Image Classification via Predictability Scores and Media Entity Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing content moderation methods are inadequate for effectively handling visual information, especially in user-generated content on platforms like YouTube and Facebook, as they rely on textual analysis and domain filtering, which are insufficient when visual content lacks text or contains misleading textual information, and are inefficient for the long tail of visual information.
Innovation Solution
A method that partitions images into media entities, generates media class descriptors, and calculates predictability scores based on probability estimations and descriptor relationships, allowing for accurate classification and moderation of visual content, including applying human body detection, face detection, and skin detection algorithms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If textual analysis and domain filtering are used for content moderation, then the solution is simple to implement, but it becomes ineffective when visual content lacks text or contains misleading textual information
Solution Approach 1:
The patent segments visual content into multiple media entities (e.g., objects, regions, features) and analyzes each entity separately using specialized detectors. This segmentation allows the system to comprehensively analyze visual content beyond simple text-based filtering, improving reliability while maintaining manageable complexity through modular processing.
Solution Approach 2:
The patent creates a unified content moderation system that handles multiple types of content (textual and visual) and multiple detection tasks (object detection, face detection, skin detection, etc.) through a single framework. This multi-functional approach ensures effectiveness across diverse content types while providing a comprehensive solution.
2Reliability
If crowd sourcing and collaborative filtering are used for visual information moderation, then human judgment can be applied, but the system becomes inefficient for the long tail of visual information
Solution Approach 1:
The patent implements automated detection algorithms (object detectors, face detectors, skin detectors, etc.) that autonomously analyze and classify visual content without requiring human intervention for each item. This self-service capability enables high-speed processing of large volumes of content, dramatically improving productivity while maintaining consistent classification standards.
Solution Approach 2:
The patent adjusts detection parameters and thresholds dynamically to optimize the balance between accuracy and processing speed. By tuning parameters such as detection confidence thresholds and processing priorities, the system can efficiently handle the long tail of visual information while maintaining reliable classification for critical content.
3Productivity
If simple image features such as skin information are used for content classification, then the processing speed is fast, but the classification accuracy is insufficient for complex visual content
Solution Approach 1:
The patent combines multiple types of image features (skin information, texture, color histograms, object detection results, face detection results) into a composite analysis framework. This composite approach integrates diverse feature types to achieve high classification accuracy for complex visual content while maintaining processing efficiency through optimized feature fusion.
Solution Approach 2:
The patent applies a hierarchical processing approach where simple features are analyzed first for quick initial classification, and more complex features are analyzed selectively for cases requiring higher accuracy. This partial application of comprehensive analysis maintains processing speed for straightforward cases while ensuring accuracy for complex content.
Data Source
AI summary
A method for determining a predictability of a media entity portion, the method includes: receiving or generating (a) reference media descriptors, and (b) probability estimations of descriptor space representatives given the reference media descriptors; wherein the descriptor space representatives are representative of a set of media entities; and calculating a predictability score of the media entity portion based on at least (a) the probability estimations of the descriptor space representatives given the reference media descriptors, and (b) relationships between the media entity portion descriptors and the descriptor space representatives. A method for processing media streams, the method may include: applying probabilistic non-parametric process on the media stream to locate media portions of interest; and generating metadata indicative of the media portions of interest.


