Adapted Object Detector for Non-Human Face Tagging
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing human face detectors are ineffective in detecting non-human faces with 'human-like' features and other regions of interest in computer-generated media content, limiting the ability to enrich multimedia content for indexing, organization, and searching.
Innovation Solution
An adapted object detector is generated by retraining an existing visual object detector using a training data set of media content items depicting the target object class, allowing for the detection of non-human faces and other regions of interest in a scalable manner, even in computer-generated content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If an existing human face detector is used to detect non-human faces with human-like features, then the detector can identify some human-like structures, but the detection accuracy and effectiveness deteriorate significantly
Solution Approach 1:
The patent applies parameter changes by retraining the object detector with modified parameters specific to non-human faces. The detector's internal parameters (weights, thresholds, feature importance) are adjusted through training on non-human face data, enabling it to detect species-specific features while maintaining human-like structure detection capabilities
Solution Approach 2:
The patent implements preliminary action by pre-training or fine-tuning the object detector on a dataset of non-human faces before deployment. This preliminary training phase prepares the detector to recognize species-specific features, so when it encounters non-human faces in media content, it can accurately detect them without requiring real-time adaptation
2Productivity
If existing object detectors are used for computer-generated content, then standard objects can be detected, but detection of non-human characters and costumed people deteriorates
Solution Approach 1:
The patent applies universality by creating a multi-functional object detector that can handle diverse object classes including non-human faces, costumed people, and standard objects. The detector is trained on a comprehensive dataset spanning multiple categories, enabling it to perform multiple detection functions with a single system, thereby maintaining both productivity and reliability across different content types
3Measurement precision
If manual tagging of regions of interest is performed, then high accuracy tagging can be achieved, but the processing time and scalability deteriorate
Solution Approach 1:
The patent implements self-service by enabling the object detector to automatically identify and tag regions of interest without human intervention. The detector autonomously processes media content, detects non-human faces and other objects, and generates tags, thereby achieving both high accuracy and scalability simultaneously by eliminating the manual tagging bottleneck
Data Source
AI summary
Disclosed herein are a system, method and architecture for media content enrichment. A visual object detector is trained using a training data set and an existing visual object detector. The newly-adapted visual object detector may be used to detect a visual object belonging to a class of visual object. The existing object detector that is used to train the adapted object detector detects a class of visual objects different from the visual object class detected by the adapted object detector. A media content item depicting a visual object detected using the adapted object detector may be associated with metadata, tag or other information about the detected visual object to enrich the media content item.


