Mobile Visual Media Tagging with Threshold-Based Bulk Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current techniques for tagging visual media, such as photos and videos, are time-consuming and cumbersome, especially when dealing with multiple persons or objects across many images, as they require manual selection and tagging, which can be inefficient.
Innovation Solution
The implementation of a system on a mobile device that uses a tagging module to determine when a manual tagging threshold is met, allowing for bulk tagging of visual media by performing facial or object recognition on multiple images, enabling quick confirmation or rejection of recognitions, and utilizing both manual and non-manual tags to improve accuracy and efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual tagging is performed for each photo individually, then tagging accuracy can be confirmed by user selection, but the time and effort required increases significantly
Solution Approach 1:
The patent segments the tagging process into two distinct phases: (1) manual tagging of a small subset of photos to establish ground truth and train the recognition system, and (2) automated bulk tagging of remaining photos using the trained system. This segmentation allows the user to invest time only in the critical accuracy-determining phase while accepting automated results for the bulk, resolving the contradiction between accuracy and time consumption.
Solution Approach 2:
The system performs preliminary manual tagging on a subset of photos before proceeding to automated tagging. This preliminary action establishes the foundation for accurate automated recognition by training the system on verified examples, thereby enabling high accuracy in the subsequent automated phase without requiring manual review of every single photo.
2Productivity
If automated facial recognition is performed on all photos, then tagging speed increases, but accuracy may decrease due to false recognitions
Solution Approach 1:
The system implements feedback by using manually tagged photos as ground truth to train and validate the automated recognition system. The manually tagged subset provides feedback signals that allow the system to learn correct recognition patterns, thereby improving the accuracy of automated tagging while maintaining high speed processing for the bulk of photos.
Solution Approach 2:
Instead of performing complete manual tagging on all photos (excessive action), the system performs partial manual tagging on a strategically selected subset. This partial action is sufficient to train the automated system effectively, achieving near-complete accuracy at much lower time cost than full manual tagging would require.
3Productivity
If bulk tagging is performed without manual threshold verification, then processing efficiency is maximized, but user confidence in results decreases
Solution Approach 1:
The system enables self-service by allowing users to define their own tagging thresholds and confidence levels. Users can automatically adjust the manual tagging threshold based on their desired balance between speed and accuracy, making the system adapt to individual user preferences without requiring constant manual intervention or verification of each tagged photo.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
This document describes techniques enabling tagging of visual media on a mobile device. In some cases the techniques determine, based on meeting a threshold of manual tagging of a person or object, to "bulk" tag visual media stored on the mobile device. Thus, the techniques can present, in rapid succession, photos and videos with the recognized person or object to enable the user to quickly and easily confirm or reject the recognition. Also, the techniques can present numerous faces for recognized persons or sub-images for recognized objects on a display at one time, thereby enabling quick and easy confirmation or rejection of the recognitions.