Real-Time Video Object Detection Using Machine Learning Detectors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional search engines are limited in their ability to accurately search and identify objects within video content, as text-based searches cannot accurately describe video content, and they struggle to process the vast amount of digital media created and shared online, making it difficult to find desired content in real-time.
Innovation Solution
A system and method that employs machine-learning detectors, such as convolutional neural networks, to identify objects, faces, or other visual features in media content, allowing users to search for specific items in videos or frames, with the ability to detect and generate confidence scores for the presence of these items, and provide user interfaces for reviewing and retraining detectors.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If text-based search engines are used to search video content, then the search process is simple and fast, but the ability to accurately identify objects within video content is limited
Solution Approach 1:
The patent introduces machine learning detectors as an intermediary between the search engine and video content. These detectors are trained to recognize specific objects, faces, or visual features, enabling accurate object identification within video content while maintaining a relatively simple search interface for users.
Solution Approach 2:
The system performs preliminary training of machine learning detectors using labeled video data before the actual search process. This preliminary action creates pre-trained models that can quickly and accurately identify objects during runtime, avoiding the need for complex real-time analysis during user searches.
2Productivity
If manual review by curators is used to process video content, then the accuracy of content identification is high, but the processing speed is too slow to handle the flood of digital media
Solution Approach 1:
The patent replaces the mechanical manual review process with automated machine learning detectors. These detectors process video content computationally, achieving both high speed (handling large volumes of media) and high accuracy (through trained recognition models), thus resolving the contradiction between productivity and precision.
Solution Approach 2:
The machine learning detectors are designed to autonomously analyze and identify objects in video content without human intervention during the search process. The system self-corrects and improves through feedback mechanisms, maintaining high accuracy while achieving automated high-speed processing.
3Measurement precision
If machine-learning detectors are trained using extensive example media content items, then the detection accuracy improves, but the training time and computational resources increase
Solution Approach 1:
The system performs detector training in advance during a preliminary phase, creating pre-trained machine learning models before deployment. This allows the detectors to be ready for immediate use with high accuracy, avoiding time-consuming training during actual search operations.
Solution Approach 2:
The system incorporates feedback mechanisms where detection results are evaluated and used to retrain or fine-tune detectors. This feedback loop progressively improves detection accuracy over time without requiring complete retraining, reducing the time loss associated with achieving high precision.
Data Source
AI summary
Described herein are systems and methods that search videos and other media content to identify items, objects, faces, or other entities within the media content. Detectors identify objects within media content by, for instance, detecting a predetermined set of visual features corresponding to the objects. Detectors configured to identify an object can be trained using a machine learned model (e.g., a convolutional neural network) as applied to a set of example media content items that include the object. The systems comprise an integrated detection unit configured to record media content, identify preferred content, and communicate the identifications of preferred content for storage in a computationally efficient manner.


