Video Object Search Using Machine Learning Detectors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional search engines are limited in their ability to effectively search for and identify objects within video content, as text-based searches cannot accurately describe visual content, making it difficult to find specific media items among the vast amount of digital media created and shared online.

Innovation Solution

The system employs machine-learning detectors, such as convolutional neural networks, to identify objects within media content items, allowing users to search for and pinpoint specific objects, faces, or actions within videos or frames, and provides user interfaces for configuring and updating these detectors based on search results and user feedback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If text-based search engines are used to search video content, then the search process is simple and fast, but the accuracy of identifying visual content is insufficient

Engineering Contradiction:
Improveaccuracy of identifying visual contentVSAvoidcomplexity of search system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces machine learning detectors as an intermediary between the user's text-based search query and the video content. These detectors are trained to recognize specific visual features (objects, faces, actions) and serve as a bridge that translates text-based search intentions into visual content identification, thereby improving accuracy while maintaining relative system simplicity

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces traditional text-matching mechanisms with machine learning-based visual recognition systems. Instead of relying on text descriptors and string matching, the system uses trained detectors that process video frames to identify visual content, substituting mechanical text-based search with intelligent visual analysis

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If machine learning detectors are employed to identify objects in video content, then the precision of search results is improved, but the processing time and computational resources increase

Engineering Contradiction:
Improveprecision of search resultsVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements preliminary action by pre-training machine learning detectors on large datasets of example media content items before actual search operations. This preprocessing step creates ready-to-use detection models that can quickly identify objects during runtime, reducing processing time during actual searches while maintaining high precision

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system applies partial action by using detectors to analyze only specific regions or features of video content that are relevant to the search query, rather than processing entire videos frame-by-frame. This selective approach reduces computational overhead and processing time while maintaining search precision

Inventive Principle:
Principle #16Partial or excessive action

3Loss of information

If conventional text descriptors are used to describe media content, then the storage requirements are low, but the ability to accurately represent visual content is limited

Engineering Contradiction:
Improveability to represent visual contentVSAvoiddata storage requirements
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The patent uses copying by creating visual representations (detectors) that capture essential features of objects in media content. Instead of storing entire videos or high-resolution images, the system stores trained detector models that can reproduce and identify visual content, efficiently representing visual information with reduced storage requirements

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11972099B2Machine learning in video classification with playback highlighting
Publication Date: 2024.04.30 MATROID INC
  • US11972099B2 patent drawing
  • US11972099B2 patent drawing
  • US11972099B2 patent drawing

AI summary

Described herein are systems and methods that search videos and other media content to identify items, objects, faces, or other entities within the media content. Detectors identify objects within media content by, for instance, detecting a predetermined set of visual features corresponding to the objects. Detectors configured to identify an object can be trained using a machine learned model (e.g., a convolutional neural network) as applied to a set of example media content items that include the object. The systems provide user interfaces that allow users to review search results, pinpoint relevant portions of media content items where the identified objects are determined to be present, review detector performance and retrain detectors, providing search result feedback, and/or reviewing video monitoring results and analytics.