Video Context Labeling via Event Graph and Zero-Shot Embeddings
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video analytics solutions for industrial settings face challenges in efficiently scaling AI solutions across new industrial settings due to the need for manual image labeling by domain experts, which is time-consuming and costly.
Innovation Solution
The proposed solution utilizes domain context computed from natural language input by domain experts to automatically label images in videos, eliminating the need for manual data engineering. This involves generating context labels by referencing a context database and an inspected event graph, allowing for scalable video analytics deployment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual image labeling by domain experts is used to train AI models for video analytics, then the accuracy and domain-specific correctness of the AI solution is improved, but the time and cost required to deploy the solution in new industrial settings increases significantly
Solution Approach 1:
The system performs preliminary action by pre-training the image labeler on general industrial imagery and pre-computing context labels from domain expert descriptions before deployment. When deployed to new settings, the system only needs to retrieve and apply pre-computed context labels rather than performing manual annotation, thus maintaining high accuracy while dramatically reducing deployment time
Solution Approach 2:
The system creates copies of context labels derived from domain expert descriptions and stores them in a context database. These context labels are then copied and applied to multiple videos and images without requiring the domain expert to re-label each individual item, enabling scalable deployment across numerous industrial scenarios while preserving annotation accuracy
2Measurement precision
If manual image labeling by domain experts is used to train AI models for video analytics, then the domain-specific correctness of the AI solution is improved, but the cost of hiring data engineer resources increases
Solution Approach 1:
The system implements self-service by enabling automatic generation of context labels through natural language processing of domain expert descriptions. The context database automatically stores and retrieves relevant context information, eliminating the need for expensive data engineer resources to manually annotate each image or video frame, thus reducing deployment costs while maintaining domain-specific accuracy
Solution Approach 2:
The system introduces an intermediary context database that bridges domain expert knowledge and AI model training. Instead of requiring direct manual labeling by domain experts for each dataset, the context database serves as an intermediary that stores pre-computed context labels, which can be automatically retrieved and applied to training data, thereby reducing the need for expensive data engineer resources
3Adaptability or versatility
If AI solutions are customized for each new industrial setting with manual labeling, then the adaptability to specific industrial scenarios is improved, but the scalability across multiple factories and assembly lines decreases
Solution Approach 1:
The system achieves universality by creating a reusable context database that can serve multiple industrial scenarios, factories, and assembly lines. The same context labels derived from domain expert descriptions can be applied across different videos and images from various industrial settings, enabling the AI solution to adapt to multiple scenarios without requiring separate manual labeling efforts for each, thus improving scalability while maintaining scenario-specific adaptability
Data Source
AI summary
Systems and method described herein involve processing a video which can include executing a zero shot image labeler on the video to generate a plurality of labels corresponding to images of the video; calculating an image embedding vector for each of the labels of the images; generating context labels to replace the each of the labels corresponding to each of the images based on context labels determined for current and previous images in time by referencing a context database with the image embedding vector to determine the context labels from an inspected event graph, wherein nodes of the event graph are indicative of events that can happen during a duration of the video; and replacing the each of the plurality of labels with the generated context labels.


