Mask R-CNN Object Recognition for Interactive Video Ads
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video processing methods struggle with precise object identification across various camera angles and lighting conditions, leading to inconsistent object recognition, and traditional advertising models are intrusive, disrupting viewer experiences in connected content environments.
Innovation Solution
The KERVdata system uses a Mask R-CNN algorithm with linear regression to generate multilevel hierarchical data for precise pixel edge boundary identification, enabling accurate object recognition across different angles and lighting conditions, and allows interactive engagement by overlaying precise boundaries defined by vertices, empowering viewers to control their advertising experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional bounding box methods are used for object identification, then the implementation is simple, but the object recognition precision is insufficient
Solution Approach 1:
The patent segments the object identification process into multiple hierarchical levels: first identifying candidate regions using bounding boxes, then applying pixel-level mask segmentation to precisely define object boundaries, and finally extracting vertex coordinates for interaction. This multi-level segmentation approach progressively refines recognition precision while managing system complexity through staged processing.
Solution Approach 2:
The patent transitions from 2D bounding box approximations to pixel-level 2D mask segmentation, and further to vertex coordinate extraction. This dimensional refinement at each stage enables progressively more precise object identification, moving from coarse spatial approximation to exact boundary definition.
2Measurement precision
If Mask R-CNN algorithm is used for precise pixel edge boundary identification, then object recognition accuracy improves, but processing time increases
Solution Approach 1:
The patent performs preliminary object candidate identification using faster bounding box methods before applying the computationally intensive Mask R-CNN algorithm. This preliminary action filters down the search space, allowing precise pixel edge boundary identification to be applied only to relevant regions, thereby reducing overall processing time while maintaining high accuracy.
Solution Approach 2:
The processing is segmented into distinct stages: rapid bounding box candidate generation, followed by selective application of Mask R-CNN for precise mask generation on identified candidates. This segmentation allows the system to balance speed and accuracy by applying computationally expensive operations only where necessary.
3Ease of operation
If traditional advertising models are used, then advertising delivery is simple, but viewer experience deteriorates
Solution Approach 1:
The patent enables viewers to self-select and interact with advertising content by allowing them to choose which objects in the video to engage with. This self-service approach gives viewers control over their advertising experience, eliminating forced exposure while maintaining advertising delivery capability through optional interaction with identified objects.
Solution Approach 2:
The system creates a universal interactive framework that can accommodate multiple advertising delivery methods: traditional non-intrusive ads, interactive object-based ads, and viewer-controlled ad selection. This multi-functional approach allows the same system to serve different advertising strategies while improving viewer experience through choice and control.
Data Source
AI summary
A method for identifying a product which appears in a video stream. The method includes playing the video stream on a video playback device, identifying key scenes in the video stream containing product images, selecting product images identified by predetermined categories of trained neural-network object identifiers stored in training datasets. Object identifiers of identified product images are stored in a database. Edge detection and masking is then performed based on at least one of shape, color and perspective of the object identifiers. A polygon annotation of the object identifiers is created using the edge detection and masking. The polygon annotation is annotated to provide correct object identifier content, accuracy of polygon shape, title, description and URL of the object identifier for each identified product image corresponding to the stored object identifiers. Also disclosed is a method for an end user to select and interact with an identified product.


