Video Object Segmentation Using Semantic Suppression and Attention
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image segmentation methods rely solely on pixel-level information, leading to erroneous segmentations when similar objects are present, causing misidentification of target objects.
Innovation Solution
An electronic device employs a method to determine and process first and second features corresponding to a target object and other regions, using semantic suppression and attention modules to enhance and suppress respective features, respectively, thereby optimizing target object segmentation accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If only pixel-level information is used for prediction, then the computation is simple and fast, but the segmentation accuracy deteriorates when similar objects are present
Solution Approach 1:
The patent segments the feature processing into two distinct pathways: pixel-level information processing and semantic-level information processing. The semantic suppressor module segments out non-target semantic information, while the attention module focuses on target-related semantic information. This segmentation allows the system to maintain computational efficiency while improving segmentation accuracy by processing different types of information through specialized channels.
Solution Approach 2:
The patent introduces semantic-level information as an intermediary between pixel-level information and final segmentation results. The semantic suppressor and attention module act as intermediaries that process semantic information and guide the pixel-level processing. This intermediary layer resolves the contradiction by adding computational steps that improve accuracy without making the entire system prohibitively complex.
2Measurement precision
If semantic-level information is added to pixel-level information, then the segmentation accuracy improves, but the computation complexity increases
Solution Approach 1:
The patent applies local quality by making different parts of the processing system have different functions: the semantic suppressor module specifically suppresses non-target semantic information, while the attention module specifically enhances target-related semantic information. This localized specialization allows the system to process semantic information efficiently by focusing computational resources where they are most needed, rather than uniformly processing all semantic information.
3Measurement precision
If the semantic suppressor module and attention module are both applied, then the target object identification accuracy improves, but the device complexity increases
Solution Approach 1:
The patent merges the semantic suppressor module and the attention module into a unified processing framework that operates on semantic-level information. Rather than treating them as completely separate systems, the patent combines their functions within the same architectural structure, allowing them to work together synergistically. The semantic suppressor removes non-target information while the attention module enhances target information, and their combined effect is integrated into the segmentation process.
Data Source
AI summary
A method and an electronic device for performing video object segmentation are provided. The method includes determining a first feature corresponding to a target object in a first image, and/or determining a second feature corresponding to other regions other than the target object in the first image, and performing a first or second processing on a mask feature corresponding to the first image, based on the determined first or second feature. A result of performing the target object segmentation on the first image is determined based on a result of the first processing and/or a result of the second processing.


