Personalized Image Segmentation via Context-Aware Semantic Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image segmentation methods, particularly those using deep learning, are limited in classifying undefined object information during the training stage, and require specific and complex user input for accurate object detection.
Innovation Solution
A personalized image segmentation device and method that includes a user input collector, a sensing information analyzer, a user semantic information generator, a multimodal foundation model, and an image-input decoder. This system converts simple user inputs into personalized semantic information, which is then used to enhance object detection performance in image segmentation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If RIS technology is used to detect object information based on user input, then the system can detect object information not learned in advance, but the user must repeatedly input the same information to detect objects accurately
Solution Approach 1:
The system performs preliminary action by collecting sensing information (location, time, device data) and generating context information in advance. This pre-processing allows the system to automatically supplement user inputs with relevant contextual data, reducing the need for repeated user input while maintaining accurate object detection capability.
2Measurement precision
If deep learning models are trained with supervised learning for image segmentation, then the models can classify defined objects, but they are limited in classifying undefined object information
Solution Approach 1:
The system introduces context information as an intermediary element that bridges supervised learning models and undefined object detection. By generating context information from sensing data and using it to supplement user inputs, the system enables multi-modal foundation models to detect objects beyond their training distribution while maintaining classification accuracy through the mediating context layer.
Data Source
AI summary
Disclosed is a personalized image segmentation device, which includes a user input collector that outputs input information in a second format based on a user input in a first format received from a first external device, a sensing information analyzer that outputs context information based on sensing information received from a second external device, a user semantic information generator that analyzes personalized semantic information based on the input information and the context information and outputs personalized user input information based on the personalized semantic information, a multimodal foundation model that encodes image data and text information respectively to generate feature information corresponding to the image data, and an image-input decoder that detects an object corresponding to the user input on the image data based on the feature information and the personalized user input information.


