Generic Visual Model Unified Instance Perception
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current machine learning models designed for specific visual tasks struggle to learn generic knowledge across tasks and domains due to diverse task definitions, leading to redundant parameters and limited collaboration between tasks, making it difficult to achieve high performance in processing various visual tasks.
Innovation Solution
A generic processing model is developed that receives visual and prompt data, extracts generic representations, and determines processing results based on these representations, allowing for unified instance perception across different tasks and domains through a universal instance perception model (UNINEXT) that reformulates tasks into a unified paradigm and jointly trains on diverse datasets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If visual tasks are subdivided into multiple branches with independent models, then processing performance for specific tasks is improved, but model complexity and redundancy increase
Solution Approach 1:
The patent applies universality by designing a single generic processing model that can handle multiple visual tasks (object detection, segmentation, tracking) through unified instance perception. The model uses generic representations that are task-agnostic, allowing the same model structure to perform different functions by varying the prompt data, thereby eliminating the need for multiple specialized models and reducing overall system complexity
Solution Approach 2:
The patent merges multiple task-specific models into one unified generic processing model. By combining the capabilities of object detection, segmentation, and tracking models into a single framework that processes visual data through generic representations, the system reduces redundancy while maintaining or improving performance across all tasks
2Measurement precision
If task-specific models are designed independently, then processing performance for individual tasks is improved, but knowledge sharing across tasks is limited
Solution Approach 1:
The generic processing model learns universal instance representations that capture fundamental object characteristics applicable across different visual tasks. These generic representations serve as shared knowledge that can be adapted to various tasks through different prompt data, enabling the model to transfer learning across tasks without requiring task-specific adaptations
Solution Approach 2:
The model performs preliminary learning of generic instance representations from diverse training data before applying them to specific tasks. By pre-training on a variety of visual tasks and domains, the model acquires transferable knowledge that can be readily applied to new tasks, improving both knowledge sharing and adaptability
3Measurement precision
If multiple independent models are trained for different visual tasks, then task-specific accuracy is improved, but computational redundancy increases
Solution Approach 1:
The patent combines multiple task-specific processing pathways into a single unified processing pipeline. The generic processing model handles detection, segmentation, and tracking in one integrated framework, eliminating redundant computation across separate models while maintaining task-specific accuracy through specialized output layers and prompt configurations
Data Source
AI summary
A method, apparatus, device, and medium are provided for processing a visual task by a generic model. In a method, visual data and prompt data associated with a visual task are received, the visual task specifying that a processing result associated with the prompt data is to be determined from the visual data. A generic prompt representation of the prompt data is obtained, the prompt data including either an image format or a language expression format. A generic visual representation of the visual data is obtained, the visual data including either an image format or a video format. The processing result is determined based on the generic prompt representation and the generic visual representation. Here, different visual tasks can be processed in a unified way, training data can be shared across a plurality of visual tasks, and the processing performance of the generic processing model can be improved.


