Multi-Task Vision Model Annotation via Specialist Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computer vision models face challenges in capturing comprehensive visual annotations at scale due to limited training data and the absence of a unified network architecture capable of addressing multiple tasks simultaneously.

Innovation Solution

A method for annotating images using multiple annotation specialist models and a data filtering and enhancement module to create a large-scale, comprehensive dataset (FLD-5B) for training a unified multi-task computer vision machine learning model (ALS-CV-MLM).

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple specialist models are used to generate annotations, then annotation quality and comprehensiveness are improved, but system complexity increases

Engineering Contradiction:
Improveannotation qualityVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The annotation system is divided into multiple specialist models, each responsible for specific annotation tasks. This segmentation allows each model to excel at its specialized function while maintaining overall system quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The framework integrates multiple specialist models into a unified annotation system that handles various annotation types simultaneously, creating a multi-functional system that improves comprehensiveness while managing complexity through modular design.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Quantity of substance

If comprehensive visual annotations are created at scale, then model training data quality is improved, but data processing time increases

Engineering Contradiction:
Improvetraining data volumeVSAvoidannotation time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

Annotations are generated in advance using automated specialist models before training data compilation. This preliminary action allows comprehensive annotations to be created at scale without delaying the training process.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Manual annotation processes are replaced with automated machine learning models. This substitution dramatically reduces annotation time while maintaining or improving annotation quality, enabling comprehensive data creation at large scale.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Adaptability or versatility

If a unified network architecture is implemented for multiple tasks, then model versatility is improved, but architectural complexity increases

Engineering Contradiction:
Improvetask versatilityVSAvoidarchitecture complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

A unified network architecture is designed to handle multiple computer vision tasks through shared representations and task-agnostic feature extraction. This universal architecture achieves versatility while managing complexity through modular task adapters.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The unified architecture incorporates dynamic task-specific adapters that can be activated or deactivated based on the required task. This dynamic approach allows the system to maintain versatility while keeping the base architecture relatively simple and manageable.

Inventive Principle:
Principle #15Dynamics

4Measurement precision

If automated annotation filtering is applied, then annotation accuracy is improved, but processing complexity increases

Engineering Contradiction:
Improveannotation accuracyVSAvoidfiltering complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

Manual quality review and filtering processes are replaced with automated machine learning-based filtering systems. These systems use learned patterns to identify and remove low-quality annotations automatically, improving accuracy while managing complexity through algorithmic approaches.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The annotation filtering system incorporates feedback loops where model predictions are continuously refined based on performance metrics and quality assessments. This feedback mechanism improves filtering accuracy over time while the system learns to handle complexity more efficiently.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250148765A1Annotating images for training computer vision models
Publication Date: 2025.05.08 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20250148765A1 patent drawing
  • US20250148765A1 patent drawing
  • US20250148765A1 patent drawing

AI summary

A method for annotating images to create a corpus for training a multi-task computer vision machine learning model is presented. The method comprises receiving, at one or more annotation specialist models, a plurality of images to be annotated. Via operation of the one or more annotation specialist models, pre-filtered annotations are generated for the plurality of images. Via operation of a data filtering and enhancement module, the pre-filtered annotations are filtered in accordance with predefined noise criteria so as to output candidate annotations for the plurality of images. The method further comprises, for each of one or more candidate annotations, selectively (1) storing the candidate annotation into the corpus as a final annotation for its associated image, or (2) adding the candidate annotation to its associated image using the one or more annotation specialist models and the data filtering and enhancement module for subsequent iterative annotation and filtering.