Target-Dataset Retraining for Cross-Domain Image Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning algorithms perform sub-optimally when applied to new or diverse image datasets, as they are typically trained on global models that fail to adapt to target domains with different image characteristics, and manual annotation of large datasets is impractical due to high prediction errors and economic costs.

Innovation Solution

The method involves generating an automatically labeled training dataset by selecting images from the original dataset that are similar to the target dataset, using image feature descriptors to measure similarity, and retraining the global model with a second training dataset that includes images predicted with high confidence, thereby locally adjusting the model to the target domain.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a global model is trained using all samples from a given training dataset, then the model can learn general patterns across the entire domain, but the model performance drops significantly when applied to diverse target images that differ from the training set

Engineering Contradiction:
Improvemodel adaptability to target domainVSAvoidprediction accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent applies local quality by transitioning from a global model trained on all training data to local models trained on specifically selected subsets of training data that are similar to the target domain. This allows each local model to specialize in particular image characteristics, improving adaptability to the target domain while maintaining reliability through focused learning on relevant patterns.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent segments the global training dataset into multiple subsets based on similarity to the target domain. Instead of using a single global model, the system divides the training data and creates multiple specialized models, each trained on a specific subset. This segmentation allows the system to select the most appropriate model for the given target images, resolving the contradiction between generalizability and domain-specific performance.

Inventive Principle:
Principle #1Segmentation

2Quantity of substance

If more images are added to the training set to increase the probability of having similar images, then the model may better cover diverse scenarios, but manual annotation of huge numbers of images becomes difficult or impossible

Engineering Contradiction:
Improvenumber of training imagesVSAvoidmanual annotation feasibility
Core Design Contradiction:
Quantity of substanceVSEase of manufacture

Solution Approach 1:

The patent applies self-service by using the machine learning model itself to identify and select relevant training images based on their similarity to the target domain. The system automatically evaluates training images and selects those most useful for the specific target application, eliminating the need for manual annotation while ensuring the selected images are highly relevant. This resolves the contradiction by enabling large-scale image selection through automated self-evaluation rather than human annotation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent introduces an intermediary mechanism (an automated selection system based on image similarity metrics) that bridges the gap between the large available training dataset and the need for domain-specific training images. This intermediary automatically identifies and selects the most relevant subset of images, replacing manual annotation processes and enabling the use of large quantities of training data without human intervention.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of manufacture

If prior-art automatic labeling methods are used to extract image data from huge numbers of images, then annotation costs are reduced, but the prediction errors are high

Engineering Contradiction:
Improveannotation costVSAvoidprediction accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent applies partial action by not attempting to annotate all available training images, but instead selectively annotating only the subset of images that are most similar to and relevant for the target domain. By focusing annotation resources on the most critical images rather than attempting comprehensive annotation, the system achieves high prediction accuracy with reduced annotation costs, resolving the contradiction between cost and precision.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20220101127A1Automatic optimization of machine learning algorithms in the presence of target datasets
Publication Date: 2022.03.31 URUGUS SA
  • US20220101127A1 patent drawing
  • US20220101127A1 patent drawing
  • US20220101127A1 patent drawing

AI summary

Methods, systems and computer program products for transferring knowledge using machine learning techniques by automatically generating training datasets are provided. New training datasets based on target datasets are automatically generated and used in machine learning techniques to perform tasks on images. One of the main benefits is the possibility to transfer the knowledge learned in one domain to another domain in which extracting data or labeling images would be costly or simply infeasible. The methods and systems also provide image training sets based on image target sets which augments data in a more efficient way and improves the content of the training set and the prediction of the machine learning techniques.