Intrinsic Modality for Domain Generalization in Image Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning models struggle with domain generalization due to domain distribution gaps, where the probability distributions of training and testing data differ, leading to performance deterioration. Additionally, collecting images from all possible domains for optimal training is expensive and impractical.

Innovation Solution

The proposed solution involves generating an intrinsic modality, a machine learning model representation that describes image portions, acting as a substitute for text modality in multi-modal networks. This intrinsic modality is combined with a visual modality obtained from a vision-only model to enhance generalization across unseen domains.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If explicit natural language text is used to describe image portions in multi-modal networks, then classification accuracy is improved, but resource consumption increases and text availability is limited across domains

Engineering Contradiction:
Improveclassification accuracyVSAvoidresource consumption
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent creates an intrinsic modality that copies the descriptive function of natural language text but in a compressed, domain-agnostic representation. Instead of using explicit text descriptions that require storage and processing of language data, the system generates intrinsic modalities that capture essential image characteristics in a compact form, reducing resource consumption while maintaining classification accuracy across domains

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system transforms the representation parameters from natural language text to intrinsic modality vectors. This parameter change allows the system to maintain the descriptive capability needed for accurate classification while significantly reducing the resource footprint. The intrinsic modality uses optimized numerical parameters instead of linguistic parameters, enabling efficient processing and domain generalization

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If images from all possible domains are collected for training, then domain generalization is improved, but cost and practicality deteriorate

Engineering Contradiction:
Improvedomain generalizationVSAvoiddata collection cost
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent extracts the essential domain-invariant features from images by generating intrinsic modalities that capture core characteristics without domain-specific variations. This extraction process removes the need to collect data from all possible domains, as the intrinsic modality representation inherently generalizes across domains by focusing on essential features rather than domain-specific details

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The intrinsic modality serves as a universal representation that functions across all domains without requiring domain-specific training data. This single representation mechanism provides multi-domain adaptability, allowing the system to generalize to unseen domains without the prohibitive cost of collecting and annotating images from every possible domain

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If text modality is used to describe high-level features, then domain agnosticism is improved, but availability of text annotations deteriorates

Engineering Contradiction:
Improvedomain agnosticismVSAvoidtext annotation availability
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The intrinsic modality acts as an intermediary between the image data and the classification task, providing domain-agnostic descriptions without requiring explicit text annotations. This intermediary representation captures the essential information needed for domain-general classification while avoiding the availability constraints of manual text annotations, bridging the gap between visual input and semantic understanding

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12340571B2Using intrinsic multimodal features of image for domain generalized
Publication Date: 2025.06.24 ADOBE INC
  • US12340571B2 patent drawing
  • US12340571B2 patent drawing
  • US12340571B2 patent drawing

AI summary

Various embodiments classify one or more portions of an image based on deriving an “intrinsic” modality. Such intrinsic modality acts as a substitute to a “text” modality in a multi-modal network. A text modality in image processing is typically a natural language text that describes one or more portions of an image. However, explicit natural language text may not be available across one or more domains for training a multi-modal network. Accordingly, various embodiments described herein generate an intrinsic modality, which is also a description of one or more portions of an image, except that such description is not an explicit natural language description, but rather a machine learning model representation. Some embodiments additionally leverage a visual modality obtained from a vision-only model or branch, which may learn domain characteristics that are not present in the multi-modal network. Some embodiments additionally fuse or integrate the intrinsic modality with the visual modality for better generalization.