Self-Supervised Transformer Pre-Training for Medical Image Domain Adaptation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning models, particularly convolutional neural networks (CNNs) face challenges in medical imaging due to overfitting and the need for costly, specialty-oriented expertise in annotation, while vision transformers (ViTs) struggle with data-hungry requirements and domain gaps between photographic and medical images, lacking effective self-supervised learning techniques.

Innovation Solution

Implement self-supervised domain-adaptive pre-training via transformers using a large-scale, unlabeled in-domain dataset to bridge the domain gap between photographic and medical images, leveraging masked image modeling (MIM) and continual pre-training to enhance model performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If self-supervised domain-adaptive pre-training is implemented via transformers using large-scale unlabeled in-domain data, then model performance and generalizability are improved, but computational resources and training time are increased

Engineering Contradiction:
Improvemodel performanceVSAvoidcomputational resources
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent applies preliminary action by implementing self-supervised domain-adaptive pre-training before the main classification task. The transformer model is pre-trained on large-scale unlabeled in-domain medical images to learn robust domain-specific representations, which are then transferred to downstream classification tasks. This preliminary training phase enables the model to adapt to the specific domain (e.g., chest x-rays) before facing the actual classification problem, improving final performance while being more efficient than training from scratch on labeled data only.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements self-service through self-supervised learning mechanisms where the model learns from unlabeled data without requiring manual annotations. The model generates its own training signals through pretext tasks such as masked image modeling or contrastive learning, enabling it to autonomously learn domain-specific features from abundant unlabeled medical images. This eliminates the need for expensive expert annotation while still achieving domain adaptation.

Inventive Principle:
Principle #25Self-service

2Quantity of substance

If vision transformers are used for medical image classification, then data utilization is improved, but domain gap between photographic and medical images persists

Engineering Contradiction:
Improvedata utilizationVSAvoiddomain adaptability
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent applies parameter changes by modifying the training regime and data characteristics. Specifically, it changes the data parameter from labeled to unlabeled (self-supervised), changes the domain parameter from generic photographic images to domain-specific medical images, and changes the learning objective parameter through domain-adaptive pre-training tasks. These parameter changes enable the transformer to better utilize medical imaging data while reducing the domain gap through continuous pre-training on domain-specific unlabeled images.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If convolutional neural networks are used for medical imaging, then annotation efficiency is improved, but overfitting occurs on small labeled datasets

Engineering Contradiction:
Improveannotation efficiencyVSAvoidmodel generalization
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies preliminary action by performing self-supervised pre-training on large-scale unlabeled medical images before the supervised fine-tuning phase. This preliminary training on abundant unlabeled data enables the model to learn robust domain-specific features and representations, which significantly reduces overfitting when subsequently trained on small labeled datasets. The pre-trained transformer serves as a better initialization point, improving generalization performance compared to training CNNs from scratch or even with supervised pre-training.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12394186B2Systems, methods, and apparatuses for implementing self-supervised domain-adaptive pre-training via a transformer for use with medical image classification
Publication Date: 2025.08.19 THE ARIZONA BOARD OF REGENTS ON BEHALF OF THE UNIV OF ARIZONA
  • US12394186B2 patent drawing
  • US12394186B2 patent drawing
  • US12394186B2 patent drawing

AI summary

Described herein are systems, methods, and apparatuses for implementing self-supervised domain-adaptive pre-training via a transformer for use with medical image classification in the context of medical image analysis. An exemplary system includes means for receiving a first set of training data having non-medical photographic images; receiving a second set of training data with medical images; pre-training an AI model on the first set of training data with the non-medical photographic images; performing domain-adaptive pre-training of the AI model via self-supervised learning operations using the second set of training data having the medical images; generating a trained domain-adapted AI model by fine-tuning the AI model against the targeted medical diagnosis task using the second set of training data having the medical images; outputting the trained domain-adapted AI model; and executing the trained domain-adapted AI model to generate a predicted medical diagnosis from an input image not present within the training data.