Medical Image Classification Using Vision-Language Model Alignment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional medical image sorting techniques rely on noisy metadata, leading to inaccurate sorting results, and conventional ML models are limited by their focus on image features, failing to utilize non-image based information and lacking versatility in recognizing unfamiliar medical devices or anatomical structures.
Innovation Solution
A machine learning-based system that pairs medical images with textual descriptions to form image-text pairs, using a vision-language model to predict similarities between images and text, thereby classifying images accurately and handling unfamiliar classes or devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional ML models focus only on image features, then the model structure remains simple, but the model cannot utilize non-image based information and lacks versatility in recognizing unfamiliar medical devices or anatomical structures
Solution Approach 1:
The patent merges image features and text features into a unified model architecture. The vision encoder processes medical images while the text encoder processes textual descriptions, and both are combined in a shared embedding space to enable the model to utilize both visual and linguistic information for improved recognition of unfamiliar medical devices and anatomical structures
Solution Approach 2:
The patent creates a universal model that can handle multiple types of inputs (images and text) and multiple classification tasks (medical device identification, anatomical structure recognition, imaging modality classification). The vision-language model architecture allows the same model to process different modalities and perform diverse classification functions, enhancing versatility without requiring separate specialized models
2Measurement precision
If rule-based techniques are used to process DICOM header information, then the processing method is simple, but the free text information is difficult and slow to process and contains noise leading to inaccurate sorting results
Solution Approach 1:
The patent replaces rule-based mechanical processing of DICOM headers with a machine learning-based vision-language model. Instead of using predefined rules to parse and categorize header information, the model learns to process both structured metadata and unstructured free text automatically, achieving higher accuracy in sorting while maintaining efficient processing speeds through parallel computation
Data Source
AI summary
Described herein are systems, methods, and instrumentalities associated with medical image classification and/or sorting. An apparatus may obtain a medical image from a medical image repository and further obtain multiple textual descriptions associated with a set of image classification labels. The apparatus may pair the medical image with one or more of the multiple text descriptions to obtain one or more corresponding image-text pairs, and classify the medical image based on a machine learning (ML) model and the one or more image-text pairs. The ML model may be configured to predict respective similarities between the medical image and the corresponding text descriptions in the one or more image-text pairs, and the apparatus may determine the class of the medical image by comparing the similarities predicted by the ML model. The apparatus may then further process the medical image based on the class of the medical image.


