Medical Image Segmentation Using Few-Shot Embeddings
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current automated segmentation techniques for medical images struggle to generalize when encountering unfamiliar or diseased anatomies, requiring the development of new models with additional data collection and labeling, which is costly and inefficient, especially for large volumetric datasets like CT and MRI.
Innovation Solution
A system configured to segment medical images using an image encoder to extract image embeddings and compute target embeddings from a set of few-shot images, which are then input to a mask decoder to associate and output a predicted segmentation mask, reducing the need for extensive user interaction and data labeling.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If task-specific neural networks are used for automated segmentation, then segmentation accuracy for predefined anatomical targets is improved, but the model fails to generalize to unfamiliar or diseased anatomies requiring new data collection and labeling
Solution Approach 1:
The patent applies universality by designing a single foundation model that can perform multiple segmentation tasks across different anatomical structures and imaging modalities. The model is trained on diverse medical imaging data during pre-training to learn universal anatomical representations, enabling it to generalize to unseen anatomies without task-specific fine-tuning, thus resolving the contradiction between specialized accuracy and general adaptability
Solution Approach 2:
The patent utilizes parameter changes by adjusting the training regime from task-specific fine-tuning to continuous pre-training on diverse data. The model parameters are optimized through pre-training on large-scale multi-modal medical images, allowing the same model to adapt to different segmentation tasks by changing the input data distribution rather than retraining parameters, thereby achieving both accuracy and generalization
2Measurement precision
If new models are developed for unfamiliar anatomies with additional data collection and labeling, then segmentation performance for specific targets is improved, but the cost and time requirements increase significantly
Solution Approach 1:
The patent applies preliminary action through pre-training the foundation model on large-scale diverse medical imaging data before deployment. This pre-training phase performs the heavy lifting of learning universal anatomical representations in advance, so that when the model encounters new anatomies, it can adapt quickly with minimal additional data and labeling, eliminating the need for extensive data collection and labeling for each new task
Solution Approach 2:
The patent utilizes copying by leveraging the pre-trained foundation model as a reusable template that can be applied to multiple segmentation tasks. Instead of creating new models from scratch for each anatomy, the same pre-trained model is copied and applied to different tasks, requiring only minimal task-specific adaptation rather than full retraining, thus reducing time and resource requirements
3Measurement precision
If extensive labeled data is collected for training, then model performance on specific tasks is improved, but the cost and complexity of data preparation increases
Solution Approach 1:
The patent applies universality by training a single foundation model on diverse multi-modal medical imaging data during pre-training. This universal training approach allows the model to learn transferable representations that work across different anatomies and modalities, reducing the need for task-specific data preparation and simplifying the overall data pipeline while maintaining high performance across multiple tasks
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Systems, computer-implemented methods, non-transitory computer readable media (computer program products), and techniques for segmenting an input medical image. An image encoder receives the input medical image (MI) and extracts image embeddings (IE) from the input medical image (MI). The image encoder receives a set of few-shot images (FIS) and computes target embeddings (TE) from the set of few-shot images (FIS). A mask decoder receives and associates the image embeddings (IE) and the target embeddings (TE) to output a predicted segmentation mask (SM) for the input medical image (IM).