Unsupervised domain adaptive medical image segmentation method based on multi-view alignment and pseudo tag optimization

By using a style transfer module with frequency domain smooth fusion and a multi-view prototype contrast learning framework, the domain offset problem in unsupervised domain adaptation medical image segmentation is solved, the quality of pseudo-labels and feature alignment accuracy are improved, and higher segmentation accuracy and boundary recognition capabilities are achieved.

CN120953610AActive Publication Date: 2025-11-14XIDIAN UNIV

Patent Information

Application Number
CN202511070025.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-14
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

Existing unsupervised domain-adaptive medical image segmentation methods suffer from domain offset in the target domain, leading to decreased model performance, poor pseudo-label quality, and affecting segmentation results.

Method used

A style transfer module with frequency domain smooth fusion is used for image-level alignment. Combined with a two-stage pseudo-label optimization and a multi-view prototype contrast learning framework, the quality of pseudo-labels and the accuracy of feature alignment are improved.

Benefits of technology

It significantly improves the accuracy and boundary precision of medical image segmentation, especially in scenarios where the foreground and background are highly similar, enabling better identification of anatomical details and boundaries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953610A_ABST
    Figure CN120953610A_ABST
Patent Text Reader

Abstract

The invention discloses an unsupervised domain adaptive medical image segmentation method based on multi-view alignment and pseudo label optimization, and aims to solve the problems of insufficient segmentation precision and low pseudo label quality caused by domain offset. According to the technical scheme, firstly, image level alignment is executed through a frequency domain smooth fusion module, and a class target domain image is generated; pre-training a segmentation network by using the image and generating an initial pseudo tag; then, through a two-stage optimization process, the integrity and the structural rationality of the pseudo tag are improved through prototype-based potential foreground completion and SAM-based structural perception enhancement in the process; and finally, on the basis of the optimized high-quality pseudo tag, constructing a multi-view prototype contrast learning framework to carry out final feature level alignment training. According to the method, the segmentation precision of the model on the label-free target domain is improved, and an effective scheme is provided for solving the challenge of scarcity of annotation data in medical image segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence technology, and specifically relates to an unsupervised domain-adaptive medical image segmentation method based on multi-view alignment and pseudo-label optimization, which can be used for automatic segmentation of unlabeled target domain medical images. Background Technology

[0002] Medical image segmentation is a crucial task in medical image processing, aiming to accurately distinguish different tissues or structures within images to improve the accuracy and efficiency of clinical diagnosis. In recent years, deep learning-based methods have demonstrated good performance in medical image segmentation. However, deep learning models heavily rely on large-scale, high-quality pixel-level labeled data. Obtaining such labeled data in the medical field presents significant challenges: due to the specialized nature of medical image data, labeling typically requires expertise in the field, resulting in a scarcity of labeled medical image datasets. Under these circumstances, supervised learning-based medical image segmentation methods struggle to achieve satisfactory segmentation results. To address this issue, numerous studies in recent years have focused on exploring ways to improve segmentation performance under conditions of unlabeled data.

[0003] Unsupervised domain adaptation (UDA) technology utilizes labeled source domain data to adapt the model to unlabeled target domain data, reducing reliance on large-scale labeled datasets and enabling knowledge transfer between source and target domains with different distributions. However, achieving effective domain adaptation faces a core challenge—domain shift. Due to differences in imaging equipment, scanning parameters, patient groups, and other factors, medical images in the source and target domains differ significantly in visual style and data distribution, leading to a sharp decline in performance when the model is directly transferred.

[0004] Existing UDA methods primarily attempt to address the domain offset problem from three aspects. First, they employ image-to-image transformation techniques to convert the style of the source domain image to the target domain style, making it visually closer to the target data. However, these methods may introduce unnatural artifacts or disrupt fine anatomical structures crucial for segmentation, leading to the loss of semantic information. Second, they force the model to learn domain-invariant feature representations by measuring differences between features in the feature space. However, these methods often perform global macro-alignment, easily ignoring subtle differences between categories, especially when foreground and background features have high similarity, resulting in poor alignment. Third, they utilize the model to generate pseudo-labels for the target domain image and use these pseudo-labels for supervised training. However, due to domain offset, the initial pseudo-labels generated by the model in the target domain are usually of poor quality, exhibiting problems such as missed foreground targets and incomplete structures. During subsequent training, these erroneous supervisory signals are continuously learned and accumulated by the model, ultimately affecting the model's performance ceiling. Summary of the Invention

[0005] To address the problems of existing technologies, this invention provides an unsupervised domain-adaptive medical image segmentation method based on multi-view alignment and pseudo-label optimization, which utilizes an improved image and feature two-layer alignment strategy and a two-stage pseudo-label optimization method to improve the segmentation accuracy of medical images in the target domain.

[0006] To achieve the above objectives, the technical solution adopted by the present invention includes the following steps:

[0007] An unsupervised domain-adaptive medical image segmentation method based on multi-view alignment and pseudo-label optimization is proposed for segmenting unlabeled target domain medical images, comprising the following steps:

[0008] S1: Perform image-level alignment process.

[0009] The style transfer module, which uses frequency domain smooth fusion, transforms the source domain image into a target domain-like image, achieving image-level alignment between the source and target domains.

[0010] S2: Use the target domain image for preliminary model training and prediction.

[0011] The target domain image is input into the image segmentation network and pre-trained using the source domain real labels as supervision; after pre-training, the real target domain image is input and the initial pseudo labels are output.

[0012] S3: Perform a two-stage pseudo-label optimization process.

[0013] The initial pseudo-labels are processed through a two-stage pseudo-label optimization process to improve their structural integrity and boundary accuracy, resulting in optimized pseudo-labels.

[0014] S4: Perform feature-level alignment.

[0015] Based on the real labels of the source domain and the optimized pseudo labels, a multi-view prototype contrastive learning framework is constructed, and the image segmentation network is trained using this framework. The trained image segmentation network is then used for medical image segmentation.

[0016] Compared with the prior art, the beneficial effects of the present invention are:

[0017] This invention addresses the lack of high-quality labeled data for medical images in the target domain by providing an unsupervised domain-adaptive medical image segmentation method based on multi-view alignment and pseudo-label optimization. First, this method solves the domain offset problem through a multi-view collaborative alignment strategy. At the image level, a style transfer module with smooth frequency-domain fusion is employed. By smoothly weighting and fusing low-frequency components in the frequency domain, it effectively transforms the image style while avoiding high-frequency artifacts caused by hard replacement, thus preserving the original image structure and semantic details crucial for segmentation to the greatest extent. At the feature level, a multi-view prototype contrastive learning framework containing independent boundary prototypes is constructed. This not only achieves intra-class compactness and inter-class separation of the main body region but also forces the network to accurately learn and distinguish blurred tissue boundaries through explicit modeling of boundary regions, significantly improving the accuracy of segmentation results at the fine structure level.

[0018] This invention designs a two-stage pseudo-label optimization process, from completion to enhancement, to improve the quality and structural rationality of pseudo-labels. In the first stage, a latent foreground completion mechanism based on reconstruction is used to back-discover regions misclassified as background using prototype-guided methods. This method proactively addresses the common problem of missed detections in pseudo-labels, significantly improving pseudo-label integrity. In the second stage, SAM is introduced as a structure-aware enhancement branch, leveraging its powerful zero-shot segmentation capabilities to perform structural correction on the completed pseudo-labels. This effectively corrects common problems in pseudo-labels such as internal holes and boundary breaks, ensuring that pseudo-labels are not only pixel-accurate but also possess anatomically sound structural rationality. Attached Figure Description

[0019] Figure 1 This is a flowchart illustrating the implementation of the present invention.

[0020] Figure 2 This is a schematic diagram of the structure of the potential prospect completion module based on reconstruction, which is the first stage of constructing the pseudo-label optimization strategy in this invention.

[0021] Figure 3 This is a schematic diagram of the structure of the second stage of constructing the pseudo-label optimization strategy in this invention, which is based on the structure-aware enhancement strategy.

[0022] Figure 4 This is a schematic diagram of the multi-view prototype comparison learning framework constructed in this invention.

[0023] Figure 5 This is a visualization of the segmentation results on the ETIS-adapted CVC-EndoScene task.

[0024] Figure 6 This is a segmentation visualization result on the ETIS adapted to the Kvasir task. Detailed Implementation

[0025] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings.

[0026] The method and process are as follows Figure 1 As shown, the implementation steps of this invention are as follows:

[0027] Step 1: Perform image-level alignment between the source domain image and the target domain image.

[0028] In unsupervised domain-adaptive medical image segmentation, the core challenge lies in the domain shift problem caused by factors such as imaging equipment and scanning parameters. This domain shift is not a single-level problem, but exists simultaneously in the low-level visual wind and the high-level feature table of the image. Due to the influence of domain shift, the performance of a model trained on the source domain will drop sharply when directly applied to the target domain.

[0029] To mitigate domain shift, several methods have been developed to transfer image styles using image transformation techniques. However, these methods, while pursuing global distribution alignment, often neglect crucial anatomical structural information in medical images. Specifically, image-level style transfer processes are prone to introducing artifacts, disrupting tissue boundary integrity, and losing fine structures, resulting in segmentation results with blurred boundaries and structural discontinuities in the target domain. Therefore, this invention proposes a frequency domain smooth fusion style transfer module that converts the source domain image into a target domain-like image, achieving image-level alignment between the source and target domains, thereby stably eliminating the visual differences between the two domains. Specifically, this includes:

[0030] (1a) Source domain image I s and target domain image I t The source domain image I is obtained by converting from the spatial domain to the frequency domain using Fast Fourier Transform. s Amplitude spectrum A s and the corresponding phase spectrum P s and target domain image I t Amplitude spectrum A t and the corresponding phase spectrum P t .

[0031] (1b) Design a center-weighted smooth mask M, whose weights decrease smoothly from the low-frequency center to the high-frequency region. The calculation formula is as follows:

[0032]

[0033] Where d(i,j) is the distance from point (i,j) to the center of the mask (c). x ,c y The Euclidean distance of ) d max To start from the center of the mask (c x ,c yThe distance from the farthest point of the mask is denoted as cos(·), where cos(·) is the cosine function.

[0034] (1c) Using the smoothing mask M to smooth the source domain image I s and target domain image I t The amplitude spectra are weighted and fused to generate a new amplitude spectrum:

[0035] A new =M⊙A t +(1-M)⊙A s

[0036] (1d) The new amplitude spectrum A new With source domain image I s Phase spectrum P s By combining these images and mapping them back to the image space using inverse Fourier transform, we obtain the final source image resembling the target, i.e., the target-like domain image I. s→t The formula is as follows:

[0037]

[0038] in, This represents the inverse Fourier transform.

[0039] Step 2: Use the target domain image for preliminary model training and prediction.

[0040] The target domain image and its corresponding source domain ground truth label are fed into the image segmentation network. Pre-training is performed using the source domain ground truth labels as supervision, and supervised by weighted cross-entropy loss and weighted intersection-union (IUU) loss. This combination of loss functions aims to simultaneously optimize pixel-level classification accuracy and the overlap of segmented regions in the overall structure, which is particularly effective in handling the common foreground-background class imbalance problem in medical images. Since the target domain image is already highly similar in style to the real target domain, this pre-training process allows the network to initially adapt to the visual characteristics of the target domain. After pre-training, real target domain images are input into the network for forward inference, and the output becomes the initial pseudo-label for the target domain.

[0041] The implementation of the image segmentation network and supervised loss in this step includes:

[0042] (2a) The encoder module consists of several convolutional blocks. The last convolutional block contains only convolution operations, while each of the remaining convolutional blocks contains both convolution operations and max pooling operations to progressively extract deep semantic features of the image.

[0043] For example, in this embodiment, the encoder module consists of 4 convolutional blocks. The last convolutional block contains only 2 consecutive 3×3 convolutional operations, and each of the remaining convolutional blocks contains 2 consecutive 3×3 convolutional operations and 1 2×2 max pooling operation.

[0044] Each time a convolutional block is passed, the feature map size is halved and the number of channels doubles; the target domain image I s→t and target domain image I t After inputting into the encoder, the output is the high-level semantic feature f of the target domain. s and target domain high-level semantic features f t .

[0045] (2b) The decoder receives the high-level semantic features output by the encoder. Its structure is symmetrical with the encoder and contains several deconvolution blocks (4 in this embodiment). Each deconvolution block doubles the size of the feature map through deconvolution operation (2×2 deconvolution in this embodiment) and fuses it with the feature map of the corresponding level of the encoder through skip connections to supplement the low-level detailed information.

[0046] (2c) The output layer receives the feature map output from the decoder and uses a 1×1 convolution operation to map its channel dimensions to the number of task-related classes. Finally, the class probability of each pixel is output through the softmax function; where the class probability of the target domain image is used as the prediction result y. s The class probability of the target domain image is used as the initial pseudo-label y. t .

[0047] (2d) The loss functions in the pre-training phase are weighted cross-entropy loss and weighted intersection-union ratio loss:

[0048]

[0049] in, This represents the weighted cross-entropy loss. This represents the weighted intersection-union loss, where N represents the number of pixels, w is the class weight, and y is the weight of the class. i p represents the value of the i-th pixel of the actual label. i The value of the i-th pixel represents the predicted probability.

[0050] Step 3: Perform a two-stage pseudo-label optimization process.

[0051] Since the domain offset problem has not been fully resolved at the feature level, the initial pseudo-labels inevitably have quality issues, particularly in the missed detections of foreground regions and the loss of details in anatomical structures. Directly using these low-quality pseudo-labels for subsequent training will lead to the accumulation of erroneous signals. Therefore, before performing the final feature alignment training, these initial pseudo-labels must be refined and optimized to improve their structural integrity and boundary accuracy, resulting in optimized pseudo-labels.

[0052] This invention proposes a two-stage pseudo-label optimization process. The first stage is based on reconstruction-based potential prospect completion, such as... Figure 2As shown. This stage aims to address the issue of missed detections by pseudo-labels. Its core idea is to utilize existing, relatively reliable foreground prototype information from the model to retrospectively examine regions initially judged as background, thereby uncovering potential foregrounds that may be misclassified. The second stage is structure-aware enhancement. After completion, the completeness of pseudo-labels is improved, but their boundaries may still be imprecise, and the overall structure may be unreasonable. To address this, a pre-trained SAM model is introduced as a structure-aware branch. Leveraging SAM's powerful zero-shot generalization ability, segmentation hypotheses with fine boundaries and complete structures are generated, such as... Figure 3 As shown, the high-fidelity results that align with the network's own predictions are filtered and retained by the consistency verification module, and used as structured supervision signals to ultimately obtain the optimized pseudo-labels.

[0053] The first stage of the two-stage optimization strategy in S3 is potential prospect completion based on reconstruction. This stage includes:

[0054] (31a) Initial pseudo-label y in the target domain t Under the guidance of step (2a), combined with the target domain high-level semantic features f t Calculate the centroid of the class as the initial prototype.

[0055]

[0056] in, Let H be the total number of pixels of class c in the target domain, H be the feature map height, W be the feature map width, and f be the number of pixels of class c in the target domain. t (i) represents the high-level semantic features f of the target domain. t The value of the i-th pixel, y t (i) represents the initial pseudo-label y t The value of the i-th pixel.

[0057] (31b) Initial pseudo-label y based on the output of the image segmentation network t Construct a background boot mask M bg The mask is defined as the region initially identified by the localization model as the background. bg =1-argmax(y t ).

[0058] (31c) Set the background guide mask M bg With the high-level semantic features f of the target domain t Feature fusion is performed to obtain a background-weighted feature representation.

[0059]

[0060] in, This is for element-wise multiplication.

[0061] (31d) calculation Local features and various prototypes The similarity between pixels is used to identify potential foreground features. When the similarity between a pixel feature and a foreground prototype exceeds a preset threshold, that pixel is identified as a potential foreground. This process can be formulated as follows:

[0062]

[0063] Where sim(·) is used to calculate cosine similarity, β is a preset threshold, and C is the total number of categories. This is the result of classifying all pixels.

[0064] (31e) Determine the category of all pixels These are integrated into the initial pseudo-labels to correct misclassified foreground elements in the original segmentation results, resulting in coarsely optimized pseudo-labels. as follows:

[0065]

[0066] Here, ∪ represents the union of the two masks.

[0067] Furthermore, the second stage of the two-stage optimization strategy in S3 is a structure-aware enhancement strategy, which includes:

[0068] (32a) Introduce a pre-trained Segment Anything Model (SAM) as a structure-aware branch.

[0069] (32b) From the coarse-optimized pseudo-labels after foreground completion In the process, a grid sparse sampling strategy is used to extract the point cue set S, as follows:

[0070] Divide the image into a k×k grid. Within each grid, select the center of the grid as either a foreground or background point based on the proportion of foreground pixels. Summarize all sampled points to form a cue set.

[0071] S = {(i,j,l)|l∈{0,1}}

[0072] Where (i,j) represents the pixel coordinates within the foreground region; l=1 indicates that the coordinates are foreground points, and l=0 indicates that the coordinates are background points.

[0073] (32c) Guided by the point cue set S, the SAM model generates a segmentation prediction map M with structural integrity. sam .

[0074] (32d) The consistency verification module filters the segmentation hypothesis, retaining the pseudo-labels from the coarse optimization. The region where consensus is reached is used to obtain the optimized pseudo-label:

[0075]

[0076] In the formula, M is the optimized pseudo-tag. sam (i) is the segmentation prediction map M sam The value of the i-th pixel, To coarsely optimize pseudo tags The value of the i-th pixel.

[0077] (32e) Pseudo-labels optimized by introducing structural loss constraints Structure and consistency with segmentation network predictions:

[0078]

[0079] in, Let SSIM(·,·) be the loss function, where SSIM(·,·) is the structural loss and λ is the weighting coefficient.

[0080] Step 4: Build a multi-view prototype contrastive learning framework and perform feature-level alignment.

[0081] Using the optimized pseudo-labels obtained in step 3, the final feature-level alignment training is performed. Existing methods typically rely on global information for construction, which inevitably dilutes key details in boundary regions, thus weakening the model's ability to model these regions. Therefore, this invention constructs a multi-view prototype contrastive learning framework, such as... Figure 4 As shown, specifically, based on the true labels from the source domain and the optimized pseudo labels from the target domain, their respective category prototypes are calculated and then fused to generate a hybrid global category prototype. Simultaneously, to address the issue of boundary information being easily diluted by main features, boundary regions are extracted through morphological operations, and independent boundary prototypes are constructed. Finally, by using a contrastive loss function, the feature representations learned by the network are forced to distinguish between different semantic categories and the transition regions between the main body and the boundary, thereby achieving compactness of intra-class features, separability of inter-class features, and explicit enhancement of boundary features.

[0082] The implementation of the multi-view prototype contrastive learning framework in this step includes:

[0083] (4a) Real labels in the source domain Pseudo-tags optimized for the target domain Under the guidance of the relevant authorities, source domain prototypes were generated respectively. and target domain prototype

[0084]

[0085] in, This represents the total number of pixels of class c in the source domain; f is the total number of pixels of class c in the target domain. s (i) represents the high-level semantic features of the target domain f s The value of the i-th pixel, f t (i) represents the high-level semantic features f of the target domain. t The value of the i-th pixel, For source domain real tags The value of the i-th pixel, Pseudo-tags optimized for the target domain The value of the i-th pixel.

[0086] (4b) For the source domain prototype and target domain prototype Perform dynamic iterative hybrid updates to generate the hybrid domain prototype for the current iteration:

[0087]

[0088] Where α is a weighting coefficient used to adjust the contribution of the source domain prototype and the target domain prototype to the global prototype update. Let be the prototype of the hybrid domain in the t-th iteration. For the prototype of the mixed domain in the (t-1)th iteration, Let be the source domain prototype for the t-th iteration. Let be the target domain prototype for the t-th iteration.

[0089] (4c) Construct a contrastive learning loss, the goal of which is to bring the global pixel features output by the network closer to the mixed domain prototypes of their corresponding classes, while increasing their distance from the mixed domain prototypes of other classes, expressed as:

[0090]

[0091] Where exp(·) is the exponential function used to convert similarity into a probability distribution, sim(·) is the cosine similarity used to calculate the similarity between feature vectors, τ is a temperature parameter used to smooth the probability distribution, and f i It represents the i-th pixel value of the semantic features of the source or target domain.

[0092] Furthermore, the multi-view prototype contrastive learning framework in S4 is further implemented by including:

[0093] (5a) Extract the boundary region mask B from the real label of the source domain and the optimized pseudo label through morphological operations:

[0094] B = dilate(y,k) - erode(y,k)

[0095] Here, dilate(y,k) is the dilation operation with a kernel size of k, and erode(y,k) is the erosion operation with a kernel size of k.

[0096] (5b) Calculate the centroid of pixel features within the boundary region to construct an independent boundary prototype.

[0097]

[0098] in, f is the total number of pixels of class c in the source domain. i B is the i-th pixel value of the semantic features in the source or target domain; i is the value of the i-th pixel in the boundary mask.

[0099] (5c) The boundary prototype A contrastive learning process is added to enhance the model’s ability to model fine boundary structures. This process complements global contrastive learning by reusing the contrastive learning loss function form in step (4c) and replacing the input prototype with the boundary prototype.

[0100] To demonstrate the advantages of this invention, the specific implementation effects can be further illustrated by the following simulation results:

[0101] I. Simulation Conditions

[0102] The segmentation method designed in the simulation was based on Python 3.10 and PyTorch 1.11 frameworks; the machine used to train the segmentation network in the simulation was equipped with an Intel(R) Core(TM), 2.5GHz CPU, 24GB RAM and one GeForce RTX 3090 GPU;

[0103] The dataset used in the simulation is a polyp image and label dataset obtained from the publicly available ETIS, CVC-EndoScene, and Kvasir datasets.

[0104] II. Simulation Content

[0105] Under the simulation conditions described above, the method of this invention was used to segment the obtained polyp images and label datasets. The segmentation evaluation metrics used were the Dice coefficient and the mean intersection-union ratio (mIoU). The Dice coefficient is a metric based on region overlap, used to quantify the similarity between the model's predicted segmentation results and the ground truth labels. A higher value indicates a more accurate segmentation result. The mean intersection-union ratio is also a metric that measures the degree of region overlap; it calculates the ratio of the intersection to the union of the predicted and ground truth regions. The value of this metric ranges from [0,1], with higher values ​​indicating better performance.

[0106] Tables 1 and 2 respectively show the performance comparison between the method of the present invention and the current mainstream unsupervised domain-adaptive medical image segmentation methods on the ETIS-adapted CVC-EndoScene task and the ETIS-adapted Kvasir task.

[0107] Table 1. Performance comparison of different UDA methods on the ETIS adaptation to the CVC-EndoScene task.

[0108] Model Dice (%) ↑ mIoU (%) ↑ No-domain adaptation 61.87 72.21 FDA 67.98 75.44 MPA-DA 72.08 75.33 GFDA 75.13 80.76 DCLPS 80.02 83.25 MLDB <![CDATA[ 81.46 ]]> <![CDATA[ 84.23 ]]> Ours 82.14 84.50

[0109] Table 2. Performance comparison of different UDA methods on the ETIS-adapted Kvasir task.

[0110]

[0111]

[0112] As shown in Tables 1 and 2, this invention achieved optimal results across all performance metrics in two different domain adaptation tasks. Specifically, in the ETIS to CVC-EndoScene adaptation task, this invention achieved a Dice coefficient of 82.14%, outperforming the second-best method by approximately 0.7%. In the ETIS to Kvasir adaptation task, the Dice coefficient reached 83.16%, an improvement of approximately 2% over the second-best method. This improvement indicates a significant advantage in the overall segmentation accuracy of the target region, enabling more accurate identification of the decision boundary between the foreground and background. Furthermore, this invention also demonstrated comprehensive performance advantages in the mIoU metric for both tasks.

[0113] To more intuitively demonstrate the performance advantages of this invention, Figure 5 and Figure 6 Visualizations of the polyp segmentation dataset are presented. In these segmentation visualization comparisons, the present invention demonstrates excellent results. Specifically, in Figure 5 and Figure 6In scenarios where the foreground target is large and its color tone is similar to the surrounding background, most contrast methods struggle to accurately define the target boundary, resulting in under-segmentation and misidentification of background areas as foreground. In contrast, this invention accurately captures the overall contour without producing erroneous segmentation results.

Claims

1. An unsupervised domain-adaptive medical image segmentation method based on multi-view alignment and pseudo-label optimization, used for segmenting unlabeled target domain medical images, characterized in that... Includes the following steps: S1: The style transfer module, which uses frequency domain smooth fusion, converts the source domain image into a target domain-like image, achieving image-level alignment between the source and target domains; S2: Input the target domain image into the image segmentation network and perform pre-training with the source domain real labels as supervision; after pre-training is completed, input the real target domain image and output the initial pseudo labels; S3: Perform a two-stage pseudo-label optimization process to process the initial pseudo-labels to improve their structural integrity and boundary accuracy, and obtain optimized pseudo-labels; S4: Perform feature-level alignment, construct a multi-view prototype contrastive learning framework based on the real labels of the source domain and the optimized pseudo labels, and use the framework to train the image segmentation network to perform medical image segmentation.

2. The method as described in claim 1, characterized in that, In S1, the processing procedure of the style transfer module for frequency domain smooth fusion includes: (1a) Source domain image I s and target domain image I t The source domain image I is obtained by converting from the spatial domain to the frequency domain using Fast Fourier Transform. s Amplitude spectrum A s and the corresponding phase spectrum P s and target domain image I t Amplitude spectrum A t and the corresponding phase spectrum P t ; (1b) Design a center-weighted smooth mask M, whose weights decrease smoothly from the low-frequency center to the high-frequency region. The calculation formula is as follows: Where d(i, is the distance from point (ij, to the mask center (c)). x ,c y The Euclidean distance of ) d max To start from the center of the mask (c x ,c y The distance from the farthest corner of the mask to the point where cos(·) is the cosine function; (1c) Using the smoothing mask M to smooth the source domain image I s and target domain image I t The amplitude spectra are weighted and fused to generate a new amplitude spectrum: A new =M⊙A t +(1-M)⊙A s (1d) The new amplitude spectrum A new With source domain image I s Phase spectrum P s By combining these images and mapping them back to the image space using inverse Fourier transform, we obtain the final source image resembling the target, i.e., the target-like domain image I. s→t The formula is as follows: in, This represents the inverse Fourier transform.

3. The method as described in claim 1, characterized in that, The image segmentation network in S2 adopts an encoder-decoder architecture, wherein: (2a) The encoder module consists of several convolutional blocks. The last convolutional block contains only convolution operations, while each of the remaining convolutional blocks contains both convolution and max pooling operations to progressively extract deep semantic features from the image. After each convolutional block, the feature map size is halved and the number of channels is doubled; the target domain image I... s→t and target domain image I t After inputting into the encoder, the output is the high-level semantic feature f of the target domain. s and target domain high-level semantic features f t ; (2b) The decoder receives the features output by the encoder. Its structure is symmetrical with the encoder and contains several deconvolution blocks. Each deconvolution block doubles the size of the feature map through deconvolution operation and fuses it with the feature map of the corresponding level of the encoder through skip connections to supplement the low-level detailed information. (2c) The output layer receives the feature map output by the decoder and uses a 1×1 convolution operation to map its channel dimensions to the number of categories related to the task. Finally, it outputs the category probability of each pixel through the softmax function; where the category probability of the target domain image is used as the prediction result y. s The class probability of the target domain image is used as the initial pseudo-label y. t .

4. The method as described in claim 3, characterized in that, The encoder module consists of four convolutional blocks. The last convolutional block contains only two consecutive 3×3 convolution operations, while each of the remaining convolutional blocks contains two consecutive 3×3 convolution operations and one 2×2 max pooling operation.

5. The method as described in claim 1, 3, or 4, characterized in that, The pre-training process employs a loss function that combines weighted cross-entropy loss and weighted intersection-over-union loss. in, This represents the weighted cross-entropy loss. This represents the weighted intersection-union loss, where N represents the number of pixels, w is the class weight, and y is the weight of the class. i p represents the value of the i-th pixel of the actual label. i The value of the i-th pixel represents the predicted probability.

6. The method as described in claim 1, characterized in that, The first stage of the two-stage optimization strategy in S3 is based on reconstruction-based potential prospect completion, including: (31a) In the initial pseudo-label y t Under the guidance of the algorithm, the centroid of the class is calculated as the initial prototype. in, Let H be the total number of pixels of class c in the target domain, H be the feature map height, W be the feature map width, and f be the number of pixels of class c in the target domain. t (i) represents the high-level semantic features f of the target domain. t The value of the i-th pixel, y t (i) represents the initial pseudo-label y t The value of the i-th pixel; (31b) Initial pseudo-label y based on the output of the image segmentation network t Construct a background boot mask M bg The region initially identified by the positioning model as the background, where M bg =1-argmax(y t ); (31c) Set the background guide mask M bg With the high-level semantic features f of the target domain t Feature fusion is performed to obtain a background-weighted feature representation. in, For element-wise multiplication; (31d) calculation Local features and various prototypes The similarity between pixels is determined by a preset threshold; when the similarity between a pixel feature and a foreground prototype exceeds a certain threshold, the pixel is identified as a potential foreground. The formula is as follows: Where sim(·) is used to calculate cosine similarity, β is a preset threshold, and C is the total number of categories. This is the result of classifying all pixels; (31e) Determine the category of all pixels These are integrated into the initial pseudo-labels to correct misclassified foreground elements in the original segmentation results, resulting in coarsely optimized pseudo-labels. as follows: Here, ∪ represents the union of the two masks.

7. The method as described in claim 6, characterized in that, The second stage of the two-stage optimization strategy in S3 is a structure-aware enhancement strategy, including: (32a) Introduce a pre-trained Segment Anything Model (SAM) as a structure-aware branch; (32b) From the coarse-optimized pseudo-labels after foreground completion In the process, a grid sparse sampling strategy is used to extract the point cue set S; (32c) Guided by the point cue set S, the SAM model generates a segmentation prediction map M with structural integrity. sam ; (32d) The consistency verification module filters the segmentation hypothesis, retaining the pseudo-labels from the coarse optimization. The region where consensus is reached is used to obtain the optimized pseudo-label: In the formula, M is the optimized pseudo-tag. sam (i) is the segmentation prediction map M sam The value of the i-th pixel, To coarsely optimize pseudo tags The value of the i-th pixel; (32e) Pseudo-labels optimized by introducing structural loss constraints Structure and consistency with segmentation network predictions: in, Let SSIM(·,·) be the loss function, where SSIM(·,·) is the structural loss and λ is the weighting coefficient.

8. The method as described in claim 7, characterized in that, The method for extracting the point cue set S using a grid sparse sampling strategy is as follows: Divide the image into a k×k grid. Within each grid, select the center of the grid as either a foreground or background point based on the proportion of foreground pixels. Summarize all sampled points to form a cue set. S = {(i,j,l)|l∈{0,1}} Where (i,j) represents the pixel coordinates within the foreground region; l=1 indicates that the coordinates are foreground points, and l=0 indicates that the coordinates are background points.

9. The method as described in claim 1, characterized in that, The multi-view prototype contrastive learning framework in S4 includes: (4a) Real labels in the source domain Pseudo-tags optimized for the target domain Under the guidance of the relevant authorities, source domain prototypes were generated respectively. and target domain prototype in, This represents the total number of pixels of class c in the source domain; f is the total number of pixels of class c in the target domain. s (i) represents the i-th pixel value of the high-level semantic feature of the target domain, f t (i) represents the high-level semantic features f of the target domain. t The value of the i-th pixel, For source domain real tags The value of the i-th pixel, Pseudo-tags optimized for the target domain The value of the i-th pixel; (4b) For the source domain prototype and target domain prototype Perform dynamic iterative hybrid updates to generate the hybrid domain prototype for the current iteration: Where α is a weighting coefficient used to adjust the contribution of the source domain prototype and the target domain prototype to the global prototype update. Let be the prototype of the hybrid domain in the t-th iteration. For the prototype of the mixed domain in the (t-1)th iteration, Let be the source domain prototype for the t-th iteration. Let be the target domain prototype for the t-th iteration. (4c) Construct a contrastive learning loss, the goal of which is to bring the global pixel features output by the network closer to the mixed domain prototypes of their corresponding classes, while increasing their distance from the mixed domain prototypes of other classes, expressed as: Where exp(·) is the exponential function used to convert similarity into a probability distribution, sim(·) is the cosine similarity used to calculate the similarity between feature vectors, τ is a temperature parameter used to smooth the probability distribution, and f i It represents the i-th pixel value of the semantic features of the source or target domain.

10. The method as described in claim 9, characterized in that, The multi-view prototype contrastive learning framework also includes: (5a) Extract the boundary region mask B from the real label of the source domain and the optimized pseudo label through morphological operations: B = dilate(y,k) - erode(y,k) Here, dilate(y,k) is the dilation operation with a kernel size of k, and erode(y,k) is the erosion operation with a kernel size of k. (5b) Calculate the centroid of pixel features within the boundary region to construct an independent boundary prototype. in, f is the total number of pixels of class c in the source domain. i B is the i-th pixel value of the semantic features in the source or target domain; i The value of the i-th pixel in the boundary mask; (5c) The boundary prototype A contrastive learning process is added to enhance the model’s ability to model fine boundary structures. This process complements global contrastive learning by reusing the contrastive learning loss function form in step (4c) and replacing the input prototype with the boundary prototype.

Citation Information

Patent Citations

  • Semi-supervised medical image segmentation method based on domain adaptation

    CN117115180A

  • Unsupervised domain adaptation method for medical image segmentation

    CN118735948A

  • Multi-source domain adaptive semantic segmentation method and apparatus based on multi-level domain correlation

    WO2025043813A1

Cited By

  • Image data processing method and device based on robust continuous learning framework

    CN121280542A

  • Target detection method based on hierarchical prototype learning and related equipment

    CN121725230A

  • Target detection method based on hierarchical prototype learning and related device

    CN121725230B