Asymmetric multi-scale discrepancy learning driven medical image segmentation method and system
Patent Information
- Application Number
- CN202610988731.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]本发明的目的是针对现有不对称双解码器所存在的不足,提供一种不对称多尺度差异学习驱动的医学影像分割方法及系统,通过同时引入不对称双解码器与多尺度差异反馈机制,强化边界细节特征学习,提升了低标注条件下的分割精度与泛化能力,为轻量化临床辅助诊断提供可行方案
[0024]一种包括计算机可读指令的计算机可读存储介质,所述计算机可读指令在被处理器执行时实现本发明所述的不对称多尺度差异学习驱动的医学影像分割方法中的步骤。
Smart Images

Figure CN122820745A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of imaging technology, and in particular to a medical image segmentation method and system driven by asymmetric multi-scale difference learning. Background Technology
[0002] Medical image segmentation is crucial for clinical diagnosis and treatment. In the image analysis of cardiac, abdominal, and lung tumors, target regions often exhibit blurred boundaries, complex structures, and significant individual differences, while high-quality labeled data is scarce and costly. Traditional supervised learning methods heavily rely on labeled data, limiting their clinical application potential; semi-supervised learning, on the other hand, utilizes both limited labeled data and abundant unlabeled data for training, effectively alleviating the problem of labeled data shortage and has become a hot research area.
[0003] Semi-supervised segmentation methods mostly employ two mainstream approaches: pseudo-label learning and consistency training. Pseudo-label learning generates high-quality pseudo-labels through model prediction, transforming unlabeled data into trainable samples and alleviating the label scarcity problem. Consistency training, by perturbing the input data, model parameters, or features, constrains the model output to maintain stability and consistency, and is the most widely used semi-supervised paradigm in medical image segmentation. Although existing pseudo-label learning and consistency training methods have made significant progress, most methods only focus on output layer consistency or pseudo-label selection strategies, failing to fully explore the multi-scale feature differences within the decoder. They have limited learning capabilities for difficult-to-segment regions such as blurred boundaries and small lesions, and still suffer from insufficient accuracy and weak generalization in low-label scenarios.
[0004] Dual-decoder structures, relying on parallel branch architectures to achieve complementary feature information, effectively enhance the model's feature representation and discrimination capabilities, and are widely used in medical image segmentation tasks. Early dual-decoder models mostly adopted symmetrical structure designs, with the two decoding branches maintaining complete consistency in network layers, channel dimensions, upsampling strategies, and supervision methods, relying only on input perturbations or parameter initialization methods to generate slight differences. This type of structure has the advantages of stable training and simple implementation, but the highly similar feature representations between branches easily lead to serious feature redundancy problems, resulting in low model computational efficiency and insufficient segmentation capabilities for key regions such as boundary details and microstructures. With further research, asymmetric dual-decoders have gradually become an improvement direction, enhancing the feature diversity between branches and reducing redundant computation through differentiated structural designs. However, existing asymmetric dual-decoder methods still have limitations. Most only complete feature fusion and output constraints in the decoding stage, lacking effective utilization of multi-scale difference information, making it difficult to jointly optimize deep feature differences with the encoder, and failing to fully leverage the structural advantages in semi-supervised, low-annotation scenarios. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of existing asymmetric dual decoders by providing an asymmetric multi-scale difference learning-driven medical image segmentation method and system. By simultaneously introducing an asymmetric dual decoder and a multi-scale difference feedback mechanism, the method enhances the learning of boundary detail features, improves segmentation accuracy and generalization ability under low-annotation conditions, and provides a feasible solution for lightweight clinical auxiliary diagnosis.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] A medical image segmentation method driven by asymmetric multi-scale difference learning includes the following steps:
[0008] S10: Obtain the medical image to be segmented and preprocess it into a uniform format;
[0009] S20: Input the preprocessed medical images into the Amld model, and the Amld model outputs the segmentation results.
[0010] The AmLd model was trained in the following way:
[0011] 1) The encoder performs multi-scale downsampling on the sample images, extracts multi-scale abstract semantic features of the sample images, and obtains multi-scale feature maps;
[0012] 2) The detail decoder performs multi-scale upsampling on the multi-scale feature map to obtain detail features; the semantic encoder performs multi-scale upsampling on the multi-scale feature map to obtain semantic features; the structure of the detail decoder is different from that of the semantic encoder.
[0013] 3) Calculate the feature difference between detail features and semantic features at each scale. This feature difference is then fed back to the encoder to guide it in feature enhancement. , , The feature map output by the detail decoder. The feature map output by the semantic decoder. For the enhanced encoder features, For encoder raw features, For feedback weighting coefficients;
[0014] 4) Calculate the total loss of the model and update the parameters of the encoder, detail decoder, and semantic decoder simultaneously through backpropagation.
[0015] The detail decoder uses the first number of channels combined with transposed convolution to complete upsampling, relying on the first receptive field to capture information; the semantic decoder uses the second number of channels combined with lightweight interpolation to complete upsampling, relying on the second receptive field to mine information; the first number of channels is greater than the second number of channels, and the first receptive field is smaller than the second receptive field.
[0016] Multi-scale feature fusion can simultaneously capture high-level semantic information and low-level detailed information, enabling the model to understand the global structure while preserving key local features such as edges, textures, and small objects. This plays a crucial role in improving the segmentation of blurred boundaries and small lesions. In a semi-supervised learning framework, the introduction of multi-scale information can further enhance the model's ability to utilize unlabeled data. By constructing a multi-scale differential guidance mechanism, the model can automatically focus on regions with higher segmentation difficulty, strengthening the learning of weak edges and low-contrast regions, thereby improving segmentation accuracy, stability, and generalization ability under low-label conditions. Combining multi-scale optimization with differential learning allows the model to learn discriminative feature representations even with limited supervised information, providing a feasible approach for efficient and accurate segmentation in clinical scenarios.
[0017] A medical image segmentation system driven by asymmetric multi-scale differential learning includes a data receiving module and an AmLd model segmentation module. The data receiving module is used to acquire the medical image to be segmented and preprocess it into a uniform format. The AmLd model segmentation module is used to perform image segmentation on the preprocessed medical image and output the segmentation result.
[0018] The AmLd model includes an encoder, an asymmetric dual decoder, a multi-scale differential feedback module, and a weighted hybrid loss module. The asymmetric dual decoder includes a detail decoder and a semantic decoder with different structures.
[0019] The encoder is used to perform multi-scale downsampling on the medical image or sample image, extract multi-scale abstract semantic features of the sample image, and obtain a multi-scale feature map.
[0020] The detail decoder is used to perform multi-scale upsampling on the multi-scale feature map to obtain detail features; the semantic encoder is used to perform multi-scale upsampling on the multi-scale feature map during the model training phase to obtain semantic features.
[0021] The multi-scale difference feedback module is used to calculate the feature difference between detail features and semantic features at each scale during the model training phase. This feature difference is then fed back to the encoder to guide it in feature enhancement. , , The feature map output by the detail decoder. The feature map output by the semantic decoder. For the enhanced encoder features, For encoder raw features, For feedback weighting coefficients;
[0022] The weighted mixed loss module is used to calculate the total loss of the model during the model training phase.
[0023] A computer program product includes computer-readable instructions that, when executed by a processor, implement the steps of the asymmetric multi-scale difference learning-driven medical image segmentation method of the present invention.
[0024] A computer-readable storage medium comprising computer-readable instructions that, when executed by a processor, implement the steps of the asymmetric multi-scale difference learning-driven medical image segmentation method of the present invention.
[0025] An electronic device includes: a memory storing program instructions; and a processor connected to the memory, executing the program instructions in the memory to implement the steps in the asymmetric multi-scale difference learning-driven medical image segmentation method of the present invention.
[0026] Compared with existing technologies, the proposed AmLd asymmetric multi-scale differential learning semi-supervised medical image segmentation model, relying on the synergistic effect of an asymmetric dual-decoder architecture and a multi-scale differential feedback mechanism, enhances the model's ability to learn features of edge details in medical images, effectively improving segmentation accuracy and generalization ability under low-annotation training conditions. Experimental results show that the model achieves Dice coefficients of 96.62%, 93.79%, and 99.33% on three datasets: left atrium of the heart, pancreas, and lung tumor, respectively, outperforming many current mainstream semi-supervised segmentation algorithms in overall segmentation performance. The model has a lightweight overall structure and can be deployed and inferenced directly in a pure CPU hardware environment. Attached Figure Description
[0027] Figure 1 This is a flowchart of the asymmetric multi-scale difference learning-driven medical image segmentation method of the present invention.
[0028] Figure 2 This is a schematic diagram of the AmLd model structure provided in the embodiment.
[0029] Figure 3 This is a comparison of the segmentation results of the left atrial MRI, pancreatic CT, and lung tumor CT datasets in the experimental case.
[0030] Figure 4 This is a comparison chart of the ablation experiment results in the experimental examples. Detailed Implementation
[0031] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0032] like Figure 1 As shown, the asymmetric multi-scale differential learning-driven medical image segmentation method provided in this embodiment includes the following steps:
[0033] S10: Obtain the medical image (or medical image) to be segmented and perform preprocessing.
[0034] Medical images first undergo a preprocessing process, which uniformly completes operations such as resolution adjustment, window width and window level standardization, and grayscale normalization. This converts the original images acquired by different devices into standard format data that the model can stably process, ensuring stable input data distribution and adapting to the grayscale distribution and anatomical structure characteristics of medical images, thereby eliminating the interference of imaging differences on segmentation performance.
[0035] S20: Input the preprocessed medical images into the AmLd model, and the AmLd model outputs the segmentation results.
[0036] like Figure 2 As shown, the AmLd (Asymmetric Multi-scale Learning Difference) model adopts an integrated encoder-asymmetric dual-decoder architecture, mainly composed of four parts: encoder, asymmetric dual decoder, multi-scale difference feedback module, and weighted hybrid loss module. The encoder extracts multi-scale abstract semantic features of medical images through multi-layer downsampling; the asymmetric dual decoder completes high-resolution detail restoration and global structure inference respectively with differentiated structures; the multi-scale difference feedback module calculates and feeds back the multi-scale feature differences of the asymmetric dual decoder to realize difference-guided encoder feature optimization; the weighted hybrid loss module jointly constrains the training process, improving convergence stability and segmentation accuracy in low-label data scenarios.
[0037] The AmLd model utilizes the multi-scale feature differences of the asymmetric dual decoder to construct feedback signals, guiding the encoder to dynamically optimize feature representations. During the training phase, it simultaneously uses the output of the asymmetric dual decoder to calculate the supervision loss and the difference feedback loss. During the inference phase, it only retains the high-resolution output of the detail decoder, balancing segmentation accuracy and computational efficiency.
[0038] Unlike general image classification networks, the encoder in this embodiment employs a lightweight residual structure customized and optimized for medical image segmentation. It uses a multi-level downsampling approach, with each level corresponding to a specific scale, progressively extracting multi-scale abstract semantic features from medical images from low to high resolution. Simultaneously, residual connections preserve detailed information at different levels. Each downsampling module consists of convolutional layers, batch normalization layers, and ReLU activation functions, effectively extracting texture, edge, and global contextual features of organs and lesions. This provides multi-scale feature support for subsequent detail restoration and semantic inference using the asymmetric dual decoder.
[0039] The asymmetric dual decoder consists of a semantic decoder and a detail decoder with different structures. The two decoding pathways are differentiated in terms of channel count, upsampling method, receptive field size, and feature learning direction, achieving functional decomposition and complementary advantages. The detail decoder uses a high channel count combined with transposed convolution for upsampling, relying on a smaller receptive field to capture fine-grained information such as local image texture and edge contours, accurately restoring high-resolution features and details of minute lesions and tissue boundaries. The semantic decoder, on the other hand, employs a simplified channel structure and lightweight interpolation upsampling method, relying on a larger receptive field to mine global context and overall anatomical structure information, focusing on high-level semantic reasoning and ensuring the integrity and rationality of organ spatial topological relationships. This asymmetric design effectively avoids the feature redundancy problem that easily occurs in traditional symmetrical dual-branch structures, allowing the two decoding branches to perform their respective functions and cooperate effectively, significantly improving the overall segmentation accuracy and edge delineation effect of the model under low-annotation training conditions.
[0040] Define I to represent the input medical image; Indicates encoder; Represents the detail decoder; This represents a semantic decoder; This represents the output of the detail decoder; If we represent the output of the semantic decoder, then we have: , .
[0041] The detail decoder employs a high-channel-count network configuration ("high" here is a relative concept, simply indicating that the detail decoder has more channels than the semantic decoder) and uses transposed convolution as the core upsampling method. This high-channel-count architecture expands the feature representation dimension, allowing the network to accommodate richer pixel-level feature information. Meanwhile, the transposed convolution possesses learnable parameters, enabling it to autonomously complete feature mapping and spatial location compensation during image size restoration, fully preserving the high-resolution spatial information of the image. This detail decoder specifically enhances the modeling of fine-grained features in medical images, such as edge textures, minute lesions, complex anatomical contours, and soft tissue boundaries, continuously improving the network's ability to express and capture subtle local features, accurately restoring details in local areas. Addressing common clinical segmentation challenges such as blurred organ boundaries, low tissue contrast, and missed detection of minute lesions, this detail decoder strengthens the feature recognition of weak boundary regions at the feature level, effectively improving the problems of blurred contours, lost details, and discontinuous edges in these difficult areas, ensuring the precision and accuracy of local segmentation results.
[0042] The detail decoder processes multi-scale feature maps obtained by the encoder through progressive downsampling. During downsampling at different levels, the encoder simultaneously preserves shallow edge texture details and deep global semantic information. These multi-scale features are passed to the detail decoder through skip connections. As the detail decoder recovers spatial resolution through progressive upsampling, it fuses its own upsampled features with the downsampled features from each level of the encoder.
[0043] The semantic decoder employs a streamlined channel layout (i.e., the semantic decoder has fewer channels than the detail decoder) and utilizes a lightweight trilinear interpolation algorithm for image upsampling. Trilinear interpolation is a parameter-free interpolation operation, requiring no additional training weights, with simple computational logic and fast processing speed, making it a lightweight upsampling solution. The semantic decoder abandons the pursuit of fitting extreme local details, focusing instead on global contextual feature inference and high-level semantic mining. It can quickly extract high-level semantic information such as the overall morphology, spatial location, and topological relationships of target organs and large lesions from the entire image, accurately controlling the integrity and rationality of human anatomical structures. Simultaneously, the streamlined channel design combined with parameter-free upsampling significantly reduces the overall number of network parameters and floating-point operations, effectively reducing hardware resource consumption during model inference. While ensuring accurate global segmentation results, it significantly improves the overall model efficiency, achieving lightweight deployment and rapid inference.
[0044] The semantic decoder processes the multi-scale feature maps obtained by the encoder through step-by-step downsampling. Both the semantic decoder and the detail decoder reuse the same set of multi-scale features output by the encoder, eliminating the need for repeated encoding and significantly improving feature utilization efficiency.
[0045] The detail decoder and semantic decoder each perform their respective functions and complement each other, fundamentally reducing feature redundancy caused by symmetrical structures and improving overall feature utilization. Combining multi-scale feature fusion, the dual decoders can simultaneously mine high-level global semantics and low-level local details, further enhancing the model's ability to extract various features. After embedding this asymmetric dual decoder into a semi-supervised learning framework, coupled with a multi-scale difference guidance mechanism, the model can proactively identify regions with high segmentation difficulty and increase the learning weight for weak edges and low-contrast regions. Even in low-label scenarios using only a small amount of labeled data, the model can still learn highly discriminative features, significantly improving segmentation accuracy, operational stability, and cross-scenario generalization ability, balancing segmentation performance and inference efficiency, and providing reliable technical support for accurate clinical segmentation of medical images.
[0046] In this embodiment, the multi-scale difference feedback module breaks the traditional mode of only unidirectional feature transmission between the encoder and decoder, and builds a bidirectional optimization path between the asymmetric dual decoder and the encoder: firstly, the feature deviation between the detail decoder and the semantic decoder is calculated at different feature scales, and then the scaled difference information is fed back to the corresponding level of the encoder layer by layer, thereby guiding the network to actively focus on blurred boundaries, weak contrast tissues, complex anatomical structures and difficult areas that are prone to missegmentation in the image, so as to realize dynamic iterative optimization of network feature expression.
[0047] Let the feature map output by the detail decoder in the l-th decoding level be... The feature map output by the semantic decoder is Then, the feature difference between the two decoders at this scale is:
[0048] , This represents the feature difference between the two decoders at the l-th decoding level (feature scale);
[0049] The calculated feature differences are then backpropagated to the l-th layer of the encoder to perform adaptive enhancement processing on the original features of the encoder. The feature enhancement formula is as follows:
[0050] , For the enhanced encoder features, For encoder raw features These are feedback weighting coefficients used to control the extent to which feature differences enhance encoder features;
[0051] To optimize details using feature differences while ensuring reasonable consistency in the overall feature distribution of the dual decoders and avoiding structural distortion caused by excessive bias, a multi-scale difference loss function is defined:
[0052] ;
[0053] In the formula: l is the feature scale index; L represents the total number of feature scales of the network as a whole, that is, the total number of layers of the encoder and decoder; The multi-scale difference loss function constrains the overall feature trend of the two decoders, ensuring the consistency of global anatomical structure prediction. MSE represents the mean squared error of the two probability maps calculated pixel by pixel, serving as the source of the consistency constraint gradient.
[0054] The entire multi-scale difference feedback mechanism forms a complete closed-loop optimization logic: the network extracts the feature bias of the dual decoders layer by layer, feeds back the feature difference information at the corresponding scale to enhance the encoder's feature representation, and allows the model to automatically shift its learning focus to regions with higher segmentation difficulty, completing the process from difference extraction to feature enhancement, and ultimately achieving end-to-end optimization for accurate segmentation. In semi-supervised learning scenarios with only a small amount of labeled data, this multi-scale difference feedback mechanism can fully exploit the effective feature information in unlabeled images, make up for the lack of supervision signals under low-label conditions, continuously enhance the model's ability to learn edge details and complex structures, and effectively improve the model's boundary segmentation accuracy, robustness, and cross-scene generalization ability.
[0055] To address issues such as training oscillations, slow convergence, and insufficient segmentation accuracy under low-label data conditions, the AmLd model employs a weighted hybrid loss combined with a deep supervision strategy for collaborative optimization. The overall loss is decomposed into four independent sub-losses: supervised loss, consistency loss, multi-scale difference loss, and deep supervision loss. These four sub-losses have clearly defined functions, respectively performing four tasks: labeled ground truth fitting, unlabeled output constraint, multi-scale feature difference regulation, and multi-level accelerated convergence. This achieves integrated collaborative optimization of labeled samples, unlabeled samples, and multi-scale features.
[0056] The expression for the model's total loss function is:
[0057] ;
[0058] In the formula, The supervised loss corresponding to the labeled samples; To adapt to the consistency loss during training with unlabeled samples; To regulate the multi-scale difference loss of feature bias in multi-scale dual decoders; The depth supervision loss for multi-level decoding output; , , , The manually set balancing weight coefficients are used to dynamically adjust the contribution ratio of the four sub-losses in the overall backpropagation, adapting to the gradient update rhythm of semi-supervised training in medical imaging.
[0059] In this model, multi-scale difference loss The hierarchical feedback optimization and the overall optimization of the total loss are nested and synergistic. They both serve the iterative update of model parameters, but they also have clear differences in their level of action, optimization path and objective focus. This is intermediate layer feature optimization applied during the training process. Its core objective is to guide the encoder to dynamically adjust the feature representations at each level through multi-scale feature difference backpropagation, making shallow features focus more on detailed textures and deep features focus more on global semantics. It belongs to "refined intermediate optimization" at the feature level and does not directly drive the overall update of model parameters. The total loss, on the other hand, is an end-to-end optimization applied to the entire model. It updates all parameters of the encoder and decoder simultaneously through backpropagation and is the "global final optimization objective" driving the overall convergence of the model. It includes the combined effects of supervised loss, consistency loss, multi-scale difference loss, and deep supervised loss. The two are nested within each other during the training process. The generated optimization signal optimizes the encoder directly through an independent feature backpropagation path, and also participates in global backpropagation as a component of the total loss, forming a dual constraint mechanism of "intermediate feature optimization + global parameter update". This ensures both the detail and semantic expressiveness of the encoder features, and achieves stable convergence of the entire model in a semi-supervised scenario.
[0060] Monitoring losses This method applies only to image samples with physician-annotated ground truth values. It fuses cross-entropy loss and Dice loss to construct a composite supervision signal, balancing pixel classification accuracy and foreground region integrity. The formula is as follows: .
[0061] Cross-entropy loss Its core function is to measure the deviation between the predicted probability and the true binary classification label pixel by pixel, finely constrain the classification result of each pixel, and improve the global pixel-level segmentation and discrimination accuracy.
[0062] ;
[0063] In the formula, N represents the total number of pixels in a single input medical image; ϵ{0, 1} represents the gold standard ground truth label for the i-th pixel, where 1 represents the foreground of the lesion / target organ and 0 represents the background tissue of the human body; This is the predicted probability value output by the model network for the i-th pixel belonging to the foreground object.
[0064] Medical imaging commonly suffers from a class imbalance where the foreground lesions / organs are significantly smaller than the background. Simple cross-entropy loss tends to favor learning the background and weaken the features of small lesions. Dice loss, relying on intersection-union logic, optimizes the foreground fitting effect, ensuring the complete and continuous segmentation contours of target organs and small lesions.
[0065] ;
[0066] In the formula, It is a local smoothing constant, with a fixed value of 10⁻⁶, used to avoid division by zero errors when the numerator and denominator values approach 0, and to ensure numerical stability during the training process.
[0067] Consistency loss Designed for massive unlabeled image samples within a semi-supervised framework, this approach specifically constrains the final segmentation output of two asymmetric decoding branches to prevent excessive bias in the dual-decoder prediction results and misalignment / distortion of the global anatomical structure. Mean Squared Error (MSE) is used to quantize the difference in the probability distributions of the two decoder outputs, forcing a uniform global segmentation contour under unlabeled data.
[0068] In the formula: It is the global pixel prediction probability distribution map output by the detail decoder; It is the global pixel prediction probability distribution map output by the semantic decoder.
[0069] Multiscale difference loss Corresponding to the multi-scale difference feedback mechanism, constraints are applied to the intermediate feature maps of the two branches at each decoding scale to stabilize the learning range of feature differences at each level, preventing drastic fluctuations in feature difference values from causing encoder feature enhancement disorder, and ensuring stable iterative training throughout the multi-scale feedback process:
[0070] ;
[0071] In the formula, L represents the total number of feature layers in the decoder; This is the output feature map of the l-th level of the detail decoder; This is the output feature map of the l-th layer of the semantic decoder; the mean square error of the two features is calculated layer by layer and summed to achieve normalized control of feature differences across the entire scale.
[0072] Deep monitoring loss This involves applying supervised constraints to the output features of each level of the dual decoder, and synchronously backpropagating gradients from shallow to deep layers, significantly accelerating the overall convergence speed of the model while improving the quality of multi-level feature extraction.
[0073] ;
[0074] In the formula, It is a supervision loss applied simultaneously to the output of the l-th layer of both the detail decoder and the semantic decoder; This is a weight coefficient specific to the l-th layer; it follows the rule of setting greater weight for deeper features and smaller weight for shallower features, matching the network characteristic that deeper features dominate the global semantics and shallower features are responsible for fine details.
[0075] To verify the feasibility of the AmLd model, a validation experiment was conducted using the publicly available Medical Segmentation Decathlon (MSD) clinical dataset. The training and test sets were uniformly divided in an 8:2 ratio. To align with low-annotation semi-supervised clinical use scenarios, only 10% of the samples in the training set were annotated with physician gold standard labels, while the remaining samples were completely unannotated, thus rigorously simulating the real-world environment of scarce annotation resources in hospitals. This study selected three image tasks with differentiated features to comprehensively test the model's adaptability in three typical segmentation scenarios: homogeneous organs, low-contrast complex organs, and small lesions.
[0076] (1) Left atrial MRI dataset (Task02)
[0077] Cardiac MRI images show clear overall atrial anatomical boundaries and uniform grayscale distribution in individual organ tissues, with an image resolution of 256×256. This data was used for benchmark testing to verify the model's completeness and basic accuracy in segmenting well-shaped, homogeneous single organs.
[0078] (2) Pancreatic CT dataset (Task07)
[0079] Abdominal CT images show minimal grayscale differences between the pancreas and surrounding gastrointestinal tract, blood vessels, and soft tissues, with blurred and intersecting tissue boundaries, indicating complex anatomical structures. The image resolution is 512×512. This study aims to examine the model's segmentation robustness in the face of weak contrast and complex adjacent structures, and to test the optimization effect of multi-scale difference feedback on blurred boundaries.
[0080] (3) Lung tumor CT dataset (Task06)
[0081] Chest CT images contain tiny tumor lesions of varying sizes and shapes. These lesions occupy a small area and have faint edges, easily confused with the lung parenchyma. The image resolution is 256×256. This study specifically evaluates the model's ability to capture details of small target lesions and faint lesion boundaries, verifying the fine-grained restoration advantages of the asymmetric detail decoder.
[0082] The experiment used three authoritative quantitative metrics recognized in the field of medical image segmentation to evaluate the performance. These three metrics comprehensively measure the model's performance from three dimensions: overlap matching degree, overall segmentation accuracy, and edge contour accuracy.
[0083] (1) Dice similarity coefficient: quantifies the degree of overlap and matching between the model segmentation prediction region and the physician's gold standard true value region. The value range is 0 ~ 1. The closer the Dice value is to 1, the higher the overall overlap of the segmentation and the better the overall segmentation accuracy of organs / lesions. It is the core overall accuracy evaluation standard for medical segmentation.
[0084] (2) Intersection over Union (IoU): The ratio of the intersection area and the union area of the predicted region and the ground truth region is calculated to intuitively reflect the overall pixel-level matching accuracy of the model. The higher the IoU value, the fewer missegmented or missed pixels of the foreground target, and the stronger the global segmentation reliability.
[0085] (3) 95% Hausdorff distance (HD95): A boundary evaluation index specifically designed for the accuracy of segmentation contour boundaries. During calculation, 5% of the extreme distance outliers are removed to avoid isolated noise interfering with the evaluation results. The smaller the HD95 value, the smaller the spatial deviation between the predicted contour and the real anatomical contour, and the better the smoothness and fit of the organ and lesion edge segmentation.
[0086] Figure 3 The visualization compares the segmentation results of the model on three test datasets: left atrial MRI, pancreatic CT, and lung tumor CT. Each row of samples, from left to right, represents the original medical image, the physician-annotated gold standard ground truth, and the segmentation prediction mask output by the AmLd model. The intuitive visualization, combined with the quantitative index data mentioned earlier, forms a complete validation, fully demonstrating that the asymmetric dual-decoder architecture and multi-scale difference feedback mechanism proposed in this embodiment can simultaneously improve the overall segmentation matching accuracy and edge detail depiction ability. Even under the stringent low-annotation semi-supervised training condition using only 10% labeled samples, the model still exhibits stable, reliable, and robust segmentation performance against interference.
[0087] The asymmetric dual-decoder architecture decouples the learning of detailed features from the learning of global semantic features: the detailed decoder built with high-channel transposed convolution focuses on mining fine-grained information such as edge texture, small lesions, and complex anatomical contours; the semantic decoder built with lightweight trilinear interpolation is responsible for quickly capturing high-level global semantics such as the overall morphology and spatial topology of organs. The two branches perform their respective functions and complement each other, reducing the feature redundancy problem that is common in symmetric dual branches from the structural root.
[0088] The multi-scale difference feedback mechanism explicitly models the hierarchical feature deviation between the two decoding branches, and propagates the scaled feature difference back to the corresponding encoder level layer by layer, dynamically enhancing the feature representation ability of difficult-to-segment regions such as blurred boundaries, low-contrast junctions, and small lesions. In semi-supervised scenarios where labeled supervision signals are scarce, this feedback link can autonomously mine learning information within unlabeled images, making up for the lack of supervision caused by a small number of labels.
[0089] The two core modules work together to enable the final model to achieve high-precision segmentation for three types of clinical images with vastly different levels of difficulty: for homogeneous organs with clear boundaries, such as the left atrium, the predicted outline fits the gold standard as a whole; for complex abdominal tissues with mixed gray levels and blurred boundaries, such as the pancreas, it can accurately distinguish the pancreas from the surrounding gastrointestinal tract, blood vessels and soft tissues; for small lung tumor lesions, it can completely detect small lesions and restore the irregular edge outline of the lesions. Whether it is a regular organ, a complex soft tissue or a small lesion, it can achieve stable and precise segmentation.
[0090] To quantitatively verify the effectiveness and synergistic gain of the two core modules—asymmetric dual decoder and multi-scale differential feedback—this paper sets up three control ablation models for comparative experiments. The configurations of the three models are as follows:
[0091] Model A: Remove the multi-scale differential feedback module and retain only the basic asymmetric dual decoder structure;
[0092] Model B: Replaced with a traditional symmetric dual-decoder structure, while retaining the complete multi-scale differential feedback module;
[0093] AmLd: The complete model proposed in this paper is equipped with both asymmetric dual decoders and multi-scale differential feedback modules.
[0094] The ablation quantification results are shown in Table 1. As can be seen from Table 1, compared to Model A and Model B, AmLd achieved optimal values for Dice and IoU indices across all three datasets, and also significantly reduced the HD95 index. Specifically: on the left atrial dataset, AmLd improved the Dice index by 2.46% compared to Model A and by 1.73% compared to Model B; on the pancreas dataset, AmLd improved the Dice index by 1.46% compared to Model A and by 0.91% compared to Model B; on the lung tumor dataset, AmLd improved the Dice index by 18.87% compared to Model A and by 18.20% compared to Model B, ultimately reaching 99.33%.
[0095] Table 1: Ablation Experiment Results
[0096]
[0097] Visualization results of ablation experiments Figure 4 As shown, combined with Figure 4More intuitively, it can be observed that in the left atrium P1 sample, both Model A and Model B exhibited local undersegmentation and incomplete contours, while AmLd's prediction results highly overlapped with the gold standard; in the pancreas P2 sample, Model A showed obvious spiculation at the segmentation boundary, while Model B had blurred edges, and AmLd completely restored the curved shape of the pancreas; in the lung tumor P3 sample, both Model A and Model B showed significant deviations in the segmentation around the complex airway, while AmLd accurately preserved the complete contour of the tumor region.
[0098] The quantitative and visualization results above jointly demonstrate that the asymmetric dual decoder and the multi-scale differential feedback mechanism have a significant synergistic gain effect: the asymmetric dual decoder, through differential channels and upsampling design, decouples detail recovery from global semantic reasoning, reducing feature redundancy; the multi-scale differential feedback mechanism, through explicit feature difference guidance, focuses the model's attention on regions with blurred boundaries and complex structures, effectively alleviating the generalization bottleneck in low-annotation scenarios. The combination of these two mechanisms jointly improves the model's segmentation accuracy and boundary detail representation ability, exhibiting a strong performance advantage, especially in tasks with complex structures and irregular boundaries, such as lung tumors.
[0099] To comprehensively and from multiple perspectives verify the overall performance of the AmLd model in low-annotation semi-supervised medical image segmentation tasks, three representative and widely used semi-supervised medical image segmentation algorithms in the industry were selected as baselines for cross-sectional comparative testing. The core principles of the three comparison methods are introduced as follows:
[0100] (1) Pseudo-Label method: As the most basic and classic foundational paradigm in the field of semi-supervised segmentation, this method uses the prediction results obtained by the model's forward inference to select prediction pixels with high confidence to generate pseudo-true labels, thereby transforming massive unlabeled samples into training samples with supervised signals, expanding the training data volume, and completing semi-supervised iterative training.
[0101] (2) Mean Teacher model: Based on the dual network architecture of teachers and students and the consistency regularization constraint, a complete semi-supervised framework is constructed. The student network updates parameters in real time, and the teacher network performs exponential mean smoothing on the parameters of the student network. The reliability and stability of the pseudo-label of unlabeled data are improved by the stable output of the teacher network. It is the current mainstream high-performance semi-supervised benchmark scheme.
[0102] (3) MT-SDD Dual-path Decoding Model: This model also introduces a dual-branch decoding structure and feature difference learning approach, but the two decoding branches adopt a completely symmetrical network configuration, which forms a direct technical contrast with the asymmetric dual decoder in this paper.
[0103] The training condition with 10% low-labeled samples was uniformly set. The AmLd medical image segmentation model was compared with three classic semi-supervised segmentation algorithms in the industry. Three different clinical image datasets were selected for the experiment: left atrial cardiac MRI dataset, pancreatic abdominal CT dataset, and lung tumor chest CT dataset. The quantitative segmentation evaluation indexes corresponding to each algorithm are shown in Table 2.
[0104] Table 2: Comparison of Segmentation Results of Different Algorithms
[0105]
[0106] Comparing the test results of the three datasets, it can be seen that the AmLd model has higher global accuracy in terms of Dice similarity coefficient and IoU than the three control algorithms. The 95% Hausdorff distance (HD95), which represents the boundary contour error, is lower than the control scheme in all cases, indicating the best overall segmentation performance. In particular, the AmLd model achieves Dice accuracy of 96.62%, 93.79%, and 99.33% on the left atrial, pancreatic, and lung tumor datasets, respectively, which is a significant performance leap compared to the existing mainstream semi-supervised segmentation schemes.
[0107] The training logic of the Pseudo-Label algorithm and the Mean Teacher algorithm heavily relies on the model's prediction results to generate high-confidence pseudo-labels, using these pseudo-labels to achieve supervised iterative training of unlabeled image samples. In the early stages of training, with only a small number of ground truth labels, the network's initial fitting effect is weak, and the generated pseudo-labels contain a large number of pixel-level mislabels. As the training progresses, the bias caused by these mislabels accumulates and expands, resulting in a significant decrease in segmentation performance when dealing with challenging regions such as low-contrast organs like the pancreas or small lesions in the lungs.
[0108] AmLd completely abandons the pseudo-label-dependent architecture and directly extracts the multi-scale feature differences of the output of each level of the asymmetric dual decoder as self-supervised constraint signals. There is no problem of pseudo-label error superposition during the entire training process. In low-label scenarios where labeled supervision signals are scarce, it has more stable training convergence characteristics and stronger segmentation robustness.
[0109] Although the MT-SDD model also adopts a dual-branch decoding structure and introduces the feature difference learning approach, the two decoding branches of this model are configured with a completely symmetrical network structure. During the feature extraction process, a large number of homogeneous redundant features are easily generated, and there is a significant shortcoming in the ability to finely model lesions and organ edges.
[0110] AmLd employs an asymmetric differential dual-decoder structure, clearly defining the functions of the two decoding pathways: the detail decoder primarily uses high-channel configuration and transposed convolutional upsampling to preserve high-resolution information and finely restore edge textures and the contours of small lesions; the semantic decoder uses simplified channels combined with trilinear interpolation for lightweight upsampling, focusing on quickly extracting the global anatomical contextual semantics of organs, reducing feature redundancy and improving boundary fitting accuracy from the root of the network structure. Simultaneously, a multi-scale differential feedback mechanism is included, which backpropagates the feature differences between the two branches at each level to the corresponding level of the encoder, adaptively enhancing the feature representation of difficult-to-segment regions such as blurred boundaries, low-contrast tissue junctions, and small tumor lesions.
[0111] The two core modules, asymmetric dual decoder and multi-scale differential feedback, work together to achieve a synergistic gain effect, ultimately enabling the AmLd model to achieve higher overall segmentation accuracy and contour boundary quality that better matches the anatomical truth in real clinical segmentation scenarios such as low annotation constraints, complex anatomical structures, and identification of small lesions.
[0112] Based on the same inventive concept, this embodiment also provides an asymmetric multi-scale differential learning-driven medical image segmentation system, including a data receiving module and an AmLd model segmentation module. The data receiving module acquires the medical image to be segmented and preprocesses it into a uniform format; the AmLd model segmentation module performs image segmentation on the preprocessed medical image and outputs the segmentation result.
[0113] See also Figure 2 The AmLd model includes an encoder, an asymmetric dual decoder, a multi-scale differential feedback module, and a weighted hybrid loss module. The asymmetric dual decoder includes a detail decoder and a semantic decoder with different structures.
[0114] The encoder is used to perform multi-scale downsampling on the medical image or sample image, extract multi-scale abstract semantic features of the sample image, and obtain a multi-scale feature map.
[0115] The detail decoder is used to perform multi-scale upsampling on the multi-scale feature map to obtain detail features; the semantic encoder is used to perform multi-scale upsampling on the multi-scale feature map during the model training phase to obtain semantic features.
[0116] The multi-scale difference feedback module is used to calculate the feature difference between detail features and semantic features at each scale during the model training phase. This feature difference is then fed back to the encoder to guide it in feature enhancement. , For the enhanced encoder features, For encoder raw features, This refers to the feedback weighting coefficient.
[0117] The weighted mixed loss module is used to calculate the total loss of the model during the model training phase.
[0118] This invention also provides a computer program product including computer-readable instructions. When the computer-readable instructions are executed in an electronic device, the program product causes the electronic device to perform the operation steps included in the method of this invention.
[0119] This invention also provides a storage medium storing computer-readable instructions that cause an electronic device to perform the operation steps included in the method of this invention.
[0120] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0121] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0122] The above specific embodiments are merely several optional embodiments of the present invention. Based on the technical solutions of the present invention and the relevant teachings of the above embodiments, those skilled in the art can make various alternative improvements and combinations to the above specific embodiments.
Claims
1. A medical image segmentation method driven by asymmetric multi-scale differential learning, characterized in that, Includes the following steps: S10: Obtain the medical image to be segmented and preprocess it into a uniform format; S20: Input the preprocessed medical images into the Amld model, and the Amld model outputs the segmentation results. The AmLd model was trained in the following way: 1) The encoder performs multi-scale downsampling on the sample images, extracts multi-scale abstract semantic features of the sample images, and obtains multi-scale feature maps; 2) The detail decoder performs multi-scale upsampling on the multi-scale feature map to obtain detail features; The semantic encoder performs multi-scale upsampling on the multi-scale feature map to obtain semantic features; the detail decoder has a different structure from the semantic encoder. 3) Calculate the feature difference between detail features and semantic features at each scale. This feature difference is then fed back to the encoder to guide it in feature enhancement. , , The feature map output by the detail decoder. The feature map output by the semantic decoder. For the enhanced encoder features, The original features of the encoder, For feedback weighting coefficients; 4) Calculate the total loss of the model and update the parameters of the encoder, detail decoder, and semantic decoder simultaneously through backpropagation.
2. The asymmetric multi-scale difference learning-driven medical image segmentation method according to claim 1, characterized in that, The encoder includes multiple levels of downsampling modules, which progressively extract multi-scale abstract semantic features of medical images from low to high, while retaining detailed information at different levels through residual connections. Each downsampling module consists of a convolutional layer, a batch normalization layer, and a ReLU activation function.
3. The asymmetric multi-scale difference learning-driven medical image segmentation method according to claim 1, characterized in that, The detail decoder uses the first number of channels combined with transposed convolution to complete upsampling, relying on the first receptive field to capture information; The semantic decoder uses a second number of channels combined with lightweight interpolation to complete upsampling and relies on the second receptive field to mine information; the number of the first channel is greater than the number of the second channel, and the first receptive field is smaller than the second receptive field.
4. The asymmetric multi-scale difference learning-driven medical image segmentation method according to claim 1, characterized in that, The multi-scale abstract semantic features obtained by the encoder downsampling are passed to the detail decoder through skip connections. During the upsampling process, the detail decoder fuses its own upsampled features with the multi-scale abstract semantic features downsampled by the encoder at each level.
5. The asymmetric multi-scale differential learning-driven medical image segmentation method according to claim 1, characterized in that, The formula for calculating the total loss is: , , , , ; In the formula, To monitor losses; This results in a loss of consistency. For multi-scale difference loss; Losses due to in-depth monitoring; , , , To balance the weighting coefficients, It is a supervision loss applied simultaneously to the output of the l-th layer of both the detail decoder and the semantic decoder; Here, L is the weight coefficient specific to level l, and L is the total number of levels. For cross-entropy loss, L Dice For Dice's loss, It is the global pixel prediction probability distribution map output by the detail decoder. It is the global pixel prediction probability distribution map output by the semantic decoder.
6. A medical image segmentation system driven by asymmetric multi-scale differential learning, characterized in that, It includes a data receiving module and an AmLd model segmentation module. The data receiving module is used to acquire the medical image to be segmented and preprocess it into a uniform format; the AmLd model segmentation module is used to perform image segmentation on the preprocessed medical image and output the segmentation result. The AmLd model includes an encoder, an asymmetric dual decoder, a multi-scale differential feedback module, and a weighted hybrid loss module. The asymmetric dual decoder includes a detail decoder and a semantic decoder with different structures. The encoder is used to perform multi-scale downsampling on the medical image or sample image, extract multi-scale abstract semantic features of the sample image, and obtain a multi-scale feature map. The detail decoder is used to perform multi-scale upsampling on the multi-scale feature map to obtain detail features; The semantic encoder is used to perform multi-scale upsampling on the multi-scale feature map during the model training phase to obtain semantic features; The multi-scale difference feedback module is used to calculate the feature difference between detail features and semantic features at each scale during the model training phase. This feature difference is then fed back to the encoder to guide it in feature enhancement. , , The feature map output by the detail decoder. The feature map output by the semantic decoder. For the enhanced encoder features, For encoder raw features, For feedback weighting coefficients; The weighted mixed loss module is used to calculate the total loss of the model during the model training phase.
7. The asymmetric multi-scale difference learning-driven medical image segmentation system according to claim 4, characterized in that, The detail decoder uses the first number of channels combined with transposed convolution to complete upsampling, relying on the first receptive field to capture information; The semantic decoder uses a second number of channels combined with lightweight interpolation to complete upsampling and relies on the second receptive field to mine information; the number of the first channel is greater than the number of the second channel, and the first receptive field is smaller than the second receptive field.
8. A computer program product comprising computer-readable instructions, characterized in that, The computer-readable instructions, when executed by a processor, implement the steps of the asymmetric multi-scale difference learning-driven medical image segmentation method according to any one of claims 1-5.
9. A computer-readable storage medium comprising computer-readable instructions, characterized in that, The computer-readable instructions, when executed by a processor, implement the steps of the asymmetric multi-scale difference learning-driven medical image segmentation method according to any one of claims 1-5.
10. An electronic device, characterized in that, include: Memory, which stores program instructions; The processor, connected to the memory, executes the program instructions in the memory to implement the steps of the asymmetric multi-scale difference learning-driven medical image segmentation method according to any one of claims 1-5.