Reliable image segmentation method and device for identification of difficult-to-segment region
By explicitly modeling the intrinsic relationship between segmentation accuracy and prediction reliability, and employing a MaxViT encoder and a dual-decoder network, combined with knowledge reconstruction and cognitive self-correction modules, the problem of low accuracy and insufficient reliability in identifying difficult-to-segment regions in existing technologies is solved, achieving accurate and reliable image segmentation of difficult-to-segment regions.
Patent Information
- Application Number
- CN202511056555.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-14
AI Technical Summary
Existing deep learning methods lack feature enhancement mechanisms when dealing with difficult-to-segment regions, resulting in low recognition accuracy. Furthermore, over-reliance on high-uncertainty regions may cause the model to be susceptible to noise interference, making it difficult to converge. They also ignore the intrinsic relationship between accuracy and reliability, affecting the model's cognitive judgment ability and generalization performance.
By explicitly modeling the intrinsic relationship between segmentation accuracy and prediction reliability, a pre-trained MaxViT encoder network, a knowledge reconstruction path, and a dual decoder network are employed. Combined with multi-stage feature loss aggregation and a cognitive self-correction module, the model is guided to focus on the expression and optimization of uncertain regions, thereby achieving accurate and reliable identification and segmentation of difficult-to-segment regions.
This improved the model's accuracy and reliability in identifying difficult-to-segment regions, enhanced the credibility of the segmentation results, and improved the model's performance in identifying complex regions and the reliability of the segmentation results.
Smart Images

Figure CN120953608A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and in particular to a reliable image segmentation method and apparatus for identifying difficult-to-segment regions. Background Technology
[0002] Image segmentation is a fundamental and crucial technique in computer vision, aiming to divide an image into semantically meaningful regions. It has wide applications in medical diagnosis, remote sensing monitoring, autonomous driving, industrial inspection, and many other scenarios. However, in real-world applications, there are often "difficult-to-segment regions" characterized by varied shapes, blurred boundaries, small scale, and high texture similarity. These regions easily lead to model recognition errors or unclear segmentation contours, severely impacting the accuracy and reliability of the segmentation results. Although existing deep learning methods have significantly improved overall segmentation accuracy, they still face significant performance bottlenecks when dealing with difficult-to-segment regions. Therefore, enhancing the model's expressive ability in these critical regions, while balancing the accuracy and reliability of the segmentation results, is of great significance for promoting the practical application of image segmentation technology.
[0003] The reliability of image segmentation models is a growing concern in the field. One mainstream approach is to assess the confidence level of segmentation results using uncertainty evaluation methods and incorporate this information into the backpropagation process to enhance model reliability. Evidence-based deep learning has gained significant attention for its ability to estimate uncertainty without multiple forward propagations. However, existing methods suffer from several shortcomings in reliability modeling: firstly, models often lack feature enhancement mechanisms for difficult-to-segment regions, leading to low accuracy in these areas; secondly, over-reliance on high-uncertainty regions for optimization can expose the model to potential noise interference, hindering convergence; and thirdly, current research largely neglects the intrinsic relationship between prediction accuracy and reliability, potentially causing a mismatch that impacts the model's cognitive judgment and generalization performance. Therefore, improving the model's ability to represent and recognize difficult-to-segment regions and deeply exploring the intrinsic coupling between accuracy and reliability not only helps improve the model's recognition performance in complex regions but also significantly enhances the credibility of segmentation results. This has significant theoretical and practical implications for building image segmentation systems with practical applications. Summary of the Invention
[0004] To achieve accurate and reliable segmentation of difficult-to-segment regions, this invention proposes a reliable image segmentation method and apparatus for identifying such regions. By explicitly modeling the intrinsic relationship between segmentation accuracy and prediction reliability, the model is guided to focus on the representation and optimization of uncertain regions, thereby achieving accurate and reliable identification and segmentation of difficult-to-segment regions in complex image backgrounds. This has significant implications for the practical application of image segmentation technology.
[0005] This application discloses a reliable image segmentation method for identifying difficult-to-segment regions, comprising the following steps: S1. Obtain the image to be segmented and preprocess it, and divide the preprocessed image to be segmented into training set, validation set and test set; S2. Construct an image segmentation model and calculate the prediction result based on the image segmentation model, wherein the image segmentation model includes: Encoder network: pre-trained MaxViT; Feature optimization network: It adopts a knowledge reconstruction path, which includes multiple interconnected knowledge reconstruction modules; Dual-decoder network: consists of a first decoder and a second decoder with identical structures; S3. Construct a dual-decoder joint loss using a multi-stage feature loss aggregation method and a cognitive self-correction module, and calculate the total loss of the image segmentation model based on the dual-decoder joint loss; the cognitive self-correction module includes uncertainty perception loss and uncertainty-controlled cross-entropy loss; S4. Train the image segmentation model, save the weights of the optimal model, and use the trained model to perform image segmentation tests.
[0006] Preferably, step S2 includes the following steps: The training set is input into the encoder network for multi-scale feature extraction; The multi-scale features are input into the first decoder to obtain the preliminary prediction results. The uncertainty map of the preliminary prediction results and the multi-scale features are input into the knowledge reconstruction path to obtain the reconstructed features. The reconstructed features are input into the second decoder to obtain the final prediction result.
[0007] Preferably, the multi-scale feature extraction includes the following steps: MaxViT is pre-trained and used as an encoder to extract features from the image to be segmented, thus obtaining multi-scale features.
[0008] Preferably, the decoding process of the first decoder and the second decoder includes the following steps: The spatial resolution of the feature map is gradually restored through upsampling operations. The upsampled feature map is then fused with the corresponding layer feature map passed by the skip connection. The fused features undergo multiple stages of convolution operations and are decoded layer by layer to restore the detailed information of the image. Finally, the first decoder generates the multi-scale output and preliminary prediction results of each stage of the first decoder, and the second decoder generates the multi-scale output and final prediction results of each stage of the second decoder.
[0009] Preferably, the knowledge reconstruction path includes the following steps: The uncertainty map of the preliminary prediction result output by the first decoder and the highest resolution feature of the multi-scale features extracted by the encoder network are input into the first layer knowledge reconstruction module, where the uncertainty map is used as a feedback signal and the multi-scale features extracted by the encoder network are used as features to be optimized. After adjustment by the first-layer knowledge reconstruction module, the reconstructed features to be optimized are passed to the next-layer knowledge reconstruction module and used as feedback signals to adjust the features at the current resolution. After layer-by-layer optimization, the final reconstructed feature map can be obtained, in which the multi-scale feature resolution of the input knowledge reconstruction module is reduced layer by layer.
[0010] Preferably, the knowledge reconstruction module reconstructs features by including the following steps: The feedback signal is downsampled by convolution operation, and the downsampled feedback signal is then concatenated with the feature to be optimized along the channel dimension. The concatenated feature maps are processed by parallel dilated convolution operations with different dilation rates. The feature maps after dilation convolution are added element-wise to the features to be optimized. The feature maps obtained by adding elements one by one are convolved to fuse multi-source information, and then the weight distribution of each channel is dynamically adjusted through a channel attention mechanism. The feature map adjusted by the channel attention mechanism is residually connected with the feature map that has not been adjusted by the channel attention mechanism to obtain the reconstructed feature to be optimized.
[0011] Preferably, step S3 includes the following steps: Based on evidence-based deep learning, the first uncertainty graph of the preliminary prediction results and the second uncertainty graph of the final prediction results are calculated. The loss of the first decoder is calculated by performing loss calculations on the multi-scale outputs of each stage of the first decoder and the multi-scale outputs of each stage of the second decoder based on a multi-stage feature fusion loss aggregation method. Multi-stage loss of the second decoder ; The first uncertainty map and the second uncertainty map are jointly input into the cognitive self-correction module to calculate the uncertainty perception loss. Cross-entropy loss due to uncertainty regulation ; Weighted aggregation , , and Total loss .
[0012] Preferably, the joint loss of the dual decoders includes: Multi-stage loss of the first decoder :
[0013] Multi-stage loss of the second decoder :
[0014] Uncertainty-perceived loss :
[0015] Cross-entropy loss due to uncertainty regulation :
[0016] in, This represents a multi-stage feature mixture loss aggregation method. and For a list, This indicates the initial prediction result of the first decoder. This represents the final prediction result of the second decoder. Indicates the true label, This represents the calculation of uncertainty-perceived loss. The calculation of cross-entropy loss represents the effect of uncertainty control. This represents the first uncertainty graph. This represents the second uncertainty plot; The total loss is as follows:
[0017] in, for The weight of the total loss. for The weight of the total loss. for The weight of the total loss. for The weight of the total loss.
[0018] Preferably, the formula for calculating the uncertainty perception loss is as follows:
[0019] in, The normalized uncertainty value, For sensitivity control parameters; The formula for calculating the cross-entropy loss under uncertainty control is as follows:
[0020] in, Represents the cross-entropy loss function. This indicates calculating the average value. Adjusting weights for uncertainty;
[0021] in, This is a hyperparameter for adjusting the intensity.
[0022] This application also discloses a reliable image segmentation apparatus for identifying difficult-to-segment regions, comprising at least one processor and a memory, the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the aforementioned reliable image segmentation method for identifying difficult-to-segment regions.
[0023] The beneficial effects of this invention are: 1. This invention proposes a reliable image segmentation method for identifying difficult-to-segment regions, which can reliably and accurately segment difficult-to-segment regions in complex image backgrounds.
[0024] 2. The knowledge reconstruction module proposed in this invention uses the uncertainty graph as a feedback signal to drive the model to adaptively focus on semantic information related to the uncertainty region, thereby improving the accurate identification of difficult-to-segment regions.
[0025] 3. The cognitive self-correction module proposed in this invention improves the reliability of the model by correcting the discrimination of high uncertainty regions and cognitive bias regions.
[0026] 4. The dual-decoder joint loss proposed in this invention achieves reliable and accurate predictions by supervising the prediction results of different decoding levels, correcting cognitive biases in the model, and improving the overall confidence of the model. Attached Figure Description
[0027] Figure 1 This invention provides a reliable image segmentation method for identifying difficult-to-segment regions. Figure 2 This is a schematic diagram of the image segmentation model and the joint loss structure of the dual decoders in an embodiment of the present invention; Figure 3 This is a knowledge reconstruction path network structure diagram of an embodiment of the present invention; Figure 4 This is a network structure diagram of the knowledge reconstruction module in an embodiment of the present invention. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided with reference to the accompanying drawings and embodiments.
[0029] This application discloses a reliable image segmentation method for identifying difficult-to-segment regions, the process of which is as follows: Figure 1 As shown, it includes the following steps: S1. Acquire and label the images to be segmented. In this embodiment, 3779 abdominal CT images from 30 patients were collected. All images were converted to a resolution of 256×256, and eight organs in the images were labeled: aorta, gallbladder, left kidney, right kidney, pancreas, spleen, stomach, and liver. Then, random rotation and random flipping operations were used to preprocess the images to be segmented, and the images were divided into training, validation, and test sets.
[0030] S2. Construct an image segmentation model, such as Figure 2 As shown, the algorithm includes an encoder network, a feature optimization network, and a dual-decoder network, and then calculates the prediction result based on an image segmentation model. The steps include: inputting the training set into the encoder network for multi-scale feature extraction; inputting the multi-scale features extracted by the encoder into the first decoder to obtain a preliminary prediction result; inputting the uncertainty map of the preliminary prediction result and the multi-scale features into the knowledge reconstruction path to obtain reconstructed features; and inputting the reconstructed features into the second decoder to obtain the final prediction result.
[0031] The encoder network employs MaxViT, a hybrid network based on CNN and Transformer. The overall architecture of MaxViT follows a typical phased design, with each phase consisting of multiple MaxViT blocks concatenated. The core of each MaxViT block lies in achieving a dynamic balance between local detail capture and global relational modeling through cascaded MBConv, block attention, and grid attention. In this embodiment, MaxViT is pre-trained using the ImageNet dataset. MaxViT comprises five phases, with the resolution halved layer by layer. The first phase maps the original input channel count to 96, the second phase maintains the same input and output channel count, and the other phases double the channel count. The trained MaxViT network is used as the encoder to extract features from the input image. Specifically, a 256×256×3 input image is fed into the pre-trained encoder. Downsampling operations are used to progressively reduce the size of the feature maps, while the CNN and Transformer in each layer are combined to gradually extract multi-layered and rich semantic information from the image, resulting in feature maps containing different levels of semantic information in phases 1 to 5. The dimensions are 128×128×96, 64×64×96, 32×32×192, 16×16×384, and 8×8×768, respectively. The above process can be expressed as formula (1). The extracted multi-scale features provide input for the subsequent feature optimization network and the first decoder.
[0032] (1) in, This represents the pre-trained encoder. Indicates the input image. Indicates the encoder's first... layer( The output feature map.
[0033] The dual-decoder network consists of a first decoder (decoder 1 in the figure) and a second decoder (decoder 2 in the figure) with identical structures. Both the first and second decoders use a SEdecoder based on a convolutional neural network to process the original feature map extracted by the encoder and the feature map reconstructed by the feature optimization network, respectively. During the decoding process, the first and second decoders gradually recover the spatial resolution of the feature map through upsampling operations. Subsequently, the upsampled feature map is fused with the corresponding layer feature map passed by the skip connection. The fused features undergo multiple stages of convolution operations, decoding layer by layer to recover the detailed information of the image. Finally, the first decoder generates the multi-scale output and preliminary prediction results of each stage of the first decoder, and the second decoder generates the multi-scale output and final prediction results of each stage of the second decoder. The above process can be expressed as formulas (2)-(4).
[0034] (2) (3) (4) in, Indicates the first decoder, This represents the initial prediction result of the first decoder, with dimensions of 256×256×9. This is a list representing the multi-scale output of each stage of the first decoder, consisting of 4 elements taken from the output of the top 4 stages of the first decoder. Each element has a dimension of 256×256×9. Indicates the knowledge reconstruction path, Indicates the encoder's first... layer( The reconstructed feature maps have dimensions of 128×128×96, 64×64×96, 32×32×192, 16×16×384, and 8×8×768, respectively. The first uncertainty graph represents the initial prediction result of the first decoder, and its dimensions are 256×256×1. Indicates the second decoder. This represents the final prediction result of the second decoder, with dimensions of 256×256×9. This is a list representing the multi-scale output of each stage of the second decoder, consisting of 4 elements taken from the output of the top 4 stages of the second decoder. Each element has a dimension of 256×256×9.
[0035] The feature optimization network employs a knowledge reconstruction path, the structure of which is as follows: Figure 3 As shown, it includes multiple cascaded knowledge reconstruction modules. In this embodiment, the knowledge reconstruction path consists of five cascaded knowledge reconstruction modules. The knowledge reconstruction path uses the uncertainty map of the preliminary prediction result of the first decoder and the multi-scale features generated by the encoder as input, and through adaptive adjustment, it passes from high-resolution features to low-resolution features layer by layer to comprehensively reconstruct information at different scales.
[0036] First, the uncertainty map of the preliminary prediction results output by the first decoder and the highest resolution feature of the multi-scale features extracted by the encoder network are input into the first layer knowledge reconstruction module. The uncertainty map serves as a feedback signal to help the model make targeted adjustments in the feature space, and the multi-scale features extracted by the encoder network serve as features to be optimized.
[0037] After adjustment by the first-layer knowledge reconstruction module, the reconstructed features to be optimized are passed to the next-layer knowledge reconstruction module and used as feedback signals to further adjust the lower-resolution features. After layer-by-layer optimization, the final reconstructed feature map is obtained, in which the multi-scale feature resolution of the input knowledge reconstruction module decreases layer by layer. Each layer of the knowledge reconstruction module optimizes the features at the current resolution level, enabling the model to adaptively learn uncertain regions at different scales. This process, through a top-down layer-by-layer transmission mechanism, ensures that all feature layers from high resolution to low resolution can be fully reconstructed at different scales. The above process can be expressed as formula (5).
[0038] (5) in, Indicating the first step in the knowledge reconstruction path A knowledge reconstruction module, As a feedback signal, in the first knowledge reconstruction module... In the In each knowledge reconstruction module, for .
[0039] like Figure 4As shown, the knowledge reconstruction module first downsamples the feedback signal through a convolution operation with a stride of 2, thereby ensuring that it is aligned with the resolution of the feature to be adjusted, providing a foundation for subsequent feature interaction. Next, the downsampled feedback signal is concatenated with the original feature map (the feature to be optimized) along the channel dimension. This concatenation operation provides sufficient information fusion for subsequent convolutional layers, enabling features from different sources to share information. This process can be expressed as formula (6).
[0040] (6) in, This represents the spliced feature map. This indicates a splicing operation along the channel. This represents a convolution operation with a stride of 2 and a kernel size of 3.
[0041] The concatenated feature maps are processed through parallel dilated convolution operations. These parallel dilated convolutions are set with different dilation rates, which in this embodiment are 2 and 4, respectively, to establish multi-level interactions between feedback signals and features at different receptive field scales, thereby mining contextual information at different scales. The feature maps after dilated convolution are then added element-wise to the original feature maps to enhance and adjust the features while preserving their original characteristics and information. This process can be expressed as formula (7).
[0042] (7) in, This represents the feature map after the addition operation. This indicates a convolution operation with a dilation rate of 2 and a kernel size of 3. This indicates a convolution operation with a dilation rate of 4 and a kernel size of 3.
[0043] The feature maps obtained by element-wise addition are further fused with multi-source information through convolution operations to strengthen the coupling relationship between local structure and semantics. On this basis, a channel attention mechanism is introduced to dynamically adjust the weight distribution of each channel to highlight key features and suppress redundant information. The feature maps adjusted by the channel attention mechanism are residually connected with the feature maps that have not been adjusted by the channel attention mechanism, aiming to retain the original information while guiding the model to focus on structurally complex or semantically ambiguous regions, thus obtaining the reconstructed features to be optimized. The above process can be expressed as formulas (8) and (9).
[0044] (8) (9) in, This represents the feature map after the convolution operation. This indicates a convolution operation with a kernel size of 3. This indicates the channel attention mechanism. This represents the reconstructed feature map.
[0045] S3. Construct a joint loss for the dual decoder using a multi-stage feature loss aggregation method and a cognitive self-correction module. Calculate the total loss of the image segmentation model based on the joint loss of the dual decoder. The process is as follows: Figure 2 As shown.
[0046] The first uncertainty graph is calculated based on the preliminary prediction results using deep learning of evidence. (Uncertainty in the figure) Figure 1 The second uncertainty plot of the final prediction results. (Uncertainty in the figure) Figure 2 The expression is as follows: (10) (11) in, Indicating evidence-based deep learning, and .
[0047] Evidence-based deep learning introduces subjective logic into deep learning, directly modeling the quality of trust and uncertainty in classification by parameterizing the network output through a Dirichlet distribution. It only requires a single forward propagation to output uncertainty, without the need for sampling or ensemble processing. Specifically, for the 9-category segmentation task (8 organ categories and 1 background category) in this embodiment, the network output is first processed through a Softplus activation function to obtain the evidence value for each category. ( ). The larger the value, the higher the value. The more substantial the evidence, the better. Next, the evidence value... It is parameterized as a Dirichlet distribution, as shown in Equation (12).
[0048] (12) ensure This satisfies the Dirichlet distribution definition. Therefore, the overall evidence quantity of the model for the current sample, i.e., the total strength parameter, is... As shown in equation (13) below.
[0049] (13) in, The larger the value, the more total evidence the model has collected, the clearer the sample characteristics, and the higher the model's confidence in the prediction results. Conversely, A smaller value indicates insufficient total evidence, which may be due to fuzzy samples, high noise levels, or samples belonging to unknown categories (out-of-distribution samples). Next, consider the trust quality. It can be calculated for the first The confidence level of the class is obtained, as shown in the following equation (14).
[0050] (14) in, The larger the value, the more likely the model considers the current sample to belong to the first position. The more compelling the evidence of the class, the better. A uniform distribution indicates that the evidence for multiple categories is similar, suggesting that the model has significant uncertainty in making decisions. This typically occurs with samples where the target and background are similar or near the classification boundary.
[0051] Evidence-based deep learning allows for a total confidence level of less than 1, preserving a "ignorance space," where the remaining portion represents the model's overall cognitive uncertainty. As shown in equations (15) and (16) below.
[0052] (15) (16) In summary, uncertainty in evidence-based deep learning is inversely proportional to the total amount of evidence. When the model has no evidence, the trust quality for each category is zero, while the uncertainty is 1. Furthermore, through... Controlling the global scale of uncertainty Refine the confidence levels for each category. Providing overall risk warnings and combining these three elements effectively avoids the misconception in traditional methods that "high probability equals reliability," making the model more robust when dealing with samples with high uncertainty.
[0053] The cognitive self-correction module includes uncertainty perception loss and uncertainty regulation cross-entropy loss. The joint loss of the dual decoder is composed of a multi-stage feature hybrid loss aggregation method and a weighted sum of the proposed uncertainty perception loss and uncertainty regulation cross-entropy loss.
[0054] The loss of the first decoder is calculated by performing multi-scale outputs at each stage of the first and second decoders based on the Multi-Stage Feature Hybrid Loss Aggregation (MUTATION) method. Multi-stage loss of the second decoder The calculation formula is as follows: (17) (18) in, This represents a multi-stage feature mixture loss aggregation method. This indicates the actual label.
[0055] The first uncertainty map corresponding to the preliminary prediction result output by the first decoder and the second uncertainty map corresponding to the final prediction result output by the second decoder are jointly input into the cognitive self-correction module to calculate the uncertainty perception loss. Cross-entropy loss due to uncertainty regulation The calculation formula is as follows: (19) (20) in, This represents the calculation of uncertainty-perceived loss. This represents the calculation of cross-entropy loss due to uncertainty control.
[0056] Weighted aggregation , , and The total loss is obtained and used as the optimization objective for backpropagation. The formula for calculating the total loss is as follows: (twenty one) in, , , and These represent the weights of each loss within the total loss. In this embodiment, Set to 1, Set to 1, Set to 0.1, Set to 1.
[0057] The multi-stage feature loss aggregation method not only calculates the loss on the final segmentation result of the decoder, but also performs deep supervision on the output of each stage of the decoder. Specifically, it first performs deep supervision on the feature maps. , , and Take a non-empty subset to obtain 15 sets. Then sum the elements in each subset to obtain 15 prediction results. Finally, the total loss is the sum of the loss of these 15 prediction results, the loss of the final output of the decoder, and the loss of the ground truth label. The above process can be expressed as the following formulas (22), (23), (24), and (25).
[0058] (twenty two) (twenty three) (twenty four) (25) in, Describes a set of 15 non-empty subsets. The function represents taking a non-empty subset of a list. The number of elements in each set. This represents a set of 15 prediction results. Dice Loss ( The weight of in the total loss is set to 0.7 in this embodiment. Represents cross-entropy loss ( The weight of in the total loss is set to 0.3 in this embodiment. This represents the final output loss value of the multi-stage feature loss aggregation method. surface The final output of the decoder, Indicates a truth value label.
[0059] Uncertainty-aware loss aims to guide the optimization of deep learning models by quantifying different levels of uncertainty, so that they can make more effective adjustments when faced with uncertain data. Its core formula is shown in the following equation (26).
[0060] (26) in, This represents the normalized uncertainty value, ranging from [0,1]. This is the sensitivity control parameter. The loss function operates under uncertainty. The expression shows a monotonically increasing trend, that is, when When the value increases, the loss value increases accordingly; conversely, when the value decreases, the loss value decreases. As the value decreases, the loss value decreases. Key parameters of the loss function. It controls the model's sensitivity to changes in uncertainty. When When it is large, The changes are more pronounced. The changes are also more drastic, therefore the loss function is more sensitive to uncertainty. Changes in [the environment] are more sensitive. Conversely, when [the changes are more sensitive] When the value is small, the change in the loss function is relatively small, resulting in a more gradual response of the model to changes in uncertainty. Therefore, the model's motivation to reduce uncertainty is relatively weak. In this embodiment, It is set to 1.
[0061] Furthermore, the uncertainty-aware loss function has a range between (0,1), and this boundedness effectively prevents numerical explosion problems that may occur during optimization, ensuring that the loss function remains numerically stable throughout the training process. Simultaneously, the smooth gradient characteristic of the loss function guarantees the stationarity of the gradient during backpropagation, avoiding gradient explosion or vanishing phenomena. A key aspect of the loss function is that... The nearest maximum gradient signal indicates that more learning and adjustments are needed. This setup aligns with human intuition: under moderate uncertainty, it is especially important to focus on and deeply analyze this data to achieve more efficient improvements.
[0062] The uncertainty-controlled cross-entropy loss function is based on a dual-decoder architecture and constructs a framework for comparing the uncertainty of the initial prediction and the final prediction. Its core idea is that when the model shows a discrepancy between prediction confidence and accuracy, the model is guided to correct its potential erroneous learning path by adjusting the weights of the cognitive bias region.
[0063] Specifically, this embodiment focuses on the following two types of "cognitive biases": First, although the model predicts correctly, its uncertainty increases, suggesting that the model's cognition is still unstable; second, the model predicts incorrectly but exhibits high confidence, a typical phenomenon of "overconfidence." For these two situations, an uncertainty adjustment weight is introduced. This is used in the weighted cross-entropy loss function. Adjusting the weights... The definition is shown in the following equation (27): (27) in, In this embodiment, a hyperparameter for adjusting the intensity is used to control the degree of influence of uncertainty changes on the loss weight. It was set to 3.
[0064] Weights based on uncertainty The definition of the cross-entropy loss under uncertainty control is shown in equation (28) below: (28) in, Represents the cross-entropy loss function. This indicates calculating the average value.
[0065] S4. Train the image segmentation model for 300 rounds. After each round of training, evaluate the model performance on the validation set and save the model weights corresponding to the highest Dice coefficients in the validation set. After training is complete, use the model with these optimal weights to segment and predict the test images.
[0066] This application also discloses a reliable image segmentation apparatus for identifying difficult-to-segment regions, including at least one processor and a memory. The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above-described reliable image segmentation method for identifying difficult-to-segment regions.
[0067] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.
Claims
1. A reliable image segmentation method for identifying difficult-to-segment regions, characterized in that, Includes the following steps: S1. Obtain the image to be segmented and preprocess it, and divide the preprocessed image to be segmented into training set, validation set and test set; S2. Construct an image segmentation model and calculate the prediction result based on the image segmentation model, wherein the image segmentation model includes: Encoder network: pre-trained MaxViT; Feature optimization network: It adopts a knowledge reconstruction path, which includes multiple interconnected knowledge reconstruction modules; Dual-decoder network: consists of a first decoder and a second decoder with identical structures; S3. Construct a dual-decoder joint loss using a multi-stage feature loss aggregation method and a cognitive self-correction module, and calculate the total loss of the image segmentation model based on the dual-decoder joint loss; the cognitive self-correction module includes uncertainty perception loss and uncertainty-controlled cross-entropy loss; S4. Train the image segmentation model, save the weights of the optimal model, and use the trained model to perform image segmentation tests.
2. The reliable image segmentation method for identifying difficult-to-segment regions according to claim 1, characterized in that, S2 includes the following steps: The training set is input into the encoder network for multi-scale feature extraction; The multi-scale features are input into the first decoder to obtain the preliminary prediction results. The uncertainty map of the preliminary prediction results and the multi-scale features are input into the knowledge reconstruction path to obtain the reconstructed features. The reconstructed features are input into the second decoder to obtain the final prediction result.
3. The reliable image segmentation method for identifying difficult-to-segment regions according to claim 2, characterized in that, The multi-scale feature extraction includes the following steps: MaxViT is pre-trained and used as an encoder to extract features from the image to be segmented, thus obtaining multi-scale features.
4. The reliable image segmentation method for identifying difficult-to-segment regions according to claim 3, characterized in that, The decoding process of the first decoder and the second decoder includes the following steps: The spatial resolution of the feature map is gradually restored through upsampling operations. The upsampled feature map is then fused with the corresponding layer feature map passed by the skip connection. The fused features undergo multiple stages of convolution operations and are decoded layer by layer to restore the detailed information of the image. Finally, the first decoder generates the multi-scale output and preliminary prediction results of each stage of the first decoder, and the second decoder generates the multi-scale output and final prediction results of each stage of the second decoder.
5. The reliable image segmentation method for identifying difficult-to-segment regions according to claim 4, characterized in that, The knowledge reconstruction path includes the following steps: The uncertainty map of the preliminary prediction result output by the first decoder and the highest resolution feature of the multi-scale features extracted by the encoder network are input into the first layer knowledge reconstruction module, where the uncertainty map is used as a feedback signal and the multi-scale features extracted by the encoder network are used as features to be optimized. After adjustment by the first-layer knowledge reconstruction module, the reconstructed features to be optimized are passed to the next-layer knowledge reconstruction module and used as feedback signals to adjust the features at the current resolution. After layer-by-layer optimization, the final reconstructed feature map can be obtained, in which the multi-scale feature resolution of the input knowledge reconstruction module is reduced layer by layer.
6. The reliable image segmentation method for identifying difficult-to-segment regions according to claim 5, characterized in that, The knowledge reconstruction module reconstructs features by including the following steps: The feedback signal is downsampled by convolution operation, and the downsampled feedback signal is then concatenated with the feature to be optimized along the channel dimension. The concatenated feature maps are processed by parallel dilated convolution operations with different dilation rates. The feature maps after dilation convolution are added element-wise to the features to be optimized. The feature maps obtained by adding elements one by one are convolved to fuse multi-source information, and then the weight distribution of each channel is dynamically adjusted through a channel attention mechanism. The feature map adjusted by the channel attention mechanism is residually connected with the feature map that has not been adjusted by the channel attention mechanism to obtain the reconstructed feature to be optimized.
7. The reliable image segmentation method for identifying difficult-to-segment regions according to claim 6, characterized in that, S3 includes the following steps: Based on evidence-based deep learning, the first uncertainty graph of the preliminary prediction results and the second uncertainty graph of the final prediction results are calculated. The loss of the first decoder is calculated by performing loss calculations on the multi-scale outputs of each stage of the first decoder and the multi-scale outputs of each stage of the second decoder based on a multi-stage feature fusion loss aggregation method. Multi-stage loss of the second decoder ; The first uncertainty map and the second uncertainty map are jointly input into the cognitive self-correction module to calculate the uncertainty perception loss. Cross-entropy loss due to uncertainty regulation ; Weighted aggregation , , and Total loss .
8. The reliable image segmentation method for identifying difficult-to-segment regions according to claim 7, characterized in that, The joint loss of the dual decoders includes: Multi-stage loss of the first decoder : Multi-stage loss of the second decoder : Uncertainty-perceived loss : Cross-entropy loss due to uncertainty regulation : in, This represents a multi-stage feature mixture loss aggregation method. and For a list, This indicates the initial prediction result of the first decoder. This represents the final prediction result of the second decoder. Indicates the true label, This represents the calculation of uncertainty-perceived loss. The calculation of cross-entropy loss represents the effect of uncertainty control. This represents the first uncertainty graph. This represents the second uncertainty plot; The total loss is as follows: in, for The weight of the total loss. for The weight of the total loss. for The weight of the total loss. for The weight of the total loss.
9. The reliable image segmentation method for identifying difficult-to-segment regions according to claim 8, characterized in that, The formula for calculating the uncertainty perception loss is as follows: in, The normalized uncertainty value, For sensitivity control parameters; The formula for calculating the cross-entropy loss under uncertainty control is as follows: in, Represents the cross-entropy loss function. This indicates calculating the average value. Adjusting weights for uncertainty; in, This is a hyperparameter for adjusting the intensity.
10. A reliable image segmentation device for identifying difficult-to-segment regions, characterized in that, It includes at least one processor and a memory, the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the reliable image segmentation method for identifying difficult-to-segment regions as described in any one of claims 1 to 9.