Uncertainty-based remote sensing image segmentation restoration method and device, and storage medium
By quantifying and adjusting repair weights based on aleatoric and epistemic uncertainties, the method enhances remote sensing image segmentation accuracy and stability by addressing noise and model limitations.
Patent Information
- Application Number
- CN202510414673.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-15
AI Technical Summary
The existing remote sensing image segmentation technology fails to effectively quantify and dynamically adjust uncertainty, resulting in noise interference and model prediction deviations, affecting segmentation accuracy and stability.
The remote sensing image segmentation repair method based on uncertainty is adopted, and through feature extraction and encoding, uncertainty error estimation and guidance repair modules, accidental and cognitive uncertainty are quantified, error correction and repair strategies are dynamically optimized, and segmentation accuracy and stability are improved.
Reduce noise interference and model prediction deviation, improve the accuracy and stability of remote sensing image segmentation, and enhance the adaptability to complex scenes.
Smart Images

Figure CN120318124A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing image processing, and particularly relates to a remote sensing image segmentation and restoration method, device and storage medium based on uncertainty. Background Art
[0002] Remote sensing image segmentation is a core technology in fields such as ground object recognition, urban planning, and environmental monitoring. It not only faces challenges such as noise interference, occlusion, and complex ground object distribution in high-resolution and multi-spectral data, but is also affected by uncertainty. Among them, the aleatoric uncertainty derived from data noise may directly lead to a decrease in the prediction confidence of the model for local pixels, especially prone to under-segmentation or over-segmentation errors in noise-intensive areas; while the epistemic uncertainty caused by insufficient model parameters or deviation in the training data distribution may cause the model to have prediction biases or segmentation errors due to cognitive limitations.
[0003] In response to the above problems, the existing technology extracts features through a convolutional neural network. Although progress has been made in semantic segmentation, there are still the following limitations: 1. The uncertainty in prediction is not quantified, resulting in the accumulation of errors and the inability to dynamically adjust the repair weights; 2. Only a single type of uncertainty is concerned, resulting in a lack of comprehensiveness in the repair strategy. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a remote sensing image segmentation and restoration method, device and storage medium based on uncertainty, which can dynamically adjust the repair weights by calculating and quantifying uncertainty, and improve the accuracy and stability of remote sensing image segmentation.
[0005] To achieve the above purpose, the present invention is implemented by adopting the following technical solutions:
[0006] In the first aspect, the present invention provides a remote sensing image segmentation and restoration method based on uncertainty, and the method includes:
[0007] Obtain a remote sensing image, and perform preprocessing on the remote sensing image to generate a remote sensing image to be repaired and an initial segmentation mask;
[0008] Input the remote sensing image and the initial segmentation mask into a trained remote sensing image segmentation and restoration model, where the remote sensing image segmentation and restoration model includes a feature extraction and encoding module, an uncertainty error estimation module, and an uncertainty-guided repair module;
[0009] Through the feature extraction and encoding module, feature extraction and fusion are performed on the remote sensing image and the initial segmentation mask to generate multi-scale high-level semantic features, and feature encoding is performed on the initial segmentation mask to generate initial segmentation features;
[0010] Through the uncertainty error estimation module, feature enhancement, mapping, and uncertainty calculation are performed on the multi-scale high-level semantic features and the initial segmentation features to obtain error estimation features, error estimation results, and total uncertainty respectively, and the error estimation results and the total uncertainty are weighted to obtain an uncertainty weighted result;
[0011] Through the uncertainty-guided repair module, fusion and correction are performed on the initial segmentation features, the error estimation features, and the uncertainty weighted result to obtain a segmentation repair result.
[0012] Combined with the first aspect, further, the method for preprocessing the remote sensing image includes: image standardization, size unification, and morphological perturbation.
[0013] Combined with the first aspect, further, the feature extraction and encoding module includes a first splicing module, an HRNet-W48 module, and an IS-Encoder module;
[0014] The feature extraction and encoding module performs the following processing steps on the input remote sensing image and initial segmentation mask:
[0015] The remote sensing image and the initial segmentation mask are spliced in the channel dimension through the first splicing module to obtain a spliced input feature;
[0016] The spliced input feature is input into the HRNet-W48 module, and multi-scale feature extraction is performed through an encoder-decoder structure to generate multi-scale features; then the multi-scale features are upsampled, and spliced in the channel dimension to fuse multi-scale information, and then the number of channels is compressed through a 1×1 first convolutional layer to generate multi-scale high-level semantic features;
[0017] The initial segmentation mask is input into the IS-Encoder module, and feature encoding is performed through a cascaded 3×3 second convolutional layer and 3×3 third convolutional layer to generate initial segmentation features.
[0018] Combined with the first aspect, further, the uncertainty error estimation module includes a difference module, a channel spatial attention module, a first Softmax function, an uncertainty module, a fusion module, and a weighting module;
[0019] The uncertainty error estimation module performs the following processing steps on the input multi-scale high-level semantic features and initial segmentation features:
[0020] Subtract the multi-scale high-level semantic features and the initial segmentation features element-wise at the same spatial scale through the difference module to obtain the attention input feature (Difference Feature):
[0021] ,
[0022] wherein, is the attention input feature or difference feature; are the multi-scale high-level semantic features; is the initial segmentation feature;
[0023] After the attention input feature is input into the channel attention module in the channel-spatial attention module, it is processed by the first average pooling layer and the first max pooling layer respectively to obtain the first average response feature and the first max response feature in the spatial dimension; then the first average response feature and the first max response feature are respectively input into a shared multi-layer perceptron for learning to obtain the average response enhanced feature and the max response enhanced feature; then the average response enhanced feature and the max response enhanced feature are fused by the first adder to obtain the fused feature; finally, after the fused feature is calculated by the first Sigmoid function to obtain the channel attention weight, the attention input feature is fused by the first multiplier to generate the enhanced feature based on the channel relationship;
[0024] After the attention input feature is input into the spatial attention module in the channel-spatial attention module, it is processed by the second average pooling layer and the second max pooling layer respectively to obtain the second average response feature and the second max response feature in the channel dimension; then after the second average response feature and the second max response feature are concatenated according to the channel dimension, they are sequentially subjected to convolution processing by a 7×7 fourth convolutional layer, calculated by the second Sigmoid function to obtain the spatial attention weight, and then the attention input feature is fused by the second multiplier to generate the enhanced feature based on the spatial relationship;
[0025] Fuse the attention input feature, the enhanced feature based on the channel relationship, and the enhanced feature based on the spatial relationship through the second adder to obtain the error estimation feature;
[0026] Map the error estimation feature through the first Softmax function to multiple channels corresponding to correctly segmented pixels, under-segmented pixels, and over-segmented pixels respectively to obtain the error estimation result;
[0027] After the error estimation result is input into the uncertainty module, different prediction categories are simulated through the Monte Carlo sampling layer, and the aleatoric uncertainty and the epistemic uncertainty are calculated respectively;
[0028] Fuse the aleatoric uncertainty and the epistemic uncertainty through the fusion module to obtain the total uncertainty;
[0029] Weight the error estimation result and the total uncertainty through the weighting module to obtain the uncertainty weighted result.
[0030] In combination with the first aspect, further, the formula for calculating the aleatoric uncertainty is:
[0031] ,
[0032] wherein, is the aleatoric uncertainty; is the predicted probability of the predicted class ; is the set of predicted classes;
[0033] The formula for calculating the epistemic uncertainty is:
[0034] ,
[0035] ,
[0036] wherein, is the epistemic uncertainty; is the error estimation result after the -th Monte Carlo sampling layer sampling; is the mean value of the -th sampling, that is, the stable prediction value; is the total number of samplings;
[0037] The formula for calculating the total uncertainty is:
[0038] ,
[0039] wherein, is the total uncertainty; and are both fusion weights, and .
[0040] In combination with the first aspect, further, the uncertainty-guided repair module includes an EGF module;
[0041] The uncertainty-guided repair module performs the following processing steps on the input initial segmentation features, error estimation features, and uncertainty weighted result:
[0042] Input the initial segmentation features, the error estimation features, and the uncertainty weighted result into the EGF module;
[0043] Splicing is performed in the channel dimension through a second splicing module to obtain an uncertainty splicing feature;
[0044] The uncertainty splicing feature is successively passed through a 3×3 fifth convolutional layer, a 5×5 depthwise separable convolutional layer, and a 1×1 sixth convolutional layer for feature extraction and correction, and is transformed through a second Softmax function to obtain a segmentation repair result.
[0045] Combined with the first aspect, further, the training method of the remote sensing image segmentation and repair model includes:
[0046] Select a remote sensing image dataset;
[0047] Preprocess the original remote sensing images in the dataset to obtain a processed dataset;
[0048] Randomly divide the processed dataset into a training set, a validation set, and a test set according to a ratio of 6:2:2;
[0049] Use the training set to train a pre-constructed remote sensing image segmentation and repair model, use the validation set to validate the trained remote sensing image segmentation and repair model, use the test set to test the validated remote sensing image segmentation and repair model, and take the remote sensing image segmentation and repair model with the optimal test result as the finally trained remote sensing image segmentation and repair model.
[0050] Combined with the first aspect, further, the training method of the remote sensing image segmentation and repair model further includes: designing a loss function and an adaptive optimization strategy to achieve end-to-end joint training and model optimization, where the loss function includes an auxiliary branch loss function, an error estimation loss function, a negative log-likelihood loss function, and a guided repair loss function;
[0051] The expression of the auxiliary branch loss function is:
[0052] ,
[0053] In the formula, is the auxiliary branch loss function; is the cross-entropy loss function; is the weight parameter for controlling the cross-entropy loss function, set to 0.5; is the Dice loss function; is the weight parameter for controlling the Dice loss function, set to 0.5; is the output result of the auxiliary branch module; is the ground truth corresponding to the output result of the auxiliary branch module;
[0054] The expression of the error estimation loss function is as follows:
[0055] ,
[0056] wherein, is the error estimation loss function; is the weight parameter for controlling the cross-entropy loss function, set to 0.5; is the weight parameter for controlling the Dice loss function, set to 0.5; is the output result of the uncertainty error estimation module; is the ground truth corresponding to the output result of the uncertainty error estimation module;
[0057] The expression of the negative log-likelihood loss function is as follows:
[0058] ,
[0059] wherein, is the negative log-likelihood loss function; is the total uncertainty at the pixel level of the th sample; is the true label; is the predicted value;
[0060] The expression of the guided repair loss function is as follows:
[0061] ,
[0062] wherein, is the guided repair loss function; is the weight parameter for controlling the cross-entropy loss function, set to 0.5; is the weight parameter for controlling the Dice loss function, set to 0.5; is the output result of the uncertainty guided repair module; is the ground truth corresponding to the output result of the uncertainty guided repair module;
[0063] The total loss function is defined as follows:
[0064] ,
[0065] wherein, is the total loss function; are all the weights of each loss function, set to 0.2, 0.3, 0.2, 0.3 in sequence.
[0066] In the second aspect, the present invention further provides a computer device, including a storage medium and a processor;
[0067] The storage medium is used to store instructions;
[0068] The processor is configured to operate according to the instructions to execute the steps of the method according to any one of the first aspect.
[0069] In a third aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method according to any one of the first aspect are implemented.
[0070] Compared with the prior art, the present invention can at least achieve the following beneficial effects:
[0071] 1. The method for segmenting and repairing remote sensing images based on uncertainty provided by the present invention introduces uncertainty weights in the segmentation error correction stage, calculates aleatoric uncertainty and epistemic uncertainty, quantifies the segmentation errors caused by data noise and model cognitive limitations in the segmentation process, and dynamically optimizes the error correction and repair strategies according to different types of uncertainty to reduce noise interference and model prediction deviation, thereby improving the accuracy and stability of remote sensing image segmentation;
[0072] 2. The method for segmenting and repairing remote sensing images based on uncertainty provided by the present invention can optimize the segmentation result in an end-to-end framework by combining an uncertainty-guided repair strategy, reduce error accumulation, and improve the adaptability of the model to complex scenes, thereby enhancing the accuracy and generalization ability of remote sensing image segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0074] Figure 1 is a flowchart of a method for segmenting and repairing remote sensing images based on uncertainty provided by an embodiment of the present invention;
[0075] Figure 2 is a schematic structural diagram of a model for segmenting and repairing remote sensing images based on uncertainty provided by an embodiment of the present invention;
[0076] Figure 3 is a schematic structural diagram of an HRNet-W48 module provided by an embodiment of the present invention;
[0077] Figure 4 is a schematic structural diagram of an IS-Encoder module provided by an embodiment of the present invention;
[0078] Figure 5 It is a schematic structural diagram of a channel spatial attention module provided by an embodiment of the present invention;
[0079] Figure 6 It is a schematic structural diagram of an uncertainty module provided by an embodiment of the present invention;
[0080] Figure 7 It is a schematic structural diagram of an EGF module provided by an embodiment of the present invention;
[0081] Figure 8 It is a visualization image of the model provided by an embodiment of the present invention during the inference phase of the WHU Building Dataset test set;
[0082] Figure 9 It is an internal structure diagram of a computer device provided by an embodiment of the present invention. Detailed implementation manners
[0083] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.
[0084] Embodiment 1:
[0085] This embodiment provides a remote sensing image segmentation and repair method based on uncertainty. As Figure 1 shown, it is a flowchart of the method provided in this embodiment, mainly including the following steps:
[0086] Step S1: Obtain a remote sensing image, and preprocess the remote sensing image to generate a remote sensing image to be repaired and an initial segmentation mask;
[0087] Step S2: Input the remote sensing image and the initial segmentation mask into a trained remote sensing image segmentation and repair model;
[0088] Step S3: Through the feature extraction and encoding module, extract and fuse the features of the remote sensing image and the initial segmentation mask to generate multi-scale high-level semantic features, and perform feature encoding on the initial segmentation mask to generate initial segmentation features;
[0089] Step S4: Through the uncertainty error estimation module, perform feature enhancement, mapping, and calculation of uncertainty on the multi-scale high-level semantic features and the initial segmentation features to obtain error estimation features, error estimation results, and total uncertainty respectively, and weight the error estimation results and the total uncertainty to obtain an uncertainty weighted result;
[0090] Step S5: Through the uncertainty-guided repair module, fuse and correct the initial segmentation features, error estimation features, and uncertainty weighted results to obtain a segmentation and repair result.
[0091] It should be noted that, as Figure 2 shown, it is a schematic structural diagram of the remote sensing image segmentation and repair model based on uncertainty provided in this embodiment, mainly including a feature extraction and encoding module, an uncertainty error estimation module, an uncertainty-guided repair module, and an auxiliary branch module; among them, the first output end of the feature extraction and encoding module is connected to the auxiliary branch module, and the second output end is sequentially connected to the uncertainty error estimation module and the uncertainty-guided repair module. In addition, the auxiliary branch module is used to calculate the auxiliary branch loss function during the model training process to enhance the stability of the model training and is not used when actually segmenting and repairing remote sensing images.
[0092] Furthermore, the feature extraction and encoding module uses HRNet-W48 as the feature extraction backbone. Referring to Figure 2 , it mainly includes a first splicing module, an HRNet-W48 module, and an IS-Encoder module; among them, after the first splicing module is connected in series with the HRNet-W48 module, it is connected in parallel with the IS-Encoder module, and the input end of the IS-Encoder module is the second input end of the first splicing module. The following combines Figure 2 to further elaborate in detail on how to extract and fuse features of the remote sensing image and the initial segmentation mask through the feature extraction and encoding module in step S3 to generate multi-scale high-level semantic features and perform feature encoding on the initial segmentation mask to generate initial segmentation features:
[0093] The remote sensing image and the initial segmentation mask are spliced in the channel dimension (4 channels) through the first splicing module to obtain the spliced input feature;
[0094] The spliced input feature is input into the HRNet-W48 module for feature extraction and fusion to generate multi-scale high-level semantic features;
[0095] The initial segmentation mask (including simulated errors) is input into the IS-Encoder module. As Figure 4 shown, it is a schematic structural diagram of the IS-Encoder module provided in this embodiment, mainly including a 3×3 second convolutional layer and a 3×3 third convolutional layer connected in series in sequence; the two convolutional layers (with a stride of 1 and a padding of 1) sequentially perform feature encoding on the initial segmentation mask to extract the spatial distribution pattern and generate the initial segmentation feature (Initial Feature), thereby retaining the position and morphological information of the error area.
[0096] Specifically, as Figure 3As shown, it is a schematic structural diagram of the HRNet-W48 module provided in this embodiment. The classifier module of the original network is removed, and it mainly includes an encoder-decoder structure and a 1×1 first convolutional layer. After inputting the concatenated input features into the HRNet-W48 module, the encoder-decoder structure will perform multi-scale feature extraction through parallel multi-resolution sub-networks to retain spatial details and semantic information at different scales, and will output four groups of multi-scale feature maps with resolutions of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the concatenated input features respectively. Then, after upsampling the four groups of multi-scale feature maps to the same resolution, they are concatenated along the channel dimension to fuse multi-scale information, and feature mapping is performed at different scales (1 / 4, 1 / 8, 1 / 16, and 1 / 32) through the 1×1 first convolutional layer to compress the number of channels, generating multi-scale and high-level semantic features, taking into account local details and global context.
[0097] In this embodiment, referring to Figure 2 , the uncertainty error estimation module mainly includes a difference module, a channel-spatial attention module, a first Softmax function, an uncertainty module, a fusion module, and a weighting module; among them, the output ends of the HRNet-W48 module and the IS-Encoder module in the feature extraction and encoding module are commonly connected to the input end of the difference module, the output end of the difference module is sequentially connected in series with the channel-spatial attention module, the first Softmax function, and the uncertainty module, the first output end and the second output end of the uncertainty module are commonly connected to the input end of the fusion module, and the output end of the fusion module and the output end of the first Softmax function are commonly connected to the input end of the weighting module.
[0098] The following combines Figure 2 , and further elaborates in detail on how to perform feature enhancement, mapping, and calculation of uncertainty on the multi-scale high-level semantic features and the initial segmentation features through the uncertainty error estimation module in step S4, respectively obtaining the error estimation features, the error estimation results, and the total uncertainty, and weighting the error estimation results and the total uncertainty to obtain the uncertainty weighted result:
[0099] The multi-scale high-level semantic features and the initial segmentation features are subtracted at the same spatial scale through the difference module to obtain the attention input features;
[0100] The attention input features are input into the channel-spatial attention module for feature enhancement to obtain the error estimator feature, which represents the error or uncertainty generated during the prediction process and is used for subsequent uncertainty calculation;
[0101] The error estimation features are passed through a first Softmax function and a 1×1 convolutional layer to map the input channels to three output channels corresponding to correctly segmented pixels, under-segmented pixels, and over-segmented pixels respectively, obtaining the error estimator prediction;
[0102] The error estimator prediction is input into the uncertainty module. As Figure 6 shown, it is the structural schematic diagram of the uncertainty module provided in this embodiment, mainly including a Monte Carlo sampling layer, a module for calculating aleatoric uncertainty, and a module for calculating epistemic uncertainty. Among them, the output end of the Monte Carlo sampling layer is respectively connected to the input ends of the module for calculating aleatoric uncertainty and the module for calculating epistemic uncertainty; after the error estimator prediction simulates different prediction categories through the Monte Carlo sampling layer, the aleatoric uncertainty and the epistemic uncertainty are respectively calculated;
[0103] The aleatoric uncertainty and the epistemic uncertainty are fused through a fusion module to obtain the total uncertainty;
[0104] The error estimator prediction and the total uncertainty are weighted and adjusted pixel by pixel through a weighting module to obtain the error estimator uncertainty-weighted prediction indicating the mislabeled areas or inaccurate areas in the model prediction, enabling the high-uncertainty areas to receive more attention during the error estimation stage, thereby improving the reliability of error repair. It can be expressed by the following formula:
[0105] ,
[0106] In the formula, is the error estimator uncertainty-weighted prediction; is the error estimator prediction; is the inverse mapping function, indicating weight decay for high-uncertainty areas and assigning higher weights to low-uncertainty areas; is the total uncertainty; is a decay coefficient greater than zero to control the influence of uncertainty on the repair weight, set to 1.5.
[0107] Specifically, aleatoric uncertainty stems from the noise in the input data itself, such as inevitable factors like image quality, lighting changes, annotation errors, etc. To model aleatoric uncertainty, in this embodiment, based on the error estimation results, entropy is used to predict the pixel-level noise variance to calculate aleatoric uncertainty:
[0108] ,
[0109] where, is the aleatoric uncertainty; is the predicted probability of the predicted class ; is the set of predicted classes.
[0110] Meanwhile, epistemic uncertainty is usually caused by insufficient data or limited model capabilities, reflecting the unknown nature of the model itself; and it is used to measure the confidence of the model, indicating whether the model lacks stability in predicting certain regions. To estimate epistemic uncertainty, in this embodiment, the MC dropout (Monte Carlo Dropout) method is adopted, enabling dropout in the inference stage, that is, randomly discarding a part of the network weights during each inference, so that the model gives different results in different inference processes. By sampling the error estimation results multiple times and calculating the variance of epistemic uncertainty for each pixel, regions where the model lacks confidence during the prediction process can be effectively identified:
[0111] ,
[0112] ,
[0113] where, is the epistemic uncertainty; is the error estimation result after the -th Monte Carlo sampling layer sampling; is the -th sampling mean, that is, the stable prediction value; is the total number of samplings.
[0114] In addition, to comprehensively measure uncertainty, in this embodiment, aleatoric uncertainty and epistemic uncertainty are weighted and fused to form the final uncertainty metric for identifying high-uncertainty regions, that is, regions where errors may exist:
[0115] ,
[0116] where, is the total uncertainty; and are both fusion weights, which can be adjusted according to the experimental results, and .
[0117] It should be noted that when the uncertainty error estimation module performs error estimation for image segmentation, it needs to analyze multi-scale high-level semantic features and initial segmentation features from both global and local scales. Among them, the global scale is used to discover large-scale annotation errors, such as error regions in large areas; the local scale is used to capture small annotation errors for fine analysis, such as errors at the edges or details of objects.
[0118] The remote sensing image segmentation and repair method based on uncertainty provided in this embodiment introduces uncertainty weights in the segmentation error correction stage, calculates aleatoric uncertainty and epistemic uncertainty, quantifies the segmentation errors caused by data noise and model cognitive limitations in the segmentation process, and dynamically optimizes the error correction and repair strategies according to different types of uncertainties, so as to reduce noise interference and model prediction bias, thereby improving the accuracy and stability of remote sensing image segmentation.
[0119] As an optional embodiment, referring to Figure 2 , in this embodiment, a refinement component - the uncertainty-guided repair module is designed to guide the correction of the prediction result through error information and uncertainty in the remote sensing image segmentation and repair model, mainly including the EGF module. As Figure 7 shown, it is the structural schematic diagram of the EGF module provided in this embodiment, mainly including a second splicing module, a 3×3 fifth convolutional layer, a 5×5 depthwise separable convolutional layer, a 1×1 sixth convolutional layer, and a second Softmax function connected in series in sequence; among them, the output end of the IS-Encoder module in the feature extraction and encoding module, the output end of the channel spatial attention module in the uncertainty error estimation module, and the output end of the weighting module are commonly connected to the input end of the second splicing module.
[0120] Next, in combination with Figure 7 , a further detailed description will be given on how to fuse and correct the initial segmentation feature, error estimation feature, and uncertainty weighting result through the uncertainty-guided repair module in step S5 to obtain the segmentation and repair result:
[0121] Input the initial segmentation feature, error estimation feature, and uncertainty weighting result into the EGF module;
[0122] Perform splicing and fusion in the channel dimension through the second splicing module, and reduce the number of channels from 2c + 3 to c to obtain the uncertainty splicing feature;
[0123] The uncertainty splicing features are successively passed through a 3×3 fifth convolutional layer, a 5×5 depthwise separable convolutional layer (Depthwise Separable Convolutions), and a 1×1 sixth convolutional layer for feature extraction and correction, which helps to extract more detailed information, and is then transformed through a second Softmax function to obtain the segmentation repair result.
[0124] The method for segmenting and repairing remote sensing images based on uncertainty provided in this embodiment can optimize the segmentation result in an end-to-end framework, reduce error accumulation, and improve the adaptability of the model to complex scenarios by combining an uncertainty-guided repair strategy, thereby improving the accuracy and generalization ability of remote sensing image segmentation.
[0125] This flowchart only shows the logical order of the method described in this embodiment. On the premise of non-conflict, in other possible embodiments of the present invention, the steps shown or described can be completed in a different order from Figure 1 the order shown. The method for segmenting and repairing remote sensing images based on uncertainty provided in this embodiment can be applied to a terminal and can be executed by a device for segmenting and repairing remote sensing images based on uncertainty. This device can be implemented in a software and / or hardware manner and can be integrated in the terminal, such as: any smart phone, tablet computer, or computer device with a communication function.
[0126] Embodiment 2:
[0127] The method for segmenting and repairing remote sensing images based on uncertainty provided in this embodiment is different from Embodiment 1 in that, in order to improve the error estimation ability of the uncertainty error estimation module, this embodiment uses a channel spatial attention module (CSAM) to enhance features of different scales, which can extract feature information from the channel dimension and the spatial dimension respectively to ensure that the model can more effectively capture the error region.
[0128] As Figure 5 shown, it is a structural schematic diagram of the channel spatial attention module provided in this embodiment, which mainly includes a channel attention module (CAM), a spatial attention module (SAM), and a second adder connected in parallel. Referring to Figure 5 , the channel spatial attention module performs the following processing steps on the input attention input features:
[0129] After the attention input feature is input into the channel attention module (CAM) in the channel-spatial attention module, it is respectively operated by the first average pooling layer (Average Pooling) and the first max pooling layer (Max Pooling) to extract global channel features in the spatial dimension, forming two different channel feature maps, namely the first average response feature and the first max response feature; then the first average response feature and the first max response feature are respectively fed into a shared multi-layer perceptron (MLP) for learning to obtain the average response enhanced feature and the max response enhanced feature; then the average response enhanced feature and the max response enhanced feature are fused by the first adder through an Add operation to obtain a fused feature; finally, after the fused feature is calculated by the first Sigmoid function to obtain the channel attention weight, the channel attention weight is multiplied element-wise with the attention input feature by the first multiplier to enhance the important channel information and generate an enhanced feature based on the channel relationship;
[0130] After the attention input feature is input into the spatial attention module (SAM) in the channel-spatial attention module, it is respectively processed by the second average pooling layer and the second max pooling layer to extract the saliency information in the spatial dimension, obtaining the second average response feature and the second max response feature in the channel dimension; then the second average response feature and the second max response feature are concatenated according to the channel dimension and then sequentially subjected to convolution processing by the 7×7 fourth convolutional layer to enhance the features in the space, calculated by the second Sigmoid function to obtain the spatial attention weight, and then the spatial attention weight is multiplied element-wise with the attention input feature by the second multiplier to enhance the response to the spatially significant region and generate an enhanced feature based on the spatial relationship;
[0131] The attention input feature, the enhanced feature based on the channel relationship, and the enhanced feature based on the spatial relationship are fused by the second adder to obtain an error estimation feature.
[0132] In some embodiments, referring to Figure 2 , an auxiliary branch module (Auxiliary Branch, AB) with multi-scale high-level semantic features as the input is designed to calculate the auxiliary branch loss function during the model training process to enhance the stability of model training; it mainly includes a cascaded auxiliary fusion module and a third Softmax function.
[0133] As an alternative embodiment, the method for preprocessing the remote sensing image in step S1 mainly includes: image normalization, size unification, and morphological perturbation, and its processing method will be specifically described in Example III for training the remote sensing image segmentation and repair model.
[0134] Example III:
[0135] This embodiment provides a method for segmenting and repairing remote sensing images based on uncertainty, and further details the method of training the remote sensing image segmentation and repair model in step S2 of the embodiment:
[0136] Step S201: Select a remote sensing image dataset;
[0137] Step S202: Preprocess the original remote sensing images in the dataset to obtain a processed dataset;
[0138] Step S203: Randomly divide the processed dataset into a training set, a validation set, and a test set according to the ratio of 6:2:2, ensuring that each subset is balanced in terms of scene types and ground object distributions;
[0139] Step S204: Use the training set to train the pre-constructed remote sensing image segmentation and repair model, use the validation set to validate the trained remote sensing image segmentation and repair model, use the test set to test the validated remote sensing image segmentation and repair model, and take the remote sensing image segmentation and repair model with the optimal test result as the finally trained remote sensing image segmentation and repair model.
[0140] It should be noted that the selected remote sensing image dataset in step S201 needs to be a high-resolution remote sensing image dataset covering multiple ground object types, ensuring that the data contains typical scenes (such as urban buildings, natural vegetation, water areas, etc.) that can provide diverse training data for the model, thereby enhancing the generalization ability of the model. In this embodiment, two publicly available remote sensing image datasets, WHU Building Dataset and Inria Building Dataset, are selected; among them, the WHU BuildingDataset dataset contains remote sensing images of urban and rural areas with a spatial resolution of 0.3 meters, which is suitable for high-resolution target recognition; while the Inria Building Dataset dataset mainly contains aerial images with a resolution range of 0.3 meters, and the images contain different urban settlements, from densely populated areas (such as the financial district of San Francisco) to alpine towns (such as Lienz in Tyrol, Austria), covering ground objects such as buildings and roads in urban environments.
[0141] Furthermore, the methods for preprocessing the original remote sensing images in the dataset in step S202 mainly include:
[0142] Image standardization (numerical normalization): By performing channel normalization on the original remote sensing images, the pixel values are mapped to the interval [-1, 1] to eliminate the illumination changes and sensor noises between different images, so as to improve the stability of the model and ensure the consistency of data distribution. The calculation formula is:
[0143] ,
[0144] In the formula, is the pixel value of the original remote sensing image; is the mean value of the WHU Building Dataset or Inria Building Dataset; is the standard deviation of the WHU Building Dataset or Inria Building Dataset.
[0145] Size unification (spatial cropping): Using the sliding window cropping technique, the large-size image is cut into local patches of 512×512 pixels, and the image size is adjusted to meet the network input requirements, thus ensuring the consistency of the input data size.
[0146] Morphological perturbation (label perturbation): By performing diversified perturbation operations on the real label, that is, applying erosion and dilation operations (kernel size from 3×3 to 7×7) to the target area, an initial segmentation mask with over-segmentation (dilation) or under-segmentation (erosion) is generated; it helps to simulate the boundary problem of urban buildings in the dataset, thereby further enhancing the learning ability of the model.
[0147] Error simulation (region erasing and forgery): By randomly erasing some target areas (maximum erasing area 20%) to simulate the missed detection phenomenon, or generating and adding false targets based on the Poisson distribution to generate false detection errors, the segmentation errors that may occur in the actual segmentation task are simulated, and the tolerance of the model to missed detection and false detection is improved.
[0148] In this embodiment, the method for training the remote sensing image segmentation and repair model further includes: adopting an online data augmentation strategy, such as random rotation, translation, and color transformation, to expand data diversity and suppress the risk of overfitting. Specifically, affine transformation (rotation , translation , scaling by 0.8 to 1.2 times) and color jitter (brightness , contrast ) are applied in real time during the training stage to enhance the robustness of the model to scale, illumination, and angle changes. For the high-resolution images of the WHU Building Dataset, it helps to enhance the learning of detailed information; while in the urban environment of the Inria Building Dataset, it can effectively improve the adaptability of the model to various scales and forms.
[0149] As an alternative embodiment, the method for training a remote sensing image segmentation and restoration model further includes: designing a loss function and an adaptive optimization strategy, seamlessly integrating the uncertainty error estimation module and the restoration mechanism into the deep learning framework to achieve end-to-end joint training and model optimization. Among them, the design of the loss function should consider both the error correction features and the uncertainty, enabling the model to achieve an effective balance between the two, mainly including the auxiliary branch loss function, the error estimation loss function, the negative log-likelihood loss function, and the guided restoration loss function.
[0150] Among them, the auxiliary branch loss function is used to directly focus on the target regions in the input samples, providing additional supervision information to better handle the problem of class imbalance. Its expression is:
[0151] ,
[0152] In the formula, is the auxiliary branch loss function; is the cross-entropy loss (CE Loss) function; is the weight parameter controlling the cross-entropy loss function, set to 0.5; is the Dice loss (Dice Loss) function; is the weight parameter controlling the Dice loss function, set to 0.5 to ensure that the model takes into account both the pixel-level classification accuracy and the consistency of the target region shape during the optimization process; is the output result of the auxiliary branch module; is the ground truth corresponding to the output result of the auxiliary branch module.
[0153] The error estimation loss function is used to directly focus on the regions where errors occur in the input samples, obtaining the detection results of the final correct segmentation, under-segmentation, and over-segmentation pixels. Its expression is:
[0154] ,
[0155] In the formula, is the error estimation loss function; is the weight parameter controlling the cross-entropy loss function, set to 0.5; is the weight parameter controlling the Dice loss function, set to 0.5; is the output result of the uncertainty error estimation module; is the ground truth corresponding to the output result of the uncertainty error estimation module.
[0156] The Negative Log-Likelihood (NLL) loss function is introduced to incorporate the total uncertainty into the optimization objective. The negative log-likelihood loss function not only considers the prediction error but also combines the uncertainty estimation, enabling the model to adaptively focus on reliable regions during training. Additionally, by penalizing the prediction deviation in high-uncertainty regions, overfitting to high-uncertainty regions is reduced, while encouraging the model to learn the reliability of uncertainty estimation, which can effectively improve the robustness of model repair and ensure the accuracy of segmentation results. Its expression is:
[0157] ,
[0158] where, is the negative log-likelihood loss function; is the total uncertainty at the pixel level of the th sample; is the ground truth; is the predicted value; is the total number of samples.
[0159] The guided repair loss function is used to directly focus on the segmentation results of the global and local parts of the image, and its expression is:
[0160] ,
[0161] where, is the guided repair loss function; is the weight parameter for controlling the cross-entropy loss function, set to 0.5; is the weight parameter for controlling the Dice loss function, set to 0.5; is the output result of the uncertainty-guided repair module; is the ground truth corresponding to the output result of the uncertainty-guided repair module;
[0162] For the joint training that coordinates segmentation accuracy and uncertainty calibration, the total loss function is defined as the weighted sum of each loss term:
[0163] ,
[0164] where, is the total loss function; are all the weights of each loss function, which are determined through experimental verification and set to 0.2, 0.3, 0.2, and 0.3 in sequence to ensure that the model can effectively utilize uncertainty information to guide the repair process while improving segmentation accuracy.
[0165] It should be noted that by jointly constraining the consistency between the segmentation output and the ground truth label using the cross-entropy loss (CE Loss) function and the Dice loss (Dice Loss) function, the accuracy of the model in pixel-level classification and boundary continuity can be ensured. Among them, the cross-entropy loss function strengthens the calibration of the class probability distribution, and the Dice loss function focuses on optimizing the overlap degree of the target region. The two cooperate to suppress the accumulation of local errors.
[0166] In some embodiments, the remote sensing image segmentation and repair model provided in this embodiment is implemented using the Pytorch framework on an NVIDIA Tesla V100 GPU, and the Adam optimizer is used. It is trained and run for 25,600 iterations, with a learning rate of 0.000125 and a batch size of 8. In addition, after the model training is completed, the weights are saved to a pth file. With the pth file, the initial segmentation results of any input model can be directly corrected and repaired, and the final result is output after visualization.
[0167] To verify the repair effect of the remote sensing image segmentation in the embodiments of the present invention, a comparative experiment was carried out, and the experimental results are shown in Table 1.
[0168]
[0169] Table 1 Performance comparison experimental results of two test sets on different models
[0170] As can be seen from Table 1, the method provided in the embodiments of the present invention has achieved the best segmentation performance on both the WHU Building Dataset and the Inria Building Dataset. Among them, on the WHU Building Dataset, the overlap degree IoU (Intersection over Union) of this method reaches 93.81%, the F1 score is 96.81%, the precision Precision is 96.96%, and the recall Recall is 96.92%, showing significant improvements compared with existing methods. On the Inria Building Dataset, this method also achieved an overlap degree IoU of 87.85%, an F1 score of 93.53%, a precision Precision of 94.13%, and a recall of 92.94%, which is overall better than existing methods.
[0171] Figure 8 It is a visualization image of the remote sensing image segmentation and repair model provided in the embodiments of the present invention in the inference stage of the WHU Building Dataset test set; among them, Figure 8From (a) to (h) are, in sequence, a remote sensing image, a label image (ground truth), an initial segmentation mask, a ground truth image of the misestimation result, a misestimation result image, an epistemic uncertainty image, an aleatoric uncertainty image, and a segmentation repair image. By Figure 8 (d) and Figure 8 (e), it can be clearly seen that the error of the initial segmentation mask is predicted by the model; by Figure 8 (f) and Figure 8 (g), it can be found that the uncertainty results are concentrated in the error areas and edge areas, which has a guiding effect on the repair process; by comparing Figure 8 (b), Figure 8 (c) and Figure 8 (h), the uncertainty-guided repair module can repair the Figure 8 (c) initial segmentation mask according to the misprediction, making the repaired Figure 8 (h) closer to the label image (ground truth) Figure 8 (b).
[0172] In summary, the uncertainty-based remote sensing image segmentation and repair method proposed in the embodiments of the present invention realizes more accurate error estimation and a more refined guiding repair strategy by combining aleatoric uncertainty and epistemic uncertainty analysis and considering uncertainty information at both the global and local scales, which can effectively improve the accuracy of remote sensing image segmentation and has stronger robustness in complex scenarios. In addition, this method can be trained end-to-end, does not rely on complex post-processing steps, but directly repairs according to the uncertainty of the image, thereby improving the processing efficiency and segmentation accuracy.
[0173] Embodiment 4:
[0174] This embodiment also provides a computer device, which can be a server, and its internal structure diagram can be as Figure 9 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface.
[0175] Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the data obtained and generated in the method for the robot to autonomously enter the packaging container. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements the method of any one of the foregoing Embodiments 1 to 3.
[0176] Those skilled in the art can understand that Figure 9 the structure shown in [figures] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figures, or combine certain components, or have different component arrangements.
[0177] The computer device provided in this embodiment can execute the method for remote sensing image segmentation and restoration based on uncertainty provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0178] Embodiment 5:
[0179] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps of the method described in any one of Embodiments 1 to 3.
[0180] The computer-readable storage medium provided in this embodiment can execute the method for remote sensing image segmentation and restoration based on uncertainty provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0181] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. In addition, terms such as "first", "second", etc. are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features.
[0182] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0183] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0184] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0185] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the present invention and the claims. All of these fall within the protection scope of the present invention.
Claims
1. A method for segmenting and repairing remote sensing images based on uncertainty, characterized in that The method includes: Obtaining a remote sensing image, and preprocessing the remote sensing image to generate a remote sensing image to be repaired and an initial segmentation mask; Inputting the remote sensing image and the initial segmentation mask into a trained remote sensing image segmentation and repair model, where the remote sensing image segmentation and repair model includes a feature extraction and encoding module, an uncertainty error estimation module, and an uncertainty-guided repair module; Through the feature extraction and encoding module, performing feature extraction and fusion on the remote sensing image and the initial segmentation mask to generate multi-scale high-level semantic features, and performing feature encoding on the initial segmentation mask to generate initial segmentation features; Through the uncertainty error estimation module, performing feature enhancement, mapping, and calculating uncertainty on the multi-scale high-level semantic features and the initial segmentation features to respectively obtain error estimation features, error estimation results, and total uncertainty, and weighting the error estimation results and the total uncertainty to obtain an uncertainty weighted result; Through the uncertainty-guided repair module, fusing and correcting the initial segmentation features, the error estimation features, and the uncertainty weighted result to obtain a segmentation and repair result.
2. The method for segmenting and repairing remote sensing images based on uncertainty according to claim 1, wherein The method for preprocessing the remote sensing image includes: image normalization, size unification, and morphological perturbation.
3. The method for segmenting and repairing remote sensing images based on uncertainty according to claim 1, wherein, The feature extraction and encoding module includes a first splicing module, an HRNet-W48 module, and an IS-Encoder module; The feature extraction and encoding module performs the following processing steps on the input remote sensing image and initial segmentation mask: Splicing the remote sensing image and the initial segmentation mask in the channel dimension through the first splicing module to obtain a spliced input feature; Inputting the spliced input feature into the HRNet-W48 module, performing multi-scale feature extraction through an encoder-decoder structure to generate multi-scale features; then upsampling the multi-scale features and splicing them in the channel dimension to fuse multi-scale information, and then compressing the number of channels through a 1×1 first convolutional layer to generate multi-scale high-level semantic features; Inputting the initial segmentation mask into the IS-Encoder module, and performing feature encoding through a cascaded 3×3 second convolutional layer and 3×3 third convolutional layer to generate initial segmentation features.
4. The method for segmenting and repairing remote sensing images based on uncertainty according to claim 1, wherein The uncertainty error estimation module includes a difference module, a channel spatial attention module, a first Softmax function, an uncertainty module, a fusion module, and a weighting module; The uncertainty error estimation module performs the following processing steps on the input multi-scale high-level semantic features and initial segmentation features: Taking the difference between the multi-scale high-level semantic features and the initial segmentation features in the same spatial scale through the difference module to obtain an attention input feature; After the attention input features are input into the channel attention module in the channel-spatial attention module, they are respectively processed by the first average pooling layer and the first max pooling layer to obtain the first average response feature and the first max response feature in the spatial dimension; then the first average response feature and the first max response feature are respectively input into a shared multi-layer perceptron for learning to obtain the average response enhanced feature and the max response enhanced feature; then, a first adder is used to fuse the average response enhanced feature and the max response enhanced feature to obtain a fused feature; Finally, after the fused feature is calculated by the first Sigmoid function to obtain the channel attention weight, the attention input features are fused through a first multiplier to generate enhanced features based on channel relationships; After the attention input features are input into the spatial attention module in the channel-spatial attention module, they are respectively processed by the second average pooling layer and the second max pooling layer to obtain the second average response feature and the second max response feature in the channel dimension; then, after the second average response feature and the second max response feature are concatenated in the channel dimension, they are successively subjected to convolution processing by a 7×7 fourth convolutional layer, the second Sigmoid function is used to calculate the spatial attention weight, and then the attention input features are fused through a second multiplier to generate enhanced features based on spatial relationships; The attention input features, the enhanced features based on channel relationships, and the enhanced features based on spatial relationships are fused through a second adder to obtain error estimation features; The error estimation features are mapped to multiple channels corresponding to correctly segmented pixels, under-segmented pixels, and over-segmented pixels respectively through the first Softmax function to obtain error estimation results; After the error estimation results are input into the uncertainty module, different prediction categories are simulated through a Monte Carlo sampling layer, and the aleatoric uncertainty and the epistemic uncertainty are calculated respectively; The aleatoric uncertainty and the epistemic uncertainty are fused through the fusion module to obtain the total uncertainty; The error estimation results and the total uncertainty are weighted through the weighting module to obtain the uncertainty weighted results.
5. The method for segmenting and restoring remote sensing images based on uncertainty according to claim 4, wherein The calculation formula for the aleatoric uncertainty is: , wherein, is accidental uncertainty; is the predicted probability of the predicted class ; is the set of predicted classes; The calculation formula for the epistemic uncertainty is: , , Wherein, is the cognitive uncertainty; is the error estimation result after the -th Monte Carlo sampling layer sampling; is the -th sampling mean value, that is, the stable prediction value; is the total number of samplings; The calculation formula for the total uncertainty is: , In the formula, is the total uncertainty; and are both fusion weights, and .
6. The method for segmenting and repairing remote sensing images based on uncertainty according to claim 1, wherein, The uncertainty-guided repair module includes an EGF module; The uncertainty-guided repair module performs the following processing steps on the input initial segmentation features, error estimation features, and uncertainty weighted results: The initial segmentation features, the error estimation features, and the uncertainty weighted results are input into the EGF module; They are concatenated in the channel dimension through a second concatenation module to obtain uncertainty concatenated features; The uncertainty concatenated features are successively subjected to feature extraction and correction through a 3×3 fifth convolutional layer, a 5×5 depthwise separable convolutional layer, and a 1×1 sixth convolutional layer, and are transformed through the second Softmax function to obtain segmentation repair results.
7. The method for segmenting and repairing remote sensing images based on uncertainty according to claim 1, wherein The training method of the remote sensing image segmentation repair model includes: Select a remote sensing image dataset; Preprocess the original remote sensing images in the dataset to obtain a processed dataset; Randomly divide the processed dataset into a training set, a validation set, and a test set according to a ratio of 6:2:2; Use the training set to train a pre-constructed remote sensing image segmentation and restoration model, use the validation set to validate the trained remote sensing image segmentation and restoration model, use the test set to test the validated remote sensing image segmentation and restoration model, and take the remote sensing image segmentation and restoration model with the optimal test result as the finally trained remote sensing image segmentation and restoration model.
8. The method for segmenting and repairing remote sensing images based on uncertainty according to claim 7, wherein The training method of the remote sensing image segmentation and restoration model further includes: designing a loss function and an adaptive optimization strategy to achieve end-to-end joint training and model optimization, where the loss function includes an auxiliary branch loss function, an error estimation loss function, a negative log-likelihood loss function, and a guided restoration loss function; The expression of the auxiliary branch loss function is: , In the formula, is the auxiliary branch loss function; is the cross-entropy loss function; is the weight parameter for controlling the cross-entropy loss function, set to 0.5; is the Dice loss function; is the weight parameter for controlling the Dice loss function, set to 0.5; is the output result of the auxiliary branch module; is the ground truth corresponding to the output result of the auxiliary branch module; The expression of the error estimation loss function is: , In the formula, is the error estimation loss function; is the weight parameter for controlling the cross-entropy loss function, set to 0.5; is the weight parameter for controlling the Dice loss function, set to 0.5; is the output result of the uncertainty error estimation module; is the ground truth corresponding to the output result of the uncertainty error estimation module; The expression of the negative log-likelihood loss function is: , In the formula, is the negative log-likelihood loss function; is the total uncertainty at the pixel level of the th sample; is the true label; is the predicted value; is the total number of samples; The expression of the guided restoration loss function is: , Wherein, is the guiding repair loss function; is the weight parameter for controlling the cross-entropy loss function, set to 0.5; is the weight parameter for controlling the Dice loss function, set to 0.5; is the output result of the uncertainty-guided repair module; is the ground truth corresponding to the output result of the uncertainty-guided repair module; The total loss function is defined as follows: , In the formula, is the total loss function; are the weights of each loss function, which are set to 0.2, 0.3, 0.2, and 0.3 in sequence.
9. A computer device, characterized in that, Including a storage medium and a processor; The storage medium is used to store instructions; The processor is used to operate according to the instructions to execute the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Cited By
Mask-guided image restoration and fusion model and training method and use method thereof
CN121010511A
Mask-guided image restoration and fusion model, its training method and usage
CN121010511B
Sample training optimization method, system and equipment for image segmentation model
CN121725012A
A sample training optimization method, system and device for an image segmentation model
CN121725012B