Image shadow removing method and system based on image filling enhancement
By constructing an image deshading method based on image filling enhancement, combining pre-trained encoder and optimized training strategies, the problem of shadow removal in complex environments is solved, efficient shadow detection and removal is achieved, and the processing capability and performance of the model is improved.
Patent Information
- Application Number
- CN202510787476.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-13
AI Technical Summary
Existing image processing technology is difficult to effectively remove shadows in complex environments, resulting in an increase in missed detection rate. The calculation efficiency of traditional methods is low, making it difficult to meet the needs of real-time inspections.
The image filling enhancement method is adopted, combined with the pre-trained graphic and text enhancement encoder, text encoder and task selection module, an image recovery model is constructed, and the training data set is constructed through random masks, and the learning rate scheduling method of cosine annealing and thermal restart is optimized to achieve end-to-end shadow detection and removal.
It significantly improves the model's processing ability for complex degraded scenarios, shortens the training cycle, improves the shadow removal effect and model performance, and adapts to dynamic lighting environments.
Smart Images

Figure CN120339135A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and more particularly, to an image de-shadowing method and system based on image filling enhancement. Background Art
[0002] In the infrastructure fields of power inspection and industrial inspection, the wide application of automated devices such as drones and robots has put forward higher requirements for image processing technology. For example, the inspection of transmission lines requires accurate identification of safety hazards such as insulator damage and conductor foreign objects, and the monitoring of substation equipment also relies on high-definition image analysis of partial discharge, rust marks, etc. However, the actual operating environment is complex and changeable, with dense shadows generated due to environmental structure occlusion or uneven illumination, dynamic illumination changes leading to image contrast imbalance, and adverse weather such as rain, fog, and dust further exacerbating image degradation. These interference factors not only obscure key defect features but also mislead detection algorithms, resulting in an increase in the miss detection rate. Traditional image processing methods are computationally inefficient when dealing with high-resolution images and are difficult to meet the real-time inspection requirements, severely restricting the deployment effect of automated terminals.
[0003] In recent years, although image de-shadowing technology has made some progress, existing solutions still have significant limitations. Physical model-based methods (such as light separation, reflectance estimation) rely on manually set parameters and have poor robustness in complex scenes and are difficult to adapt to dynamic environmental changes. Deep learning methods (such as GAN, CNN) have improved the shadow removal efficiency through end-to-end training, but their performance highly depends on the quality of labeled data, and there is a lack of large-scale professional datasets in the field of power inspection, resulting in limited model generalization ability. In addition, mainstream technologies mostly focus on single-task optimization and ignore the synergy with multi-tasks such as image inpainting and super-resolution, leading to the fact that the processing strategies for a single task can no longer meet the requirements of more complex application environments.
[0004] Regarding the problems in the related art, no effective solution has been proposed yet. Summary of the Invention
[0005] In view of this, the present invention provides an image de-shadowing method and system based on image filling enhancement to solve the above-mentioned problems.
[0006] To solve the above problems, the specific technical solutions adopted by the present invention are as follows:
[0007] According to one aspect of the present invention, there is provided an image de-shadowing method based on image filling enhancement, including the following steps:
[0008] S1. Based on a single-task image enhancement model, combined with a pre-trained text-image enhancement encoder, a text encoder, and a task selection module, construct an image restoration model;
[0009] S2. Based on the ground truth images of the de-shadowing dataset, construct an image inpainting task training dataset using random masks, train the image restoration model using the image inpainting task training dataset, and after training is completed, save the model weight file;
[0010] S3. Based on the model weight file, use the image inpainting task training dataset to retrain the trained image restoration model, and combine the learning rate scheduling method of cosine annealing and warm restart to optimize the image restoration model during the training process. Based on the optimized image restoration model, determine and save the model optimized weight file;
[0011] S4. Based on the model optimized weight file, initialize the image restoration model, and use the initialized image restoration model to perform de-shadowing processing on the to-be-processed shadowed image to obtain the shadow-removed image.
[0012] Preferably, the construction of the image restoration model by combining the pre-trained image-text enhancement encoder and the text encoder based on the single-task image enhancement model includes the following steps:
[0013] S11. According to the pre-configured single-task image enhancement model, obtain the initial model structure by introducing the image enhancement encoder in the pre-trained image-text enhancement encoder.
[0014] S12. Based on the initial model structure, obtain the multi-level model structure by introducing the text encoder.
[0015] S13. According to the multi-level model structure, obtain the image restoration model by introducing the task selection module and the hybrid attention module.
[0016] Among them, the image enhancement encoder is used to extract image features corresponding to the weights of the image encoder;
[0017] The text encoder is used to encode the task description text into an embedding vector to obtain task semantic features, so as to realize the conversion from text to semantic feature vectors.
[0018] The task selection module is used to fuse the image features and task semantic features using the cross-attention mechanism to obtain a task-guided feature map.
[0019] The hybrid attention module is used to extract features at different levels and generate scale features by concatenation, and combine the task-guided feature map to dynamically adjust the spatial and channel weights.
[0020] Preferably, the fusion of the image features and task semantic features using the cross-attention mechanism to obtain the task-guided feature map includes the following steps:
[0021] Input the task semantic features into a preset linear layer, and based on the weight matrix learned by the linear layer during training, perform a linear transformation on the task semantic features to map the task semantic features into a projection vector that matches the dimension of the image features;
[0022] According to the image features and the projection vector, use the cross-attention mechanism to calculate the fused features to obtain a task-guided feature map;
[0023] Among them, the calculation formula of the task-guided feature map is:
[0024] ;
[0025] In the formula, F ca represents the task-guided feature map, represents the cross-attention calculation, F img represents the image features, and E task represents the projection vector.
[0026] Preferably, the steps of extracting features at different levels and generating scale features by splicing, and dynamically adjusting the spatial and channel weights in combination with the task-guided feature map include the following:
[0027] Extract features at different levels from the image enhancement encoder in the text-image enhancement encoder, and generate multi-scale features by upsampling and splicing to obtain a task semantic vector;
[0028] Based on the self-attention mechanism, fuse the task semantic vector with the multi-scale features to obtain a self-attention fused feature;
[0029] Use a feed-forward network to perform a residual connection on the self-attention fused feature and the task-guided feature map to obtain a feature with enhanced non-linear expression ability after passing through the feed-forward network;
[0030] Based on the feature with enhanced non-linear expression ability after passing through the feed-forward network, adjust the weights of the attention channels;
[0031] Among them, the calculation formula of the task semantic vector is:
[0032] ;
[0033] In the formula, F mul represents the task semantic vector, Concat represents splicing, UpSample represents upsampling, and F low , F mid , F high represent features at different levels respectively;
[0034] The calculation formula of the self-attention fused feature is:
[0035] ;
[0036] In the formula, F refined represents the self-attention fusion feature, SelfAttention(·) represents the self-attention mechanism, and F mul represents the task semantic vector, and E task represents the projected text semantic vector;
[0037] The calculation formula for the feature after enhancing the non-linear expression ability through the feed-forward network is:
[0038] ;
[0039] In the formula, F fined represents the feature after enhancing the non-linear expression ability through the feed-forward network, FFN(·) represents the feed-forward neural network, and F ca represents the task-guided feature map, and F refined represents the self-attention fusion feature.
[0040] Preferably, the steps of constructing an image filling task training dataset by using a random mask according to the real image of the de-shadowing dataset, training an image restoration model by using the image filling task training dataset, and saving the model weight file after the training is completed are as follows:
[0041] S21. Collect a de-shadowing dataset containing real images, preprocess the de-shadowing dataset by using a shadow detection model, generate a mask by using the random mask method according to the preprocessing result, and generate an image filling task training dataset based on the mask;
[0042] S22. Configure task prompt data according to the image filling task training dataset;
[0043] S23. Based on the task prompt data and combined with the image filling task training dataset, perform iterative training on the image restoration model until the image restoration model converges;
[0044] S24. For the trained image restoration model, extract and save the weight file trained by the image restoration model.
[0045] Preferably, the steps of collecting a de-shadowing dataset containing real images, preprocessing the de-shadowing dataset by using a shadow detection model, generating a mask by using the random mask method according to the preprocessing result, and generating an image filling task training dataset based on the mask are as follows:
[0046] S211. Collect a de-shadowing dataset containing real images, and perform shadow area recognition processing on the de-shadowing dataset by using a shadow detection model;
[0047] S212. For a real image containing a shadow area, map it to a real image without a shadow area to obtain the corresponding spatial coordinate positions;
[0048] S213. Perform operations within the mapped shadow area, randomly generate masks of several geometric shapes, and determine whether the masks meet the prediction threshold range by calculating the size of the masks on the real image. If they meet the requirements, execute step S214; if not, regenerate the masks of geometric shapes;
[0049] S214. Based on the generated masks, pair them with the real images in the de-shadowing dataset to generate a training dataset for the image filling task for training.
[0050] Preferably, the steps of re-training the trained image restoration model based on the model weight file using the training dataset for the image filling task and optimizing the image restoration model during the training process by combining the cosine annealing and warm restart learning rate scheduling method, and determining and saving the model optimization weight file according to the optimized image restoration model include:
[0051] S31. Load the model weight file into the image restoration model and perform weight initialization processing on the image restoration model;
[0052] S32. Use the training dataset for the image filling task as the input pictures for training the image restoration model, and configure task prompt data for the image restoration model;
[0053] S33. Use the training dataset for the image filling task and the task prompt data to train the image restoration model, and optimize the learning rate of the image restoration model by combining the cosine annealing and warm restart learning rate scheduling method;
[0054] S34. Determine whether the image restoration model reaches the convergence condition in the de-shadowing task. If so, stop the training and save the model optimization weight file after the image restoration model is trained. If not, return to step S33.
[0055] Preferably, the optimizing the learning rate of the image restoration model by combining the cosine annealing and warm restart learning rate scheduling method includes:
[0056] Gradually reduce the learning rate of the image restoration model through the cosine function, and combine the introduction of a learning rate peak decay mechanism to control the learning rate of the image restoration model to form a decreasing warm restart;
[0057] Among them, the expression for gradually reducing the learning rate of the image restoration model through the cosine function and combining the introduction of a learning rate peak decay mechanism to control the learning rate of the image restoration model to form a decreasing warm restart is:
[0058] ;
[0059] ;
[0060] wherein, lr t represents the learning rate of the current training times, lr min , lr max respectively represent the minimum value and the maximum value of the learning rate, T cycle represents the length of a complete training cycle, t represents the current training times, represents the maximum learning rate of the k-th training cycle.
[0061] Preferably, the method for initializing the image restoration model based on the model-optimized weight file and using the initialized image restoration model to perform shadow removal processing on the to-be-processed shadowed image to obtain the shadow-removed image includes the following steps:
[0062] S41. Based on the model-optimized weight file, perform secondary weight initialization processing on the image restoration model completed in step S3;
[0063] S42. Obtain the to-be-processed shadowed image and input it into the image restoration model after secondary weight initialization processing;
[0064] S43. The image restoration model automatically detects the shadow area through the learned multi-task processing module and generates the shadow-removed image.
[0065] According to another aspect of the present invention, there is provided an image shadow removal system based on image filling enhancement, including:
[0066] A model construction module, configured to construct an image restoration model based on a single-task image enhancement model, in combination with a pre-trained image-text enhancement encoder, a text encoder, and a task selection module;
[0067] A primary training module, configured to construct an image filling task training data set by using a random mask according to the real images of the shadow removal data set, train the image restoration model by using the image filling task training data set, and save the model weight file after the training is completed;
[0068] A secondary training module, configured to re-train the trained image restoration model by using the image filling task training data set based on the model weight file, optimize the image restoration model during the training process in combination with the learning rate scheduling method of cosine annealing and warm restart, and determine and save the model-optimized weight file according to the optimized image restoration model;
[0069] The shadow processing module is used to initialize the image restoration model based on the model optimization weight file, and use the initialized image restoration model to perform shadow removal on the to-be-processed shadowed image to obtain the shadow-removed image.
[0070] The beneficial effects of the present invention are as follows:
[0071] 1. By introducing a guidance-based multi-task learning mechanism, the present invention realizes end-to-end shadow detection and removal, and at the same time improves the model's comprehensive processing ability for complex degradations. Based on a single-task image enhancement model, a multi-modal encoder pre-trained based on CLIP is added, and cross-attention mechanism is used to fuse text guidance and image features to achieve dynamic adaptation of tasks, significantly improving the model's processing ability to handle complex degradation scenarios.
[0072] 2. The present invention constructs an image inpainting task using a shadow removal dataset, pre-trains the model through a controllable mask generation mechanism, strengthens the model's global understanding of image structure and texture, and also provides high-quality initial weights for the fine-tuning of subsequent shadow removal tasks, significantly shortening the training cycle and improving the model performance.
[0073] 3. Through transfer learning optimization, the present invention transfers the pre-trained weights to the shadow removal task for model training, and combines the task prompt mechanism to ensure the rapid convergence of model parameters and the performance improvement of shadow removal effect in complex lighting environments. Description of the Drawings
[0074] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the following described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings. In the drawings:
[0075] Figure 1 is a flowchart of an image shadow removal method based on image inpainting enhancement according to an embodiment of the present invention;
[0076] Figure 2 is a schematic block diagram of an image shadow removal system based on image inpainting enhancement according to an embodiment of the present invention;
[0077] Figure 3 is an overall structure diagram of shadow removal in an image shadow removal method based on image inpainting enhancement according to an embodiment of the present invention;
[0078] Figure 4It is a flowchart of high-quality feature map output based on a pre-trained CLIP image encoder in an image shadow removal method based on image filling enhancement according to an embodiment of the present invention;
[0079] Figure 5 It is a multitasking flowchart in an image shadow removal method based on image filling enhancement according to an embodiment of the present invention.
[0080] In the figure:
[0081] 1. Model construction module; 2. First training module; 3. Second training module; 4. Shadow processing module. Detailed implementation manners
[0082] In order to enable those skilled in the art of this technology to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the protection scope of this application.
[0083] According to an embodiment of the present invention, an image shadow removal method and system based on image filling enhancement are provided.
[0084] Now, the present invention will be further described in conjunction with the accompanying drawings and specific implementation manners, as Figure 1 and Figures 3 - 5 shown, according to an embodiment of the present invention, an image shadow removal method based on image filling enhancement is provided, including the following steps:
[0085] S1. Based on a single-task image enhancement model, in combination with a pre-trained text-image enhancement encoder, a text encoder, and a task selection module, construct an image restoration model;
[0086] As a preferred implementation manner, the construction of the image restoration model based on the single-task image enhancement model in combination with the pre-trained text-image enhancement encoder and the text encoder includes the following steps:
[0087] S11. According to the pre-configured single-task image enhancement model, by introducing the image enhancement encoder in the pre-trained text-image enhancement encoder, obtain an initial model structure,
[0088] S12. Based on the initial model structure, by introducing the text encoder, obtain a multi-level model structure;
[0089] S13. According to the multi-level model structure, by introducing the task selection module and the hybrid attention module, obtain the image restoration model;
[0090] Among them, the image enhancement encoder is used to extract image features corresponding to the weights of the image encoder;
[0091] The text encoder is used to encode the task description text into an embedding vector to obtain task semantic features, so as to realize the conversion from text to semantic feature vectors;
[0092] The task selection module is used to fuse the image features and task semantic features by using the cross-attention mechanism to obtain a task-guided feature map;
[0093] As a preferred implementation manner, the step of fusing the image features and task semantic features by using the cross-attention mechanism to obtain a task-guided feature map includes the following steps:
[0094] Input the task semantic features into a preset linear layer, and based on the weight matrix learned by the linear layer during training, perform a linear transformation on the task semantic features to map the task semantic features into a projection vector matching the dimension of the image features;
[0095] According to the image features and the projection vector, use the cross-attention mechanism to calculate the fusion features to obtain a task-guided feature map;
[0096] Among them, the calculation formula of the task-guided feature map is:
[0097] ;
[0098] In the formula, F ca represents the task-guided feature map, represents the cross-attention calculation, F img represents the image features, E task represents the projection vector.
[0099] The hybrid attention module is used to extract features at different levels and generate scale features through splicing, and combine with the task-guided feature map to dynamically adjust the spatial and channel weights.
[0100] As a preferred implementation manner, the step of extracting features at different levels and generating scale features through splicing, combining with the task-guided feature map, and dynamically adjusting the spatial and channel weights includes the following steps:
[0101] Extract features at different levels from the image enhancement encoder in the text-image enhancement encoder, and generate multi-scale features through upsampling and splicing to obtain task semantic vectors;
[0102] Based on the self-attention mechanism, fuse the task semantic vectors with the multi-scale features to obtain self-attention fusion features;
[0103] A residual connection is made between the self-attention fusion features and the task-guided feature maps using a feed-forward network to obtain features with enhanced non-linear expression ability after passing through the feed-forward network;
[0104] Based on the features with enhanced non-linear expression ability after passing through the feed-forward network, the weights of the attention channels are adjusted.
[0105] It should be noted that, based on the general single-task image enhancement model, an image enhancement encoder in the pre-trained text-image enhancement encoder (CLIP) is introduced (the weights of the image enhancement encoder are from the pre-trained CLIP model, and the network structure of the image enhancement encoder is exactly the same as that of the CLIP visual encoder, so the pre-trained CLIP weights are directly loaded as initial parameters), and the image features corresponding to the weights of the image encoder are extracted:
[0106] ;
[0107] In the formula, x represents the image data, and Encoder image (x) represents the image encoder, and F img represents the image features based on the weights (the context information obtained after feature extraction).
[0108] It should be noted that the weights are the "parameters" of the encoder, the context is the "output" of the encoder, and the context information has a "guiding and optimizing" effect on the weights. Specifically: the weights are the "knowledge carriers" of the encoder, determining how to extract context information from the input image; the context information is the "inference result" of the weights for a specific image, and the two have a dependency relationship of "model parameters" and "feature output".
[0109] Among them, the extracted context information (image features) and the internal features of the model are used through self-attention and cross-attention to dynamically enhance the feature response values in the shadow-related regions).
[0110] In addition, using the text encoder Encoder text to encode the task description text Task prompt into a semantic vector Task embedding , the conversion from text to semantic feature vector can be achieved. The calculation formula for the conversion from text to semantic feature vector is:
[0111] ;
[0112] In the formula, Task embedding represents the semantic vector, Encoder text (·) represents the text encoder, and Task prompt represents the task description text.
[0113] Then, the image features (here, the image feature map is obtained by inputting the original image into the CLIP image encoder) and the text vector are simultaneously fed into the task selection module. First, the text encoding is mapped (projected) into the same dimension as the feature map to align the text and image feature spaces and achieve cross-modal interaction. The formula for mapping (projecting) the text encoding into the same dimension as the feature map is:
[0114] ;
[0115] In the formula, E task represents the projection vector matching the image feature dimension, and Linear(·) represents the linear layer. Task embedding represents the task embedding vector (semantic vector).
[0116] It should be noted that the weight matrix of the linear layer is learned during the training process and can adaptively generate different weight distributions according to the semantic features of the input task text (such as "removing shadows" or "repairing the insulator area").
[0117] The task semantic features are extracted through the text encoder and then mapped through the linear layer to make the weight values task-specific, laying the foundation for dynamically adjusting the feature map weights. Specifically: Task embedding is linearly transformed through the linear layer Linear. The weight matrix of the linear layer is learned during the training process. According to the semantic features of Task embedding , it is mapped into a projection vector matching the image feature map dimension. Different Tasks embedding (corresponding to different task texts) will output different Es task when passing through this linear layer due to different input features.
[0118] In addition, when dynamically fusing text semantics and image features to generate a task-guided feature map, Q, K, and V in the cross-attention mechanism are first defined. Among them, Q represents the image feature F img ; K and V are the projected text semantic vectors E task , and the fused feature is calculated based on the cross-attention mechanism:
[0119] ;
[0120] Specifically, the original information can be retained through residual connection, and the feed-forward network (FFN) enhances the non-linear expression ability, specifically including:
[0121] Extract different levels of features {F low , F mid , F high} from the CLIP encoder, and generate multi-scale features through upsampling and concatenation:
[0122] ;
[0123] Wherein, F mul represents the task semantic vector, Concat represents concatenation, UpSample represents upsampling, F low , F mid , F high respectively represent features at different levels;
[0124] Then, the task semantic vector F mul is fused with the multi-scale features to enhance the response of the key regions:
[0125] ;
[0126] Wherein, F refined represents the self-attention fusion feature, SelfAttention(·) represents the self-attention mechanism, F mul represents the task semantic vector, E task represents the projected text semantic vector;
[0127] The original information is retained through the residual connection, and the feed-forward network (FFN) enhances the non-linear expression ability. The calculation formula for the feature after enhancing the non-linear expression ability through the feed-forward network is:
[0128] ;
[0129] Wherein, F final represents the feature after enhancing the non-linear expression ability through the feed-forward network, FFN(·) represents the feed-forward neural network, F ca represents the task-guided feature map, F refined represents the self-attention fusion feature.
[0130] Specifically, by introducing the hybrid attention module, the spatial and channel weights are dynamically adjusted:
[0131] ;
[0132] ;
[0133] ;
[0134] Wherein, M spatial represents the spatial attention weight map, M channel represents the channel attention weight map, SpatialAttention represents the spatial attention module, F final represents the feature after enhancing the non-linear expression ability through the feed-forward network, F outputIt represents the features after spatial and channel attention adjustment, and ChannelAttention represents the channel attention module.
[0135] Then, according to the task requirements, the final shadow-removed image is generated through a lightweight convolutional layer:
[0136] ;
[0137] In the formula, OutPut represents the enhanced feature map after task semantic guidance, Conv(·) represents the convolutional layer, Decoder represents the decoder module, and F output represents the features after spatial and channel attention adjustment.
[0138] S2. According to the real images in the shadow-removed dataset, an image inpainting task training dataset is constructed using a random mask. The image restoration model is trained using the image inpainting task training dataset, and after the training is completed, the model weight file is saved.
[0139] As a preferred implementation manner, the step of constructing an image inpainting task training dataset using a random mask according to the real images in the shadow-removed dataset, training the image restoration model using the image inpainting task training dataset, and saving the model weight file after the training is completed includes the following steps:
[0140] S21. Collect the shadow-removed dataset containing real images, preprocess the shadow-removed dataset using a shadow detection model, and according to the preprocessing results, generate a mask using the random mask method, and generate an image inpainting task training dataset based on the mask;
[0141] As a preferred implementation manner, the step of collecting the shadow-removed dataset containing real images, preprocessing the shadow-removed dataset using a shadow detection model, generating a mask using the random mask method according to the preprocessing results, and generating an image inpainting task training dataset based on the mask includes the following steps:
[0142] S211. Collect the shadow-removed dataset containing real images, and perform shadow area recognition processing on the shadow-removed dataset using a shadow detection model;
[0143] S212. For the real images containing shadow areas, map them to the real images without shadow areas to obtain the corresponding spatial coordinate positions;
[0144] S213. Perform operations within the mapped shadow areas, randomly generate masks of several geometric shapes, and judge whether the masks meet the prediction threshold range by calculating the sizes of the masks on the real images. If they meet, execute step S214. If they do not meet, regenerate the geometric shape masks;
[0145] S214. Based on the generated mask, pair it with the real images in the de-shadowing dataset to generate a training dataset for the image inpainting task for training.
[0146] S22. Configure task prompt data according to the training dataset for the image inpainting task;
[0147] S23. Based on the task prompt data and combined with the training dataset for the image inpainting task, iteratively train the image restoration model until the image restoration model converges;
[0148] S24. For the trained image restoration model, extract and save the weight file trained by the image restoration model.
[0149] It should be noted that when using the images in the de-shadowing task dataset as materials, the idea of using a random mask (Mask) is used to create an image inpainting dataset (in the method for constructing the image inpainting dataset, for each 512×512 high-quality power equipment image, a mask is generated in the area corresponding to the shadow with a probability of 60% (pre-generated by the shadow detection model).
[0150] The shape of the mask is randomly rectangular (40%), elliptical (30%), or polygonal (30%), and the area is controlled within 1 / 12 to 1 / 4 of the image area. The generated masked image is obtained by setting the pixels in the masked area to 0, and the original image is used as the label, meeting the data requirements for the pre-training of image inpainting in step S22). The positions where the mask is generated are divided into the shadow area and the non-shadow area (the coordinates of the real shadow area are located through the shadow mask output by the shadow detection model, combined with a probability control strategy (6:4 ratio) and logical constraint operations to ensure that the mask is only generated within the target area). Among them, the ratio of the mask generated in the shadow area to the non-shadow area is 6:4.
[0151] First, use the shadow detection model to detect the position of the shadow area in the low-quality images (images with shadows) in the de-shadowing dataset. Second, generate a mask on the high-quality image (a clear image without shadows). If a mask needs to be generated in the shadow area, the mask will be generated in the detected shadow area. Otherwise, the generated mask will avoid the shadow area. Among them, the area of the mask is controlled between 1 / 4 and 1 / 12 of the image size, which not only ensures the challenge of the training task but also avoids the loss of image details due to an overly large mask. Finally, pair the image with the added mask and the original image to form a training data pair, generating the image inpainting dataset D1.
[0152] Specifically, the shadow detection model adopts a U-Net architecture. The encoder is a pre-trained ResNet50, and the decoder contains 4 layers of transposed convolutions. The input is three channels of RGB + HSL (6 channels in total), and the output generates a binary mask after Otsu threshold segmentation and 3×3 kernel morphological operations. The training data uses the ISTD dataset (1500 pairs of images) and the self-built power inspection dataset (800 pairs of images). The loss function is BCE + Dice (α = 0.75), and the optimizer is AdamW (learning rate 1e-4). After 100 rounds of training, the Dice coefficient on the validation set reaches 0.92, meeting the accuracy requirements for shadow area localization.
[0153] In addition, when generating a mask on a high-quality image (a clear image without shadows), the shadow detection model is used to process a low-quality shadowed image to determine the position of the shadow area, and a binary mask is output (1 for the shadow area and 0 for the non-shadow area), which is then mapped to the corresponding spatial coordinate position of the high-quality image to ensure pixel-level alignment.
[0154] Operations are performed within the mapped shadow area (element-wise multiplication of the target area output by the shadow detection model and a randomly generated area mask to ensure that the mask is only valid within the target area). A geometric shape mask is randomly generated, and the shape area is calculated to ensure that it is between 1 / 4 and 1 / 12 of the image size. If not satisfied, the geometric shape is regenerated.
[0155] It should be noted that the accuracy and generalization ability of shadow area mask generation can be improved through an end-to-end data processing flow that combines self-supervised contrast learning and a dynamic deformable convolutional network, including: data augmentation and positive and negative sample construction. By randomly cropping, rotating, and perturbing the lighting of unlabeled power equipment images, enhanced image pairs are generated; positive sample pairs are composed of different enhanced versions of the same image; negative sample pairs are composed of enhanced versions of different images.
[0156] Among them, a positive sample pair means that after the same image is enhanced, the semantics of the key components are ensured to be consistent (the intersection over union IoU of the mask is calculated through a pre-trained segmentation model, and it is only accepted as a positive sample pair when IoU > 80%), but the shadow areas are different (such as changes in shadow concentration and range); a negative sample pair means that device images with similar semantics but different types (such as transformers of different models) are selected, and after enhancement, a contrast pair is constructed to force the model to learn fine-grained feature differences.
[0157] In addition, the prediction confidence of the computational model for each sample is calculated. Samples with high confidence but incorrect predictions are defined as "hard negative samples". During training, a sample difficulty evaluation module is introduced. Based on the cross-entropy between the model prediction probability and the true label, samples with cross-entropy > the set threshold are selected as hard negative samples, and the sample weights are dynamically adjusted. The weight of hard negative samples in contrastive learning is increased by 1.5 times. At the same time, their corresponding "adversarial enhanced versions" are generated (for example, adding slight perturbations to the shadow areas of hard negative samples) to improve the robustness of the model to error-prone samples.
[0158] In addition, when using the deep neural network VIT to extract image features, a high-dimensional feature representation F1 is obtained. Among them, dynamic deformable convolution is introduced in the feature extraction layer, and by learning the offset Δp k the sampling position of the convolution kernel is adaptively adjusted to accurately capture the irregular boundaries and details of the shadow. For the input feature map, the output feature is:
[0159] ;
[0160] where P represents the current position, k represents the sampling point index of the convolution kernel, and p k represents the original convolution kernel position, w k represents the weight, and Δm k represents the modulation factor.
[0161] Specifically, by cascading deformable convolution layers with different dilation rates, a multi-scale receptive field feature pyramid is constructed. Combining the outputs of dynamic deformable convolutions at different scales, high-level semantics and low-level details are fused to improve the representation ability for complex shadow structures. The decoding part of the U-Net segmentation architecture is adopted to upsample the features output by the dynamic deformable convolution network to restore the image resolution and output the shadow area mask.
[0162] During model training, the contrastive learning loss L contrastive and the cross-entropy segmentation loss L seg are optimized simultaneously.
[0163] ;
[0164] In the formula, s(·) represents the cosine similarity score, N represents the number of positive samples participating in the calculation, f i , f j both represent the feature vectors of positive sample images, f n represents the features of other images (negative samples), y is the true label (0 or 1), p represents the probability of the model predicting the positive class, τ∈(0,1) is the temperature coefficient, and λ1 represents the hyperparameter.
[0165] It should be noted that after obtaining the generated dataset D1, the current task is specified as an image inpainting task by setting the task prompt (Prompt) to "Please help me fill in the missing parts of the image". Through the task prompt, the model can dynamically adjust the feature extraction and restoration strategies. (The task prompt enables the model to dynamically optimize the feature extraction direction and restoration logic for a specific task (such as "filling in the missing parts of the image") through the process of "semantic encoding → weight mapping generation → feature modulation → strategy adjustment", improving the accuracy and pertinence of task execution and focusing on the repair of the missing area.)
[0166] Finally, using the set task prompt, pre-train on the multi-task model structure (image restoration model). The training is carried out for 50 iterations until the model converges, and then the training of the model on the image inpainting task is ended. After the training is completed, save the weight file W1 trained by the model as the initial weight for the subsequent shadow removal task.
[0167] S3. Based on the model weight file, use the training dataset for the image inpainting task to retrain the trained image restoration model, and combine the learning rate scheduling method of cosine annealing and warm restart to optimize the image restoration model during the training process. According to the optimized image restoration model, determine and save the model optimized weight file;
[0168] As a preferred implementation, the step of using the training dataset for the image inpainting task to retrain the trained image restoration model based on the model weight file, and combining the learning rate scheduling method of cosine annealing and warm restart to optimize the image restoration model during the training process, and determining and saving the model optimized weight file according to the optimized image restoration model includes the following steps:
[0169] S31. Load the model weight file into the image restoration model and perform weight initialization processing on the image restoration model;
[0170] S32. Use the training dataset for the image inpainting task as the input pictures for the training of the image restoration model, and configure task prompt data for the image restoration model;
[0171] S33. Use the training dataset for the image inpainting task and the task prompt data to train the image restoration model, and combine the learning rate scheduling method of cosine annealing and warm restart to optimize the learning rate of the image restoration model;
[0172] S34. Judge whether the image restoration model reaches the convergence condition in the shadow removal task. If so, stop the training and ensure the model optimized weight file after the training of the image restoration model. If not, return to step S33.
[0173] It should be noted that when using the de-shadowed dataset as the image input for model (image restoration model) training, it includes low-quality images with shadows and their corresponding high-quality shadow-free images. At the same time, the task prompt is set as "Please help me remove the shadows in the picture", and through text encoding, the model is guided to focus on the feature extraction and restoration of the shadow area, clarifying the current task goal.
[0174] Use the saved pre-trained weights W1 as the weight parameters for model initialization, and use the set dataset and task prompt to train the model until the model reaches the convergence condition in the de-shadowing task, and save the final optimized weights W2.
[0175] In addition, combining cosine annealing with warm restart, a dynamic learning rate scheduling method is proposed. The learning rate is gradually decreased through the cosine function to help the model converge more stably in the later stage of training. And a learning rate peak decay mechanism is introduced, that is, after each restart, the learning rate peak decays by 80%, and the peak decay is 80% of the previous cycle, forming a decreasing warm restart.
[0176] As a preferred implementation, the learning rate scheduling method combining cosine annealing with warm restart for optimizing the learning rate of the image restoration model includes:
[0177] Gradually decrease the learning rate of the image restoration model through the cosine function, and combine the introduction of the learning rate peak decay mechanism to control the learning rate of the image restoration model, forming a decreasing warm restart;
[0178] Among them, the expression for gradually decreasing the learning rate of the image restoration model through the cosine function and combining the introduction of the learning rate peak decay mechanism to control the learning rate of the image restoration model, forming a decreasing warm restart is:
[0179] ;
[0180] ;
[0181] In the formula, lr t represents the learning rate of the current training iteration, lr min , lr max represent the minimum and maximum values of the learning rate respectively, T cycle represents the length of a complete training cycle, t represents the current training iteration, represents the maximum learning rate of the k-th training cycle. The initial parameters are set as: lr max = 1e-6, lr min = 1e-6, T cycle = 20.
[0182] S4. Initialize the image restoration model based on the model-optimized weight file, and use the initialized image restoration model to perform shadow removal on the to-be-processed shadowed image to obtain the shadow-removed image.
[0183] As a preferred implementation, the step of initializing the image restoration model based on the model-optimized weight file and using the initialized image restoration model to perform shadow removal on the to-be-processed shadowed image to obtain the shadow-removed image includes the following steps:
[0184] S41. Based on the model-optimized weight file, perform secondary weight initialization on the image restoration model completed in step S3;
[0185] S42. Obtain the to-be-processed shadowed image and input it into the image restoration model after secondary weight initialization;
[0186] S43. The image restoration model automatically detects the shadow area through the learned multi-task processing module and generates the shadow-removed image.
[0187] It should be noted that the specific implementation steps of initializing the image restoration model based on the model-optimized weight file and using the initialized image restoration model to perform shadow removal on the to-be-processed shadowed image to obtain the shadow-removed image are as follows:
[0188] Model loading: Load the trained weight file W2 into the image restoration model for image filling enhancement to complete model initialization.
[0189] Input the to-be-processed image: Input the to-be-processed low-quality shadowed image into the model, supporting single-image or batch processing; combined with the text encoder, input the task prompt (such as "Please help me remove the shadow in the picture"), and the model can dynamically adjust the feature extraction strategy based on text guidance, focusing on the restoration of the shadow area without manual intervention.
[0190] End-to-end inference: The model automatically detects the shadow area through the learned multi-task processing module (the model extracts the global features of the image through the CLIP encoder and combines the cross-attention mechanism to fuse the text guidance information, and the multi-task processing module dynamically adjusts the feature map according to the task weights, giving priority to restoring the texture and structure of the shadow area), and combines the structure and texture understanding ability pre-trained for the image filling task to generate a high-quality shadow-removed image.
[0191] Output the result: Directly output the clear shadow-removed image for subsequent tasks such as power inspection defect recognition and industrial inspection.
[0192] Such as Figure 2As shown, according to another embodiment of the present invention, an image de-shadowing system based on image filling enhancement is provided, including:
[0193] A model construction module 1, configured to construct an image restoration model based on a single-task image enhancement model, in combination with a pre-trained image-text enhancement encoder, a text encoder, and a task selection module;
[0194] A primary training module 2, configured to construct an image filling task training dataset by using a random mask according to the real images in the de-shadowing dataset, use the image filling task training dataset to train the image restoration model, and after the training is completed, save the model weight file;
[0195] A secondary training module 3, configured to re-train the trained image restoration model based on the model weight file by using the image filling task training dataset, and optimize the image restoration model during the training process in combination with the learning rate scheduling method of cosine annealing and warm restart, and determine and save the model optimization weight file according to the optimized image restoration model;
[0196] A shadow processing module 4, configured to perform initialization processing on the image restoration model based on the model optimization weight file, and use the initialized image restoration model to perform de-shadowing processing on the to-be-processed shadowed image to obtain the shadow-removed image.
[0197] In summary, by means of the above technical solutions of the present invention, the present invention realizes end-to-end shadow detection and removal by introducing a guidance-based multi-task learning mechanism, and at the same time improves the model's comprehensive processing ability for complex degradations. On the basis of a single-task image enhancement model, a multi-modal encoder pre-trained based on CLIP is added, and the cross-attention mechanism is used to fuse text guidance and image features to realize the dynamic adaptation of tasks, significantly improving the model's processing ability to handle complex degradation scenarios. The present invention constructs an image filling task by using a de-shadowing dataset, pre-trains the model through a controllable mask generation mechanism, strengthens the model's global understanding of image structure and texture, and also provides high-quality initial weights for the fine-tuning of subsequent de-shadowing tasks, greatly shortening the training cycle and improving the model performance. The present invention optimizes through transfer learning, migrates the pre-trained weights to the de-shadowing task for model training, and combines the task prompt mechanism to ensure the rapid convergence of model parameters and the performance improvement of the shadow removal effect in a complex illumination environment.
[0198] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) containing computer-usable program code.
[0199] The specific embodiments described above further elaborate on the objectives, technical solutions and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. An image de-shadowing method based on image filling enhancement, characterized in that, It includes the following steps: S1. Based on a single-task image enhancement model, combined with a pre-trained text-image enhancement encoder, a text encoder, and a task selection module, construct an image restoration model; S2. According to the real images in the de-shadow dataset, use random masks to construct an image inpainting task training dataset, use the image inpainting task training dataset to train the image restoration model, and after the training is completed, save the model weight file; S3. Based on the model weight file, use the image inpainting task training dataset to train the trained image restoration model again, and combine the learning rate scheduling method of cosine annealing and warm restart to optimize the image restoration model during the training process. According to the optimized image restoration model, determine and save the model optimized weight file; S4. Based on the model optimized weight file, perform initialization processing on the image restoration model, and use the initialized image restoration model to perform de-shadow processing on the to-be-processed shadowed image to obtain an image after removing the shadow.
2. The method for removing shadow from an image based on image filling enhancement according to claim 1, wherein The construction of the image restoration model by combining the pre-trained text-image enhancement encoder and the text encoder based on the single-task image enhancement model includes the following steps: S11. According to the pre-configured single-task image enhancement model, obtain the initial model structure by introducing the image enhancement encoder in the pre-trained text-image enhancement encoder; S12. Based on the initial model structure, obtain a multi-level model structure by introducing the text encoder; S13. According to the multi-level model structure, obtain the image restoration model by introducing the task selection module and the hybrid attention module; Among them, the image enhancement encoder is used to extract image features corresponding to the weights of the image encoder; The text encoder is used to encode the task description text into an embedding vector to obtain task semantic features, so as to realize the conversion from text to semantic feature vectors; The task selection module is used to fuse the image features and task semantic features by using the cross-attention mechanism to obtain a task-guided feature map; The hybrid attention module is used to extract features at different levels and generate scale features by splicing, and combine the task-guided feature map to dynamically adjust the spatial and channel weights.
3. A method for removing shadows from an image based on image filling enhancement according to claim 2, wherein, The fusion of the image features and task semantic features by using the cross-attention mechanism to obtain a task-guided feature map includes the following steps: Input the task semantic features into a preset linear layer, and based on the weight matrix learned by the linear layer during training, perform a linear transformation on the task semantic features to map the task semantic features into a projection vector matching the dimension of the image features; According to the image features and the projection vector, use the cross-attention mechanism to calculate the fusion features to obtain a task-guided feature map; Among them, the calculation formula of the task-guided feature map is: ; Where, F ca represents the task-guided feature map, represents cross-attention calculation, and F img represents the image feature, and E task represents the projection vector.
4. The image shadow removal method based on image filling enhancement according to claim 2, characterized in that, The extraction of features at different levels and the generation of scale features by splicing, and the combination of the task-guided feature map to dynamically adjust the spatial and channel weights include the following steps: Extract features at different levels from the image enhancement encoder in the text-image enhancement encoder, and generate multi-scale features through upsampling and splicing to obtain task semantic vectors; Based on the self-attention mechanism, the task semantic vector is fused with multi-scale features to obtain the self-attention fusion feature; The feed-forward network is used to perform residual connection on the self-attention fusion feature and the task-guided feature map to obtain the feature with enhanced non-linear expression ability after passing through the feed-forward network; Based on the feature with enhanced non-linear expression ability after passing through the feed-forward network, the weights of the attention channels are adjusted; Among them, the calculation formula of the task semantic vector is: ; In the formula, F mul represents the task semantic vector, Concat represents concatenation, UpSample represents upsampling, and F low , F mid , F high represent features at different levels respectively; The calculation formula of the self-attention fusion feature is: ; Where, F refined represents the self-attention fusion feature, SelfAttention(·) represents the self-attention mechanism, F mul represents the task semantic vector, and E task represents the projected text semantic vector; The calculation formula of the feature with enhanced non-linear expression ability after passing through the feed-forward network is: ; where, F final represents the feature after enhancing the non - linear expression ability through the feed - forward network, FFN(·) represents the feed - forward neural network, and F ca represents the task - guided feature map, and F refined represents the self - attention fusion feature.
5. A method for removing shadows from an image based on image filling enhancement according to claim 1, wherein According to the real images of the de-shadowing dataset, using a random mask to construct an image inpainting task training dataset, using the image inpainting task training dataset to train the image restoration model, and after the training is completed, saving the model weight file includes the following steps: S21. Collect the de-shadowing dataset containing real images, preprocess the de-shadowing dataset using the shadow detection model, and according to the preprocessing results, generate a mask using the random mask method, and generate an image inpainting task training dataset based on the mask; S22. Configure the task prompt data according to the image inpainting task training dataset; S23. Based on the task prompt data and combined with the image inpainting task training dataset, perform iterative training on the image restoration model until the image restoration model converges; S24. For the trained image restoration model, extract and save the weight file trained by the image restoration model.
6. The method for removing shadow from an image based on image filling enhancement according to claim 5, wherein The step of collecting the de-shadowing dataset containing real images, preprocessing the de-shadowing dataset using the shadow detection model, and according to the preprocessing results, generating a mask using the random mask method, and generating an image inpainting task training dataset based on the mask includes the following steps: S211. Collect the de-shadowing dataset containing real images, and perform shadow area recognition processing on the de-shadowing dataset using the shadow detection model; S212. For the real image containing the shadow area, map it to the real image without the shadow area to obtain the corresponding spatial coordinate position; S213. Perform operations within the mapped shadow area, randomly generate masks of several geometric shapes, and judge whether the mask meets the prediction threshold range by calculating the size of the mask on the real image. If it meets the requirement, execute step S214. If it does not meet the requirement, regenerate the geometric shape mask; S214. Based on the generated mask, pair it with the real image in the de-shadowing dataset to generate an image inpainting task training dataset for training.
7. A method for removing shadows from an image based on image filling enhancement according to claim 1, characterized in that Based on the model weight file, using the image inpainting task training dataset to train the trained image restoration model again, and combining the learning rate scheduling method of cosine annealing and warm restart to optimize the image restoration model during the training process. According to the optimized image restoration model, determine and save the model optimized weight file includes the following steps: S31. Load the model weight file into the image restoration model and perform weight initialization processing on the image restoration model; S32. Use the training dataset of the image filling task as the input pictures for training the image restoration model, and configure task prompt data for the image restoration model; S33. Use the training dataset of the image filling task and the task prompt data to train the image restoration model, and optimize the learning rate of the image restoration model by combining the learning rate scheduling method of cosine annealing and warm restart; S34. Determine whether the image restoration model converges on the shadow removal task. If so, stop training and save the model optimization weight file after the image restoration model is trained. If not, return to step S33.
8. A method for removing shadows from an image based on image filling enhancement according to claim 7, characterized in that, The optimization process of the learning rate of the image restoration model by combining the learning rate scheduling method of cosine annealing and warm restart includes: Gradually reduce the learning rate of the image restoration model through the cosine function, and combine the introduction of the learning rate peak decay mechanism to control the learning rate of the image restoration model to form a decreasing warm restart; Among them, the expression for gradually reducing the learning rate of the image restoration model through the cosine function and combining the introduction of the learning rate peak decay mechanism to control the learning rate of the image restoration model to form a decreasing warm restart is: ; ; where lr t represents the learning rate of the current training iteration, lr min , lr max represent the minimum and maximum values of the learning rate respectively, T cycle represents the length of a complete training cycle, t represents the current training iteration, represents the maximum learning rate of the k-th training cycle.
9. A method for removing shadows from an image based on image filling enhancement according to claim 1, characterized in that, The steps for initializing the image restoration model based on the model optimization weight file and using the initialized image restoration model to remove the shadow from the to-be-processed shadowed image to obtain the shadow-removed image are as follows: S41. Based on the model optimization weight file, perform a secondary weight initialization process on the image restoration model trained in step S3; S42. Obtain the to-be-processed shadowed image and input it into the image restoration model after the secondary weight initialization process; S43. The image restoration model automatically detects the shadow area through the learned multi-task processing module and generates the shadow-removed image.
10. An image de-shadowing system based on image filling enhancement is used to implement the image de-shadowing method based on image filling enhancement described in any one of claims 1-9, characterized in that Including: A model construction module for constructing an image restoration model based on a single-task image enhancement model, in combination with a pre-trained text-image enhancement encoder, a text encoder, and a task selection module; A primary training module for constructing a training dataset for the image filling task using a random mask based on the real images of the shadow removal dataset, training the image restoration model using the training dataset of the image filling task, and saving the model weight file after the training is completed; A secondary training module for re-training the trained image restoration model using the training dataset of the image filling task based on the model weight file, optimizing the image restoration model during the training process by combining the learning rate scheduling method of cosine annealing and warm restart, and determining and saving the model optimization weight file according to the optimized image restoration model; A shadow processing module for initializing the image restoration model based on the model optimization weight file and using the initialized image restoration model to remove the shadow from the to-be-processed shadowed image to obtain the shadow-removed image.
Citation Information
Patent Citations
Image completion method and system based on semantic edge fusion
CN112184585A
Image shadow removal model and construction method, device and application thereof
CN115375589A
Image shadow removal method based on mask refinement
CN118314038A
Image shadow removing method and device based on diffusion model, equipment and medium
CN118691497A
Shadow removing method and system based on diffusion, segmentation and super-resolution model
CN119048357A
Cited By
Image restoration method, system and equipment based on multi-modal large model driving
CN121458592A
Image restoration methods, systems, and devices based on multimodal large model-driven approaches.
CN121458592B