A method for restoring images of bad weather in real scenes based on a vision-language model
By adopting visual language models and semi-supervised learning frameworks in severe weather image recovery technology, combining teacher-student networks and weather prompt learning, the problems of insufficient generalization ability and semantic context neglect in the existing technology are solved, and more efficient image restoration performance and semantic recovery quality are achieved.
Patent Information
- Application Number
- CN202410454101.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-16
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-04-16
AI Technical Summary
Existing inclement weather image recovery technology shows significant performance gaps in real-world applications and ignores the semantic context of the image, resulting in insufficient generalization capability and low recovery quality.
The semi-supervised learning and teacher-student network structure based on visual language model are adopted to optimize the image restoration model by combining labeled synthetic images and label-free real-world images, and visual language models are used for visibility evaluation and pseudo-label generation, combining weather prompt learning and semantic regularization loss.
It significantly improves the generalization ability and image restoration performance of the model under harsh weather conditions in real world, improves image clarity and semantic recovery quality, and enables the model to restore harsh weather images in the real world more accurately.
Smart Images

Figure CN118537264B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and artificial intelligence, and in particular, to a method for restoring images of real - world bad weather based on a vision - language model. Background Art
[0002] Images captured under bad weather conditions are usually affected by various factors, such as raindrops, haze, and snowflakes, which significantly reduce the quality of the images. Image restoration under bad weather conditions is an important research direction in the field of computer vision, aiming to restore the clarity and visibility of images through various algorithms. Traditional methods usually focus on a single weather phenomenon (such as rain removal, haze removal, or snow removal), and use specific algorithms to process images captured under specific weather conditions. In recent years, researchers have tried to develop integrated solutions to handle image restoration problems under multiple bad weather conditions through a single model. Although these methods have shown certain effectiveness on synthetic datasets, in real - world application scenarios, due to the significant domain differences between the training data and the actual application scenarios, their generalization ability is limited.
[0003] In some existing technologies, a single model is used to handle image restoration techniques under multiple bad weather conditions, and image restoration for different bad weather conditions is achieved through joint training and unified model weights. In addition, through techniques such as knowledge distillation and contrast learning, some methods attempt to handle image restoration under multiple weather conditions. Although these methods have made progress in handling image restoration under multiple bad weather conditions, they usually ignore the semantic context of the images, and most methods still rely on synthetic data for training, which limits their effectiveness in real - world applications.
[0004] In summary, there are two main problems in the existing bad - weather image restoration technologies: First, most of the existing methods are trained based on synthetic data, resulting in significant performance gaps when dealing with real - world scenarios; Second, while these methods restore image clarity, they often ignore the semantic context of the scene and fail to fully utilize the semantic information of the images to improve the restoration quality. Due to these limitations, the existing technologies are difficult to effectively handle complex bad - weather conditions in the real world, thus affecting the practicality and accuracy of image restoration models in key applications such as urban surveillance and autonomous driving. Summary of the Invention
[0005] The purpose of the present invention is to overcome the above - mentioned defects of the existing technologies and provide a method for restoring images of real - world bad weather based on a vision - language model to improve the image restoration performance in real - world scenarios under multiple bad weather conditions.
[0006] The purpose of the present invention can be achieved through the following technical solutions:
[0007] In one aspect of the present invention, a method for restoring images of bad weather in real scenes based on a vision-language model is provided. The input images are processed using a pre-trained image restoration framework to achieve image restoration. Among them, the training process of the image restoration framework includes:
[0008] Obtain samples of bad weather images with labels and input them into the student network. Based on the restored images and the corresponding labels, calculate the supervision loss;
[0009] Obtain samples of unlabeled bad weather images and input them into the teacher network and the student network respectively. Based on the images restored by the student network, use the vision-language model to evaluate the visibility. Based on the pseudo-labels and the images restored by the student network, calculate the pseudo-label loss.
[0010] Based on the images restored by the student network, use the pre-trained CLIP model to calculate the corresponding image embeddings. Based on the image embeddings and the text embeddings of the preset clear weather prompt vectors, calculate the weather prompt learning loss.
[0011] Based on the images restored by the student network and the scene descriptions of the preset unlabeled bad weather image samples under good weather conditions, calculate the semantic regularization loss.
[0012] Calculate the feature similarity loss for aligning the images restored by the student network, the pseudo-labels, and the unlabeled bad weather image samples.
[0013] Based on the supervision loss, pseudo-label loss, weather prompt learning loss, semantic regularization loss, and feature similarity loss, calculate the comprehensive loss to achieve the training of the image restoration framework.
[0014] As a preferred technical solution, the generation process of the pseudo-labels includes:
[0015] Before training, use the vision-language model to perform visibility scoring on the unlabeled bad weather image samples to generate pseudo-labels.
[0016] During training, for the images restored by the student network, use the vision-language model to perform visibility scoring. If the scoring result is higher than the visibility score of the corresponding unlabeled bad weather image sample, update the pseudo-labels.
[0017] As a preferred technical solution, the pseudo-label loss is calculated using the following formula:
[0018]
[0019] where respectively represent the score and pseudo-label for evaluating the visibility of the image restored by the student network using the vision-language model, represents the appearance loss function.
[0020] As a preferred technical solution, the training process of the CLIP model includes:
[0021] Construct weather prompt vectors for capturing imaging features related to weather conditions in four weather conditions: clear, rainy, foggy, and snowy;
[0022] Initialize the weather prompt vectors in the embedding space of the text encoder of the CLIP model, and use real images in different weather conditions to extract reference image embeddings through the image encoder of the CLIP model, and achieve the training of the CLIP model with the goal of minimizing the weather classification loss.
[0023] As a preferred technical solution, the weather prompt learning loss is calculated using the following formula:
[0024]
[0025] where cos(·,·) represents the cosine similarity, respectively represent the text embeddings of the weather prompt vectors in four weather conditions: clear, rainy, foggy, and snowy, represents the image restored by the student network embedding.
[0026] As a preferred technical solution, the process of obtaining the scene description of the unlabeled adverse weather image sample under good weather conditions includes:
[0027] For the unlabeled adverse weather image sample, use another vision-language model to generate a negative scene description of the image weather condition, and convert the corresponding positive scene description based on the negative scene description.
[0028] As a preferred technical solution, the semantic regularization loss is calculated using the following formula:
[0029]
[0030] where cos(·,·) represents the cosine similarity, respectively represent the embeddings of the positive scene description and the negative scene description of the image weather condition, represents the image restored by the student network embedding.
[0031] As a preferred technical solution, the comprehensive loss is calculated using the following formula:
[0032]
[0033] Among them, are the supervision loss, the pseudo-label loss, the weather prompt learning loss, the semantic regularization loss, and the feature similarity loss respectively. w1, w2, w3, and w4 are weighting parameters.
[0034] Another aspect of the present invention provides an electronic device, including: one or more processors and a memory. The memory stores one or more programs, and the one or more programs include instructions for executing the foregoing real-scene bad weather image restoration method based on a vision-language model.
[0035] Another aspect of the present invention provides a computer-readable storage medium, including one or more programs for execution by one or more processors of an electronic device. The one or more programs include instructions for executing the foregoing real-scene bad weather image restoration method based on a vision-language model.
[0036] Compared with the prior art, the present invention has at least one of the following beneficial effects:
[0037] (1) Strong model generalization ability: By introducing an image restoration framework based on semi-supervised learning and a teacher-student network structure, and training by combining labeled synthetic images and unlabeled real-world images, the generalization ability of the model under real-world bad weather conditions is significantly improved. This solves or partially solves the problem of insufficient generalization ability caused by the prior art relying on synthetic data training, enabling the model to more accurately restore bad weather images in the real world.
[0038] (2) Strong adaptability of image restoration under different weather conditions: The present invention adopts a weather prompt learning method. By giving specific weather condition prompts to the vision-language model, the model is trained to accurately identify and handle image features in various weather situations, which not only improves the restoration quality of image clarity but also enhances the adaptability and accuracy of the model under diverse weather conditions.
[0039] (3) Strong image semantic restoration ability: The present invention realizes the training of the image restoration framework based on description-assisted semantic enhancement. By converting the negative descriptions of bad weather images into positive scene descriptions to guide the image semantic restoration process, it can effectively restore the image semantic content distorted by weather-related noise while maintaining the improvement of image visual clarity, ensuring that the restored image is not only visually close to the ideal state but also more real and accurate in content, which is crucial for improving the performance of downstream tasks such as object detection and semantic segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 Schematic diagram of the image restoration framework in the embodiment;
[0041] Figure 2 Schematic diagram of the image visibility evaluation in the embodiment;
[0042] Figure 3 Schematic diagram of the weather prompt learning in the embodiment;
[0043] Figure 4 Schematic diagram of the semantic enhancement of the description assistance in the embodiment. Specific implementation manners
[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0045] Embodiment 1
[0046] The existing image restoration technologies for bad weather mainly have the following disadvantages:
[0047] (1) Insufficient generalization ability: Most of the existing methods are mainly trained based on synthetic data, resulting in significant performance gaps when they process bad weather images in real-world scenarios. This method is difficult to adapt to the changing real-world conditions, limiting its practicality.
[0048] (2) Ignoring semantic information: The existing technologies have made certain progress in improving the visual clarity of images, but often ignore the semantic context of the restored images. This results in the fact that although the images look clearer, the restored images may not accurately reflect the true intention and content of the scene, affecting the comprehensiveness of the image restoration quality.
[0049] (3) Lack of an effective actual data training mechanism: Since most of the existing methods rely on synthetic data, they lack an effective training mechanism for real-world data. This limits their effects in practical applications because there are obvious differences between synthetic data and real-world data.
[0050] Therefore, the existing methods are usually trained on synthetic images. When processing images under real-world bad weather conditions, there are often problems of insufficient generalization ability, and at the same time, the semantic context related to the weather is often ignored, resulting in limited performance.
[0051] To address the above drawbacks, this embodiment provides a method for restoring images in real-world adverse weather conditions based on a vision-language model, which adopts an image restoration framework based on semi-supervised learning and a teacher-student learning network, and utilizes the vision-language model to improve the image restoration performance in real-world scenarios under various adverse weather conditions, aiming to make up for these deficiencies in the prior art.
[0052] See Figure 1 FIG. for a schematic diagram of the training process of the image restoration framework. The image restoration framework includes a trainable student network (i.e., the restoration network (student) in Figure 1 and a non-trainable teacher network (i.e., the restoration network (teacher) in Figure 1 . The teacher network is the exponential moving average of the student network. During training, labeled data (i.e., adverse weather image samples with labels) are input into the student network, and the supervised loss is calculated based on the restored image (i.e., the restoration value (labeled) in Figure 1 ) and the label. In addition, unlabeled data (i.e., unlabeled adverse weather image samples) are input into the student network and the teacher network respectively, and the restoration values (student) and the restoration value (teacher) are obtained respectively. Based on the restoration value (teacher), a vision-language model is used for visibility evaluation, the pseudo-label database is updated, and the pseudo-label loss with the restoration value (student) is calculated. Using another vision-language model, the weather cue learning loss and the semantic regularization loss (i.e., the cue & semantic loss in Figure 1 ) are calculated based on the restoration value (student). Finally, based on the supervised loss, pseudo-label loss, weather cue learning loss, and semantic regularization loss calculated above, the parameter training of the student network is achieved. The framework adopts multiple vision-language models (VLMs) to improve the clarity and semantic information of the image during the process of removing weather-related noise.
[0053] The key points of this method are as follows:
[0054] (1) Image Assessment and Pseudo-Labeling: By introducing a vision-language model to comprehensively evaluate images in adverse weather, this technique does not rely on traditional labeled data training. By generating high-quality pseudo-labels, unlabeled real-world images can be effectively utilized, thereby improving the generalization ability and restoration performance of the model in real scenarios.
[0055] To ensure high-quality pseudo-labels, a pseudo-label database is established to evaluate the effect of adverse weather image restoration by leveraging the zero-shot ability of large vision-language models. Subsequently, paired clear real images are not required for model training.
[0056] Image Evaluation: For the prediction of real adverse weather images and corresponding methods for removing weather effects, a key issue is how to evaluate the quality of restored images. Existing low-level image quality assessment methods mainly focus on technical distortions, including noise, blur, and compression artifacts. However, there is a situation where images affected by adverse weather, although free from common noise interference, have a significant decrease in visibility due to rain, fog, and snow noise. Therefore, it becomes particularly important to find an effective method to automatically evaluate the quality of images under the background of removing adverse weather noise. This method uses a vision-language model to prompt based on weather-related image quality problems and quantifies the response of the vision-language model into a numerical score.
[0057] As Figure 2 shown in the schematic diagram of image visibility evaluation, first, a conversion template is designed to ask the vision-language model to evaluate the image. Then, a five-level scoring system in the study of Mean Opinion Score (MOS) is adopted, namely excellent, good, medium, poor, and extremely poor, corresponding to scores from one to five. After that, based on the score r vlm predicted by the vision-language model, the probabilities of these five vocabulary labels are converted into a numerical score. Thus, we obtain the visibility evaluation of each restored image.
[0058] Pseudo-label Generation: For unlabeled adverse weather image samples Assign and update pseudo-labels for them according to the image evaluation based on the vision-language model Select ideal pseudo-label images without weather-related artifacts The study found that r vlm can obtain better pseudo-labels with fewer weather-related artifacts.
[0059] Initially, a pseudo-label database is constructed to store the current best pseudo-labels of unlabeled images. During the model training process, the score of the vision-language model based on image visibility will be evaluated, and the predictions of the model and the recorded pseudo-labels will be evaluated. If the model achieves a better restoration effect, that is, the image restored by the student network obtains a higher quality score, then the image restored by this student network will be used as the updated ideal pseudo-label image, and the pseudo-label database will be updated accordingly. In practice, the prediction of the teacher network, which is the exponential moving average of the student network, is used for comparison. Finally, the pseudo-label loss of the online model is calculated using the updated pseudo-labels: where and are the prediction and the corresponding pseudo-label respectively, is any type of appearance loss, such as
[0060] (2) Weather Prompt Learning: The weather prompt learning method is adopted. By giving specific weather condition prompts to the vision-language model, the model is trained to accurately identify and handle image features under various weather conditions. This not only improves the restoration quality of image clarity but also enhances the adaptability and accuracy of the model under diverse weather conditions.
[0061] Leveraging the rich knowledge embedded in large pre-trained vision-language models, this framework can understand the concepts of images under favorable and adverse weather conditions. In particular, the CLIP model is expected to identify weather conditions of images, such as sunny, rainy, foggy, or snowy. Subsequently, the learned concept of "clarity" is used to guide the model learning to achieve clear restoration results. To enhance the ability of CLIP to accurately distinguish weather conditions in multiple scenarios, the prompt learning method is adopted to obtain prompt embeddings for image features of each weather condition, as Figure 3 shown, where (a) is a schematic diagram of the prompt representation learning process and (b) is a schematic diagram of the restoration model optimization process. It should be noted that specific weather prompts are used for learning in this method. As an alternative, more diverse weather condition prompts can be introduced, or dynamically generated prompts can be adopted to adapt to a wider range of weather scenarios and image features.
[0062] Prompt embedding learning: Referring to (a) in Figure 3 , the CLIP model integrates image and text encoders in a shared feature space and demonstrates excellent performance in visual classification and natural language. The prompt learning method is adopted instead of relying on labor-intensive and error-prone prompt engineering, such as handcrafted text prompts like "rain" or "a photo of rain". Specifically, while keeping the parameters of the pre-trained CLIP model fixed, a set of four weather prompts {t c , t r , t h , t s} representing sunny, rain, fog, and snow conditions are used as learnable vectors. These different context prompts capture the imaging features related to each weather condition.
[0063] Initializing weather prompts in the embedding space of the CLIP text encoder. Meanwhile, real images y under different weather conditions (sunny, rain, fog, snow) are collected to extract reference image embeddings through the CLIP image encoder The training objective is to minimize the classification loss, i.e., the cross-entropy loss, by classifying the weather prompts into their respective weather categories c: p(c = i|y) = σ(z i ), where σ represents the Softmax function and cos(·, ·) represents the cosine similarity.
[0064] Restoration model optimization: Refer to Figure 3 In (b) of, the knowledge obtained from the learned weather cues is used to guide the training of the restoration model to generate images with enhanced clarity. Specifically, in the restoration model optimization stage, the weather cue learning loss Maximizes the image embedding of the model-predicted restored image (i.e., the predicted output of the student network, corresponding to Figure 1 The restoration value (student) in ) and the text embedding of the clear weather cue The similarity between.
[0065]
[0066] The knowledge obtained from the learned weather cues guides the training of the restoration model to generate images with better clarity. In this stage, by maximizing the similarity between the image embedding of the restored image predicted by the model And the text embedding of the clear weather cue to optimize the weather cue learning loss
[0067] In the preliminary study, only The prediction of the optimized model is used to reduce weather-related artifacts, but the resulting images show significant noise. Assume that there is a feasible solution within the space of minimizing the weather cue learning loss. To address this issue and enhance the regularization of model learning, the feature similarity loss Is used to align the model's prediction with the pseudo-label y ps And the input x u . This method uses the visual encoder of Depth Anything for feature extraction to ensure robustness in multiple scenarios.
[0068] (3) Description-assist Semantic Enhancement: Utilize the text generation ability of the vision-language model to transform the negative description of the bad weather image into a positive scene description to guide the semantic restoration process of the image. This highlights that while maintaining the improvement of image visual clarity, it can also effectively restore the semantic content of the image, ensuring that the restored image is not only visually close to the ideal state but also more realistic and accurate in content.
[0069] Restoring images under adverse weather conditions not only requires enhancing image clarity but also recovering the image semantics distorted by weather-related noise, which is crucial for improving the effectiveness of downstream tasks. Existing image restoration methods trained on synthetic data often overlook the potential of restoring image semantics. Therefore, in this method, a method is introduced that utilizes the image-text understanding ability of large vision-language models to enhance the semantics of adverse weather image restoration.
[0070] When restoring images with degraded quality, this task becomes challenging due to the ill-posed nature of the image restoration task itself. In contrast, using natural language to describe the appearance and content of adverse weather images is relatively straightforward. In this method, first, vision-language models (VLMs) are utilized to provide natural language descriptions containing rich semantic information about the scene and weather conditions for adverse weather images. As Figure 4 shown, where (a) is the semantic loss calculation process and (b) is the negative description - positive description conversion process. For a given input image, a (negative) caption describing the image's weather condition is generated using a VLM. For example, "A person is walking along the street in heavy rain..." describes a scene containing the object "person" in rainy weather. The description also provides additional environmental context, such as the surrounding environment "street". This text information, together with the image, provides a more intuitive high-level understanding of the scene.
[0071] Subsequently, the negative scene description d neg associated with the image with degraded quality under adverse weather conditions is converted into a pseudo-clear representation. This conversion is achieved by prompting large language models (LLMs) to generate a positive description d pos corresponding to the restored image. For example, given the above negative description of adverse weather, we can imagine its positive, clearly restored image, such as "The weather looks great. A person is walking..." Intuitively, d pos and d neg should have similar descriptions of the image content such as objects and the environment, but different descriptions of the weather and visibility, i.e., good weather vs. bad weather. Different from the aforementioned weather cues, d pos and d neg are customized for a specific image. To achieve this, refer to Figure 4 (b), and use large VLMs, such as LLaVA, to generate d neg . After that, by prompting LLMs, such as Llama, to generate d pos .
[0072] The training of the model integrates semantic-aware regularization to promote the prediction results to be consistent with the positive description indicating good weather conditions. Given the positive and negative descriptions, a description-assisted semantic regularization loss
[0073]
[0074] It should be noted that the described semantic enhancement of the auxiliary in this method can be further expanded through more complex natural language processing techniques. For example, sentiment analysis can be introduced to evaluate the impact of weather on the scene atmosphere, or a knowledge graph can be used to enrich the semantic understanding of the image content.
[0075] In the initial experiment, it was observed that LLMs were occasionally difficult to generate weather change descriptions with unchanged content. To solve this problem, some examples of negative-to-positive description conversion were manually marked and introduced into the in-context learning method. Finally, the total loss is the supervised appearance loss semi-supervised pseudo-label loss weather prompt loss description-assisted semantic loss and feature similarity loss which is a weighted combination of, where w1, w2, w3, w4 are weighted parameters used to balance the loss values:
[0076]
[0077] Based on the above comprehensive loss the network parameters of the student network are updated. It should be noted that the loss function is not limited to the above formula. Alternative solutions include exploring new combinations of loss functions or automated methods for adjusting the weights of loss terms to achieve more optimized training results.
[0078] The combination of the above technologies not only significantly improves the restoration quality of images under adverse weather conditions, but also ensures the accuracy and richness of semantic information during the image restoration process, making this method show superior performance in terms of image clarity and semantic restoration, fully meeting the requirements for high-quality image restoration in real-world applications.
[0079] During the prediction process, the pre-trained student network is used to process the input real-scene adverse weather image to obtain the restored image.
[0080] To verify the effectiveness of this method, a feasibility verification was carried out in the form of experiments and user studies, and the results show the superiority and practicality of this technical solution. Specifically:
[0081] Quantitative comparison: This method was compared with a variety of the latest all-weather image restoration methods, including Restormer, TransWeather, TKL, WeatherDiff, WGWS-Net, MWDT, PromptIR, and DA-CLIP, etc. Multiple no-reference image quality assessment metrics were used for quantitative evaluation, including NIMA, MUSIQ, CLIP-IQA, LIQE, and Q-Align, as well as the proposed image visibility assessment method VLM-Vis based on vision-language models. On all image quality assessment metrics, this method ranked first on average and performed best under almost all weather conditions. In addition, this method obtained the best VLM-Vis score under different weather conditions. These results demonstrate the superiority of this method over existing state-of-the-art all-weather image restoration methods (mainly focused on synthetic data evaluation) on real data.
[0082] Qualitative comparison: A qualitative evaluation was conducted on the real-world evaluation dataset. Compared with other methods, this method is more effective in processing real-world all-weather images, especially in reducing artifacts of rain, fog, and snow. It is worth noting that our method can effectively eliminate the fog effect in rain and snow scenes, significantly improving image visibility.
[0083] User study: A user study was conducted to evaluate visual quality. Ten real-world images were randomly selected for comparison for each weather scene. A total of 32 participants were invited to evaluate, half of whom had relevant image processing experience. Two factors were considered, namely image visibility and quality, focusing on the degree of removal of weather-related artifacts and the authenticity of the restored images. Overall, this method demonstrated obvious advantages over other competing methods in terms of visibility and quality under different weather conditions.
[0084] In summary, through a series of experiments and user studies, this method has proven its feasibility and effectiveness in improving the performance of image restoration under all-weather conditions, especially in enhancing image clarity, semantic restoration, and superiority in real-world application scenarios.
[0085] In summary, compared with the prior art, this method has at least one of the following advantages:
[0086] (1) Improving generalization ability: The semi-supervised learning framework adopted utilizes labeled synthetic images and unlabeled real-world images. By combining labeled synthetic images and unlabeled real-world images for training, it significantly enhances the model's generalization ability and image restoration performance under real-world all-weather conditions. This solves or partially solves the problem of insufficient generalization ability caused by the prior art relying on synthetic data training, enabling the model to more accurately restore all-weather images in the real world.
[0087] (2) It can remove weather-related noise: By learning specific weather condition prompt embeddings through a vision-language model, the model's ability to recognize and process image features in different weather conditions is enhanced, thereby improving image clarity and the effect of removing weather-related noise.
[0088] (3) Enhance image semantic restoration: Use the vision-language model to generate natural language descriptions of images in bad weather and convert them into positive descriptions representing clear weather conditions to guide the image semantic restoration process, effectively restoring the semantic content of images distorted by weather-related noise. This is often overlooked in the prior art. Especially when restoring image clarity, it is crucial to retain and recover the semantic information of the image for improving the performance of downstream tasks such as object detection and semantic segmentation.
[0089] (4) Optimize training efficiency and effect: Through image evaluation and pseudo-label generation techniques, unlabeled real-world images can be effectively utilized for model training. At the same time, through iterative training and model optimization strategies, the model parameters are continuously fine-tuned to further improve the quality and accuracy of the restored images.
[0090] (5) Improve user experience: By enhancing image clarity and semantic accuracy, the visual quality of the restored images is significantly improved, providing users with a more satisfactory visual experience.
[0091] In summary, through its unique technical solution, this method not only solves the limitations of the prior art technically but also brings significant improvements in application effects, including improving efficiency, reducing costs, and enhancing user experience, demonstrating superior performance and application value compared to the prior art.
[0092] It should be noted that although specific vision-language models (such as CLIP and LLaVA) are used in this method, other existing or future-developed vision-language models can be considered as alternative solutions to meet different image restoration requirements and optimization goals.
[0093] This method is not limited to the restoration of images in bad weather and can also be used for:
[0094] Other types of image restoration tasks: In addition to the restoration of images in bad weather, this technical solution can also be applied to other types of image restoration tasks, such as low-light image enhancement, image denoising, super-resolution, etc.
[0095] Video restoration and enhancement: Extend this technical solution to the video field to restore and enhance videos captured under bad weather conditions, improving video quality and visibility.
[0096] Augmented Reality (AR): In AR applications, enhancing the visual quality and realism of the virtual environment is a major challenge. This technical solution can be used to enhance the image quality in the environment and provide a more immersive user experience.
[0097] Embodiment 2
[0098] This embodiment provides an electronic device, including: one or more processors and a memory. The memory stores one or more programs, and the one or more programs include instructions for executing the real-scene bad weather image restoration method based on the vision language model as described in Embodiment 1.
[0099] Embodiment 3
[0100] This embodiment provides a computer-readable storage medium, including one or more programs for execution by one or more processors of an electronic device. The one or more programs include instructions for executing the real-scene bad weather image restoration method based on the vision language model as described in Embodiment 1.
[0101] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for restoring a real scene severe weather image based on a visual language model, characterized in that: The input image is processed using a pre-trained image restoration framework to restore the image, wherein the training process of the image restoration framework includes: Obtain labeled severe weather image samples and input them into the student network, and calculate the supervision loss based on the restored images and corresponding labels; Obtain unlabeled severe weather image samples and input them into the teacher network and the student network respectively, perform visibility evaluation using a visual language model based on the image restored by the student network, and calculate the pseudo-label loss based on the pseudo-label and the image restored by the student network. Based on the image restored by the student network, the corresponding image embedding is calculated using the pre-trained CLIP model, and the weather prompt learning loss is calculated based on the image embedding and the text embedding of the preset sunny weather prompt vector. Based on the image restored by the student network and the scene description of the preset unlabeled bad weather image samples under good weather conditions, the semantic regularization loss is calculated. Compute the feature similarity loss used to align the restored images from the student network, the pseudo-labels, and the unlabeled severe weather image samples; Based on the supervision loss, pseudo label loss, weather prompt learning loss, semantic regularization loss and feature similarity loss, the comprehensive loss is calculated to realize the training of the image restoration framework. Among them, the weather prompts learning loss The calculation is done using the following formula: , in, represents the cosine similarity, Text embeddings of weather prompt vectors for four weather conditions: sunny, rainy, foggy, and snowy. Represents the image after the student network is restored Embedding.
2. According to claim 1, a method for restoring a real scene severe weather image based on a visual language model is characterized in that: The pseudo-label generation process includes: Before training, the visual language model is used to score the visibility of unlabeled severe weather image samples and generate pseudo labels; During the training process, the image restored by the student network is scored for visibility using a visual language model. If the score is higher than the visibility score of the corresponding unlabeled severe weather image sample, the pseudo label is updated.
3. The method for restoring a real scene severe weather image based on a visual language model according to claim 1, characterized in that: The pseudo-label loss The calculation is done using the following formula: , in, , They respectively represent the score and pseudo label of the visibility evaluation of the image restored by the student network using the visual language model, represents the appearance loss function.
4. The method for restoring a real scene severe weather image based on a visual language model according to claim 1, characterized in that: The training process of the CLIP model includes: Construct weather cue vectors for capturing imaging features related to weather conditions under four weather conditions: clear, rainy, foggy and snowy. A weather prompt vector is initialized in the embedding space of the text encoder of the CLIP model, and reference image embeddings are extracted through the image encoder of the CLIP model using real images under different weather conditions, so as to train the CLIP model with the goal of minimizing weather classification loss.
5. The method for restoring a real scene severe weather image based on a visual language model according to claim 1, characterized in that: The process of obtaining the scene description of the unlabeled bad weather image sample under good weather conditions includes: For unlabeled severe weather image samples, another visual language model is used to generate negative scene descriptions describing the weather conditions of the images, and the corresponding positive scene descriptions are converted based on the negative scene descriptions.
6. The method for restoring a real scene severe weather image based on a visual language model according to claim 1, characterized in that: The semantic regularization loss The calculation is done using the following formula: , in, represents the cosine similarity, , Represent the embeddings of the positive and negative scene descriptions of the image weather conditions, respectively. Represents the image after the student network is restored Embedding.
7. The method for restoring a real scene severe weather image based on a visual language model according to claim 1, characterized in that: The comprehensive loss The calculation is done using the following formula: , in, , , , , They are supervision loss, pseudo label loss, weather prompt learning loss, semantic regularization loss, and feature similarity loss. , , , is the weighting parameter.
8. An electronic device, characterized in that: include: One or more processors and a memory, wherein the memory stores one or more programs, and the one or more programs include instructions for executing the real scene severe weather image restoration method based on the visual language model as described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that: It includes one or more programs for execution by one or more processors of an electronic device, and the one or more programs include instructions for executing the real scene severe weather image restoration method based on the visual language model as described in any one of claims 1-7.
Citation Information
Patent Citations
Image restoration method under various severe weather conditions
CN116681615A
Non-ideal supervision-based severe weather image restoration system
CN117391969A