Image restoration method and device for various interferences, equipment and medium

By introducing a prompt generation network into the image restoration model, generating and fusing a variety of prompt information to assist the image restoration network in dealing with a variety of interference factors, the complexity and resource requirements in the prior art are solved, and a more efficient image restoration effect is achieved.

CN120147192AActive Publication Date: 2025-06-13BEIJING UNIV OF POSTS & TELECOMM
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510226143.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-13
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with the image degradation problem caused by multiple interference factors, and requires the deployment of multiple independent models, which increases complexity and resource requirements.

Method used

An image restoration model based on multiple prompts is adopted, including an image restoration network and a prompt generation network. The prompt generates content prompts, style prompts and learning prompts, and generates target prompt information through fusion encoding, and combines the image restoration network feature layer to assist in image restoration.

Benefits of technology

It realizes flexible response to various interference factors, improves image restoration effect and performance, and reduces the number of models and computing resource requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147192A_ABST
    Figure CN120147192A_ABST
Patent Text Reader

Abstract

The invention provides an image restoration method and device for various interferences, equipment and a medium, and relates to the technical field of computer vision. The method at least comprises the following steps: processing a sample degradation image through a to-be-trained prompt generation network, and respectively generating content prompt information, style prompt information and to-be-learned prompt information; performing fusion coding on the multiple pieces of prompt information through a to-be-trained prompt generation network to obtain target prompt information; in the process of carrying out image restoration on the sample degraded image by the to-be-trained image restoration network, combining the target prompt information with a feature layer of the image restoration network to obtain a sample restoration image; training a to-be-trained image restoration model based on the sample restoration image and the sample clean image to obtain a trained image restoration model; and inputting the to-be-restored image into the trained image restoration model to obtain a restored image, thereby improving the image restoration effect and performance for various interferences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular, to an image restoration method, apparatus, device and medium for various interferences. Background Art

[0002] Computer Vision (CV) is playing an increasingly important role in various fields of real life, widely applied in various scenarios such as recognition and detection, and has become one of the indispensable key technologies for building smart cities, autonomous driving systems, and intelligent security systems. However, in the actual application process, the quality of image acquisition has a decisive impact on the accuracy of subsequent processing tasks. Factors such as haze, rain and snow weather, camera shake, thermal noise, shadow occlusion, low light conditions, and image loss will significantly reduce the image quality and generate damaged images. Furthermore, it seriously affects the performance of downstream tasks. Image restoration technology is precisely born to solve the above problems. It aims to restore damaged or degraded images to a state close to their original state through algorithms. In this context, how to efficiently handle various types of image damage problems to achieve high-quality image restoration has become a key topic of great concern in the field of computer vision research. With the continuous progress of technology and the increasing richness of application scenarios, this challenge has also prompted researchers to continuously explore more advanced solutions, striving to further improve the overall performance and reliability of image processing.

[0003] Image restoration technology is a comprehensive method that combines denoising and enhancement. It uses mathematical models and computer vision algorithms to optimize image quality and enhance detail clarity. This technology aims to recover clearer and more accurate information from images affected by noise, rain, snow, fog, or degraded image quality due to problems such as low light and shadows. By improving the overall quality of the image, image restoration can not only enhance the visual effect of the image itself but also significantly improve the performance of subsequent processing steps such as semantic segmentation, object detection, and tracking, thereby greatly enhancing the accuracy and stability of the results when analyzing these images. Therefore, in the face of complex and variable interference factors, developing high-performance image restoration algorithms has become one of the indispensable key links in promoting the development of the computer vision field.

[0004] Currently, image restoration usually only focuses on solving a certain type of degradation problem, and its performance is often unsatisfactory when encountering multiple interference factors. In addition, in order to process various types of damaged images, multiple independent models need to be deployed, which greatly increases the complexity in practical applications. Although multi-task image restoration algorithms have gradually attracted attention in recent years, the types of tasks they can effectively handle are still limited and the restoration effects are often unsatisfactory. Therefore, how to develop an image restoration method that can flexibly cope with multiple interference factors to improve the image restoration effect for multiple interferences is a technical problem to be solved in the present invention. Summary of the invention

[0005] Based on the above technical problems, the present invention provides an image restoration method, device, equipment and medium for multiple interferences, aiming to improve the image restoration effect and performance for multiple interferences.

[0006] A first aspect of the present invention provides an image restoration method for multiple interferences, the method comprising: Inputting the sample degraded image into the image restoration network to be trained and the hint generation network to be trained in the image restoration model to be trained respectively; The sample degraded image is processed by the prompt generation network to be trained to generate content prompt information, style prompt information and prompt information to be learned corresponding to the sample degraded image, respectively; the content prompt information represents the image content prompt corresponding to the sample degraded image, the style prompt information represents the degradation type prompt corresponding to the sample degraded image, and the prompt information to be learned represents the semantic prompt information corresponding to the sample degraded image and having a higher dimension than the content prompt information and the style prompt information; fusing and encoding the content prompt information, the style prompt information and the prompt information to be learned through the prompt generation network to be trained to obtain target prompt information; In the process of restoring the sample degraded image by the image restoration network to be trained, combining the target prompt information with the feature layer of the image restoration network to obtain a sample restored image output by the image restoration network; Based on the sample restored image and the sample clean image corresponding to the sample degraded image, training the image restoration model to be trained until a trained image restoration model is obtained; The image to be restored is input into the trained image restoration model to obtain a restored image.

[0007] A second aspect of the present invention provides an image restoration device for multiple interferences, the device comprising: An image input module, configured to input a sample degraded image into a to-be-trained image restoration network and a to-be-trained prompt generation network in the to-be-trained image restoration model respectively; A prompt generation module, configured to process the sample degraded image through the to-be-trained prompt generation network to generate content prompt information, style prompt information, and to-be-learned prompt information corresponding to the sample degraded image respectively; the content prompt information represents the image content prompt corresponding to the sample degraded image, the style prompt information represents the degradation type prompt corresponding to the sample degraded image, and the to-be-learned prompt information represents semantic prompt information of a higher dimension than the content prompt information and the style prompt information corresponding to the sample degraded image; A prompt encoding module, configured to perform fusion encoding on the content prompt information, the style prompt information, and the to-be-learned prompt information through the to-be-trained prompt generation network to obtain target prompt information; An image output module, configured to combine the target prompt information with a feature layer of the image restoration network during the process of the to-be-trained image restoration network restoring the sample degraded image, to obtain a sample restored image output by the image restoration network; A model training module, configured to train the to-be-trained image restoration model based on the sample restored image and the sample clean image corresponding to the sample degraded image until a trained image restoration model is obtained; An image restoration module, configured to input an image to be restored into the trained image restoration model to obtain a restored image.

[0008] The third aspect of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the computer program is executed by the processor, it implements the image restoration method for multiple interferences according to the first aspect of the embodiments of the present invention.

[0009] The fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, where when the computer program is executed by a processor, it implements the image restoration method for multiple interferences according to the first aspect of the embodiments of the present invention.

[0010] In the image restoration method for multiple interferences provided by the present invention, an image restoration model based on multiple cues is trained. The image restoration model includes two parts: an image restoration network and a cue generation network. The image restoration network is used to complete the image restoration task, and the cue generation network is responsible for generating content cue information, style cue information, and cue information to be learned, so as to respectively represent the image content cue, degradation type cue, and higher-dimensional semantic cue information of the degraded image, and integrate the multiple cues to realize the guidance of the image restoration network and assist in completing the image restoration. Thus, through the image restoration method of the present invention, the style information cue, content information cue, and cue information to be learned are respectively realized in the way of multiple cues, and the multiple cue information is fused with the image restoration network to assist in image restoration, completing the image restoration for multiple interferences (i.e., multiple tasks), solving more types of image degradation problems at one time, being able to flexibly cope with multiple interference factors for image restoration, and improving the image restoration effect and performance for multiple interferences. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0012] Figure 1 is a flowchart of the steps of an image restoration method for multiple interferences shown in an embodiment of the present invention; Figure 2 is a training schematic diagram of a content perceptron shown in an embodiment of the present invention; Figure 3 is a training schematic diagram of a style perceptron shown in an embodiment of the present invention; Figure 4 is a structural schematic diagram of a perceptron to be learned shown in an embodiment of the present invention; Figure 5 is a structural schematic diagram of a cue encoder shown in an embodiment of the present invention; Figure 6 is an overall network structural schematic diagram of an image restoration model shown in an embodiment of the present invention; Figure 7 is a training flowchart of an image restoration model shown in an embodiment of the present invention; Figure 8 is an inference flowchart of an image restoration model shown in an embodiment of the present invention; Figure 9 is a comparison diagram of visualization effects under three task scenarios proposed in an embodiment of the present invention; Figure 10 It is a comparison diagram of visualization effects under five task scenarios proposed in an embodiment of the present invention; Figure 11 It is a structural block diagram of an image restoration device for multiple interferences provided in an embodiment of the present invention. Specific embodiments

[0013] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0014] Please refer to Figure 1 , Figure 1 It is a step flowchart of an image restoration method for multiple interferences shown in an embodiment of the present invention. As Figure 1 shown, the image restoration method for multiple interferences provided in this embodiment at least includes the following steps: Step S11: Input the sample degraded images into the image restoration network to be trained and the prompt generation network to be trained in the image restoration model to be trained respectively.

[0015] In this embodiment, the image restoration model to be trained consists of two parts: the image restoration network to be trained and the prompt generation network to be trained. Among them, the image restoration network is used to complete the image restoration task, and the prompt generation network is responsible for generating various prompts to assist the image restoration network to complete the image restoration. In this embodiment, a sample data set is prepared for training the image restoration model. The sample data set includes: sample degraded images and sample clean images corresponding to the sample degraded images. The sample clean image is the label corresponding to the sample degraded image, which is the clear image after removing the interference corresponding to the sample degraded image.

[0016] Among them, the sample degraded image is a degraded image sample for training the image restoration model. The degraded image in this embodiment refers to an image with one or more interferences, such as a damaged image generated due to interference factors such as noise, haze, rain and snow, occlusion, and / or low light, and there is no limitation thereto; the sample data set includes multiple sample degraded images with different interference factors.

[0017] In this embodiment, the sample degraded images are input into the image restoration network to be trained and the prompt generation network to be trained respectively, and the sample degraded images are processed in parallel by the image restoration network to be trained and the prompt generation network to be trained.

[0018] Step S12: Process the sample degraded image through the to-be-trained prompt generation network to respectively generate content prompt information, style prompt information, and learnable prompt information corresponding to the sample degraded image.

[0019] In this embodiment, after inputting the sample degraded image into the to-be-trained prompt generation network, the to-be-trained prompt generation network processes the sample degraded image to respectively generate content prompt information corresponding to the sample degraded image, style prompt information corresponding to the sample degraded image, and learnable prompt information corresponding to the sample degraded image. Among them, the content prompt information (Content Prompt) represents the image content prompt corresponding to the sample degraded image, the style prompt information (Style Prompt) represents the degradation type prompt corresponding to the sample degraded image, and the learnable prompt information (Learnable Prompt) represents the semantic prompt information with a higher dimension than the content prompt information and the style prompt information corresponding to the sample degraded image.

[0020] Step S13: Perform fusion encoding on the content prompt information, the style prompt information, and the learnable prompt information through the to-be-trained prompt generation network to obtain target prompt information.

[0021] In this embodiment, after the to-be-trained prompt generation network generates the content prompt information, style prompt information, and learnable prompt information corresponding to the sample degraded image, in order to integrate these prompts, the to-be-trained prompt generation network also performs fusion encoding on the content prompt information, style prompt information, and learnable prompt information to obtain target prompt information, which is used to guide the image restoration network to complete the image restoration task.

[0022] Step S14: During the process of the to-be-trained image restoration network performing image restoration on the sample degraded image, combine the target prompt information with the feature layer of the image restoration network to obtain the sample restored image output by the image restoration network.

[0023] In this embodiment, after inputting the sample degraded image into the to-be-trained image restoration network, the to-be-trained image restoration network performs image restoration on the sample degraded image. During the process of the sample degraded image being restored, the target prompt information output by the to-be-trained prompt generation network is combined with the feature layer in the to-be-trained image restoration network to achieve the restoration of a clear image, and the sample restored image output by the to-be-trained image restoration network is obtained.

[0024] Step S15: Train the to-be-trained image restoration model based on the sample restored image and the sample clean image corresponding to the sample degraded image until a trained image restoration model is obtained.

[0025] In this embodiment, the image restoration model to be trained may be trained based on the sample restored image corresponding to the sample degraded image and the sample clean image corresponding to the sample degraded image, until a trained image restoration model is obtained.

[0026] In an optional implementation, the loss between the sample restored image and the sample clean image can be calculated (for example, calculating the L1 loss, etc.), and the model parameters of the image restoration model to be trained are updated based on the calculated loss until a trained image restoration model is obtained.

[0027] Step S16: input the image to be restored into the trained image restoration model to obtain a restored image.

[0028] In this embodiment, the trained image restoration model is used to restore degraded images with multiple interferences. In the application process of the trained image restoration model, the image to be restored can be input into the trained image restoration model to obtain the restored image output by the trained image restoration model. The image to be restored is an image that needs to be restored for image interference, and the restored image is an image after the image to be restored is restored. This embodiment does not impose any restrictions on the interference factors involved in the image to be restored. The trained image restoration model of this embodiment can realize image restoration under single interference and compound interference conditions.

[0029] In this embodiment, the prompt generation network uses multiple prompts to realize the three parts of style information prompts, content information prompts and prompts to be learned respectively, and the multiple prompt information is integrated with the image restoration network to assist in image restoration, completing the image restoration for multiple interferences (i.e., multiple tasks), so as to solve more types of image degradation problems at one time, and can flexibly respond to multiple interference factors for image restoration, further improving the image restoration effect and image restoration performance for multiple interferences. In this way, it also avoids the design and training of special models for each specific degradation situation in the related art, reduces the demand for computing resources and storage space, and reduces the workload of maintaining multiple independent systems.

[0030] In combination with the above embodiments, in one implementation, the present invention further provides an image restoration method for multiple interferences, wherein the prompt generation network to be trained includes at least: a pre-trained content sensor, a pre-trained style sensor, and a sensor to be learned to be trained; in the method, the above step S12 may specifically include steps S21 to S23: Step S21: inputting the sample degraded image into the pre-trained content sensor to obtain the content prompt information.

[0031] In this embodiment, the pre-trained content perceptron and the pre-trained style perceptron in the prompt generation network to be trained are perceptrons with pre-trained and fixed parameters, which are not integrated with the image restoration network to be trained and do not participate in the training of the entire image restoration model. Moreover, both the pre-trained content perceptron and the pre-trained style perceptron in this embodiment are retrained by the image encoder in the pre-trained visual text network.

[0032] After the sample degraded image is input into the prompt generation network to be trained, three perceptrons in the prompt generation network to be trained (the pre-trained content perceptron, the pre-trained style perceptron, and the perceptron to be learned to be trained) will generate corresponding prompts: content prompt information, style prompt information, and prompt information to be learned. Specifically, the sample degraded image is input into the pre-trained content perceptron to obtain the content prompt information output by the pre-trained content perceptron.

[0033] Step S22: Input the sample degraded image into the pre-trained style perceptron to obtain the style prompt information.

[0034] In this embodiment, after the sample degraded image is input into the prompt generation network to be trained, the sample degraded image is input into the pre-trained style perceptron to obtain the style prompt information output by the pre-trained style perceptron.

[0035] Step S23: Input the sample degraded image into the perceptron to be learned to be trained to obtain the prompt information to be learned.

[0036] Since the content perceptron and the style perceptron are trained separately and do not participate in the training of the entire image restoration by being integrated with the image restoration network to be trained, in order to provide learnable parameters for the network during the entire image restoration process, this embodiment proposes a perceptron to be learned, which is used to infer that there are similar situations between some scenes in the image. For example, there is often a haze condition on rainy days, hoping that the prompt generation network can find the connection between the scenes.

[0037] After the sample degraded image is input into the prompt generation network to be trained, the sample degraded image is input into the perceptron to be learned to be trained to obtain the prompt information to be learned output by the perceptron to be learned to be trained.

[0038] It should be noted that this embodiment does not impose any limitation on the execution order between step S21 and step S23. For example, step S21 to step S23 can be executed in any order, or can be executed simultaneously.

[0039] Combined with the above embodiments, in one implementation, the present invention also provides an image restoration method for multiple interferences. In this method, the training steps of the pre-trained content perceptron may include steps S31 to S34: Step S31: Input the sample clean image corresponding to the sample degraded image into the pre-trained text-image multi-modal network to obtain the clean image description text corresponding to the sample degraded image.

[0040] Since the sample data set used to train the image restoration model often lacks a large number of text-image pairs, in this embodiment, by inputting these images into the pre-trained text-image multi-modal network to generate content descriptions, the text information of the clean image is generated.

[0041] Specifically, in this embodiment, the sample clean image corresponding to the sample degraded image is input into the pre-trained text-image multi-modal network to obtain the clean image description text corresponding to the sample degraded image output by the pre-trained text-image multi-modal network. Among them, the text-image multi-modal network may include: the BLIP network, which is a multi-modal encoder-decoder architecture and can generate corresponding text according to the image. The sample degraded image and the sample clean image used to train the content perception device in this embodiment are the same as the sample degraded image and the sample clean image used to train the image restoration model, belonging to the same sample data set.

[0042] Step S32: Input the clean image description text into the text encoder in the pre-trained visual text network to obtain the first text feature.

[0043] Since the scale of the image restoration data set is limited, especially the number of degraded image-clean image descriptions is very small, it is difficult to directly use these data to train the network to converge. To solve this problem, in this embodiment, a pre-trained visual text network is used. The pre-trained visual text network depends on a large number of image-text pair data for pre-training to achieve the alignment between image features and text features, and it already contains rich image and text encoding information. The pre-trained visual text network includes: a text encoder and an image encoder. And when training the content perception device, the text encoder of the pre-trained visual text network is frozen, and only the image encoder of the pre-trained visual text network is re-trained.

[0044] In this embodiment, after obtaining the clean image description text corresponding to the sample degraded image, the clean image description text can be input into the text encoder in the pre-trained visual text network to obtain the first text feature output by the text encoder in the pre-trained visual text network. This first text feature is the text feature output by the text encoder during the training of the content perception device.

[0045] Step S33: Input the sample degraded image into the image encoder in the pre-trained visual text network to obtain the first image feature.

[0046] In this embodiment, the sample degraded image is also input into the image encoder in the pre-trained vision-text network to obtain the first image feature output by the image encoder in the pre-trained vision-text network. The first image feature is the image feature output by the image encoder during the training of the content perceptron. Among them, when training the image encoder, the weights (network parameters) of the image encoder in the pre-trained vision-text network are used as the weights (network parameters) of the initialized image encoder.

[0047] Step S34: Compare and learn the first text feature and the first image feature to obtain a first contrastive learning loss, and update the network parameters of the image encoder based on the first contrastive learning loss until the updated first image encoder is obtained. Take the updated first image encoder as the pre-trained content perceptron.

[0048] In this embodiment, the content perceptron uses contrastive learning for training: after obtaining the first text feature and the first image feature, the first text feature and the first image feature are compared and learned to obtain a first contrastive learning loss. Based on the first contrastive learning loss, the network parameters of the image encoder in the pre-trained vision-text network are updated until the updated first image encoder is obtained, so as to take the updated first image encoder as the trained content perceptron, that is, the pre-trained content perceptron.

[0049] That is to say, in this embodiment, by inputting the sample degraded image-clean image description text, the image encoder in the pre-trained vision-text network is trained to obtain an image encoder for determining the image content, which is used as the pre-trained content perceptron.

[0050] In an alternative embodiment, the vision-text network can be a CLIP network. CLIP (Contrastive Language-Image Pretraining) is a model trained by a contrastive learning method that can map images and texts to the same latent space. It jointly trains image and text data, enabling the model to understand the relationship between text and images, and is widely used in tasks such as image generation, search, and classification.

[0051] In one embodiment, as Figure 2 shown, Figure 2 is a training schematic diagram of the content perceptron shown in an embodiment of the present invention. In Figure 2 the graphic-text multimodal network is a BLIP network, and the vision-text network can be a CLIP network. Figure 2On the left is the sample clean image corresponding to the sample degraded image. The sample clean image is input into the pre-trained BLIP network to obtain the clean image description text corresponding to the sample degraded image. Then, the clean image description text is input into the text encoder of the pre-trained CLIP network to obtain the first text feature. Figure 2 On the right is the sample degraded image. The sample degraded image is input into the image encoder of the pre-trained CLIP network to obtain the first image feature. Based on the comparison between the first text feature and the first image feature, the network parameters of the image encoder of the pre-trained CLIP network are updated again until the trained image encoder of the CLIP network is obtained to be used as the pre-trained content-aware detector.

[0052] Combined with the above embodiments, in one implementation, the present invention also provides an image restoration method for multiple interferences. In this embodiment, the training steps of the pre-trained style-aware detector may include steps S41 to S44: Step S41: Determine the dataset corresponding to the sample degraded image and determine the scene description text corresponding to the dataset.

[0053] For the generation of style-related prompt text, this embodiment adopts context creation based on the dataset to determine the scene description: determine the dataset corresponding to the sample degraded image, and based on the dataset corresponding to the sample degraded image, determine the scene description text corresponding to the dataset, that is, the scene description text corresponding to the sample degraded image. The sample degraded image used to train the style-aware detector in this embodiment is the same as the sample degraded image used to train the image restoration model and belongs to the same sample dataset. The dataset corresponding to the sample degraded image is the sample dataset. It can be understood that the scene description text is the description text of the degradation type corresponding to the sample degraded image. For example, the description texts of degradation types such as fog, noise, and occlusion.

[0054] In this embodiment, each dataset corresponds to a scene description. For example, in a fog dataset, the description is "A photo taken on a hazy day". By way of example, the scene description texts corresponding to each scene dataset are shown in Table 1: Table 1 Correspondence table of scenes and scene description texts

[0055] Step S42: Input the scene description text into the text encoder of the pre-trained vision-text network to obtain the second text feature.

[0056] This embodiment is the same as the training content perceptron. It uses a pre-trained visual-text network. The pre-trained visual-text network depends on a large amount of image-text pair data for pre-training to achieve alignment between image features and text features, and it already contains rich image and text encoding information. The pre-trained visual-text network includes: a text encoder and an image encoder. When training the style perceptron, freeze the text encoder of the pre-trained visual-text network and only retrain the image encoder of the pre-trained visual-text network again.

[0057] In this embodiment, after obtaining the scene description text corresponding to the sample degraded image, the scene description text can be input into the text encoder in the pre-trained visual-text network to obtain the second text feature output by the text encoder in the pre-trained visual-text network. This second text feature is the text feature output by the text encoder during the training of the style perceptron.

[0058] Step S43: Input the sample degraded image into the image encoder in the pre-trained visual-text network to obtain a second image feature.

[0059] In this embodiment, the sample degraded image is also input into the image encoder in the pre-trained visual-text network to obtain the second image feature output by the image encoder in the pre-trained visual-text network. This second image feature is the image feature output by the image encoder during the training of the style perceptron. Among them, when training this image encoder, the weights (network parameters) of the image encoder in the pre-trained visual-text network are used as the weights (network parameters) of the initialized image encoder.

[0060] Step S44: Perform contrastive learning on the second text feature and the second image feature to obtain a second contrastive learning loss, and update the network parameters of the image encoder based on the second contrastive learning loss until the updated second image encoder is obtained. Take the updated second image encoder as the pre-trained style perceptron.

[0061] In this embodiment, the style perceptron also uses the method of contrastive learning for training: after obtaining the second text feature and the second image feature, perform contrastive learning on the second text feature and the second image feature to obtain a second contrastive learning loss, and update the network parameters of the image encoder in the pre-trained visual-text network based on the second contrastive learning loss until the updated second image encoder is obtained, so as to take the updated second image encoder as the trained style perceptron, that is, the pre-trained style perceptron.

[0062] That is to say, in this embodiment, by inputting the sample degraded image - scene description text, the image encoder in the pre - trained vision - text network is trained to obtain an image encoder for determining the image degradation type, which serves as the pre - trained style perceptron. In an alternative embodiment, the vision - text network can be a CLIP (Contrastive Language - Image Pretraining) network.

[0063] In one embodiment, as Figure 3 shown, Figure 3 is a schematic diagram of the training of the style perceptron shown in an embodiment of the present invention. In Figure 3 , the text - image multi - modal network is a BLIP network, and the vision - text network can be a CLIP network. Figure 3 On the left is the scene description text corresponding to the sample degraded image, belonging to datasets of different dotted - line types (i.e., squares of different dotted - line types in the figure). Then, the scene description text is input into the text encoder of the pre - trained CLIP network to obtain the second text feature. Figure 3 On the right is the sample degraded image. The sample degraded image is input into the image encoder of the pre - trained CLIP network to obtain the second image feature. Based on the contrastive learning between the second text feature and the second image feature, the network parameters of the image encoder of the pre - trained CLIP network are updated again until the trained image encoder of the CLIP network is obtained, which serves as the pre - trained style perceptron.

[0064] Combining the above embodiments, in one implementation, the present invention also provides an image restoration method for multiple interferences. In this embodiment, the to - be - learned perceptron to be trained at least includes: a pre - trained vision Transformer structure, a to - be - trained first multi - layer perceptron, a to - be - trained second multi - layer perceptron, a to - be - trained self - attention layer, and a to - be - trained cross - attention layer; the above step S23 can specifically include steps S51 to S55: Step S51: Input the sample degraded image into the pre - trained vision Transformer structure to obtain a plurality of image feature vectors.

[0065] In this embodiment, after the sample degraded image is input into the to - be - learned perceptron to be trained, first, the sample degraded image is input into the pre - trained vision Transformer (Vision Transformer, ViT) structure to obtain a plurality of image feature vectors. At this stage, the pre - trained vision Transformer structure remains frozen, and there is no gradient propagation or parameter update. For example, the pre - trained vision Transformer structure in this embodiment can be a ViT - b16 model pre - trained on MAGENET1K.

[0066] Step S52: Input the multiple image feature vectors into the first multi-layer perceptron to be trained for processing, obtain multiple first feature vectors, and perform self-attention operation among the multiple first feature vectors through the self-attention layer to be trained, so as to obtain a self-attention result.

[0067] In this embodiment, input the multiple image feature vectors output by the pre-trained Vision Transformer structure into the first multi-layer perceptron to be trained for processing, obtain multiple first feature vectors, and then perform self-attention operation (Self-Attention operation) among the multiple first feature vectors through the self-attention layer to be trained, so as to obtain the self-attention result output by the self-attention layer to be trained, and update the parameters of the first multi-layer perceptron during this process.

[0068] Step S53: Initialize multiple vectors to be learned, and perform cross-attention operation on the multiple vectors to be learned and the self-attention result through the cross-attention layer to be trained, so as to obtain a cross-attention result.

[0069] In this embodiment, before the model training of the image restoration model to be trained starts, randomly initialize multiple vectors to be learned, and then input the multiple vectors to be learned and the obtained self-attention result into the cross-attention layer to be trained. Through the cross-attention layer to be trained, perform cross-attention operation (Cross-Attention operation) on the multiple vectors to be learned and the self-attention result, so as to obtain the cross-attention result output by the cross-attention layer to be trained.

[0070] Step S54: Input the cross-attention result into the second multi-layer perceptron to be trained for processing, so as to obtain updated vectors to be learned.

[0071] In this embodiment, after obtaining the cross-attention result, input the cross-attention result into the second multi-layer perceptron to be trained for processing, so as to obtain the updated vectors to be learned output by the second multi-layer perceptron to be trained.

[0072] It should be noted that in the subsequent model training process, the updated learnable vector is used as the initialized learnable vector to perform cross-attention operation with the subsequent self-attention result, so as to obtain the subsequent cross-attention result. Moreover, the first multi-layer perceptron and the second multi-layer perceptron in this embodiment are different multi-layer perceptrons, corresponding to different network parameters and not sharing weights. Among them, the multi-layer perceptron (Multi-layer Perceptron, MLP) is also called an artificial neural network. During the training process of the learnable perceptron, the parameters in the first multi-layer perceptron, the second multi-layer perceptron, the self-attention layer, and the cross-attention layer will be updated, and the first multi-layer perceptron, the second multi-layer perceptron, the self-attention layer, and the cross-attention layer each have independent parameters to be trained and do not share weights.

[0073] Step S55: Obtain the learnable prompt information based on the updated learnable vector and the cross-attention result.

[0074] In this embodiment, after obtaining the updated learnable vector, the learnable prompt information output by the learnable perceptron to be trained can be obtained based on the updated learnable vector and the cross-attention result. In an optional example, it may be to perform a residual connection based on the updated learnable vector and the cross-attention result to obtain the learnable prompt information.

[0075] In one embodiment, as Figure 4 shown, Figure 4 is a schematic structural diagram of the learnable perceptron shown in an embodiment of the present invention. In Figure 4 the left side is the sample degraded image. The sample degraded image is first processed by a pre-trained vision Transformer structure (i.e., Figure 4 the ViT in Figure 4 ), generating 196 image feature vectors, and the dimension of each image feature vector is 768. At this stage, the ViT remains frozen and no gradient propagation or parameter update occurs. These 196 image feature vectors pass through the first multi-layer perceptron ( Figure 4 the upper row of MLP in Figure 4The MLP in the lower middle line) processes the cross-attention results, thereby generating a set of 32 updated learnable vectors, and each updated learnable vector has 768 dimensions. Finally, an element-wise addition is performed on these updated learnable vectors and the cross-attention results to obtain a 768*1 vector as the final learnable prompt information.

[0076] Currently, using prompt information to guide another network to generate results has become a common method. However, there is no general formula for how to combine prompt information with the eigenvalues of the network. Generally speaking, the method of combining prompt information with the eigenvalues of the original network depends on the amount of information contained in the prompt information. When the prompt information is a scalar with less information, such as using time T to guide the diffusion model to generate, usually T is directly embedded into the network features through multiplication. When the prompt information contains more information, such as using the output of the object detection network as the prompt information, the Embedding layer is used to align the prompt information with the shape of the network feature map. After alignment, the connection operation is used to combine the prompt information and the network eigenvalues. When using the eigenvalues of other backbone networks, encoders and other neural networks as encoded information, the information content is the richest, and the Embedding layer can be used to align the prompt information with the shape of the network feature map, and then connect it through the self-attention layer.

[0077] The embodiments of the present invention use three types of prompt information, and their information densities are different: among them, the content prompt information contains the most information, the learnable prompt information is the second, and the style prompt information contains the least information. How to combine these three types of prompt information with the feature map in the middle of the image restoration network becomes a problem. To solve this problem, in combination with the above embodiments, in one implementation, the present invention also provides an image restoration method for multiple interferences. In this method, the to-be-trained prompt generation network further includes: a to-be-trained prompt encoder, and the to-be-trained prompt encoder at least includes: three to-be-trained third multi-layer perceptrons and three to-be-trained fourth multi-layer perceptrons; and, the above step S13 may specifically include steps S61 to S63: Step S61: Input the content prompt information, the style prompt information, and the learnable prompt information into their respective corresponding to-be-trained third multi-layer perceptrons for processing and then perform a linear transformation to obtain the first prompt feature, the second prompt feature, and the third prompt feature corresponding to the content prompt information, the style prompt information, and the learnable prompt information respectively.

[0078] In this embodiment, three third multi-layer perceptrons and three fourth multi-layer perceptrons are respectively used to process the content prompt information, style prompt information, and to-be-learned prompt information. That is, each type of prompt information corresponds to a third multi-layer perceptron and a fourth multi-layer perceptron respectively. The third multi-layer perceptron and the fourth multi-layer perceptron in this embodiment are different multi-layer perceptrons, and the six multi-layer perceptrons in the prompt encoder to be trained correspond to different network parameters and do not share weights.

[0079] In this embodiment, the content prompt information, style prompt information, and to-be-learned prompt information are respectively input into their corresponding third multi-layer perceptrons to be trained for processing and then linearly transformed to obtain the first prompt feature, second prompt feature, and third prompt feature corresponding to the content prompt information, style prompt information, and to-be-learned prompt information respectively.

[0080] Step S62: Input the first prompt feature, the second prompt feature, and the third prompt feature into their corresponding fourth multi-layer perceptrons to be trained for processing to obtain the fourth prompt feature, fifth prompt feature, and sixth prompt feature corresponding to the content prompt information, style prompt information, and to-be-learned prompt information respectively.

[0081] In this embodiment, then, the first prompt feature, the second prompt feature, and the third prompt feature are respectively input into their corresponding fourth multi-layer perceptrons to be trained for processing to obtain the fourth prompt feature, fifth prompt feature, and sixth prompt feature corresponding to the content prompt information, style prompt information, and to-be-learned prompt information respectively.

[0082] Step S63: After connecting the fourth prompt feature, the fifth prompt feature, and the sixth prompt feature, perform processing through self-attention and a feed-forward network in sequence to obtain the target prompt information output by the prompt encoder to be trained.

[0083] In this embodiment, after obtaining the fourth prompt feature, the fifth prompt feature, and the sixth prompt feature, connect the fourth prompt feature, the fifth prompt feature, and the sixth prompt feature to obtain a connection feature. Then, perform self-attention processing on the connection feature to obtain a seventh prompt feature, and then process the seventh prompt feature through a feed-forward network (Feed Forward Neural Network, FFN) to obtain the target prompt information output by the prompt encoder to be trained.

[0084] In this embodiment, in order to effectively fuse prompt information, a prompt encoder is proposed. This is an effective fusion module for multiple prompts, which can fuse multiple promptings and be used by downstream modules.

[0085] In one embodiment, as Figure 5 shown, Figure 5The following is a schematic structural diagram of a prompt encoder shown in an embodiment of the present invention. In Figure 5 each piece of prompt information (style prompt information, content prompt information, and prompt information to be learned) is first processed by a third multi-layer perceptron ( Figure 5 the MLP in the second-to-last row), and then undergoes a linear transformation and is further processed by a fourth multi-layer perceptron ( Figure 5 the MLP in the fourth-to-last row), reducing it from the original 768 dimensions to 256 dimensions. Then, the three 256-dimensional prompts are concatenated to form a 768-dimensional prompt again. This final 768-dimensional prompt is successively operated on through self-attention and a feed-forward network to obtain the final target prompt information for input into the image restoration network.

[0086] In related technologies, image restoration methods have shown good performance on some datasets, but not all datasets have been conquered. When facing relatively complex datasets, the performance is often unsatisfactory and cannot effectively restore damaged images. To solve this problem, in combination with the above embodiments, in one implementation manner, the present invention also provides an image restoration method for multiple interferences. In this method, the image restoration network to be trained is a diffusion model to be trained; and, specifically, the above step S14 may include step S71 and step S72: Step S71: The sample degraded image is successively subjected to forward diffusion and reverse diffusion through the image restoration network to be trained.

[0087] In this embodiment, the image restoration network to be trained first performs forward diffusion on the sample degraded image to add noise to the sample degraded image, and then performs reverse diffusion on the noisy sample image to denoise the noisy sample degraded image, so as to perform image restoration of the sample degraded image.

[0088] Step S72: During the reverse diffusion process of the reverse diffusion, the target prompt information is combined with the feature layer in the reverse diffusion process through a cross-attention mechanism to obtain the sample restored image.

[0089] In this embodiment, during the reverse diffusion process of the reverse diffusion of the image restoration network to be trained, through a cross-attention mechanism, the target prompt information output by the prompt generation network to be trained is combined with the feature layer in the reverse diffusion process of the image restoration network to be trained to obtain the sample restored image output by the image restoration network to be trained. In an optional implementation manner, the diffusion model to be trained in this embodiment can be obtained by improving the structure based on Stable Diffusion.

[0090] In view of the problem of insufficient absolute performance when dealing with complex data sets in related technologies, this embodiment uses a diffusion model to solve this problem. The diffusion model has excellent capabilities in the field of image generation and is very suitable for solving this problem after being modified.

[0091] In one embodiment, in order to more effectively solve the problem of restoring degraded images of multiple tasks, this embodiment proposes an image restoration model based on prompts and a diffusion model, constituting a new paradigm in the field of image generation based on pre-training of large vision-language models, as Figure 6 shown. Figure 6 FIG. is a schematic diagram of the overall network structure of an image restoration model shown in an embodiment of the present invention. In Figure 6 it, the image restoration model includes: an image restoration network (Image Restoration Part, IRP) and a prompt generation network (Prompt Generation Part, PGP). Among them, IRP processes the task of image restoration, while PGP is responsible for generating prompts.

[0092] Among them, the prompt generation network includes four components: a content perceptron, a style perceptron, a to-be-learned perceptron, and a prompt encoder. When a degraded image is input into the network, three perceptrons in the prompt generation network will generate corresponding prompts: a content prompt, a style prompt, and a to-be-learned prompt. Then, these prompts are uniformly encoded by the prompt encoder to generate a unified prompt used by the image restoration network. Finally, the unified prompt is combined with the feature layer in the image restoration network to achieve the restoration of a clear image. Experimental results show that this method is superior to many mainstream methods in both single-interference and composite-interference cases. Thus, in view of the problem that most related technologies can only solve a few concentrated images, this embodiment uses the method of multi-prompt information to solve this problem; in the face of different degraded images, the image restoration model of this embodiment can solve more types of network degradation problems at one time, and is an image restoration model with stronger adaptability.

[0093] Combined with the above embodiments, in one implementation manner, the present invention also provides an image restoration method for multiple interferences. In this method, step S15 above may specifically include the following steps: Based on the sample restored image and the sample clean image corresponding to the sample degraded image, train the to-be-trained image restoration network, the to-be-trained to-be-learned perceptron, and the to-be-trained prompt encoder until a trained image restoration network, a trained to-be-learned perceptron, and a trained prompt encoder are obtained; based on the trained image restoration network, the pre-trained content perceptron, the pre-trained style perceptron, the trained to-be-learned perceptron, and the trained prompt encoder, obtain the trained image restoration model.

[0094] In an optional embodiment, the image restoration method for various interferences shown in any of the foregoing embodiments can be deployed and implemented as a preprocessing part in a vision system to assist in completing subsequent vision tasks. For example, it can be applied to a variety of different vision application scenarios, including but not limited to: autonomous driving vehicles, security monitoring systems, industrial automation detection, and medical image analysis, etc. For example, it has broad application prospects in autonomous driving, and accurate road sign recognition and obstacle detection are crucial. By deploying the image restoration preprocessing module (i.e., the image restoration model proposed in the present invention), the vision systems in these fields can obtain more accurate data support and thus make more informed decisions.

[0095] Image restoration often participates in the entire vision network as part of image preprocessing. For example, as the upstream of object detection and object tracking tasks, it assists in better detection and tracking tasks. One practical scenario is as a preprocessing module during real-time detection tasks, reading in damaged images and restoring them. The deployment of the system is divided into a training stage and an inference stage. The flowchart of the training stage is as Figure 7 shown Figure 7 which is the training flowchart of the image restoration model shown in an embodiment of the present invention. The training process is as follows: Construct a dataset: For the one-by-one (a test method for evaluating the network restoration ability, training a separate model for each specific restoration task) image dehazing task, use the RESIDE dataset. For the one-by-one image deraining task, use the Rain100L dataset, which contains 200 pairs of clean rain images for training and 100 pairs of images for testing. For the one-by-one image denoising task, use the BSD400 dataset. BSD400 contains 400 training images. In this embodiment, their corresponding noise versions are generated by adding Gaussian noise with different levels σ∈(15, 25, 50), and then the denoising task is evaluated on the BSD68 dataset. In the all-in-one (a test method for evaluating the network restoration ability, training a single model for image restoration of multiple degradation types) setting, a single model is trained on the combined set of the above training datasets and tested on multiple restoration tasks. To further illustrate the excellent performance of the network proposed in this embodiment, this embodiment also adds a deblurring task and uses the GoPro dataset. For the low-light restoration task, use the LOL dataset.

[0096] Read image data: Read the corresponding image data and its true label data from the training dataset, and the batch size of loading training image data for each GPU is 1.

[0097] Configure training parameters: Build the model using PyTorch 2.0 and train it on a system with 4 NVIDIA A100 GPUs. The specific training parameters are shown in Table 2. First, train the content perceptron and style perceptron on the above dataset. Use the BLIP-ViT-L network to generate image-text pairs. Use the CLIP network implemented based on OpenCLIP and the ViT-B-32 model pre-trained on the laion2b_s34b_b79k pre-training data. During training, CLIP freezes the text encoder and text embedding completely and only trains the image encoder. The batch_size is 256 and the learning rate is 1e-4. After that, the content perceptron and style perceptron are completely frozen during the training process.

[0098] The model outputs the prediction result: The model infers the image restoration result, which is an RGB image identical to the input degraded image.

[0099] Calculate the loss and perform backpropagation: During training, calculate the loss between the ground truth label image corresponding to the input degraded image and the output restored image. Use the L1 loss to calculate the difference between the model output and the expectation.

[0100] Save the training result: After training, obtain the final image restoration model and save it offline, such as in HDF5, TensorFlow SavedModel, ONNX, etc., and export the model through the corresponding APIs or tools. At the same time, the model can be verified on the validation set.

[0101] Table 2 System hardware configuration and software version information table

[0102] The flowchart of the inference stage is as Figure 8 shown Figure 8 This is the inference flowchart of the image restoration model shown in an embodiment of the present invention. The inference process is as follows: Load the model: Load the offline image restoration model using methods such as ONNX and TensorRT. Using tools such as ONNX and TensorRT can significantly improve the model loading and inference efficiency and support cross-platform flexibility.

[0103] Data Reading: Read and decode images from the camera in real-time. During the data reading process, the program will initialize the camera connection, capture image frames in real-time and perform decoding processing, converting them into digital information that can be analyzed for subsequent image processing and analysis algorithms. Initialize the camera through OpenCV or a similar computer vision library. Then, use the camera's API to set parameters such as frame rate and resolution, and start capturing the image stream in real-time. After each frame is captured, use the built-in decoder to convert the original image data (such as H.264 encoded) into matrix data in RGB or grayscale format for subsequent processing. This process is usually continuously executed in a loop to ensure the continuity and real-time nature of the image data.

[0104] Send the read image into the model for inference: The output of the image restoration model is the restored image. Further send the restored image to the next step to complete subsequent vision tasks.

[0105] In addition, in another embodiment, the image restoration model proposed in the above embodiment can be deployed and implemented as part of an image processing software. Whether it is professional graphic design software or lightweight applications for ordinary users, it can provide users with a powerful tool to improve the quality of photos and other visual content. Especially in fields such as post-processing of photography, preservation of historical documents, and medical image analysis, the image restoration function will play an irreplaceable role. At this time, for the deployment and inference stages of the image restoration model: When loading the model, there are some specific considerations when the image restoration model is deployed as part of an image processing software: The trained model and its inference logic can be seamlessly integrated into the existing image processing software framework to ensure good cooperation with other functional modules.

[0106] Data Loading: Real-time data: Read image frames from the camera, and after necessary preprocessing (such as scaling, normalization), send them into the model. Custom data: Users upload image or video files, and the system automatically parses and converts them into a format acceptable to the model, ensuring the same preprocessing standards as the real-time data.

[0107] Save the results: Ensure that the inference results are presented to users in an intuitive way, display the detection results or classification labels through a graphical interface, and enhance the overall user experience.

[0108] Through the above deployment and implementation steps, the inference module can be effectively integrated into the image processing software, not only maintaining consistency with the training stage, but also fully considering the needs of different data sources and actual application scenarios. This deployment method enhances the functionality and flexibility of the software, making it more adaptable to diverse user needs.

[0109] In one embodiment, to verify the effectiveness of the image restoration method for multiple interferences proposed in the embodiments of the present invention, this embodiment strictly follows the previous research and conducts experiments in two different environments: (1) All-In-One task and (2) One-By-One task. In the All-In-One setting, a model is trained to perform image restoration of multiple degradation types. In the One-By-One task setting, a separate model is trained for each specific restoration task.

[0110] I. Results for All-In-One: Three Degradation Types To be consistent with the effect tests of mainstream methods, first, three hybrid image restoration tasks of defogging, haze removal, and denoising were tested. For comparison, this embodiment selected the networks that had been tested on the same tasks, including PromptIR, AirNet, and DL, etc. The quantitative analysis results are shown in Table 3: Table 3 Comparison with Mainstream All-In-One Networks on Three Tasks

[0111] As can be seen from Table 3, in this scenario of the three tasks, the average effect of multiple restoration tasks has been improved, and there are also improvements to varying degrees in specific tasks. Compared with PromptIR, the method (Ours) proposed in the present invention has improved the PSNR index of the defogging task by 0.11 dB. As Figure 9 shown, Figure 9 is the visualization effect comparison diagram under the three task scenarios proposed in one embodiment of the present invention. From Figure 9 the visualized image restoration effect in, GT represents the labeled image. The method proposed in the present invention effectively eliminates the blur in the image and produces a cleaner result than PromptIR.

[0112] II. Results for All-In-One: Five Degradation Types To further illustrate the obvious advantages of the method proposed in the present invention in the face of complex scenarios, in addition to the above three tasks, more task scenario mixtures were carried out. These include five scenarios of haze elimination, denoising, noise reduction, deblurring, and low-light image enhancement. The scenario settings are consistent with the comparison network, and the comparison network also has 5 scenarios. For this reason, 5 datasets were combined to complete the network training. These datasets include the datasets for the above three task settings and other datasets: the GoPro dataset for motion deblurring and the LOL dataset for low-light image enhancement.

[0113] The results are shown in Table 4. Compared with the networks that have also completed the above five-scenario image restoration tasks, the method of the present invention (Ours in Table 4) is still at the leading level. Compared with IDR, the method of the present invention leads by 1.08 dB in the average restoration task of the five scenarios. In each specific task, there are also varying degrees of leads. At the same time, it can be observed that due to the use of the diffusion model, there is an obvious advantage in the denoising task, and the advantage is 2.73 dB. As Figure 10 shown, Figure 10 Figure 4 is a comparison diagram of the visualization effects under the five task scenarios proposed in an embodiment of the present invention. Figure 10 It shows the restoration effect of the damaged image by the image restoration model proposed by the present invention when facing five mixed scenarios, where GT represents the labeled image.

[0114] Table 4 Analysis table for comparison with the mainstream All-In-One network on five tasks

[0115] III. One-By-One Results This embodiment also evaluates the performance of the image restoration model proposed by the present invention in a one-by-one setting, where a separate model is trained for each different restoration task. A series of experiments show that the method proposed by the present invention has achieved good results in the One-By-One scenario. For the defogging task, the method of the present invention is trained on the RESIDE OTS and tested on the OTS-outdoor dataset. The experimental results are shown in Table 5. The results show that compared with PromptIR, the PSNR index of the method of the present invention has increased by 0.25 dB. Compared with the traditional single-task defogging network, it also has certain advantages. For the deraining task, the present invention is trained and tested on the Rain100L dataset. Comparisons are made with PromptIR and other networks. The experimental results are shown in Table 6. It can be observed that the method proposed by the present invention has obvious advantages. In this task, the method of the present invention leads PromptIR by 0.24 dB and 0.004 in the PSNR and SSIM indexes respectively.

[0116] Table 5 Separate table of One-By-One results on the SOTS-Outdoor dataset

[0117] Table 6 One-By-One results on the RAIN100L dataset

[0118] It should be noted that for method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the embodiments of the present invention are not limited by the described action sequences, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.

[0119] Based on the same inventive concept, an embodiment of the present invention provides an image restoration device for multiple interferences. Refer to Figure 11 , Figure 11 which is a structural block diagram of an image restoration device for multiple interferences provided by an embodiment of the present invention. As shown in Figure 11 , the image restoration device for multiple interferences in this embodiment may include: An image input module, configured to input a sample degraded image into a to-be-trained image restoration network and a to-be-trained prompt generation network in the to-be-trained image restoration model respectively; A prompt generation module, configured to process the sample degraded image through the to-be-trained prompt generation network to generate content prompt information, style prompt information, and to-be-learned prompt information corresponding to the sample degraded image respectively; the content prompt information represents the image content prompt corresponding to the sample degraded image, the style prompt information represents the degradation type prompt corresponding to the sample degraded image, and the to-be-learned prompt information represents the semantic prompt information of a higher dimension than the content prompt information and the style prompt information corresponding to the sample degraded image; A prompt encoding module, configured to fuse and encode the content prompt information, the style prompt information, and the to-be-learned prompt information through the to-be-trained prompt generation network to obtain target prompt information; An image output module, configured to combine the target prompt information with the feature layer of the image restoration network during the process of the to-be-trained image restoration network restoring the sample degraded image to obtain the sample restored image output by the image restoration network; A model training module, configured to train the to-be-trained image restoration model based on the sample restored image and the sample clean image corresponding to the sample degraded image until a trained image restoration model is obtained; An image restoration module, configured to input an image to be restored into the trained image restoration model to obtain a restored image.

[0120] Optionally, the prompt generation network to be trained at least includes: a pre-trained content perceptron, a pre-trained style perceptron, and a perceptron to be learned to be trained; the pre-trained content perceptron and the pre-trained style perceptron are re-trained by an image encoder in a pre-trained vision-text network; The prompt generation module includes: The first prompt generation module is used to input the sample degraded image into the pre-trained content perceptron to obtain the content prompt information; The second prompt generation module is used to input the sample degraded image into the pre-trained style perceptron to obtain the style prompt information; The third prompt generation module is used to input the sample degraded image into the perceptron to be learned to be trained to obtain the to-be-learned prompt information.

[0121] Optionally, the image restoration device further includes: a first training module for training the pre-trained content perceptron. The first training module includes: The first text generation module is used to input the sample clean image corresponding to the sample degraded image into the pre-trained text-image multi-modal network to obtain the clean image description text corresponding to the sample degraded image; The first text processing module is used to input the clean image description text into the text encoder in the pre-trained vision-text network to obtain the first text feature; The first image processing module is used to input the sample degraded image into the image encoder in the pre-trained vision-text network to obtain the first image feature; The first perceptron training module is used to perform contrastive learning on the first text feature and the first image feature to obtain the first contrastive learning loss, and update the network parameters of the image encoder based on the first contrastive learning loss until the updated first image encoder is obtained, and use the updated first image encoder as the pre-trained content perceptron.

[0122] Optionally, the image restoration device further includes: a second training module for training the pre-trained style perceptron. The second training module includes: The second text generation module is used to determine the dataset corresponding to the sample degraded image and determine the scene description text corresponding to the dataset; The second text processing module is used to input the scene description text into the text encoder in the pre-trained vision-text network to obtain the second text feature; The second image processing module is used to input the sample degraded image into the image encoder in the pre-trained vision-text network to obtain the second image feature; The second perceptron training module is used to perform contrastive learning on the second text feature and the second image feature to obtain a second contrastive learning loss, and update the network parameters of the image encoder based on the second contrastive learning loss until the updated second image encoder is obtained. The updated second image encoder is used as the pre-trained style perceptron.

[0123] Optionally, the perceptron to be trained for learning at least includes: a pre-trained Vision Transformer structure, a first multi-layer perceptron to be trained, a second multi-layer perceptron to be trained, a self-attention layer to be trained, and a cross-attention layer to be trained; The third prompt generation module includes: The first processing module is used to input the sample degraded image into the pre-trained Vision Transformer structure to obtain a plurality of image feature vectors; The second processing module is used to input the plurality of image feature vectors into the first multi-layer perceptron to be trained for processing to obtain a plurality of first feature vectors, and perform self-attention operations among the plurality of first feature vectors through the self-attention layer to obtain a self-attention result; The third processing module is used to initialize a plurality of vectors to be learned, and perform cross-attention operations on the plurality of vectors to be learned and the self-attention result through the cross-attention layer to obtain a cross-attention result; The fourth processing module is used to input the cross-attention result into the second multi-layer perceptron to be trained for processing to obtain updated vectors to be learned; The fifth processing module is used to obtain the to-be-learned prompt information based on the updated vectors to be learned and the cross-attention result.

[0124] Optionally, the prompt generation network to be trained further includes: a prompt encoder to be trained, and the prompt encoder to be trained at least includes: three third multi-layer perceptrons to be trained and three fourth multi-layer perceptrons to be trained; The prompt encoding module includes: The sixth processing module is used to input the content prompt information, the style prompt information, and the to-be-learned prompt information into their respective corresponding third multi-layer perceptrons to be trained for processing and then perform a linear transformation to obtain a first prompt feature, a second prompt feature, and a third prompt feature corresponding to the content prompt information, the style prompt information, and the to-be-learned prompt information respectively; A seventh processing module, configured to input the first hint feature, the second hint feature, and the third hint feature into respective corresponding fourth multi-layer perceptrons to be trained for processing, so as to obtain a fourth hint feature, a fifth hint feature, and a sixth hint feature corresponding to the content hint information, the style hint information, and the to-be-learned hint information respectively; An eighth processing module, configured to connect the fourth hint feature, the fifth hint feature, and the sixth hint feature, and then sequentially perform processing through self-attention and a feed-forward network to obtain the target hint information output by the to-be-trained hint encoder.

[0125] Optionally, the to-be-trained image restoration network is a to-be-trained diffusion model; the image output module includes: A diffusion processing module, configured to sequentially perform forward diffusion and reverse diffusion on the sample degraded image through the to-be-trained image restoration network; A hint combination module, configured to combine the target hint information with a feature layer in the reverse diffusion process through a cross-attention mechanism during the reverse diffusion process of the reverse diffusion to obtain the sample restored image.

[0126] Based on the same inventive concept, another embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the image restoration method for multiple interferences as described in any of the above embodiments of the present invention are implemented.

[0127] Based on the same inventive concept, another embodiment of the present invention provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes, the steps in the image restoration method for multiple interferences as described in any of the above embodiments of the present invention are implemented.

[0128] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, refer to the partial description of the method embodiment.

[0129] Each embodiment in this specification is described in a progressive manner, and the key point of each embodiment is the difference from other embodiments. For the same and similar parts among the embodiments, refer to each other.

[0130] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, an apparatus, or a computer program product. Therefore, the embodiments of the present invention can take the form of an all-hardware embodiment, an all-software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0131] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or a plurality of flows and / or blocks

[0132] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one or more of the flows Figure 1 or a plurality of flows and / or blocks

[0133] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal devices, such that a series of operation steps are executed on the computer or other programmable terminal devices to generate a computer-implemented process, so that the instructions executed on the computer or other programmable terminal devices provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or a plurality of flows and / or blocks

[0134] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.

[0135] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising said element.

[0136] The above has introduced in detail a method, apparatus, device and medium for image restoration against multiple interferences provided by the present invention. Specific examples are used in this text to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for image restoration against multiple interferences, characterized in that: The method comprises: Inputting the sample degraded image into the image restoration network to be trained and the hint generation network to be trained in the image restoration model to be trained respectively; The sample degraded image is processed by the prompt generation network to be trained to generate content prompt information, style prompt information and prompt information to be learned corresponding to the sample degraded image, respectively; the content prompt information represents the image content prompt corresponding to the sample degraded image, the style prompt information represents the degradation type prompt corresponding to the sample degraded image, and the prompt information to be learned represents the semantic prompt information corresponding to the sample degraded image and having a higher dimension than the content prompt information and the style prompt information; fusing and encoding the content prompt information, the style prompt information and the prompt information to be learned through the prompt generation network to be trained to obtain target prompt information; In the process of restoring the sample degraded image by the image restoration network to be trained, combining the target prompt information with the feature layer of the image restoration network to obtain a sample restored image output by the image restoration network; Based on the sample restored image and the sample clean image corresponding to the sample degraded image, training the image restoration model to be trained until a trained image restoration model is obtained; The image to be restored is input into the trained image restoration model to obtain a restored image.

2. The image restoration method for multiple interferences according to claim 1, characterized in that: The prompt generation network to be trained includes at least: a pre-trained content sensor, a pre-trained style sensor and a to-be-trained learning sensor; the pre-trained content sensor and the pre-trained style sensor are obtained by retraining the image encoder in the pre-trained visual text network; The sample degraded image is processed by the prompt generation network to be trained to generate content prompt information, style prompt information and prompt information to be learned corresponding to the sample degraded image, including: Inputting the sample degraded image into the pre-trained content sensor to obtain the content prompt information; Inputting the sample degraded image into the pre-trained style sensor to obtain the style prompt information; The sample degraded image is input into the to-be-learned perceptron to be trained to obtain the to-be-learned prompt information.

3. The image restoration method for multiple interferences according to claim 2, characterized in that: The training steps of the pre-trained content sensor include: Inputting a sample clean image corresponding to the sample degraded image into a pre-trained image-text multimodal network to obtain a clean image description text corresponding to the sample degraded image; Inputting the clean image description text into a text encoder in a pre-trained visual text network to obtain a first text feature; Inputting the sample degraded image into the image encoder in the pre-trained visual text network to obtain a first image feature; The first text feature and the first image feature are subjected to comparative learning to obtain a first comparative learning loss, and the network parameters of the image encoder are updated based on the first comparative learning loss until an updated first image encoder is obtained, and the updated first image encoder is used as the pre-trained content sensor.

4. The image restoration method for multiple interferences according to claim 2, characterized in that: The training steps of the pre-trained style sensor include: Determine a data set corresponding to the sample degraded image, and determine a scene description text corresponding to the data set; Inputting the scene description text into a text encoder in a pre-trained visual text network to obtain a second text feature; Inputting the sample degraded image into the image encoder in the pre-trained visual text network to obtain a second image feature; The second text feature and the second image feature are subjected to comparative learning to obtain a second comparative learning loss, and the network parameters of the image encoder are updated based on the second comparative learning loss until an updated second image encoder is obtained, and the updated second image encoder is used as the pre-trained style perceptron.

5. The image restoration method for multiple interferences according to claim 2, characterized in that: The perceptron to be trained and learned includes at least: a pre-trained visual Transformer structure, a first multi-layer perceptron to be trained, a second multi-layer perceptron to be trained, a self-attention layer to be trained, and a cross-attention layer to be trained; Inputting the sample degraded image into the to-be-learned perceptron to be trained to obtain the to-be-learned prompt information includes: Inputting the sample degraded image into the pre-trained visual Transformer structure to obtain multiple image feature vectors; Input the plurality of image feature vectors into the first multi-layer perceptron to be trained for processing to obtain a plurality of first feature vectors, and perform a self-attention operation between the plurality of first feature vectors through the self-attention layer to be trained to obtain a self-attention result; Initializing a plurality of vectors to be learned, and performing a cross attention operation on the plurality of vectors to be learned and the self-attention result through the cross attention layer to be trained to obtain a cross attention result; Inputting the cross attention result into the second multi-layer perceptron to be trained for processing to obtain an updated vector to be learned; Based on the updated vector to be learned and the cross-attention result, the prompt information to be learned is obtained.

6. The image restoration method for multiple interferences according to claim 2, characterized in that: The prompt generation network to be trained further includes: a prompt encoder to be trained, and the prompt encoder to be trained includes at least: three third multi-layer perceptrons to be trained and three fourth multi-layer perceptrons to be trained; The content prompt information, the style prompt information and the prompt information to be learned are fused and encoded by the prompt generation network to be trained to obtain target prompt information, including: Inputting the content prompt information, the style prompt information and the prompt information to be learned into the corresponding third multi-layer perceptron to be trained for processing and then performing linear transformation to obtain the first prompt feature, the second prompt feature and the third prompt feature corresponding to the content prompt information, the style prompt information and the prompt information to be learned respectively; Inputting the first prompt feature, the second prompt feature and the third prompt feature into the corresponding fourth multi-layer perceptron to be trained for processing, respectively, to obtain the fourth prompt feature, the fifth prompt feature and the sixth prompt feature corresponding to the content prompt information, the style prompt information and the prompt information to be learned, respectively; After connecting the fourth prompt feature, the fifth prompt feature and the sixth prompt feature, they are processed by self-attention and feedforward networks in turn to obtain the target prompt information output by the prompt encoder to be trained.

7. The image restoration method for multiple interferences according to any one of claims 1 to 6, characterized in that: The image restoration network to be trained is a diffusion model to be trained; in the process of restoring the sample degraded image by the image restoration network to be trained, the target prompt information is combined with the feature layer of the image restoration network to obtain a sample restored image output by the image restoration network, including: Performing forward diffusion and backward diffusion on the sample degraded image in sequence through the image restoration network to be trained; In the reverse diffusion process of the reverse diffusion, the target prompt information is combined with the feature layer in the reverse diffusion process through a cross attention mechanism to obtain the sample restored image.

8. An image restoration device for multiple interferences, characterized in that: The device comprises: An image input module, used for inputting the sample degraded image into the image restoration network to be trained and the hint generation network to be trained in the image restoration model to be trained respectively; a prompt generation module, configured to process the sample degraded image through the prompt generation network to be trained, and respectively generate content prompt information, style prompt information and prompt information to be learned corresponding to the sample degraded image; the content prompt information represents the image content prompt corresponding to the sample degraded image, the style prompt information represents the degradation type prompt corresponding to the sample degraded image, and the prompt information to be learned represents the semantic prompt information corresponding to the sample degraded image and having a higher dimension than the content prompt information and the style prompt information; A prompt encoding module, used for fusing and encoding the content prompt information, the style prompt information and the prompt information to be learned through the prompt generation network to be trained to obtain target prompt information; An image output module, used for combining the target prompt information with the feature layer of the image restoration network during the process of the image restoration network to be trained performing image restoration on the sample degraded image, so as to obtain a sample restored image output by the image restoration network; A model training module, used for training the image restoration model to be trained based on the sample restored image and the sample clean image corresponding to the sample degraded image, until a trained image restoration model is obtained; The image restoration module is used to input the image to be restored into the trained image restoration model to obtain a restored image.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is executed by the processor, the image restoration method for multiple interferences as described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the image restoration method for multiple interferences as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Multi-degradation-factor image restoration method based on collaborative semantic comparison

    CN118096569A

  • Face image restoration method based on generation diffusion prior

    CN118333866A

  • Image restoration method and device based on text prompt

    CN118505571A

  • Self-supervised image denoising method based on diffusion model guidance

    CN118691494A

  • Image super-resolution method and system based on information guide diffusion model

    CN118799188A