Adaptive single image reflection removal method, device, equipment and storage medium
Patent Information
- Application Number
- CN202311311862.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-11
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-10-11
AI Technical Summary
[0005]本申请的主要目的在于提供一种自适应单张图像反光去除方法、装置、设备及存储介质,旨在解决现有的基于深度学习的反光去除方法无法针对每张图片进行针对性的反光去除的技术问题
[0039]本申请提供一种自适应单张图像反光去除方法、装置、设备及存储介质,自适应单张图像反光去除方法包括:基于图像训练集对预设的初始模型进行训练以获得训练后的预训练模型;将所述图像训练集中的各训练图像输入至所述预训练模型中,根据所述预训练模型输出的各第一中间图像对所述预训练模型进行迭代调整以获得目标模型;根据所述目标模型对待去除反光的目标图像进行迭代预测,以得到所述目标图像的反光去除结果。
Smart Images

Figure CN117372304B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to an adaptive single-image reflection removal method, apparatus, device, and storage medium. Background Technology
[0002] During image capture, reflections often occur in the image due to the light source and the material of the subject, which can have a certain impact on subsequent image processing.
[0003] For single-image deglazing, both the reflected image and noise interference within the reflected image are unknown; only the reflected image itself is known. Therefore, single-image deglazing is an ill-posed problem. Existing single-image deglazing techniques utilize prior information about the reflected image, transmission and reflection patterns, and the reflected image, and add constraints to each image to address this ill-posed problem. However, this prior information cannot fully describe the inherent features of the image, and therefore cannot effectively solve the deglazing problem. Existing techniques include deep learning-based methods for single-image deglazing, but these methods model images with various reflection contaminations by directly learning the statistical correlation between the transmission and reflection layers on a large-scale training dataset, and then perform deglazing based on the model.
[0004] However, the above methods use the same de-glare model for different images, ignoring the internal characteristics of the images and failing to perform targeted de-glare removal for each image. As a result, the accuracy of the de-glare images obtained cannot meet the user's needs. Summary of the Invention
[0005] The main objective of this application is to provide an adaptive single-image reflection removal method, apparatus, device, and storage medium, aiming to solve the technical problem that existing deep learning-based reflection removal methods cannot perform targeted reflection removal for each image.
[0006] To achieve the above objectives, this application provides an adaptive single-image reflection removal method, which includes:
[0007] The initial model is trained based on the image training set to obtain the pre-trained model.
[0008] Each training image in the image training set is input into the pre-trained model, and the pre-trained model is iteratively adjusted according to each first intermediate image output by the pre-trained model to obtain the target model;
[0009] The target image to be de-reflected is iteratively predicted based on the target model to obtain the de-reflection result of the target image.
[0010] Optionally, in one feasible embodiment, the step of training a preset initial model based on an image training set to obtain a pre-trained model includes:
[0011] The image training set is input into a preset initial model to obtain the result set output by the initial model;
[0012] Calculate the total loss between the result set and the image training set, and adjust the weight parameters in the initial model according to the total loss to obtain the pre-trained model.
[0013] Optionally, in one feasible embodiment, the step of calculating the total loss between the result set and the image training set includes:
[0014] Calculate the pixel loss, similarity loss, perceptual loss, adversarial loss, exclusion loss, and reconstruction loss between the result set and the image training set;
[0015] The pixel loss, similarity loss, perceptual loss, adversarial loss, exclusion loss, and reconstruction loss are weighted according to a preset weight set to obtain the total loss between the result set and the image training set.
[0016] Optionally, in one feasible embodiment, the step of iteratively adjusting the pre-trained model based on each first intermediate image output by the pre-trained model to obtain the target model includes:
[0017] Based on each first intermediate image output by the pre-trained model, the pre-trained model corresponding to each training image is iteratively adjusted to obtain the iterative model corresponding to each training image.
[0018] Each training image is input into the corresponding iterative model, and the final training result output by each iterative model is obtained.
[0019] The pre-trained model is iteratively adjusted based on the final training results to obtain the target model.
[0020] Optionally, in a feasible embodiment, the step of iteratively adjusting the pre-trained model corresponding to each of the training images based on each first intermediate image output by the pre-trained model includes:
[0021] Obtain the first intermediate image output by each of the pre-trained models;
[0022] Calculate the reconstruction loss between each first intermediate image and the corresponding training image for each first intermediate image;
[0023] Each of the pre-trained models is adjusted according to the reconstruction loss to obtain the first intermediate model corresponding to each of the training images.
[0024] Each of the training images is input into the first intermediate model corresponding to each of the training images;
[0025] The output of each first intermediate model is used as the new first intermediate image. After using each first intermediate model as the new pre-trained model, the process returns to the step of calculating the reconstruction loss between each first intermediate image and the corresponding training image, until the number of iterations reaches a preset first number and then stops. The current first intermediate model is used as the iterative model corresponding to each training image.
[0026] Optionally, in one feasible embodiment, the step of obtaining the first intermediate image output by each of the pre-trained models includes:
[0027] Obtain the transmission layer and reflection layer output by each of the pre-trained models;
[0028] The transmission layer and the reflection layer corresponding to each of the training images are reconstructed to obtain the first intermediate image corresponding to each of the training images.
[0029] Optionally, in a feasible embodiment, the step of iteratively predicting the target image to be de-reflected based on the target model to obtain the reflection removal result of the target image includes:
[0030] The target image to be de-reflected is input into the target model to obtain a second intermediate image;
[0031] Calculate the reconstruction loss between the second intermediate image and the target image, and adjust the target model according to the reconstruction loss to obtain the second intermediate model;
[0032] The first intermediate model is used as the new target model, and the step of inputting the target image to be de-reflected into the target model is returned to be executed until the number of iterations reaches the preset second number and then stops. The current output of the target model is the reflection removal result of the target image.
[0033] Furthermore, to achieve the above objectives, this application also provides an adaptive single-image reflection removal device, which is a virtual device, and includes:
[0034] The pre-training module is used to train a pre-set initial model based on an image training set to obtain a pre-trained model.
[0035] An iterative training module is used to input each training image in the image training set into the pre-trained model, and to iteratively adjust the pre-trained model according to each first intermediate image output by the pre-trained model to obtain the target model;
[0036] The iterative prediction module is used to iteratively predict the target image to be de-reflected based on the target model, so as to obtain the de-reflection result of the target image.
[0037] In addition, to achieve the above objectives, this application also provides an adaptive single-image reflection removal device, which includes: a memory, a processor, and an adaptive single-image reflection removal program stored in the memory and executable on the processor. When the adaptive single-image reflection removal program is executed by the processor, it implements the steps of the adaptive single-image reflection removal method as described above.
[0038] This application also provides a storage medium storing an adaptive single-image reflection removal program, which, when executed by a processor, implements the steps of the adaptive single-image reflection removal method described above.
[0039] This application provides an adaptive single-image reflection removal method, apparatus, device, and storage medium. The adaptive single-image reflection removal method includes: training a preset initial model based on an image training set to obtain a pre-trained model; inputting each training image in the image training set into the pre-trained model; iteratively adjusting the pre-trained model according to each first intermediate image output by the pre-trained model to obtain a target model; and iteratively predicting the target image to be reflected to be removed according to the target model to obtain the reflection removal result of the target image.
[0040] Compared to existing technologies that use a fixed anti-reflection model to train and predict all images, the adaptive single-image anti-reflection removal method of this application first trains a preset initial model based on an image training set to obtain a pre-trained model. Then, each training image in the image training set is input into the pre-trained model. During training, iterative training is performed for each training image. The pre-trained model is adjusted by combining the training results of each training image to obtain the target model. When predicting the target image to be anti-reflected, iterative prediction is also performed on the target image based on the target model to obtain the final prediction result, i.e., the anti-reflection removal result of the target image.
[0041] Thus, the method of iteratively predicting the target image to be de-reflected based on the above-mentioned iterative training model, compared with the traditional method of training and predicting all images using a fixed de-reflection model, the adaptive single-image de-reflection removal method of this application iteratively trains for each training image when training the model, adjusts the model parameters by combining the iterative results of each training image to obtain the target model, and iteratively predicts for the target image when de-reflecting it. This allows for refinement of the prediction results based on the internal features of the target image, improving the targeting of image de-reflection and thus improving the accuracy of the prediction results. Attached Figure Description
[0042] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0043] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without creative effort.
[0044] Figure 1 This is a schematic diagram of the adaptive single-image reflection removal device structure for the hardware operating environment of the device involved in the embodiments of this application;
[0045] Figure 2 This is a schematic diagram illustrating the implementation process of an embodiment of the adaptive single-image reflection removal method of this application;
[0046] Figure 3 This is a schematic diagram of the model framework of an embodiment of the adaptive single-image reflection removal method of this application;
[0047] Figure 4 This is a schematic diagram illustrating the implementation process of an embodiment of the adaptive single-image reflection removal method of this application;
[0048] Figure 5 This is a schematic diagram of the functional modules of the adaptive single-image reflection removal device involved in the embodiments of this application.
[0049] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0050] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0051] It should be noted that during image capture, due to the light source and the material of the subject, reflections often occur in the image, which can have a certain impact on subsequent image processing.
[0052] For single-image deglazing, both the reflected image and noise interference within the reflected image are unknown; only the reflected image itself is known. Therefore, single-image deglazing is an ill-posed problem. Existing single-image deglazing techniques utilize prior information about the reflected image, transmission and reflection patterns, and the reflected image, and add constraints to each image to address this ill-posed problem. However, this prior information cannot fully describe the inherent characteristics of the image, and therefore cannot effectively solve the deglazing problem. Among existing techniques, there are deep learning-based methods for single-image deglazing. These methods model images with various reflection contaminations by directly learning the statistical correlation between the transmission and reflection layers on a large-scale training dataset, and then perform deglazing based on the model. While these methods achieve state-of-the-art performance, their main drawback is that they use the same training weights for all test images. Each image with reflection contamination has its own unique characteristics. If all test images are given uniform weights obtained from an external training set, the unique internal characteristics of the images are ignored. Therefore, satisfactory results cannot be obtained. This is because cascading can refine the results. Some methods in image restoration employ cascading structures and achieve good results. However, these methods blindly cascade modules with the same weights, only improving model performance to a limited extent.
[0053] However, the above methods use the same de-glare model for different images, ignoring the internal characteristics of the images and failing to perform targeted de-glare removal for each image. As a result, the accuracy of the de-glare images obtained cannot meet the user's needs.
[0054] To address the aforementioned problems, this application provides an adaptive single-image reflection removal method, apparatus, device, and storage medium. The adaptive single-image reflection removal method includes: training a preset initial model based on an image training set to obtain a pre-trained model; inputting each training image from the image training set into the pre-trained model; iteratively adjusting the pre-trained model based on each first intermediate image output by the pre-trained model to obtain a target model; and iteratively predicting the target image to be reflected to be removed based on the target model to obtain the reflection removal result of the target image.
[0055] Compared to existing technologies that use a fixed anti-reflection model to train and predict all images, the adaptive single-image anti-reflection removal method of this application first trains a preset initial model based on an image training set to obtain a pre-trained model. Then, each training image in the image training set is input into the pre-trained model. During training, iterative training is performed for each training image. The pre-trained model is adjusted by combining the training results of each training image to obtain the target model. When predicting the target image to be anti-reflected, iterative prediction is also performed on the target image based on the target model to obtain the final prediction result, i.e., the anti-reflection removal result of the target image.
[0056] Thus, the method of iteratively predicting the target image to be de-reflected based on the above-mentioned iterative training model, compared with the traditional method of training and predicting all images using a fixed de-reflection model, the adaptive single-image de-reflection removal method of this application iteratively trains for each training image when training the model, adjusts the model parameters by combining the iterative results of each training image to obtain the target model, and iteratively predicts for the target image when de-reflecting it. This allows for refinement of the prediction results based on the internal features of the target image, improving the targeting of image de-reflection and thus improving the accuracy of the prediction results.
[0057] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0058] Please refer to Figure 1 , Figure 1 This is a schematic diagram of the adaptive single-image reflection removal device structure for the hardware operating environment of the device involved in the embodiments of this application.
[0059] The terminal device in this application embodiment can be a computer or a personal computer or other terminal device with data processing capabilities.
[0060] like Figure 1 As shown, the adaptive single-image reflection removal device may include: a processor 1001, such as a CPU, a memory 1005, and a communication bus 1002. The communication bus 1002 is used to establish communication between the processor 1001 and the memory 1005. The memory 1005 may be a high-speed RAM or a stable, non-volatile memory, such as a disk drive. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0061] Optionally, the adaptive single-image reflection removal device may also include a user interface 1003, a network interface 1004, a camera, RF (Radio Frequency) circuitry, sensors, audio circuitry, a WiFi module, etc. The user interface may include a display screen and an input submodule such as a keyboard. Optionally, the user interface may also include a standard wired interface or a wireless interface. The network interface may optionally include a standard wired interface or a wireless interface (such as a WiFi interface).
[0062] Those skilled in the art will understand that Figure 1 The adaptive single-image reflection removal device structure shown does not constitute a limitation on the adaptive single-image reflection removal device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0063] like Figure 1 As shown, the memory 1005, which serves as a computer storage medium, may include an operating system, a network communication module, and an adaptive single-image reflection removal program. The operating system is a program that manages and controls the hardware and software resources of the adaptive single-image reflection removal device, supporting the operation of the adaptive single-image reflection removal program and other software and / or programs. The network communication module is used to enable communication between the various components within the memory 1005, as well as communication with other hardware and software in the adaptive single-image reflection removal device.
[0064] exist Figure 1 In the adaptive single-image reflection removal device shown, the processor 1001 is used to execute the adaptive single-image reflection removal program stored in the memory 1005 and perform the following operations:
[0065] The initial model is trained based on the image training set to obtain the pre-trained model.
[0066] Each training image in the image training set is input into the pre-trained model, and the pre-trained model is iteratively adjusted according to each first intermediate image output by the pre-trained model to obtain the target model;
[0067] The target image to be de-reflected is iteratively predicted based on the target model to obtain the de-reflection result of the target image.
[0068] Furthermore, the processor 1001 can call the adaptive single-image reflection removal program stored in the memory 1005 and also perform the following operations:
[0069] The image training set is input into a preset initial model to obtain the result set output by the initial model;
[0070] Calculate the total loss between the result set and the image training set, and adjust the weight parameters in the initial model according to the total loss to obtain the pre-trained model.
[0071] Furthermore, the processor 1001 can call the adaptive single-image reflection removal program stored in the memory 1005 and also perform the following operations:
[0072] Calculate the pixel loss, similarity loss, perceptual loss, adversarial loss, exclusion loss, and reconstruction loss between the result set and the image training set;
[0073] The pixel loss, similarity loss, perceptual loss, adversarial loss, exclusion loss, and reconstruction loss are weighted according to a preset weight set to obtain the total loss between the result set and the image training set.
[0074] Furthermore, the processor 1001 can call the adaptive single-image reflection removal program stored in the memory 1005 and also perform the following operations:
[0075] Based on each first intermediate image output by the pre-trained model, the pre-trained model corresponding to each training image is iteratively adjusted to obtain the iterative model corresponding to each training image.
[0076] Each training image is input into the corresponding iterative model, and the final training result output by each iterative model is obtained.
[0077] The pre-trained model is adjusted based on the final training results to obtain the target model.
[0078] Furthermore, the processor 1001 can call the adaptive single-image reflection removal program stored in the memory 1005 and also perform the following operations:
[0079] Obtain the first intermediate image output by each of the pre-trained models;
[0080] Calculate the reconstruction loss between each first intermediate image and the corresponding training image for each first intermediate image;
[0081] Each of the pre-trained models is adjusted according to the reconstruction loss to obtain the first intermediate model corresponding to each of the training images.
[0082] Each of the training images is input into the first intermediate model corresponding to each of the training images;
[0083] The output of each first intermediate model is used as the new first intermediate image. After using each first intermediate model as the new pre-trained model, the process returns to the step of calculating the reconstruction loss between each first intermediate image and the corresponding training image, until the number of iterations reaches a preset first number and then stops. The current first intermediate model is used as the iterative model corresponding to each training image.
[0084] Furthermore, the processor 1001 can call the adaptive single-image reflection removal program stored in the memory 1005 and also perform the following operations:
[0085] Obtain the transmission layer and reflection layer output by each of the pre-trained models;
[0086] The transmission layer and the reflection layer corresponding to each of the training images are reconstructed to obtain the first intermediate image corresponding to each of the training images.
[0087] Furthermore, the processor 1001 can call the adaptive single-image reflection removal program stored in the memory 1005 and also perform the following operations:
[0088] The target image to be de-reflected is input into the target model to obtain a second intermediate image;
[0089] Calculate the reconstruction loss between the second intermediate image and the target image, and adjust the target model according to the reconstruction loss to obtain the second intermediate model;
[0090] The first intermediate model is used as the new target model, and the step of inputting the target image to be de-reflected into the target model is returned to be executed until the number of iterations reaches the preset second number and then stops. The current output of the target model is the reflection removal result of the target image.
[0091] This application provides an adaptive single-image reflection removal method. In the first embodiment of this adaptive single-image reflection removal method, please refer to... Figure 2 The adaptive single-image reflection removal method includes:
[0092] Step S10: Train the preset initial model based on the image training set to obtain the pre-trained model.
[0093] It should be noted that in this embodiment, the user-preset initial model is constructed by improving upon the existing IBCLN (IterativeBoost Convolutional LSTM Network) network, and is referred to as PNACR (Personalized Single Image Reflection Removal Network through Adaptive Cascade Refinement). The structure of PNACR is as follows: Figure 3 As shown, similar to IBCLN, PNACR also constructs two subnetworks with identical structures f. t and f r f t Used to predict the transmission layer, f r Used to predict the reflection layer. The image of the reflection contamination, the transmission layer, and the reflection layer are concatenated as input to two sub-networks. The transmission layer and the reflection layer serve as auxiliary information to each other; the transmission layer can help predict the reflection layer, and vice versa.
[0094] In terms of model structure, PNACR makes the following two improvements compared to the basic network of IBCLN:
[0095] (1) To ensure that the model can extract more useful features from the input image, PNACR added 6 CBAM-Residual (attention residual) blocks to the original 11 Conv-relu (feature extraction) modules. The channel attention and spatial attention mechanisms in CBAM (attention) allow the model to learn feature representations with a specific focus;
[0096] (2) In recent years, residual learning techniques have demonstrated excellent performance in low-level image processing tasks. Therefore, PNACR replaces the four dilated convolutional layers in the original network with a single dilated SE-Residual (attention residual) module. This module contains six sequentially connected dilated SE-Residual blocks. This dilated SE-Residual block is a modification of the SE-Residual block, eliminating the normalization layer to save memory and preserve image features. Furthermore, to improve the stability of the network during training, 0.1 is empirically chosen as the scaling factor for scaling the residuals. Unlike these methods, we employ dilated convolutions in these SE-Residual blocks to obtain a larger receptive field. The dilation rates in these six dilated SE-Residual blocks are 2, 2, 4, 4, 4, and 4, respectively.
[0097] Regarding cascading strategies, IBCLN blindly cascades modules with equal weights to refine the results. PNACR, on the other hand, proposes an adaptive cascading strategy. This strategy adjusts the model's weights for the next iteration based on the model's output at the current iteration.
[0098] In this embodiment, the preset initial model is first trained in the first stage. The images in the training set are input into the initial model, and the transmitted image and the reflected image are used as outputs to train the initial model to obtain the pre-trained model.
[0099] Further, in a feasible embodiment, step S10 above, the step of training a preset initial model based on an image training set to obtain a trained pre-trained model, includes:
[0100] Step S101: Input the image training set into the preset initial model to obtain the result set output by the initial model;
[0101] Step S102: Calculate the total loss between the result set and the image training set, and adjust the weight parameters in the initial model according to the total loss to obtain the pre-trained model after training.
[0102] In this embodiment, when training the initial model, the images of the image training set are first output to the initial model to obtain the result image set output by the initial model. Then, the total loss between the result image and the training image is calculated. The weight parameters of the initial model are then adjusted according to the calculated total loss to obtain the pre-trained model.
[0103] Further, in a feasible embodiment, step S102 above, the step of calculating the total loss between the result set and the image training set, includes:
[0104] Step S1021: Calculate the pixel loss, similarity loss, perceptual loss, adversarial loss, exclusion loss, and reconstruction loss between the result set and the image training set;
[0105] Step S1022: The pixel loss, similarity loss, perceptual loss, adversarial loss, exclusion loss and reconstruction loss are weighted according to a preset weight set to obtain the total loss between the result set and the image training set.
[0106] It should be noted that the total loss includes pixel loss, similarity loss, perceptual loss, adversarial loss, exclusion loss, and reconstruction loss. Pixel loss measures the difference between two images at the pixel level. Similarity loss, also known as SSIM loss, measures the similarity between two images in terms of brightness, contrast, and structure. Perceptual loss compares the low-level and high-level semantic features of the image using the existing VGG19 network. Adversarial loss measures the realism of the predicted transmission image. Exclusion loss measures the degree of image segmentation. Reconstruction loss measures the difference between the reconstructed image based on the transmission and reflection images and the original image. Through self-supervised reconstruction loss, the initial model can learn the internal information of the input image within the test time, improving its relevance to the image.
[0107] In this embodiment, when calculating the total loss between the result set and the training image set, it is necessary to first calculate the pixel loss, similarity loss, perceptual loss, adversarial loss, exclusion loss and reconstruction loss between the result image and the training image, and then sum them according to the weights of each loss to obtain the total loss between the result image and the training image.
[0108] Specifically, during calculation, the pixel loss is defined as follows:
[0109]
[0110] Where T is the transmitted image, i.e., the image without reflection, I is the image containing reflection, and R is the reflected image. Based on experience, the parameters in the formula are defined as a = 0.3, b = 0.6, c = 1.0, d = 0.2. The SSIM loss is defined as follows:
[0111]
[0112] The definition of perceived loss is as follows:
[0113]
[0114] Where, φ l λ represents the output of layer l in the VGG19 network. lThe weights representing the results of layer l, and the definition of adversarial loss are as follows:
[0115]
[0116]
[0117] Where G represents the proposed network, D represents the discriminator, and the definition of exclusive loss is as follows:
[0118]
[0119]
[0120] Where n represents the predicted downsampling of the transmission and reflection layers by 2. n-1 Next. λ t and λ R represents the normalization factor, ° represents element-wise multiplication, and the reconstruction loss is defined as follows:
[0121]
[0122] The total loss of the model is defined as follows:
[0123]
[0124] Based on experience, the weight parameters are set to λ1 = 1.0, λ2 = 0.3, λ3 = 0.1, λ4 = 0.01 and λ5 = 1.0. Based on the above formula, the total loss of the model output image can be calculated.
[0125] Step S20: Input each training image in the image training set into the pre-trained model, and iteratively adjust the pre-trained model according to each first intermediate image output by the pre-trained model to obtain the target model;
[0126] In this embodiment, after the first stage of training, the terminal uses the training images in the graphics training set to perform the second stage of iterative training on the pre-trained model. This allows the pre-trained model to adapt to the internal features of each training image during the iteration process, and to adaptively adjust the weight parameters of the pre-trained model based on the internal features. Finally, by combining the training results of all training images, the weight parameters of the pre-trained model are adjusted to obtain the target model that has been trained.
[0127] Specifically, after the first phase of training, in order to explore the unique internal features of the images input to the model during model testing and to quickly adapt to new breakthroughs based on these internal features, the engineers performed meta-mutually assisted learning on PNACR and refined the image prediction results through an adaptive cascade structure.
[0128] Step S30: Iteratively predict the target image to be de-reflected based on the target model to obtain the de-reflection removal result of the target image.
[0129] In this embodiment, after obtaining the target model, the target image to be de-reflected is input into the target model, and the target image is iteratively predicted. During the iteration process, the model can adapt to the internal features of the target image to adjust the model parameters, thereby improving the targeting of the target image and thus improving the accuracy of the prediction results.
[0130] Specifically, after the technicians set up the PNACR model, the terminal inputs the training set into the model. The PNACR model is trained in the first stage based on the total loss between the output result set and the training set to obtain the pre-trained model. Then, the training images in the training set are input into the pre-trained model for iterative training. The total loss between the transmission image obtained after iterative training and the training image is calculated. The total loss of all training images in the training set is used to adjust the pre-trained model to obtain the target model. Finally, the target image to be de-reflected is input into the target model for iterative prediction. The final result is the projected image of the target image.
[0131] Thus, the method of iteratively predicting the target image to be de-reflected based on the above-mentioned iterative training model, compared with the traditional method of training and predicting all images using a fixed de-reflection model, the adaptive single-image de-reflection removal method of this application iteratively trains for each training image when training the model, adjusts the model parameters by combining the iterative results of each training image to obtain the target model, and iteratively predicts for the target image when de-reflecting it. This allows for refinement of the prediction results based on the internal features of the target image, improving the targeting of image de-reflection and thus improving the accuracy of the prediction results.
[0132] Furthermore, based on the first embodiment of the adaptive single-image reflection removal method of this application described above, a second embodiment of the adaptive single-image reflection removal method of this application is proposed.
[0133] In the second embodiment of the adaptive single-image reflection removal method of this application, the step S20 above, which iteratively adjusts the pre-trained model based on each first intermediate image output by the pre-trained model to obtain the target model, includes:
[0134] Step S201: Based on each first intermediate image output by the pre-trained model, iteratively adjust the pre-trained model corresponding to each training image to obtain the iterative model corresponding to each training image.
[0135] In this embodiment, after the training images of the training set are input into the pre-trained model, the pre-trained model is iteratively adjusted according to the first intermediate image output by the pre-trained model for each training image to obtain the iterative model corresponding to each training image.
[0136] Further, in a feasible embodiment, step S201 above, the step of iteratively adjusting the pre-trained model corresponding to each training image based on each first intermediate image output by the pre-trained model, includes:
[0137] Step S2011: Obtain the first intermediate image output by each of the pre-trained models;
[0138] Step S2012: Calculate the reconstruction loss between each first intermediate image and the corresponding training image;
[0139] Step S2013: Adjust each of the pre-trained models according to each of the reconstruction losses to obtain the first intermediate model corresponding to each of the training images;
[0140] Step S2014: Input each of the training images into the first intermediate model corresponding to each of the training images respectively;
[0141] Step S2015: The output of each first intermediate model is used as the new first intermediate image, and after using each first intermediate model as the new pre-trained model, the process returns to the step of calculating the reconstruction loss between each first intermediate image and the corresponding training image, until the number of iterations reaches a preset first number and then stops. The current first intermediate model is used as the iterative model corresponding to each training image.
[0142] In this embodiment, for each training image in the training set, after inputting the training image into the pre-trained model, the first intermediate image output by the pre-trained model is obtained. The first intermediate image includes a transmission layer and a reflection layer. The first intermediate image can be reconstructed based on the transmission layer and the reflection layer. Then, the reconstruction loss between the first intermediate image and the training image is calculated. After adjusting the parameters of the model based on the reconstruction loss, the training image is input into the model again. This process is repeated until the number of iterations set by the technician is reached. Finally, the parameters of the model are adjusted to obtain the iterative model for the training image.
[0143] Further, in a feasible embodiment, step S2011 above, the step of obtaining the first intermediate image output by each of the pre-trained models, includes:
[0144] Step S20111: Obtain the transmission layer and reflection layer output by each of the pre-trained models;
[0145] Step S20112: Reconstruct the transmission layer and reflection layer corresponding to each of the training images to obtain the first intermediate image corresponding to each of the training images.
[0146] In this embodiment, after the training image is pre-trained on the input value, the pre-trained model outputs a transmission layer and a reflection layer, and the first intermediate image can be reconstructed based on the transmission layer and the reflection layer.
[0147] Step S202: Input each of the training images into the iterative model corresponding to each of the training images, and obtain the final training result output by each of the iterative models.
[0148] Step S203: Adjust the pre-trained model according to the final training results to obtain the target model.
[0149] In this embodiment, after obtaining the iterative model corresponding to each training model, the training model is then input into its corresponding iterative model to obtain the final training result of each training model. Based on the total loss between each training result and each training image, the pre-trained model is adjusted by combining the total loss of all training images to obtain the target model.
[0150] Specifically, in the PNACR setup, removing reflections from an image is considered a task because each image contaminated with reflections has its own characteristics. The purpose of meta-assisted learning is to obtain a good set of initial parameters from a series of related tasks and then quickly adapt to a new task—a new image—using limited information. In the meta-assisted training phase, the transmission and reflection layers are first predicted using a pre-trained model. Then, the original image is reconstructed using the predicted transmission and reflection layers. The model is adjusted based on the calculated reconstruction loss. Next, the adjusted model refines the previous prediction results. After m iterations, the final prediction result is obtained. Then, the total loss for this training image can be calculated. After calculating the total loss for all images in the meta-training task, the model parameters can be globally updated. At this point, the meta-training phase is complete, a good set of initial parameters is obtained, and the final target model is derived.
[0151] Thus, the method of iteratively predicting the target image to be de-reflected based on the above-mentioned iterative training model in this application, compared with the traditional method of training and predicting all images using a fixed de-reflection model, the adaptive single-image de-reflection removal method of this application iteratively trains for each training image when training the model. During the training process, it can adapt to the internal features of the image and adjust the model parameters by combining the iterative results of each training image to obtain the target model, thereby improving the adaptive ability of the model.
[0152] Furthermore, based on the first and second embodiments of the adaptive single-image reflection removal method of this application described above, a third embodiment of the adaptive single-image reflection removal method of this application is proposed.
[0153] In the third embodiment of the adaptive single-image reflection removal method of this application, step S30, which iteratively predicts the target image to be reflected to be removed based on the target model to obtain the reflection removal result of the target image, includes:
[0154] Step S301: Input the target image to be de-reflected into the target model to obtain a second intermediate image;
[0155] Step S302: Calculate the reconstruction loss between the second intermediate image and the target image, and adjust the target model according to the reconstruction loss to obtain the second intermediate model;
[0156] Step S303: Use the first intermediate model as the new target model, and return to the step of inputting the target image to be de-reflected into the target model until the number of iterations reaches the preset second number, and then stop. Use the current output transmission layer of the target model as the reflection removal result of the target image.
[0157] In this embodiment, as Figure 4 As shown, after training the target model, reflection removal can be performed on the target image to be de-reflected. After inputting the target image into the target model, the predicted reflection layer and transmission layer can be obtained, and a second intermediate image can be reconstructed. Then, the reconstruction loss between the second intermediate image and the target image is calculated, and the parameters of the target model are adjusted according to the reconstruction loss to obtain the second intermediate model. The second intermediate model is then used as the target model. If the number of iterations i is less than m, the step of inputting the target image to be de-reflected into the target model is returned until the number of iterations is greater than or equal to m, that is, the prediction stops after the number of iterations preset by the user. At this time, the transmission layer output by the target model is the final prediction result of the target image.
[0158] Specifically, in the meta-mutually assisted testing phase, we use the model to predict new images. This allows us to explore the unique internal characteristics of the images, i.e., by adjusting the model using a self-supervised reconstruction loss. Then, we can refine the results using the updated model. After m iterations, we obtain the final prediction for the image.
[0159] The model described above for processing new images is a personalized model that uses adaptive cascaded refinement. We can generate such a model for each new image.
[0160] Thus, the method of iteratively predicting the target image to be de-reflected based on the above-mentioned iterative training model, compared with the traditional method of training and predicting all images using a fixed de-reflection model, the adaptive single-image de-reflection removal method of this application iteratively trains for each training image when training the model, adjusts the model parameters by combining the iterative results of each training image to obtain the target model, and iteratively predicts for the target image when de-reflecting it. This allows for refinement of the prediction results based on the internal features of the target image, improving the targeting of image de-reflection and thus improving the accuracy of the prediction results.
[0161] In addition, please refer to Figure 5 , Figure 5 This is a functional block diagram of the adaptive single-image reflection removal device of this application. This application also provides an adaptive single-image reflection removal device, which includes:
[0162] The pre-training module 10 is used to train a preset initial model based on an image training set to obtain a pre-trained model.
[0163] The iterative training module 20 is used to input each training image in the image training set into the pre-trained model, and to iteratively adjust the pre-trained model according to each first intermediate image output by the pre-trained model to obtain the target model.
[0164] The iterative prediction module 30 is used to perform iterative prediction on the target image to be de-reflected based on the target model, so as to obtain the de-reflection removal result of the target image.
[0165] Optionally, the pre-trained module includes:
[0166] The first training unit is used to input the image training set into a preset initial model to obtain the result set output by the initial model; calculate the total loss between the result set and the image training set, and adjust the weight parameters in the initial model according to the total loss to obtain the pre-trained model after training.
[0167] Optionally, the first training unit includes:
[0168] The total loss calculation subunit is used to calculate the pixel loss, similarity loss, perceptual loss, adversarial loss, exclusion loss, and reconstruction loss between the result set and the image training set; the pixel loss, similarity loss, perceptual loss, adversarial loss, exclusion loss, and reconstruction loss are weighted according to a preset weight set to obtain the total loss between the result set and the image training set.
[0169] Optionally, the iterative training module includes:
[0170] The model adjustment unit is configured to iteratively adjust the pre-trained model corresponding to each training image based on each first intermediate image output by the pre-trained model, so as to obtain an iterative model corresponding to each training image; input each training image into the iterative model corresponding to each training image, and obtain the final training result output by each iterative model; and adjust the pre-trained model according to the final training results to obtain the target model.
[0171] Optionally, the model adjustment unit includes:
[0172] An iterative adjustment subunit is used to obtain the first intermediate image output by each of the pre-trained models; calculate the reconstruction loss between each first intermediate image and the corresponding training image; adjust each pre-trained model according to the reconstruction loss to obtain the first intermediate model corresponding to each training image; input each training image into the first intermediate model corresponding to each training image; use the output of each first intermediate model as the new first intermediate image, and after using each first intermediate model as the new pre-trained model, return to the step of calculating the reconstruction loss between each first intermediate image and the corresponding training image, until the number of iterations reaches a preset first number, and stop, and use the current first intermediate model as the iterative model corresponding to each training image.
[0173] Optionally, the iterative adjustment sub-unit includes:
[0174] The image reconstruction subunit is used to obtain the transmission layer and reflection layer output by each of the pre-trained models; and to reconstruct the transmission layer and reflection layer corresponding to each of the training images to obtain the first intermediate image corresponding to each of the training images.
[0175] Optionally, the iterative prediction module includes:
[0176] An iterative prediction unit is used to input the target image to be de-reflected into the target model to obtain a second intermediate image; calculate the reconstruction loss between the second intermediate image and the target image, and adjust the target model according to the reconstruction loss to obtain a second intermediate model; use the first intermediate model as a new target model, and return to execute the step of inputting the target image to be de-reflected into the target model until the number of iterations reaches a preset second number, and then stop, and use the current output transmission layer of the target model as the final prediction result of image de-reflection.
[0177] The specific implementation of the adaptive single-image reflection removal device of this application is basically the same as the embodiments of the adaptive single-image reflection removal method described above, and will not be repeated here.
[0178] Furthermore, this application also proposes a storage medium storing an adaptive single-image reflection removal program, which, when executed by a processor, implements the steps of the adaptive single-image reflection removal method of this application as described above.
[0179] The specific embodiments of the computer storage medium in this application are basically the same as the embodiments of the adaptive single-image reflection removal method described above, and will not be repeated here.
[0180] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0181] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0182] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0183] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. An adaptive single-image reflection removal method, characterized in that, The adaptive single-image reflection removal method includes: The initial model is trained based on the image training set to obtain the pre-trained model. Each training image in the image training set is input into the pre-trained model to obtain the first intermediate image output by each pre-trained model. The first intermediate image corresponding to each training image is obtained by reconstructing the transmission layer and reflection layer corresponding to each training image. Calculate the reconstruction loss between each first intermediate image and the corresponding training image for each first intermediate image; Each of the pre-trained models is adjusted according to the reconstruction loss to obtain the first intermediate model corresponding to each of the training images. Each of the training images is input into the first intermediate model corresponding to each of the training images; The output of each first intermediate model is used as the new first intermediate image, and after using each first intermediate model as the new pre-trained model, the step of calculating the reconstruction loss between each first intermediate image and the corresponding training image is returned until the number of iterations reaches a preset first number and then stops. The current first intermediate model is used as the iterative model corresponding to each training image. Each training image is input into the corresponding iterative model, and the final training result output by each iterative model is obtained. The pre-trained model is adjusted based on the final training results to obtain the target model; The target image to be de-reflected is iteratively predicted according to the target model to obtain the de-reflection result of the target image; The step of training a preset initial model based on an image training set to obtain a pre-trained model includes: The image training set is input into a preset initial model to obtain the result set output by the initial model; Calculate the total loss between the result set and the image training set, and adjust the weight parameters in the initial model according to the total loss to obtain the pre-trained model after training. The total loss includes pixel loss, similarity loss, perceptual loss, adversarial loss, exclusion loss and reconstruction loss. The initial model is based on the structure of PNACR, which is an improvement on the existing IBCLN network. PNACR has two identical sub-networks, ft and fr, where ft is used to predict the transmission layer and fr is used to predict the reflection layer. PNACR adds 6 CBAM-Residual blocks to the 11 Conv-ReLU modules. PNACR replaces the 4 dilated convolutional layers in the original network with 1 dilated SE-Residual module, which contains 6 dilated SE-Residual blocks connected in sequence.
2. The adaptive single-image reflection removal method according to claim 1, characterized in that, The step of calculating the total loss between the result set and the image training set includes: Calculate the pixel loss, similarity loss, perceptual loss, adversarial loss, exclusion loss, and reconstruction loss between the result set and the image training set; The pixel loss, similarity loss, perceptual loss, adversarial loss, exclusion loss, and reconstruction loss are weighted according to a preset weight set to obtain the total loss between the result set and the image training set.
3. The adaptive single-image reflection removal method according to claim 1, characterized in that, The step of obtaining the first intermediate image output by each of the pre-trained models includes: Obtain the transmission layer and reflection layer output by each of the pre-trained models; The transmission layer and the reflection layer corresponding to each of the training images are reconstructed to obtain the first intermediate image corresponding to each of the training images.
4. The adaptive single-image reflection removal method according to claim 1, characterized in that, The step of iteratively predicting the target image to be de-reflected based on the target model to obtain the de-reflection removal result of the target image includes: The target image to be de-reflected is input into the target model to obtain a second intermediate image; Calculate the reconstruction loss between the second intermediate image and the target image, and adjust the target model according to the reconstruction loss to obtain the second intermediate model; The second intermediate model is used as the new target model, and the step of inputting the target image to be de-reflected into the target model is returned to be executed until the number of iterations reaches the preset second number and then stops. The current output of the target model is the reflection removal result of the target image.
5. An adaptive single-image reflection removal device, characterized in that, The adaptive single-image reflection removal device includes: The pre-training module is used to train a pre-set initial model based on an image training set to obtain a pre-trained model. An iterative training module is used to input each training image in the image training set into the pre-trained model, obtain a first intermediate image output by each pre-trained model, wherein the first intermediate image corresponding to each training image is obtained by reconstructing the transmission layer and reflection layer corresponding to each training image; calculate the reconstruction loss between each first intermediate image and the corresponding training image; adjust each pre-trained model according to the reconstruction loss to obtain a first intermediate model corresponding to each training image; input each training image into the first intermediate model corresponding to each training image; and input each first intermediate image into the first intermediate model corresponding to each training image. The results output by the intermediate model are used as new first intermediate images. After using each first intermediate model as a new pre-trained model, the process returns to the step of calculating the reconstruction loss between each first intermediate image and its corresponding training image, continuing until the number of iterations reaches a preset first number. The current first intermediate models are then used as the iterative models corresponding to each training image. Each training image is input into its corresponding iterative model, and the final training result output by each iterative model is obtained. The pre-trained models are adjusted based on the final training results to obtain the target model. The iterative prediction module is used to iteratively predict the target image to be de-reflected based on the target model, so as to obtain the de-reflection result of the target image; The pre-training module includes: The first training unit is used to input the image training set into a preset initial model to obtain the result set output by the initial model; calculate the total loss between the result set and the image training set, and adjust the weight parameters in the initial model according to the total loss to obtain the pre-trained model after training. The total loss includes pixel loss, similarity loss, perceptual loss, adversarial loss, exclusion loss and reconstruction loss. The initial model is based on the structure of PNACR, which is an improvement on the existing IBCLN network. PNACR includes two identical subnetworks, ft and fr, where ft is used to predict the transmission layer and fr is used to predict the reflection layer. PNACR adds 6 CBAM-Residual blocks to the 11 Conv-ReLU modules. PNACR replaces the 4 dilated convolutional layers in the original network with 1 dilated SE-Residual module, which contains 6 sequentially connected dilated SE-Residual blocks.
6. An adaptive single-image reflection removal device, characterized in that, The adaptive single-image reflection removal device includes a memory and a processor, wherein the memory stores an adaptive single-image reflection removal program, and when the adaptive single-image reflection removal program is executed by the processor, it implements the steps of the adaptive single-image reflection removal method as described in any one of claims 1 to 4.
7. A storage medium, characterized in that, The storage medium stores an adaptive single-image reflection removal program, which, when executed by a processor, implements the steps of the adaptive single-image reflection removal method as described in any one of claims 1 to 4.