Multi-image anti-reflection method and system based on space-time information
Through a multi-image dereflection method based on spatiotemporal information, the para-LAP algorithm is used to align the transmission images and sparsely constrain the reflection layer gradient, which solves the problem of poor dereflection effect in the existing technology and achieves efficient image dereflection effect.
Patent Information
- Application Number
- CN202311434745.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-31
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2043-10-31
AI Technical Summary
Existing technologies are not effective in removing image reflections. In particular, single-image dereflection technology is ill-posed and relies on insufficient prior information. Multi-image dereflection technology is not applicable to moving objects, and deep learning methods rely on high-performance equipment and training data.
A multi-image de-reflection method based on spatiotemporal information is adopted, and the motion information and brightness relationship between images are calculated using the para-LAP algorithm. The spatial smoothness of the image is achieved by iteratively aligning the transmission image and sparsely constraining the reflection layer gradient.
It improves the de-reflection effect, reduces the dependence on high-performance equipment and training data, and realizes cost-effective de-reflection of multiple images.
Smart Images

Figure CN117218040B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a method and system for removing reflections from multiple images based on spatiotemporal information. Background Art
[0002] With the continuous advancement of image processing technology, people are demanding higher quality images captured by imaging devices such as mobile phones and tablets. However, if reflective objects such as glass are present in the scene, the reflected image will be superimposed on the image captured by the imaging device, reducing the clarity of the transmitted image and resulting in poor image quality. Therefore, removing reflections from images is one of the keys to improving image quality.
[0003] Existing image dereflection techniques are generally categorized as single-image dereflection techniques and multi-image dereflection techniques. Single-image dereflection techniques are based on the mathematical model I = T + R + N, where I refers to the image containing reflections, T refers to the transmission image (i.e., the image without reflections), R refers to the reflection image, and N refers to noise. The task of image dereflection is to recover the reflection-free image T from the input image I. For single-image dereflection, T, R, and N are unknown, and only I is known. The number of unknown parameters exceeds the number of known parameters, making single-image dereflection an ill-posed problem. Existing single-image dereflection techniques utilize prior information on I, T, and R and impose constraints on these parameters to address this ill-posed problem. However, this prior information cannot fully describe the inherent characteristics of the image, so existing single-image dereflection techniques cannot effectively address the dereflection problem and are generally ineffective.
[0004] Existing multi-image de-reflection technology requires capturing two or more images at different angles and then utilizing motion information between the images or device information to remove reflections. Compared to single-image de-reflection technology, multi-image de-reflection utilizes more information and generally achieves better de-reflection results. Currently, commonly used multi-image de-reflection algorithms utilize motion information between a series of images to decompose the image into projected and reflective layers. However, this method is generally not suitable for videos or image sequences containing moving objects.
[0005] In recent years, with the rapid development of deep learning, a growing number of deep learning-based image de-reflection techniques have been proposed, achieving excellent de-reflection results. However, these techniques require a large amount of additional training data and time, and the training process relies on high-performance computing equipment, limiting their practical application in real life and production. Summary of the Invention
[0006] To address the deficiencies of the above-mentioned prior art, the present invention provides a method and system for dereflection of multiple images based on spatiotemporal information. Based on the motion information between the two images and the spatial information of a single image, an accurate and fast image registration method, the parametric local all-pass (para-LAP) method, is adopted to align the transmission layers of the two images and perform sparse constraints on the gradient of the reflection layer between the two images to ensure the spatial smoothness of the dereflected image, thereby completing the dereflection task of multiple images and obtaining multiple dereflected images with better effects.
[0007] In a first aspect, the present invention provides a method for de-reflecting multiple images based on spatiotemporal information.
[0008] A method for removing reflections from multiple images based on spatiotemporal information, comprising:
[0009] Acquire an image sequence including a plurality of images, extract the preceding and following images from the image sequence, use the two extracted images as input images, and initialize a transmission image of the input image;
[0010] Based on the transmission image initialized by the two input images, the para-LAP algorithm is used to calculate the deformation between the two transmission images, obtain the motion information and brightness relationship between the two transmission images, and then align the two transmission images;
[0011] Fix the previous transmission image, calculate the optimal solution of the subsequent transmission image according to the optimization equation, and update the subsequent transmission image; based on the updated subsequent transmission image, calculate the optimal solution of the previous transmission image according to the optimization equation, and update the previous transmission image;
[0012] Based on the two updated transmission images, the alignment and update operations are continuously and iteratively performed until the maximum number of loop iterations is reached, and the final updated transmission image is output, that is, the final multiple de-reflected images are output.
[0013] In a second aspect, the present invention provides a multiple image de-reflection system based on spatiotemporal information.
[0014] A multi-image de-reflection system based on spatiotemporal information, comprising:
[0015] An image acquisition and preprocessing module, configured to acquire an image sequence comprising a plurality of images, extract the preceding and following images from the image sequence, use the two extracted images as input images, and initialize a transmission image of the input image;
[0016] The image alignment module is used to calculate the deformation between the two transmission images based on the transmission images initialized by the two input images using the para-LAP algorithm, obtain the motion information and brightness relationship between the two transmission images, and then align the two transmission images;
[0017] An image processing module is used to fix the previous transmission image, calculate the optimal solution of the subsequent transmission image according to the optimization equation, and update the subsequent transmission image; based on the updated subsequent transmission image, calculate the optimal solution of the previous transmission image according to the optimization equation, and update the previous transmission image;
[0018] The dereflection image output module is used to continuously and iteratively perform alignment and update operations based on the two updated transmission images until the maximum number of loop iterations is reached, and output the final updated transmission image, that is, output the final multiple dereflection images.
[0019] In a third aspect, the present disclosure further provides an electronic device comprising a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the steps of the method described in the first aspect are completed.
[0020] In a fourth aspect, the present disclosure further provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, complete the steps of the method described in the first aspect.
[0021] One or more of the above technical solutions have the following beneficial effects:
[0022] 1. The present invention provides a method and system for de-reflection of multiple images based on spatiotemporal information. According to the motion information between the two images and the spatial information of a single image, an accurate and fast image registration method - a parameterized all-pass filtering method is adopted to align the transmission layers of the two images and perform sparse constraints on the gradient of the reflection layer between the two images to ensure the spatial smoothness of the de-reflected image. Through continuous iterative solution, the de-reflection task of multiple images is completed, and multiple de-reflected images with better effects are obtained. Compared with the traditional single-image deblurring algorithm, the present invention utilizes more information in multiple images, such as temporal and spatial information, and has a better de-reflection effect.
[0023] 2. The present invention combines the para-LAP algorithm to accurately estimate the motion information between images even if there is a brightness difference between the two input images. Moreover, this method does not rely on training data and high-performance computing equipment, does not require additional training time, and is cost-effective. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0025] Figure 1 This is an overall flow chart of the method for de-reflecting multiple images based on spatiotemporal information according to an embodiment of the present invention. DETAILED DESCRIPTION
[0026] It should be noted that the following detailed descriptions are exemplary only and are intended to describe specific embodiments and provide further explanation of the present invention, and are not intended to limit the exemplary embodiments according to the present invention. Unless otherwise indicated, all technical and scientific terms used herein have the same meanings as those commonly understood by those of ordinary skill in the art to which the present invention belongs. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0027] Example 1
[0028] In view of the problems in existing image dereflection methods, such as poor dereflection effect of single-image dereflection methods, unsuitability of multi-image dereflection methods for videos and image sequences including moving objects, and many restrictions and strong dependencies of dereflection methods based on deep learning, this embodiment provides a multi-image dereflection method based on spatiotemporal information, such as Figure 1 As shown, the following steps are included:
[0029] Step S1, obtaining an image sequence including a plurality of images, extracting the preceding and following images in the image sequence, using the two extracted images as input images, and initializing a transmission image of the input image;
[0030] Step S2: Based on the transmission images initialized from the two input images, the para-LAP algorithm is used to calculate the deformation between the two transmission images, obtain the motion information and brightness relationship between the two transmission images, and then align the two transmission images;
[0031] Step S3, fixing the front transmission image, calculating the optimal solution of the rear transmission image according to the optimization equation, and updating the rear transmission image; based on the updated rear transmission image, calculating the optimal solution of the front transmission image according to the optimization equation, and updating the front transmission image;
[0032] Step S4: Based on the two updated transmission images, the alignment and update operations are continuously and iteratively performed until the maximum number of iterations is reached, and the final updated transmission image is output, that is, the final multiple de-reflected images are output.
[0033] The following content introduces the de-reflection method of multiple images based on spatiotemporal information proposed in this embodiment in more detail.
[0034] In step S1, an image sequence consisting of multiple images is acquired. The preceding and following images are extracted from the sequence. These two extracted images are used as input images to initialize the transmission image of the input image. For a single image, this can be represented by the mathematical model: I = T + R + N, where I is the image containing reflections, i.e., the input image; T is the transmission image, i.e., the image after de-reflection; R is the reflection image; and N is random noise. The purpose of this embodiment is to extract the de-reflection image T from the input image I.
[0035] In this embodiment, an image sequence including multiple images is obtained, and the images in the image sequence all include reflections. The two images before and after the image sequence are extracted and recorded as input images I1 and I2. The transmission images of input images I1 and I2 are initialized and recorded as transmission images T1 and T2 respectively. Initialization here means directly assigning the input image to the transmission image, that is, input image I1 is used as the initialized transmission image T1, and input image I2 is used as the initialized transmission image T2. At the same time, the parameter λ is initialized. s and λ t , set the maximum number of iterations β max .
[0036] In step S2, based on the transmission image initialized from the two input images, the para-LAP algorithm is used to calculate the deformation between the two transmission images, obtain the motion information and brightness relationship between the two transmission images, and then align the two transmission images.
[0037] Specifically, for the initialized transmission images T1 and T2, the para-LAP algorithm is used to calculate the deformation between the two transmission images to obtain the motion information between the two transmission images (u x (x,y),u y (x, y)) and brightness relationship F, where motion information refers to the displacement information between each corresponding pixel in the two transmission images, u x (x,y) represents the horizontal coordinate displacement of the pixel point (x,y), u y (x,y) represents the vertical coordinate displacement of the pixel point (x,y).
[0038] The para-LAP algorithm is an image registration algorithm that uses image registration to determine the deformation between two images. The para-LAP algorithm works as follows: when there is no grayscale difference between the reference image and the floating image, the local transformation in the smooth spatial transformation can be approximated by a translation transformation, which is equivalent to performing an all-pass filter on the image block. This estimation then generates a local all-pass filter, from which the local deformation is extracted. This filter is then used to obtain an elastic deformation field through a sliding window. Finally, this elastic deformation field is fitted using a linear equation with a small number of parameters to obtain a smooth deformation field.
[0039] Furthermore, based on the calculated motion information, the two transmission images are aligned, aligning transmission image T2 with transmission image T1. This accurate and rapid image registration method aligns the projection layers of the two images. Furthermore, the para-LAP algorithm is used for calculation, taking image brightness into account, ensuring accurate alignment even when the brightness of the two images differs.
[0040] In step S3, the previous transmission image is fixed, the optimal solution of the subsequent transmission image is calculated according to the optimization equation, and the subsequent transmission image is updated; based on the updated subsequent transmission image, the optimal solution of the previous transmission image is calculated according to the optimization equation, and the previous transmission image is updated.
[0041] Specifically, one of the two images is fixed. In this embodiment, the first transmission image T1 is fixed. That is, T1 is used as a known quantity. The optimal solution of the second transmission image T2 is calculated according to the optimization equation. The optimization equation is:
[0042] E=E d +λ s E s +λ t E t (1)
[0043]
[0044]
[0045]
[0046] Among them, T represents the transmission image, T1 represents the front transmission image, T2 represents the back transmission image, and E d is the data fidelity term, E s is the spatial term, E t is the time term, represents the first-order derivative, L represents the second-order derivative, F represents the brightness relationship between the two transmission images T1 and T2, (u x (x,y),u y(x,y)) represents the deformation between images, and E represents the total energy.
[0047] Taking T1 as a known quantity and combining the input images I1 and I2, the first-order derivatives and second-order derivatives of the input images I1 and I2 and the previous transmission image T1 are calculated (i.e., the first-order derivatives and second-order derivatives of each pixel in the image are calculated) to minimize the total energy E in the optimization equation. In this way, the optimal solution of the subsequent transmission image is obtained, and this optimal solution is used as the new subsequent transmission image T2. That is, the subsequent transmission image T2 is updated based on this optimal solution.
[0048] Next, based on the updated subsequent transmission image T2, fix image T2 and repeat the above steps to recalculate and solve the optimal solution of the preceding transmission image T1, and update the subsequent transmission image T1 based on the optimal solution.
[0049] During this iteration, an optimization equation is used to solve and calculate the transmission image initialized from the two input images. This optimization equation not only considers the motion information between the two images, but also the spatial information in a single image. At the same time, a sparse constraint is imposed on the gradient of the reflection layer between the two images, ensuring the spatial smoothness of the de-reflected image, thereby obtaining a more effective de-reflected image.
[0050] On this basis, step S4 is finally executed, that is, based on the two updated transmission images, the alignment and update operations are continuously and iteratively performed until the maximum number of loop iterations is reached, and the final updated transmission image is output, that is, the final multiple de-reflected images are output.
[0051] Specifically, in the current iteration process, the two updated transmission images T1 and T2 are obtained through the above calculation, and it is determined whether the current iteration number β is greater than or equal to the preset maximum iteration number β max If yes, stop the iteration and output two transmission images T1 and T2, that is, output the final de-reflection images T1 and T2; otherwise, if no, update β to β+1, repeat the above steps S2 to S3, and iterate until the number of iterations is greater than the preset maximum number of iterations β max .
[0052] The multi-image dereflection method based on spatiotemporal information proposed in this embodiment not only utilizes the spatial information of each image, but also the motion information between the two images. Compared with the single-image deblurring algorithm, it utilizes more information from multiple images and has a better dereflection effect. Combined with the para-LAP algorithm, the motion information between the two input images can be accurately estimated even if there is a brightness difference between the two input images. The method proposed in this embodiment does not rely on training data and high-performance computing equipment, does not require additional training time, and is cost-effective.
[0053] Example 2
[0054] This embodiment provides a multi-image de-reflection system based on spatiotemporal information, including:
[0055] An image acquisition and preprocessing module, configured to acquire an image sequence comprising a plurality of images, extract the preceding and following images from the image sequence, use the two extracted images as input images, and initialize a transmission image of the input image;
[0056] The image alignment module is used to calculate the deformation between the two transmission images based on the transmission images initialized by the two input images using the para-LAP algorithm, obtain the motion information and brightness relationship between the two transmission images, and then align the two transmission images;
[0057] An image processing module is used to fix the previous transmission image, calculate the optimal solution of the subsequent transmission image according to the optimization equation, and update the subsequent transmission image; based on the updated subsequent transmission image, calculate the previous transmission image according to the optimization equation and update the previous transmission image;
[0058] The dereflection image output module is used to continuously and iteratively perform alignment and update operations based on the two updated transmission images until the maximum number of loop iterations is reached, and output the final updated transmission image, that is, output the final multiple dereflection images.
[0059] Example 3
[0060] This embodiment provides an electronic device, including a memory and a processor, and computer instructions stored in the memory and executed on the processor. When the computer instructions are executed by the processor, the steps of the above-mentioned method for de-reflecting multiple images based on spatiotemporal information are completed.
[0061] Example 4
[0062] This embodiment further provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the steps of the above-mentioned method for de-reflecting multiple images based on spatiotemporal information are completed.
[0063] The steps involved in the above embodiments 2 to 4 correspond to those in the method embodiment 1. For detailed implementation, please refer to the relevant description of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media that includes one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and cause the processor to perform any method of the present invention.
[0064] Those skilled in the art will appreciate that the modules or steps of the present invention described above can be implemented using a general-purpose computer device. Alternatively, they can be implemented using program code executable by a computing device, which can then be stored in a storage device and executed by the computing device. Alternatively, they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0065] The above description is only a preferred embodiment of the present invention. Although the specific implementation of the present invention is described in conjunction with the accompanying drawings, it does not limit the scope of protection of the present invention. Those skilled in the art should understand that on the basis of the technical solution of the present invention, various modifications or variations that can be made by those skilled in the art without creative work are still within the scope of protection of the present invention.
Claims
1. A method for removing reflections from multiple images based on spatiotemporal information, characterized in that: include: Acquire an image sequence including a plurality of images, extract the preceding and following images from the image sequence, use the two extracted images as input images, and initialize a transmission image of the input image; Based on the transmission image initialized by the two input images, the para-LAP algorithm is used to calculate the deformation between the two transmission images, obtain the motion information and brightness relationship between the two transmission images, and then align the two transmission images; Fix the previous transmission image, calculate the optimal solution of the subsequent transmission image according to the optimization equation, and update the subsequent transmission image; based on the updated subsequent transmission image, calculate the optimal solution of the previous transmission image according to the optimization equation, and update the previous transmission image; Based on the two updated transmission images, the alignment and update operations are continuously and iteratively performed until the maximum number of loop iterations is reached, and the final updated transmission image is output, that is, the final multiple de-reflected images are output.
2. The method for removing reflections from multiple images based on spatiotemporal information according to claim 1, wherein: Initializing the transmission image of the input image means directly assigning the input image to the transmission image, that is, using the input image I1 as the initialized transmission image T1 and using the input image I2 as the initialized transmission image T2.
3. The method for removing reflections from multiple images based on spatiotemporal information according to claim 1, wherein: The optimization equation is: E=E d +λ s AND s +λ t AND t Where T represents the transmission image, T1 represents the previous transmission image, T2 represents the next transmission image, I1 represents the previous input image, I2 represents the next input image, and E d is the data fidelity term, E s is the spatial term, E t is the time term, represents the first-order derivative, L represents the second-order derivative, F represents the brightness relationship between the two transmission images T1 and T2, (u x (x,y),u y (x,y)) represents the deformation between the two transmission images, and E represents the total energy.
4. The method for removing reflections from multiple images based on spatiotemporal information according to claim 3, wherein: Fix the previous transmission image, calculate the optimal solution of the subsequent transmission image according to the optimization equation, and update the subsequent transmission image, including: Taking the previous transmission image T1 as a known quantity and combining it with the two input images I1 and I2, the first-order derivatives and second-order derivatives of the two input images I1 and I2 and the previous transmission image T1 are calculated, and the total energy E in the optimization equation is minimized. In this way, the optimal solution of the subsequent transmission image is obtained, and this optimal solution is used as the new subsequent transmission image T2.
5. The method for removing reflections from multiple images based on spatiotemporal information according to claim 4, wherein: Calculating the first-order derivative and second-order derivative of an image refers to calculating the first-order derivative and second-order derivative of each pixel in the image.
6. A multi-image de-reflection system based on spatiotemporal information, characterized in that: include: An image acquisition and preprocessing module, configured to acquire an image sequence comprising a plurality of images, extract the preceding and following images from the image sequence, use the two extracted images as input images, and initialize a transmission image of the input image; The image alignment module is used to calculate the deformation between the two transmission images based on the transmission images initialized by the two input images using the para-LAP algorithm, obtain the motion information and brightness relationship between the two transmission images, and then align the two transmission images; An image processing module is used to fix the previous transmission image, calculate the optimal solution of the subsequent transmission image according to the optimization equation, and update the subsequent transmission image; based on the updated subsequent transmission image, calculate the optimal solution of the previous transmission image according to the optimization equation, and update the previous transmission image; The dereflection image output module is used to continuously and iteratively perform alignment and update operations based on the two updated transmission images until the maximum number of loop iterations is reached, and output the final updated transmission image, that is, output the final multiple dereflection images.
7. The multiple image de-reflection system based on spatiotemporal information according to claim 6, wherein: Initializing the transmission image of the input image means directly assigning the input image to the transmission image, that is, using the input image I1 as the initialized transmission image T1 and using the input image I2 as the initialized transmission image T2.
8. The multiple image de-reflection system based on spatiotemporal information according to claim 6, wherein: The optimization equation is: E=E d +λ s AND s +λ t AND t Where T represents the transmission image, T1 represents the previous transmission image, T2 represents the next transmission image, I1 represents the previous input image, I2 represents the next input image, and E d is the data fidelity term, E s is the spatial term, E t is the time term, represents the first-order derivative, L represents the second-order derivative, F represents the brightness relationship between the two transmission images T1 and T2, (u x (x,y),u y (x,y)) represents the deformation between the two transmission images, and E represents the total energy.
9. An electronic device, characterized in that: The invention comprises a memory and a processor, and computer instructions stored in the memory and executed on the processor. When the computer instructions are executed by the processor, the steps of the method for de-reflecting multiple images based on spatiotemporal information are completed as described in any one of claims 1 to 5.
10. A computer-readable storage medium, characterized in that: Used to store computer instructions, which, when executed by a processor, complete the steps of the method for de-reflecting multiple images based on spatiotemporal information according to any one of claims 1 to 5.
Citation Information
Patent Citations
Video light reflection removing method based on time and space
CN115424173A
Method, device and equipment for removing light reflection of single image and storage medium
CN116612027A