Image glasses reflection elimination method and device based on multistage refining network
By building an adaptive reflective perception network and a multi-level refinement network, the glasses reflection in the image is eliminated, and the poor visual quality caused by the constrained data set is solved, and higher quality image recovery is achieved.
Patent Information
- Application Number
- CN202510123729.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-06-20
AI Technical Summary
In the prior art, the data set of the reflection elimination method of image glasses is limited, resulting in poor image visual quality and hindering the development of research and application in related fields.
Using a multi-level refinement network method, by collecting the initial glasses reflective images and glasses-free reflective images, an adaptive reflective perception network and a multi-level refinement reflective elimination network are built, and the network is trained using the target loss function to simulate a wider range of reflective scenes, and through the coarse and refinement network, eliminating reflections to restore details around the human eye.
It improves the image visual quality, solves the problem of data set limitations, and promotes the research and application development in related fields.
Smart Images

Figure CN120182140A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of face beautification, and particularly relates to an image glasses reflection elimination method and device based on a multi-level refinement network. Background Art
[0002] Reflections are very common in our daily life. Especially when taking pictures through transparent media such as glass and resin, the reflections often obscure the subject behind the transparent medium, thus reducing the quality and visibility of the image. This is a common optical phenomenon. Eliminating these reflections to enhance the visibility of the main object has become an important task in the field of computer vision. Glasses reflections can be regarded as a special subtask of this task, and their elimination is of great significance for subsequent computer vision tasks such as face recognition and face key point detection.
[0003] In related technologies, the generation network and discriminant network in the adversarial neural network can be alternately trained using the obtained training sample set to obtain a trained adversarial neural network, and then the glasses reflection area in the image to be processed can be eliminated through the trained adversarial neural network to obtain an image without glasses reflections; alternatively, the image to be processed can be mosaicked and then input into a target detection model for target detection, and the image to be processed can be cropped according to the target detection result to obtain a high-definition image; the high-definition image is input into an anti-generation network model to eliminate shadows and reflections to obtain a restored image.
[0004] However, in related technologies, the datasets used have limitations in terms of reflection types and scenarios, mainly concentrated in indoor environments, resulting in poor visual quality of the obtained images, hindering the research and application development in related fields, and there is an urgent need for improvement. Summary of the Invention
[0005] This application provides an image glasses reflection elimination method and device based on a multi-level refinement network to solve problems in related technologies such as limited datasets, resulting in poor visual quality of the obtained images, and hindering the research and application development in related fields.
[0006] The first aspect embodiment of the present application provides an image glasses specular reflection elimination method based on a multi-level refinement network, which is applied to the model training stage. Among them, the method includes the following steps: collecting at least one initial glasses specular reflection image to be eliminated and the corresponding initial glasses-free specular reflection image of the at least one initial glasses specular reflection image to be eliminated; performing data processing on the at least one initial glasses specular reflection image and the initial glasses-free specular reflection image to respectively obtain the glasses specular reflection image to be eliminated and the glasses-free specular reflection image after data processing of the at least one initial glasses specular reflection image data processing and the initial glasses-free specular reflection image data processing; based on the glasses specular reflection image to be eliminated, the glasses-free specular reflection image and the pre-trained ResNet-34 network, constructing an adaptive specular reflection perception network suitable for the glasses specular reflection image to be eliminated; based on the specular reflection perception result output by the adaptive specular reflection perception network, a coarse-level network with different scale convolutional kernels and a refinement network, constructing a multi-level refinement specular reflection elimination network suitable for the glasses specular reflection image to be eliminated; based on the glasses specular reflection image to be eliminated and the glasses-free specular reflection image, training the adaptive specular reflection perception network and the multi-level refinement specular reflection elimination network by using a target loss function to obtain a trained adaptive specular reflection perception network and a multi-level refinement specular reflection elimination network for eliminating the glasses specular reflection in the glasses specular reflection image to be eliminated.
[0007] Optionally, in an embodiment of the present application, the constructing a multi-level refinement specular reflection elimination network suitable for the glasses specular reflection image to be eliminated based on the specular reflection perception result output by the adaptive specular reflection perception network, a coarse-level network with different scale convolutional kernels and a refinement network includes: based on the specular reflection perception result, selecting convolutional kernels of different scales to determine the coarse-level network by using the convolutional kernels; constructing the refinement network based on the coarse-level result output by the coarse-level network and the specular reflection perception result; constructing the multi-level refinement specular reflection elimination network based on the refinement network, the coarse-level result and the specular reflection perception result.
[0008] Optionally, in an embodiment of the present application, the based on the specular reflection perception result, selecting convolutional kernels of different scales to determine the coarse-level network by using the convolutional kernels includes: based on the specular reflection perception result and the convolutional kernels of different scales, obtaining at least one initial depth convolutional feature corresponding to the glasses specular reflection image to be eliminated; passing the at least one initial depth convolutional feature through a 1×1 convolutional layer to obtain a depth convolutional feature map corresponding to the at least one initial depth convolutional feature; passing the depth convolutional feature map through at least one convolutional layer to obtain a spatial attention map corresponding to the depth convolutional feature map; based on the spatial attention map and a target mask, obtaining an attention feature corresponding to the spatial attention map; determining the coarse-level network based on the attention feature.
[0009] Optionally, in an embodiment of the present application, the expression of the target loss function can be but is not limited to:
[0010] L loss = λ1L residual + λ2L pixel + λ3L MP + λ4L adv ,
[0011] where the hyperparameters λ1, λ2, λ3, λ4 are the weights of each loss, and are respectively set to 1, 10, 1, 1; L residual is the residual reconstruction loss, L pixel is the pixel loss, L MP is the multi-scale perception loss, L adv is the adversarial loss.
[0012] Optionally, in an embodiment of the present application, the expression of the image with glasses reflection to be eliminated can be but is not limited to:
[0013] I = T + (1 - W)·R,
[0014] where I is the input image, T and R are the face image without reflection and the reflection layer respectively, and W is a three-channel reflection weight map to be predicted by the network, aiming to accurately represent information such as the position, size, and intensity of the reflection.
[0015] An embodiment of the second aspect of the present application provides an image glasses reflection elimination method based on a multi-level refinement network. Using the image glasses reflection elimination method based on the multi-level refinement network as described above, and applied to the model application stage. The method includes the following steps: obtaining at least one actual initial image with glasses reflection to be eliminated and the corresponding actual initial image without glasses reflection of the at least one actual initial image with glasses reflection to be eliminated; performing data processing on the at least one actual initial image with glasses reflection to be eliminated and the actual initial image without glasses reflection to respectively obtain the actual image with glasses reflection after data processing of the at least one actual initial image with glasses reflection to be eliminated and the actual image without glasses reflection after data processing; inputting the actual image with glasses reflection to be eliminated and the actual image without glasses reflection into the trained adaptive reflection perception network and the trained multi-level refinement reflection elimination network to output the image with glasses reflection eliminated after eliminating the glasses reflection of the actual image with glasses reflection to be eliminated. The trained adaptive reflection perception network is constructed from the actual image with glasses reflection to be eliminated, the actual image without glasses reflection, and a pre-trained ResNet-34 network. The trained multi-level refinement reflection elimination network is constructed from the reflection perception result output by the trained adaptive reflection perception network, a coarse network with convolutional kernels of different scales, and a refinement network.
[0016] In the third aspect of the embodiments of the present application, an image glasses reflection elimination device based on a multi-level refinement network is provided. The method for eliminating image glasses reflection based on the multi-level refinement network as described above is adopted and applied to the model training stage. The device includes: an acquisition module, configured to acquire at least one initial image to be eliminated of glasses reflection and an initial image without glasses reflection corresponding to the at least one initial image to be eliminated of glasses reflection; a first data processing module, configured to perform data processing on the at least one initial image to be eliminated of glasses reflection and the initial image without glasses reflection, so as to respectively obtain an image to be eliminated of glasses reflection and an image without glasses reflection after data processing of the at least one initial image to be eliminated of glasses reflection and the initial image without glasses reflection; a first construction module, configured to construct an adaptive reflection perception network applicable to the image to be eliminated of glasses reflection based on the image to be eliminated of glasses reflection, the image without glasses reflection, and a pre-trained ResNet-34 network; a second construction module, configured to construct a multi-level refinement reflection elimination network applicable to the image to be eliminated of glasses reflection based on the reflection perception result output by the adaptive reflection perception network, a coarse-level network, and a refinement network with convolution kernels of different scales; a training module, configured to train the adaptive reflection perception network and the multi-level refinement reflection elimination network based on the image to be eliminated of glasses reflection and the image without glasses reflection by using a target loss function, so as to obtain a trained adaptive reflection perception network and a multi-level refinement reflection elimination network for eliminating glasses reflection in the image to be eliminated of glasses reflection.
[0017] Optionally, in an embodiment of the present application, the second construction module includes: a determination unit, configured to select convolution kernels of different scales based on the reflection perception result, so as to determine the coarse-level network by using the convolution kernels; a first construction unit, configured to construct the refinement network based on the coarse-level result output by the coarse-level network and the reflection perception result; a second construction unit, configured to construct the multi-level refinement reflection elimination network based on the refinement network, the coarse-level result, and the reflection perception result.
[0018] Optionally, in an embodiment of the present application, the determining unit includes: a first generating subunit, configured to obtain at least one initial depth convolution feature corresponding to the image of the glasses reflection to be eliminated based on the reflection perception result and the convolution kernels of different scales; a second generating subunit, configured to pass the at least one initial depth convolution feature through a 1×1 convolution layer to obtain a depth convolution feature map corresponding to the at least one initial depth convolution feature; a third generating subunit, configured to pass the depth convolution feature map through at least one convolution layer to obtain a spatial attention map corresponding to the depth convolution feature map; a fourth generating subunit, configured to obtain an attention feature corresponding to the spatial attention map based on the spatial attention map and the target mask; a determining subunit, configured to determine the coarse-level network based on the attention feature.
[0019] Optionally, in an embodiment of the present application, the expression of the target loss function may be, but is not limited to:
[0020] L loss = λ1L residual + λ2L pixel + λ3L MP + λ4L adv ,
[0021] wherein, the hyperparameters λ1, λ2, λ3, λ4 are the weights of each loss, which are respectively set to 1, 10, 1, 1; L residual is the residual reconstruction loss, L pixel is the pixel loss, L MP is the multi-scale perception loss, L adv is the adversarial loss.
[0022] Optionally, in an embodiment of the present application, the expression of the image of the glasses reflection to be eliminated may be, but is not limited to:
[0023] I = T + (1 - W)·R,
[0024] wherein, I is the input image, T and R are the non-reflective face image and the reflection layer respectively, and W is a three-channel reflection weight map to be predicted by the network, which is designed to accurately represent information such as the position, size, and intensity of the reflection.
[0025] In the embodiment of the fourth aspect of the present application, an image glasses specular reflection elimination device based on a multi-level refinement network is provided. The method for eliminating image glasses specular reflection based on the multi-level refinement network as described above is adopted and applied to the model application stage. Among them, the device includes: an acquisition module, configured to acquire at least one actual initial glasses specular reflection image to be eliminated and the actual initial glasses-free specular reflection image corresponding to the at least one actual initial glasses specular reflection image to be eliminated; a second data processing module, configured to perform data processing on the at least one actual initial glasses specular reflection image to be eliminated and the actual initial glasses-free specular reflection image, respectively obtaining the actual glasses specular reflection image to be eliminated and the actual glasses-free specular reflection image after data processing of the at least one actual initial glasses specular reflection image to be eliminated and the actual initial glasses-free specular reflection image; an output module, configured to input the actual glasses specular reflection image to be eliminated and the actual glasses-free specular reflection image into the trained adaptive specular reflection perception network and the trained multi-level refinement specular reflection elimination network, so as to output the glasses specular reflection elimination image after eliminating the glasses specular reflection of the actual glasses specular reflection image to be eliminated. Among them, the trained adaptive specular reflection perception network is constructed by the actual glasses specular reflection image to be eliminated, the actual glasses-free specular reflection image, and the pre-trained ResNet-34 network, and the trained multi-level refinement specular reflection elimination network is constructed by the specular reflection perception result output by the trained adaptive specular reflection perception network, a coarse network with convolution kernels of different scales, and a refinement network.
[0026] In the embodiment of the fifth aspect of the present application, an electronic device is provided, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the program to implement the method for eliminating image glasses specular reflection based on the multi-level refinement network as described in the above embodiments.
[0027] In the embodiment of the sixth aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the program is executed by a processor, it implements the method for eliminating image glasses specular reflection based on the multi-level refinement network as described above.
[0028] In the embodiment of the seventh aspect of the present application, a computer program product is provided, including a computer program, and when the program is executed, it implements the method for eliminating image glasses specular reflection based on the multi-level refinement network as described above.
[0029] The embodiments of the present application can perform data processing on the collected initial image with glasses reflection to be eliminated and the corresponding initial image without glasses reflection, and then obtain the image with glasses reflection to be eliminated and the image without glasses reflection respectively. Based on the image with glasses reflection to be eliminated, the image without glasses reflection, and the pre-trained ResNet-34 network, an adaptive reflection perception network and a multi-level refinement reflection elimination network suitable for the image with glasses reflection to be eliminated are constructed. Then, the adaptive reflection perception network and the multi-level refinement reflection elimination network are trained using the target loss function to obtain a trained adaptive reflection perception network for eliminating the glasses reflection in the image with glasses reflection to be eliminated and a trained multi-level refinement reflection elimination network. By synthesizing a dataset of glasses reflection images and mixing it with the image pairs without glasses reflection for training to simulate a wider range of reflection scenarios, and through the trained adaptive reflection perception network to perceive information such as the position, area, and intensity of the reflection, eliminating the reflection through the coarse-level network, and then further eliminating the remaining reflection and restoring the details around the human eyes through the refinement network. In addition, the embodiments of the present application can also dynamically adjust the receptive field according to the size of the reflection area to restore the texture under different reflection areas. Thus, the problems in the related art, such as limited dataset, resulting in poor visual quality of the obtained images and hindering the research and application development in related fields, are solved.
[0030] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, where:
[0032] Figure 1 FIG. is a flowchart of an image glasses reflection elimination method based on a multi-level refinement network according to an embodiment of the present application;
[0033] Figure 2 FIG. is a schematic block diagram of the design of a dynamic kernel selection module when constructing a coarse-level network according to an embodiment of the present application;
[0034] Figure 3 FIG. is a flowchart of the working principle of an image glasses reflection elimination method based on a multi-level refinement network according to an embodiment of the present application;
[0035] Figure 4 FIG. is a schematic block diagram of an image glasses reflection elimination device based on a multi-level refinement network according to an embodiment of the present application;
[0036] Figure 5 FIG. is a flowchart of an image glasses reflection elimination method based on a multi-level refinement network according to another embodiment of the present application;
[0037] Figure 6 Flowchart of the working principle of the image glasses reflection elimination method based on a multi-level refinement network provided in another embodiment of the present application;
[0038] Figure 7 Block diagram of the image glasses reflection elimination device based on a multi-level refinement network provided in another embodiment of the present application;
[0039] Figure 8 Schematic structural diagram of an electronic device provided according to an embodiment of the present application. Detailed implementation manners
[0040] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application and should not be construed as a limitation to the present application.
[0041] The image glasses reflection elimination method and device based on a multi-level refinement network according to the embodiments of the present application will be described below with reference to the accompanying drawings. Aiming at the problem that the dataset is limited in the above-mentioned background technology, resulting in poor visual quality of the obtained images and hindering the research and application development in related fields, the present application provides an image glasses reflection elimination method based on a multi-level refinement network. In this method, data processing can be performed on the collected initial image to be eliminated of glasses reflection and the corresponding initial image without glasses reflection, and then the image to be eliminated of glasses reflection and the image without glasses reflection can be obtained respectively. Based on the image to be eliminated of glasses reflection, the image without glasses reflection and the pre-trained ResNet-34 network, an adaptive reflection perception network and a multi-level refinement reflection elimination network suitable for the image to be eliminated of glasses reflection are constructed. Then, the adaptive reflection perception network and the multi-level refinement reflection elimination network are trained using the target loss function to obtain a trained adaptive reflection perception network for eliminating the glasses reflection in the image to be eliminated of glasses reflection and a trained multi-level refinement reflection elimination network. By synthesizing a glasses reflection image dataset and mixing it with the image pairs without glasses reflection for training to simulate a wider range of reflection scenarios, and through the trained adaptive reflection perception network to perceive information such as the position, area and intensity of the reflection, the reflection is eliminated by the coarse-level network, and then the remaining reflection is further eliminated by the refinement network and the details around the human eyes are restored. In addition, the embodiments of the present application can also dynamically adjust the receptive field according to the size of the reflection area to restore the texture under different reflection areas. Thus, the problems in the related technology, such as the limited dataset resulting in poor visual quality of the obtained images and hindering the research and application development in related fields, are solved.
[0042] Specifically, Figure 1A flowchart of an image glasses reflection elimination method based on a multi-level refinement network provided according to an embodiment of the present application.
[0043] As Figure 1 shown, the image glasses reflection elimination method based on a multi-level refinement network is applied to the model training stage, and the method includes the following steps:
[0044] In step S101, at least one initial image to eliminate glasses reflection and at least one initial image without glasses reflection corresponding to the at least one initial image to eliminate glasses reflection are collected.
[0045] As a possible implementation manner, embodiments of the present application can respectively collect an initial image to eliminate glasses reflection when the light is on and an initial image without glasses reflection when the light is off for a single person wearing glasses in front of the camera, so as to obtain a training set applied to the model training stage.
[0046] Exemplarily, embodiments of the present application can fix the camera at a certain position, let the person wearing glasses sit in front of the camera, turn on the light, and adjust the position of the light so that there is reflection on the glasses in front of the camera, take an initial image to eliminate glasses reflection, and then turn off the light and take an initial image without glasses reflection.
[0047] In step S102, data processing is performed on at least one initial image to eliminate glasses reflection and the initial image without glasses reflection to respectively obtain a glasses reflection elimination image and a non-glasses reflection image after data processing of the at least one initial image to eliminate glasses reflection and the initial image without glasses reflection. Among them, the expression of the glasses reflection elimination image can be but is not limited to:
[0048] I = T + (1 - W)·R,
[0049] where I is the input image, T and R are the non-reflective face image and the reflection layer respectively, and W is a three-channel reflection weight map to be predicted by the network, which is designed to accurately represent information such as the position, size, and intensity of the reflection.
[0050] In some embodiments, embodiments of the present application perform data processing on at least one initial image to eliminate glasses reflection and the initial image without glasses reflection, and then obtain the corresponding glasses reflection elimination image and non-glasses reflection image after data processing. Among them, the expression of the glasses reflection elimination image can be but is not limited to:
[0051] I = T + (1 - W)·R,
[0052] where I is the input image, T and R are the non-reflective face image and the reflection layer respectively, and W is a three-channel reflection weight map to be predicted by the network, which is designed to accurately represent information such as the position, size, and intensity of the reflection.
[0053] Exemplarily, in the embodiments of the present application, the initial image with glasses reflection to be eliminated and the initial image without glasses reflection can be adjusted to a resolution of 512X512 through cropping and scaling, and the existing and mature tone transfer algorithm can be used to transfer the tone of the initial image with glasses reflection to the image without glasses reflection.
[0054] In step S103, an adaptive reflection perception network applicable to the image with glasses reflection to be eliminated is constructed based on the image with glasses reflection to be eliminated, the image without glasses reflection, and the pre-trained ResNet-34 network.
[0055] In the actual execution process, the embodiments of the present application can use the pre-trained ResNet-34 network as the basis of the adaptive reflection perception network, adaptively perceive the useful information of the reflection, and combine the image with glasses reflection to be eliminated and the image without glasses reflection to complete the construction of the adaptive reflection perception network.
[0056] In step S104, a multi-level refined reflection elimination network applicable to the image with glasses reflection to be eliminated is constructed based on the reflection perception result output by the adaptive reflection perception network, the coarse-level network with convolutional kernels of different scales, and the refinement network.
[0057] Those skilled in the art can understand that when constructing the multi-level refined reflection elimination network in the embodiments of the present application, a network framework structure based on encoding and decoding can be used. Specifically, for example, a coarse-level network and a refinement network with convolutional kernels of different scales, and combined with the reflection perception result output by the adaptive reflection perception network.
[0058] Exemplarily, the embodiments of the present application can construct a multi-level refined reflection elimination network based on the network framework structure of encoding and decoding, and use the reflection perception result of the adaptive reflection perception network as a guide to remove the glasses reflection.
[0059] Optionally, in an embodiment of the present application, a multi-level refined reflection elimination network applicable to the image with glasses reflection to be eliminated is constructed based on the reflection perception result output by the adaptive reflection perception network, the coarse-level network with convolutional kernels of different scales, and the refinement network, including: based on the reflection perception result, selecting convolutional kernels of different scales to determine the coarse-level network by using the convolutional kernels; constructing the refinement network based on the coarse-level result output by the coarse-level network and the reflection perception result; and constructing the multi-level refined reflection elimination network based on the refinement network, the coarse-level result, and the reflection perception result.
[0060] In some embodiments, when constructing the multi-level refinement anti-reflection network in the embodiments of the present application, the main content is as follows: First, the reflection layer and the transmission layer are respectively predicted by the coarse-level network, and a dynamic kernel selection module with convolution kernels of different scales is designed in the coarse-level network to dynamically select an appropriate convolution kernel according to the size of the reflective area to obtain different receptive fields; through the refinement network, the remaining reflections around the human eyes are further eliminated, and the area around the eyes is refined at the same time.
[0061] Among them, in the embodiments of the present application, the refinement network mainly has three functions: (1) further eliminating the strong reflections that are not completely eliminated by the coarse-level network; (2) eliminating the artifacts left after the elimination by the coarse-level network; (3) refining the details of the position of the eyes on the human face to make its structural texture clearer.
[0062] Exemplarily, in the embodiments of the present application, the reflections on the glasses in the image of the glasses to be eliminated can be initially eliminated by the coarse-level network first, and then the result of the initial elimination and the reflection perception result output by the adaptive reflection perception network are input into the refinement network to further eliminate the reflections on the glasses.
[0063] Optionally, in an embodiment of the present application, based on the reflection perception result, convolution kernels of different scales are selected to determine the coarse-level network by using the convolution kernels, including: based on the reflection perception result and convolution kernels of different scales, obtaining at least one initial depth convolution feature corresponding to the image of the glasses to be eliminated; passing the at least one initial depth convolution feature through a 1×1 convolution layer to obtain a depth convolution feature map corresponding to the at least one initial depth convolution feature; passing the depth convolution feature map through at least one convolution layer to obtain a spatial attention map corresponding to the depth convolution feature map; based on the spatial attention map and the target mask, obtaining the attention feature corresponding to the spatial attention map; determining the coarse-level network based on the attention feature.
[0064] In some embodiments, when constructing the coarse-level network in the embodiments of the present application, it mainly lies in the design of the dynamic kernel selection module, and its process is as Figure 2 shown, and its main content is as follows:
[0065] First, the input feature X is processed by multiple parallel convolution kernels with different scales to generate the initial depth convolution feature U i , and its expression can be but is not limited to:
[0066]
[0067] Among them, is the convolution with the kernel k i and the dilation d iFor the depth convolution, i represents several branches, that is, different convolutional kernels. In the embodiments of the present application, i can be taken as 2, that is, there are two branches. The small kernel k1 is a dilated convolution of 3x3, and the large kernel k2 is a dilated convolution of 7x7.
[0068] Furthermore, in the embodiments of the present application, it is assumed that there are N different convolutional kernels, and the feature maps corresponding to each kernel pass through a 1×1 convolutional layer for further processing to obtain the depth convolution feature map, and its expression can be but is not limited to:
[0069]
[0070] In order to enhance the network's ability to focus on the most relevant spatial context regions to eliminate specular reflections of different areas, the embodiments of the present application design a spatial selection mechanism to select appropriate parts from the feature maps generated by convolutional kernels of different scales, and then connect the features from different receptive fields. Its expression can be but is not limited to:
[0071]
[0072] Then, for the connected feature map apply channel-based average and max pooling operations (denoted as P avg (·) and P max (·)) to effectively extract spatial relationships.
[0073] Next, pass these pooled feature maps through a convolutional layer to convert them into N spatial attention maps, and its expression can be but is not limited to:
[0074]
[0075] Among them, for each spatial attention map in the embodiments of the present application apply the sigmoid activation function to obtain a single spatial selection target mask for each kernel, and its expression can be but is not limited to:
[0076]
[0077] Among them, σ(·) represents the sigmoid function.
[0078] Then, use these spatial selection target masks to weight the feature maps of each convolutional kernel, and through a convolutional layer fuse these weighted features to generate the final attention feature S, and its expression can be but is not limited to:
[0079]
[0080] Finally, the final output of the MLSK module is the element-wise product of the input feature X and the attention feature S, and its expression can be but is not limited to:
[0081] Y = X · S,
[0082] Exemplarily, in the embodiments of the present application, the input feature X can be first input and processed by multiple parallel convolutional kernels with different scales to generate an initial depth convolutional feature U i , and further processed through a 1×1 convolutional layer to obtain a depth convolutional feature map. Among them, in the embodiments of the present application, a spatial selection mechanism is designed to connect features from different receptive fields, and then for the connected feature map apply channel-based average and max pooling operations, and then pass these pooled feature maps through a convolutional layer to convert them into N spatial attention maps. Apply the sigmoid activation function to each spatial attention map to obtain a single spatial selection target mask for each kernel. Then use these spatial selection masks to weight the feature maps of each convolutional kernel, and pass through a convolutional layer to fuse these weighted features to generate the final attention feature S. Finally, the final output is the element-wise product of the input feature X and the attention feature S.
[0083] In step S105, based on the image with glasses reflection to be eliminated and the image without glasses reflection, use the target loss function to train the adaptive reflection perception network and the multi-level refinement reflection elimination network to obtain the trained adaptive reflection perception network and multi-level refinement reflection elimination network for eliminating the glasses reflection in the image with glasses reflection to be eliminated. Among them, the expression of the target loss function can be but is not limited to:
[0084] L loss = λ1L residual + λ2L pixel + λ3L MP + λ4L adv ,
[0085] where the hyperparameters λ1, λ2, λ3, λ4 are the weights of each loss, which are respectively set to 1, 10, 1, 1; L residual is the residual reconstruction loss, L pixel is the pixel loss, L MP is the multi-scale perception loss, and L adv is the adversarial loss.
[0086] In some embodiments, when training the adaptive specular reflection perception network and the multi-level refinement specular reflection elimination network in the embodiments of the present application, a target loss function can be constructed using the residual reconstruction loss, pixel loss, multi-scale perception loss, and adversarial loss. Then, the adaptive specular reflection perception network and the multi-level refinement specular reflection elimination network are continuously iteratively optimized using the target loss function, the image with specular reflection to be eliminated, and the image without specular reflection, and the iterative training is continued until the value of the target loss function is small enough. Among them, the expression of the target loss function can be, but is not limited to:
[0087] L loss = λ1L residual + λ2L pixel + λ3L MP + λ4L adv ,
[0088] where the hyperparameters λ1, λ2, λ3, λ4 are the weights of each loss, and are respectively set to 1, 10, 1, 1; L residual is the residual reconstruction loss, L pixel is the pixel loss, L MP is the multi-scale perception loss, L adv is the adversarial loss.
[0089] Furthermore, in the embodiments of the present application, the residual reconstruction loss can be used to ensure that the linear combination of the reflection layer and the transmission layer is close to the input image I. That is, if the effect of specular reflection elimination is good enough, the linear combination of the reflection layer and the transmission layer will be close to the input image I. The expression of its loss function can be, but is not limited to:
[0090]
[0091] where L2 is the loss representation of the L2 criterion, I is the input image, is the reconstruction result.
[0092] The pixel loss can be used to calculate the pixel difference between the ground truth image T and the outputs of the coarse network and the refinement network, and the L1 criterion loss is used to calculate the absolute difference. The expression of its loss function can be, but is not limited to:
[0093] L pixel = L1(T - T coarse ) + L1(T - T refine ),
[0094] where L1 is the loss representation of the L1 criterion, T is the ground truth, T coarse is the output of the coarse network, and T refine is the output of the refinement network.
[0095] The multi-scale perceptual loss can extract features from different layers of the decoder, input them into the convolutional layer, generate outputs of different resolutions, and compare these outputs with the ground truth image using the perceptual distance instead of the LMSE distance, thereby capturing more context information at different scales. This loss takes into account both low-level and high-level information, and the expression of its loss function can be but is not limited to:
[0096]
[0097] where, and are the outputs of the penultimate layer and the fifth-to-last layer respectively, with sizes 1 / 2 and 1 / 4 of the original image size. T 3 and T 5 correspond to the ground truth images matching these outputs. Layers with smaller sizes are not considered due to relatively less information. Among them, in the embodiments of this application, γ3 = 0.8 and γ5 = 0.6 are set.
[0098] Furthermore, in the embodiments of this application, all images are input into the VGG19 network, and the outputs of the "conv1_2" layer and the "conv2_2" layer in VGG19 are compared.
[0099] The adversarial loss is to further improve the quality of the restored image. Among them, in the embodiments of this application, a multi-layer discriminator network D can be used to evaluate the image quality, and the expression of its loss function can be but is not limited to:
[0100] L adv = ∑ T∈D -logD(T, T refine ),
[0101] Next, in combination with Figure 3 the working principle of the image glasses reflection elimination method based on the multi-level refinement network proposed in the embodiments of this application will be introduced.
[0102] where, Figure 3 is the flowchart of the working principle of the image glasses reflection elimination method based on the multi-level refinement network provided according to an embodiment of this application.
[0103] Step S301: Image acquisition.
[0104] Among them, in the embodiments of this application, the initial image with glasses reflection to be eliminated and the initial image without glasses reflection when the light is off of a single person wearing glasses in front of the camera can be collected respectively, and then the training set for the model training stage can be obtained.
[0105] Step S302: Data processing.
[0106] Among them, in the embodiments of the present application, the initial anti-glare glasses image and the initial non-glare glasses image can be adjusted to a resolution of 512X512 through cropping and scaling, and the existing and mature tone transfer algorithm is used to transfer the tone of the initial anti-glare glasses image to the non-glare glasses image.
[0107] Step S303: Construct an adaptive glare perception network.
[0108] Step S304: Construct a multi-level refinement glare elimination network.
[0109] Step S305: Design an objective loss function.
[0110] Among them, in the embodiments of the present application, the objective loss function can be constructed by using residual reconstruction loss, pixel loss, multi-scale perception loss, and adversarial loss.
[0111] Step S306: Train the network.
[0112] According to the image glasses glare elimination method based on a multi-level refinement network proposed in the embodiments of the present application, the initial anti-glare glasses image and the corresponding initial non-glare glasses image collected can be processed, and then the anti-glare glasses image and the non-glare glasses image can be obtained respectively. Based on the anti-glare glasses image, the non-glare glasses image, and the pre-trained ResNet-34 network, an adaptive glare perception network and a multi-level refinement glare elimination network suitable for the anti-glare glasses image are constructed. Then, the adaptive glare perception network and the multi-level refinement glare elimination network are trained by using the objective loss function to obtain a trained adaptive glare perception network for eliminating the glare in the anti-glare glasses image and a trained multi-level refinement glare elimination network. By synthesizing a glare glasses image dataset and mixing it with the non-glare glasses image pair for training to simulate a wider range of glare scenarios, and through the trained adaptive glare perception network to perceive information such as the position, area, and intensity of the glare, the glare is eliminated by the coarse-level network, and then the remaining glare is further eliminated by the refinement network and the details around the human eyes are restored. In addition, in the embodiments of the present application, the receptive field can also be dynamically adjusted according to the size of the glare area to restore the texture under different glare areas. Thus, the problems in the related art that the dataset is limited, resulting in poor visual quality of the obtained images and hindering the research and application development in related fields are solved.
[0113] Next, a description is given of an image glasses glare elimination device based on a multi-level refinement network proposed in the embodiments of the present application with reference to the accompanying drawings.
[0114] Figure 4 It is a block diagram of an image glasses glare elimination device provided according to an embodiment of the present application.
[0115] AsFigure 4 As shown, the image glasses glare elimination device 40 based on a multi-level refinement network is applied to the model training stage. Among them, the device 40 includes: an acquisition module 401, a first data processing module 402, a first construction module 403, a second construction module 404, and a training module 405.
[0116] Among them, the acquisition module 401 is used to acquire at least one initial image with glasses glare to be eliminated and at least one initial image without glasses glare corresponding to the initial image with glasses glare to be eliminated.
[0117] The first data processing module 402 is used to perform data processing on at least one initial image with glasses glare to be eliminated and the initial image without glasses glare, so as to obtain the image with glasses glare to be eliminated and the image without glasses glare after data processing of the at least one initial image with glasses glare to be eliminated and the initial image without glasses glare respectively.
[0118] The first construction module 403 is used to construct an adaptive glare perception network suitable for the image with glasses glare to be eliminated based on the image with glasses glare to be eliminated, the image without glasses glare, and the pre-trained ResNet-34 network.
[0119] The second construction module 404 is used to construct a multi-level refinement glare elimination network suitable for the image with glasses glare to be eliminated based on the glare perception result output by the adaptive glare perception network, the coarse-level network with convolutional kernels of different scales, and the refinement network.
[0120] The training module 405 is used to train the adaptive glare perception network and the multi-level refinement glare elimination network based on the image with glasses glare to be eliminated and the image without glasses glare by using the target loss function, so as to obtain the trained adaptive glare perception network and multi-level refinement glare elimination network for eliminating the glasses glare in the image with glasses glare to be eliminated.
[0121] Optionally, in an embodiment of the present application, the second construction module 404 includes: a determination unit, a first construction unit, and a second construction unit.
[0122] Among them, the determination unit is used to select convolutional kernels of different scales based on the glare perception result, so as to determine the coarse-level network by using the convolutional kernels.
[0123] The first construction unit is used to construct a refinement network based on the coarse-level result output by the coarse-level network and the glare perception result.
[0124] The second construction unit is used to construct a multi-level refinement glare elimination network based on the refinement network, the coarse-level result, and the glare perception result.
[0125] Optionally, in an embodiment of the present application, the determination unit includes: a first generation subunit, a second generation subunit, a third generation subunit, a fourth generation subunit, and a determination subunit.
[0126] Among them, the first generation subunit is configured to obtain at least one initial depth convolution feature corresponding to the image of the glasses reflection to be eliminated based on the reflection perception result and convolution kernels of different scales.
[0127] The second generation subunit is configured to pass at least one initial depth convolution feature through a 1×1 convolution layer to obtain a depth convolution feature map corresponding to at least one initial depth convolution feature.
[0128] The third generation subunit is configured to pass the depth convolution feature map through at least one convolution layer to obtain a spatial attention map corresponding to the depth convolution feature map.
[0129] The fourth generation subunit is configured to obtain attention features corresponding to the spatial attention map based on the spatial attention map and the target mask.
[0130] The determination subunit is configured to determine the coarse-level network based on the attention features.
[0131] Optionally, in an embodiment of the present application, the expression of the target loss function may but is not limited to:
[0132] L loss = λ1L residual + λ2L pixel + λ3L MP + λ4L adv ,
[0133] Among them, the hyperparameters λ1, λ2, λ3, λ4 are the weights of each loss, which are respectively set to 1, 10, 1, 1; L residual is the residual reconstruction loss, L pixel is the pixel loss, L MP is the multi-scale perception loss, L adv is the adversarial loss.
[0134] Optionally, in an embodiment of the present application, the expression of the image of the glasses reflection to be eliminated may but is not limited to:
[0135] I = T + (1 - W)·R,
[0136] Among them, I is the input image, T and R are the non-reflective face image and the reflection layer respectively, and W is a three-channel reflection weight map to be predicted by the network, aiming to accurately represent information such as the position, size, and intensity of the reflection.
[0137] It should be noted that the foregoing explanatory description of the embodiment of the method for eliminating image glasses reflection based on a multi-level refinement network is also applicable to the device for eliminating image glasses reflection based on a multi-level refinement network in this embodiment, and will not be elaborated here.
[0138] The device for eliminating image glasses reflection based on a multi-level refinement network proposed according to an embodiment of the present application can perform data processing on the collected initial image with glasses reflection to be eliminated and the corresponding initial image without glasses reflection, and then obtain the image with glasses reflection to be eliminated and the image without glasses reflection respectively. Based on the image with glasses reflection to be eliminated, the image without glasses reflection, and the pre-trained ResNet-34 network, an adaptive reflection perception network and a multi-level refinement reflection elimination network suitable for the image with glasses reflection to be eliminated are constructed. Then, the adaptive reflection perception network and the multi-level refinement reflection elimination network are trained using the target loss function to obtain a trained adaptive reflection perception network for eliminating the glasses reflection in the image with glasses reflection to be eliminated and a trained multi-level refinement reflection elimination network. By synthesizing a dataset of glasses reflection images and mixing it with the image pairs without glasses reflection for training to simulate a wider range of reflection scenarios, and through the trained adaptive reflection perception network to perceive information such as the position, area, and intensity of the reflection, eliminating the reflection through the coarse-level network, and then further eliminating the remaining reflection and restoring the details around the human eyes through the refinement network. In addition, the embodiment of the present application can also dynamically adjust the receptive field according to the size of the reflection area to restore the texture under different reflection areas. Thus, the problems in the related art that the dataset is limited, resulting in poor visual quality of the obtained images and hindering the research and application development in related fields are solved.
[0139] The above embodiments describe the model training stage. The following describes the embodiments of the model application stage.
[0140] Figure 5 It is a flowchart of a method for eliminating image glasses reflection based on a multi-level refinement network according to another embodiment of the present application.
[0141] As Figure 5 shown, this method for eliminating image glasses reflection based on a multi-level refinement network is applied to the model application stage, and the method includes the following steps:
[0142] In step S501, at least one actual initial image with glasses reflection to be eliminated and at least one corresponding actual initial image without glasses reflection are obtained.
[0143] In step S502, data processing is performed on at least one actual initial image with glasses reflection to be eliminated and at least one actual initial image without glasses reflection, respectively obtaining an actual image with glasses reflection after data processing of the at least one actual initial image with glasses reflection to be eliminated and an actual image without glasses reflection.
[0144] In step S503, the actual image with glasses reflection to be eliminated and the actual image without glasses reflection are input into the trained adaptive reflection perception network and the trained multi-level refinement reflection elimination network to output an image with glasses reflection eliminated from the actual image with glasses reflection to be eliminated. Among them, the trained adaptive reflection perception network is constructed from the actual image with glasses reflection to be eliminated, the actual image without glasses reflection, and the pre-trained ResNet-34 network, and the trained multi-level refinement reflection elimination network is constructed from the reflection perception result output by the trained adaptive reflection perception network, a coarse-level network with convolutional kernels of different scales, and a refinement network.
[0145] As a possible implementation manner, embodiments of the present application may first obtain at least one actual initial image with glasses reflection to be eliminated and the corresponding actual initial image without glasses reflection, and then perform data processing on the at least one actual initial image with glasses reflection to be eliminated and the actual initial image without glasses reflection to obtain an actual image with glasses reflection after data processing and an actual image without glasses reflection, and input the actual image with glasses reflection to be eliminated and the actual image without glasses reflection into the trained adaptive reflection perception network and the trained multi-level refinement reflection elimination network to output an image with glasses reflection eliminated from the actual image with glasses reflection to be eliminated.
[0146] The following combines Figure 6 to introduce the working principle of the image glasses reflection elimination method based on a multi-level refinement network proposed in embodiments of the present application.
[0147] Among them, Figure 6 is a flowchart of the working principle of the image glasses reflection elimination method based on a multi-level refinement network provided in another embodiment of the present application.
[0148] Step S601: Input the actual image with glasses reflection to be eliminated and the actual image without glasses reflection after data processing.
[0149] Step S602: Obtain the trained adaptive reflection perception network.
[0150] Step S603: Obtain the trained multi-level refinement reflection elimination network.
[0151] Step S604: Output an image with glasses reflection eliminated.
[0152] According to the image glasses specular reflection elimination method based on a multi-level refinement network proposed in an embodiment of the present application, the processed actual specular reflection image to be eliminated and the actual specular reflection-free image of glasses can be input into a trained adaptive specular reflection perception network and a trained multi-level refinement specular reflection elimination network, and then an image of glasses with specular reflection eliminated from the actual specular reflection image to be eliminated is output. The trained adaptive specular reflection perception network senses information such as the position, area, and intensity of the specular reflection. The specular reflection is eliminated through a coarse-level network, and then the residual specular reflection is further eliminated and the details around the human eyes are restored through a refinement network. In addition, the embodiment of the present application can also dynamically adjust the receptive field according to the size of the specular reflection area to restore the texture under different specular reflection areas. Thus, the problems in the related art that the dataset is limited, resulting in poor visual quality of the obtained images and hindering the research and application development in related fields are solved.
[0153] Next, a device for eliminating specular reflection of glasses in an image based on a multi-level refinement network proposed in an embodiment of the present application is described with reference to the accompanying drawings.
[0154] Figure 7 It is a block diagram of a device for eliminating specular reflection of glasses in an image based on a multi-level refinement network provided in another embodiment of the present application.
[0155] As Figure 7 shown, the device 70 for eliminating specular reflection of glasses in an image based on a multi-level refinement network is applied to the model application stage. Among them, the device 70 includes: an acquisition module 701, a second data processing module 702, and an output module 703.
[0156] Among them, the acquisition module 701 is used to acquire at least one actual initial specular reflection image to be eliminated of glasses and at least one actual initial specular reflection-free image of glasses corresponding to the at least one actual initial specular reflection image to be eliminated of glasses.
[0157] The second data processing module 702 is used to perform data processing on at least one actual initial specular reflection image to be eliminated of glasses and the actual initial specular reflection-free image of glasses to respectively obtain an actual specular reflection image to be eliminated and an actual specular reflection-free image of glasses after data processing of the at least one actual initial specular reflection image to be eliminated of glasses and data processing of the actual initial specular reflection-free image of glasses.
[0158] An output module 703 is configured to input the actual image with glasses reflection to be eliminated and the actual image without glasses reflection into the trained adaptive reflection perception network and the trained multi-level refined reflection elimination network, so as to output an image with glasses reflection eliminated from the actual image with glasses reflection to be eliminated. The trained adaptive reflection perception network is constructed from the actual image with glasses reflection to be eliminated, the actual image without glasses reflection, and a pre-trained ResNet-34 network. The trained multi-level refined reflection elimination network is constructed from the reflection perception result output by the trained adaptive reflection perception network, a coarse-level network with convolutional kernels of different scales, and a refinement network.
[0159] It should be noted that the foregoing explanation of the embodiment of the method for eliminating glasses reflection in an image based on a multi-level refinement network also applies to the device for eliminating glasses reflection in an image based on a multi-level refinement network in this embodiment, and will not be elaborated here.
[0160] The device for eliminating glasses reflection in an image based on a multi-level refinement network according to an embodiment of the present application can input the processed actual image with glasses reflection to be eliminated and the actual image without glasses reflection into the trained adaptive reflection perception network and the trained multi-level refined reflection elimination network, and then output an image with glasses reflection eliminated from the actual image with glasses reflection to be eliminated. The trained adaptive reflection perception network senses information such as the position, area, and intensity of the reflection, the coarse-level network eliminates the reflection, and the refinement network further eliminates the remaining reflection and restores the details around the human eyes. In addition, the embodiment of the present application can also dynamically adjust the receptive field according to the size of the reflection area to restore the texture under different reflection areas. Thus, the problems in the related art that the dataset is limited, resulting in poor visual quality of the obtained images and hindering the research and application development in related fields are solved.
[0161] Figure 8 FIG. is a schematic structural diagram of an electronic device according to an embodiment of the present application. The electronic device may include:
[0162] A memory 801, a processor 802, and a computer program stored on the memory 801 and executable on the processor 802.
[0163] When the processor 802 executes the program, it implements the method for eliminating glasses reflection in an image based on a multi-level refinement network provided in the foregoing embodiment.
[0164] Further, the electronic device further includes:
[0165] A communication interface 803 for communication between the memory 801 and the processor 802.
[0166] The memory 801 is used to store a computer program executable on the processor 802.
[0167] The memory 801 may include high-speed RAM memory and may also include non-volatile memory, such as at least one disk memory.
[0168] If the memory 801, the processor 802, and the communication interface 803 are implemented independently, the communication interface 803, the memory 801, and the processor 802 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity in representation, Figure 8 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0169] Optionally, in a specific implementation, if the memory 801, the processor 802, and the communication interface 803 are integrated on a single chip, the memory 801, the processor 802, and the communication interface 803 can communicate with each other through an internal interface.
[0170] The processor 802 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0171] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned method for eliminating image glasses reflection based on a multi-level refinement network is implemented.
[0172] The embodiments of the present application also provide a computer program product, including a computer program, and when the program is executed, the above-mentioned method for eliminating image glasses reflection based on a multi-level refinement network is implemented.
[0173] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0174] In addition, the terms "first" and "second" are used only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0175] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or N executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of this application belong.
[0176] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part (electronic device) having one or N wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.
[0177] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0178] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0179] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, may exist separately as individual physical units, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0180] The above-mentioned storage medium may be a read-only memory, a magnetic disk or an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present application.
Claims
1. A method for eliminating reflections from image glasses based on a multi-level refinement network, characterized in that: Applied to the model training stage, wherein the method comprises the following steps: Collecting at least one initial image of glasses reflection to be eliminated and an initial image without glasses reflection corresponding to the at least one initial image of glasses reflection to be eliminated; Performing data processing on the at least one initial image with glasses reflection to be eliminated and the initial image without glasses reflection, so as to obtain the image with glasses reflection to be eliminated and the image without glasses reflection after the data processing of the at least one initial image with glasses reflection to be eliminated and the data processing of the initial image without glasses reflection respectively; Based on the image with glasses reflection to be eliminated, the image without glasses reflection and a pre-trained ResNet-34 network, an adaptive reflection perception network suitable for the image with glasses reflection to be eliminated is constructed; Based on the reflection perception results output by the adaptive reflection perception network, the coarse-level network and the refined network with convolution kernels of different scales, a multi-level refined reflection elimination network suitable for the glasses reflection image to be eliminated is constructed; Based on the glasses reflection image to be eliminated and the image without glasses reflection, the adaptive reflection perception network and the multi-level refined reflection elimination network are trained using a target loss function to obtain a trained adaptive reflection perception network and a multi-level refined reflection elimination network for eliminating glasses reflection in the glasses reflection image to be eliminated.
2. The method according to claim 1, characterized in that The method of constructing a multi-level refined reflection elimination network suitable for the glasses reflection image to be eliminated based on the reflection perception result output by the adaptive reflection perception network, a coarse-level network and a refined network with convolution kernels of different scales, comprises: Based on the reflection perception result, selecting convolution kernels of different scales to determine the coarse-level network using the convolution kernels; Constructing the refined network based on the coarse-level result output by the coarse-level network and the reflection perception result; Based on the refined network, the coarse-level result and the reflection perception result, the multi-level refined reflection elimination network is constructed.
3. The method according to claim 2, characterized in that The selecting convolution kernels of different scales based on the reflection perception result to determine the coarse-level network by using the convolution kernels includes: Based on the reflection perception result and the convolution kernels of different scales, obtaining at least one initial deep convolution feature corresponding to the glasses reflection image to be eliminated; Passing the at least one initial deep convolutional feature through a 1×1 convolutional layer to obtain a deep convolutional feature map corresponding to the at least one initial deep convolutional feature; Passing the deep convolutional feature map through at least one convolutional layer to obtain a spatial attention map corresponding to the deep convolutional feature map; Based on the spatial attention map and the target mask, obtaining an attention feature corresponding to the spatial attention map; The coarse-level network is determined based on the attention feature.
4. The method according to claim 1, characterized in that: The expression of the objective loss function is: L loss =λ1L residual +λ2L pixel +λ3L MP +λ4L adv , Among them, the hyperparameters λ1, λ2, λ3, and λ4 are the weights of each loss, which are set to 1, 10, 1, and 1 respectively; L residual is the residual reconstruction loss, L pixel is the pixel loss, L MP is the multi-scale perceptual loss, L adv For confrontational losses.
5. The method according to claim 1, characterized in that The expression of the glasses reflection image to be eliminated is: I=T+(1-W)·R, Among them, I is the input image, T and R are the non-reflective face image and the reflection layer respectively, and W is a three-channel reflection weight map to be predicted by the network, which aims to accurately represent the position, size, strength and other information of the reflection.
6. A method for eliminating reflections from image glasses based on a multi-level refinement network, characterized in that: The method for eliminating glare from image glasses based on a multi-level refinement network as described in any one of claims 1 to 5 is applied to the model application stage, wherein the method comprises the following steps: Acquire at least one actual initial image to be cleared of glasses reflections and an actual initial image without glasses reflections corresponding to the at least one actual initial image to be cleared of glasses reflections; Performing data processing on the at least one actual initial image with glasses reflection to be eliminated and the actual initial image without glasses reflection, to obtain the actual image with glasses reflection to be eliminated and the actual image without glasses reflection after the data processing of the at least one actual initial image with glasses reflection to be eliminated and the data processing of the actual initial image without glasses reflection respectively; The actual image of glasses reflection to be eliminated and the actual image without glasses reflection are input into the trained adaptive reflection perception network and the trained multi-level refined reflection elimination network to output the glasses reflection eliminated image after the glasses reflection is eliminated from the actual image of glasses reflection to be eliminated, wherein the trained adaptive reflection perception network is constructed by the actual image of glasses reflection to be eliminated, the actual image without glasses reflection and a pre-trained ResNet-34 network, and the trained multi-level refined reflection elimination network is constructed by the reflection perception result output by the trained adaptive reflection perception network, a coarse-level network with convolution kernels of different scales, and a refined network.
7. An image glasses reflection elimination device based on a multi-level refinement network, characterized in that: Applied to the model training stage, wherein the device comprises: A collection module, used for collecting at least one initial image of glasses reflection to be eliminated and an initial image without glasses reflection corresponding to the at least one initial image of glasses reflection to be eliminated; A first data processing module is used to perform data processing on the at least one initial image to be eliminated with glasses reflection and the initial image without glasses reflection, so as to obtain the image to be eliminated with glasses reflection and the image without glasses reflection after the data processing of the at least one initial image to be eliminated with glasses reflection and the data processing of the initial image without glasses reflection respectively; A first construction module is used to construct an adaptive reflection perception network suitable for the image with glasses reflection to be eliminated based on the image with glasses reflection to be eliminated, the image without glasses reflection and a pre-trained ResNet-34 network; The second construction module is used to construct a multi-level refined reflection elimination network suitable for the glasses reflection image to be eliminated based on the reflection perception result output by the adaptive reflection perception network, the coarse-level network and the refined network with convolution kernels of different scales; A training module is used to train the adaptive reflection perception network and the multi-level refined reflection elimination network based on the glasses reflection image to be eliminated and the image without glasses reflection using a target loss function, so as to obtain a trained adaptive reflection perception network and a multi-level refined reflection elimination network for eliminating glasses reflection in the glasses reflection image to be eliminated.
8. An image glasses reflection elimination device based on a multi-level refinement network, characterized in that: The method for eliminating glare from image glasses based on a multi-level refinement network as described in any one of claims 1 to 5 is applied to the model application stage, wherein the device comprises: An acquisition module, used for acquiring at least one actual initial image with glasses reflection to be eliminated and an actual initial image without glasses reflection corresponding to the at least one actual initial image with glasses reflection to be eliminated; A second data processing module is used to perform data processing on the at least one actual initial image with glasses reflection to be eliminated and the actual initial image without glasses reflection, to obtain the actual image with glasses reflection to be eliminated and the actual image without glasses reflection after the data processing of the at least one actual initial image with glasses reflection to be eliminated and the data processing of the actual initial image without glasses reflection respectively; The output module is used to input the actual image of glasses reflection to be eliminated and the actual image without glasses reflection into the trained adaptive reflection perception network and the trained multi-level refined reflection elimination network, so as to output the glasses reflection eliminated image after the glasses reflection is eliminated from the actual image of glasses reflection to be eliminated, wherein the trained adaptive reflection perception network is constructed by the actual image of glasses reflection to be eliminated, the actual image without glasses reflection and a pre-trained ResNet-34 network, and the trained multi-level refined reflection elimination network is constructed by the reflection perception result output by the trained adaptive reflection perception network, a coarse-level network with convolution kernels of different scales, and a refined network.
9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for eliminating glare from image glasses based on a multi-level refinement network as described in any one of claims 1 to 5 or the method for eliminating glare from image glasses based on a multi-level refinement network as described in claim 6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the method for eliminating glare from image glasses based on a multi-level refinement network as described in any one of claims 1 to 5 or the method for eliminating glare from image glasses based on a multi-level refinement network as described in claim 6.