Cascade Surface Specular Highlights Removal Method Based on Conditional Generative Adversarial Network

By applying a multi-scale gradient cascade method based on condition generation adversarial networks in the robot vision system, the problem of grab failure caused by high reflection and position deviation of the step-palletizing is solved, efficient highlight removal and step-recognition are achieved, and the accuracy of robot grabbing is improved.

CN114841877BActive Publication Date: 2025-06-13HANGZHOU DIANZI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210449142.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-26
Publication Date
2025-06-13
Estimated Expiration
2042-04-26

AI Technical Summary

Technical Problem

In manufacturing factories, the high reflection and position deviation of the rung palletizing result in the failure of the robot's grasp, and the prior art is difficult to effectively remove the highlights and accurately identify the rung posture and position.

Method used

Using a multi-scale gradient cascade method based on conditional generation adversarial network, a highlight removal model that can process images end-to-end by combining the U-net structure generator and PatchGAN discriminator, combined with spatial context-intensive blocks and multi-scale gradient cascade modules.

Benefits of technology

It realizes effective removal of the highlights on the step surface, improves the accuracy of step recognition and grasping point positioning, and enhances the robot's processing ability of the highly reflective metal material steps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114841877B_ABST
    Figure CN114841877B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for removing specular highlights on stepped surfaces based on conditional generative adversarial networks. A multi-scale conditional generative adversarial network MSDGC-GAN is constructed, aiming to obtain specular highlight removal images in an end-to-end manner. The main structure of the network includes a generator and a discriminator. The basic architecture of the generator adopts a U-net structure with spatial context dense blocks as basic modules, and uses its encoding-decoding structure characteristics to extract deep structure information of images, and fully extracts the spatial context background information features between feature pixels. The connection between encoding and decoding adopts a SOS enhancement strategy structure to achieve refined enhancement processing of images, realize the conduction of gradients, and improve the problem of unstable gradients during network training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of machine vision and is used for quickly and accurately identifying and grasping the stepped palletizing of factory robots. Through the camera installed on the robotic arm, the stepped palletizing on the production line is photographed, and the conditional generative adversarial network is used to process the highly reflective part of the stepped surface, so as to accurately identify the posture and position of the steps, providing guarantee for the subsequent grasping of the robotic arm. Background Art

[0002] With the rise of the new round of scientific and technological revolution and industrial transformation wave, more and more enterprises have begun to notice the unique advantages of the intelligent integration of enterprise production in future competition. As the representative of intelligent equipment, robots have increasingly become the power booster for enterprises to transform and upgrade, the multiplier for enterprise production efficiency, and the amplifier for enterprise competitive advantages. Naturally, they have become the high points for major enterprises to compete for economic development and are widely used in various production scenarios. Palletizing robots, as an important branch of industrial robots, play an outstanding role in various industrial fields. Among them, the traditional palletizing robots based on teach pendants are currently adopted by many domestic factories for related operations due to their simple operation and easy maintenance. However, such robots require the objects to be palletized to be strictly placed at fixed teach points. Once the palletized objects deviate from the original teach points due to certain factors, the robot will be unable to locate the target object, resulting in grasping failure. And currently, the scale, intensification, and automation degree of manufacturing factories are getting higher and higher. Once a problem occurs in a certain production link, it will seriously affect its production efficiency and cause a series of troubles. As Figure 1 shown in the automatic production line of escalator step palletizing of an elevator manufacturing enterprise. In this production line, each step group is first stacked into a stepped stack with a fixed structure through the assembly production line. Subsequently, the neatly stacked stepped stack will be transported from the assembly production line to the designated palletizing position through roller transmission. Finally, the palletizing robot will perform the de-palletizing and palletizing operations on each step group in the stepped stack according to the program through the pre-set teach points. However, due to the long-distance transportation of the stepped stack, it is inevitable that there will be collisions with other steps or objects during the transportation process. At the same time, there will also be errors in the assembly process of the stepped stack, which will cause the step group to deviate from the teach point position, resulting in grasping failure.

[0003] Moreover, different from ordinary material objects, on the actual escalator step palletizing production line, the material of the steps to be grasped by the robot is usually aluminum metal, so it has the characteristics of high reflectivity and complex background under natural light. The existence of high-light coverage in some areas of the step surface will bring troubles to subsequent image processing steps in visual guidance, such as threshold segmentation, edge straight line detection, etc., and is prone to problems such as uneven segmentation and detection failure, which will further affect subsequent step recognition, grasping point positioning and extraction, etc. Therefore, the removal of high light from the images of escalator steps made of metal materials is of great significance for the robot palletizing system guided by vision. Summary of the Invention

[0004] In view of the current research status, the present invention proposes a stepped high-light removal model with strong robustness. In this model, first, an input stepped image I with highlights to be removed is taken H , and a corresponding clear image I with highlight suppression is generated without any other relevant information assistance S . Therefore, a multi-scale conditional generative adversarial network MSDGC-GAN is constructed, aiming to obtain a high-light removal image in an end-to-end manner. The main structure of the network consists of a generator and a discriminator. The basic architecture of the generator adopts the U-net structure, and its encoding-decoding structure characteristics are used to extract the deep structure information of the image. At the same time, in order to fully extract the spatial context background information features between feature pixels, a spatial context dense block is proposed as the basic module of the generator. To solve the problem that the network is prone to lose some scale feature information in the downsampling pooling operation, a multi-scale gradient cascade method is proposed. By cascading and outputting from the bottom-layer features in turn to make up for the downsampling feature loss between adjacent modules, and connecting the discriminators and generators at each scale with their respective cascade outputs, the network enhances the ability to process image details and has multi-scale discrimination ability, realizing the conduction of gradients and improving the problem of unstable gradients during network training. In the loss function stage, the dichromatic reflection model is analyzed, and the diffuse reflection component estimation of the image is applied to the loss function. At the same time, the adversarial loss function and the feature matching loss are combined as the total target loss

[0005] The specific steps of the method of the present invention are as follows

[0006] A stepped image high-light removal method based on a conditional generative adversarial network, the steps of which are as follows

[0007] Step (1), construct a training set

[0008] By obtaining the stepped image I with highlights to be removed H1 and the corresponding stepped image I without highlights H2 to construct a one-to-one stepped surface high-light comparison data set, and perform the same cutting on the images to select the optimal image pair; generate a corresponding clear image I with highlight suppression without any other relevant information assistance S ;

[0009] Step (2), build a stepped image high-light removal model MSDGC-GAN based on a conditional generative adversarial network, and use the training data set processed by data cutting for training

[0010] The stepped image high-light removal model based on a conditional generative adversarial network includes a main network and a multi-scale gradient cascade module

[0011] The main network includes a generator and a discriminator;

[0012] The generator adopts the basic architecture of the U-net network, and the U-net network adopts an encoding-decoding structure;

[0013] In the downsampling process of the encoder, the first to fifth spatial context dense layers (Encode-SCFDB) in cascade are adopted;

[0014] The first to fifth spatial context dense layers (Encode-SCFDB) are all used to extract the spatial context background information features between feature pixels, extract and transfer the image pixel background feature information based on multi-path parallel slice-by-slice convolution, and obtain the semantic feature information of the image through a deep dense network; it includes a plurality of cascaded spatial context dense blocks SCF and a first transition layer;

[0015] The spatial context dense block SCF can fully obtain the spatial context background information between the pixels in each row and column of the feature map. Specifically, it first uses the input feature map to apply a 1X1 convolutional bottleneck layer for dimensionality reduction, InstanceNorm instance normalization, and LeakyReLu activation to generate a new feature map Q. Subsequently, the feature map Q is divided into upper and lower branches within the SCF layer, and different direction spatial slice-by-slice convolution strategies in the left, right, up, down order and up, down, left, right order are adopted respectively; then, after normalizing and activating the slice information of each path, the two branches are weighted and fused, and then passed through a 3X3 convolution, InstanceNorm instance normalization, and LeakyReLu activation to output the feature map;

[0016] The input of each spatial context dense block is the original feature map and the output feature map of the spatial context dense block cascaded in the previous stage within the same spatial context dense layer;

[0017] The first transition layer uses 1x1 convolution to compress the feature channels, and then realizes the downsampling operation of the feature map through average pooling, reducing the source input feature map to 1 / 2 size;

[0018] The decoding layer includes the sixth to tenth spatial context dense blocks (Decode-SCFDB) cascaded in sequence. They correspond one-to-one to the cascaded feature output maps of the encoder at each scale, output the feature maps enhanced by each initial cascade, and send the feature images to the discriminator at the corresponding scale after a 1x1 convolution. Among them, the structure of the tenth spatial context dense block is the same as that of the first spatial context dense block, the structure of the ninth spatial context dense block is the same as that of the second spatial context dense block, the structure of the eighth spatial context dense block is the same as that of the third spatial context dense block, the structure of the seventh spatial context dense block is the same as that of the fourth spatial context dense block, and the structure of the sixth spatial context dense block is the same as that of the fifth spatial context dense block;

[0019] The first output end of the first spatial context dense block is connected to the first input end of the first connection layer, and the second output end is connected to the input end of the second spatial context dense block; the second input end of the first connection layer is connected to the output end of the second 1×1 convolution layer, and the output end is connected to the input end of the first 1×1 convolution layer; the output end of the first 1×1 convolution layer is connected to the input end of the third 3×3 convolution layer; the output end of the third 3×3 convolution layer is connected to the input end of the tenth spatial context dense block;

[0020] The first output end of the second spatial context dense block is connected to the first input end of the second connection layer, and the second output end is connected to the input end of the third spatial context dense block; the second input end of the second connection layer is connected to the output end of the third 1×1 convolution layer, and the output end is connected to the input end of the first upsampling layer; the output end of the first upsampling layer is connected to the input ends of the second 1×1 convolution layer and the second 3×3 convolution layer; the output end of the second 3×3 convolution layer is connected to the input end of the ninth spatial context dense block;

[0021] The first output end of the third spatial context dense block is connected to the first input end of the third connection layer, and the second output end is connected to the input end of the fourth spatial context dense block; the second input end of the third connection layer is connected to the output end of the fourth 1×1 convolution layer, and the output end is connected to the input end of the second upsampling layer; the output end of the second upsampling layer is connected to the input ends of the third 1×1 convolution layer and the first 3×3 convolution layer; the output end of the first 3×3 convolution layer is connected to the input end of the eighth spatial context dense block;

[0022] The first output end of the fourth spatial context dense block is connected to the first input end of the fourth connection layer, and the second output end is connected to the input ends of the fifth spatial context dense block and the seventh spatial context block; the second input end of the fourth connection layer is connected to the output end of the fourth upsampling layer, and the output end is connected to the input end of the third upsampling layer; the output end of the third upsampling layer is connected to the input end of the fourth 1×1 convolution layer;

[0023] The connections between the first to fourth spatial context dense layers and the tenth to seventh spatial context dense blocks respectively adopt the SOS enhancement strategy structure to replace the long-distance skip connection method of the traditional U-net; therefore, the decoder realizes the refinement and enhancement processing of the image. Specifically:

[0024]

[0025] In the formula, ψ k (x) represents the output feature map of the decoder, represents the result after processing by the current spatial context dense block in the decoder, and Up() represents upsampling;

[0026] The discriminator adopts a PatchGAN discriminator, which includes three discriminators D1, D2, and D3 with the same structure, and respectively receives the output images from different scales of the decoder;

[0027] Step (3): Use the trained cascade image specular highlight removal model based on the conditional generative adversarial network to achieve the specular highlight processing of the cascade surface;

[0028] During the model training process, the total loss function consists of three parts: the adversarial loss function, the diffuse reflection component loss function, and the feature matching loss;

[0029] Adversarial loss function is as follows:

[0030]

[0031] In the formula, G represents the generator, and D k represents the output result of each discriminator, and min G () represents taking the minimum of the loss gap between the generator and the discriminator, represents the conditional loss function:

[0032]

[0033] In the formula, E (x,y) represents the minimum consumption cost of deriving the image y from the image x, and E (x,z) represents the minimum consumption cost of deriving the image z from the image x.

[0034] Diffuse reflection component loss function is as follows:

[0035]

[0036] In the formula, z k and y k are respectively the predicted and true diffuse reflection components of the k-th scale image, and c krepresents the number of image channels at the k-th scale (RGB 3 channels), w k and h k are the size of the image at the k-th scale (ranging from 256*256 to 32*32); ∑ represents the accumulation of the loss results of the first k images.

[0037] Feature matching loss is defined as follows:

[0038]

[0039] In the formula, N i represents the number of elements in the i-th layer of the discriminator, T is the total number of layers, represents the i-th layer feature of the discriminator D k and x k represents the real image corresponding to the k-th scale.

[0040] Therefore, the total loss function is as follows:

[0041]

[0042] where λ 1 , λ 2 are the weights of the feature matching loss function and the diffuse reflection loss respectively.

[0043] Another object of the present invention is to provide a stepped image specular highlight removal device based on a conditional generative adversarial network, including:

[0044] An image acquisition module for acquiring a stepped image I with specular highlights to be removed H ;

[0045] A specular highlight removal module that uses the trained stepped image specular highlight removal model based on a conditional generative adversarial network to remove the specular highlights from the stepped image I with specular highlights to be removed H for specular highlight removal.

[0046] Advantages of the present invention:

[0047] The present invention designs a method based on a conditional generative adversarial network to solve the problem of removing specular highlights on the surface of a single stepped image. By replacing the spatial context dense block with the basic module of the traditional U-net architecture, the feature extraction ability of the network for deep image information is enhanced, and by replacing the long-distance skip connection method of the traditional U-net with the SOS enhancement strategy structure, the network is endowed with multi-scale discrimination ability and the training gradient of the network is stabilized. A custom dataset for network training and testing is established by simulating specular highlight illumination on the stepped surface.

[0048] The present invention provides a more effective solution, which fully extracts the texture and background feature information of the image, can effectively generate high-quality images, effectively restore the surface features, and make the feature restoration closer to the actual picture. Brief Description of the Drawings

[0049] Figure 1 It is the architecture diagram of the cascaded image specular highlight removal model based on the conditional generative adversarial network of the present invention;

[0050] Figure 2 It is the architecture diagram of the spatial context dense layer Encode - SCFDB;

[0051] Figure 3 It is the schematic diagram of Slice - by - Slice convolution;

[0052] Figure 4 It is the feature concatenation method;

[0053] Figure 5 It is the production of the cascaded dataset;

[0054] Figure 6 It is the comparison of cascaded surface specular highlight removal, where (a) is the specular highlight image, (b) is Ramos, (c) is Yi, (d) is Ye, (e) is pix2pixHD, (f) is the present invention, (g) is the non - specular highlight image; the first row is the vertical specular highlight processing effect, and the second row is the horizontal specular highlight processing effect.

[0055] Figure 7 It is the comparison of cascaded image specular highlight removal in the actual work station, where (a) is the specular highlight image, (b) is Ramos, (c) is Ye, (d) is Yamamoto, (e) is pix2pixHD, (f) is the present invention; the first row is the processing effect of the local specular highlight area, the second row is the processing effect of the large - area specular highlight, and the third row is the effect diagram of processing the background metal specular highlight area.

[0056] Figure 8 It is the comparison of classical specular highlight images, where (a) is the specular highlight image, (b) is Yamamoto, (c) is Yi, (d) is Ye, (e) is pix2pixHD, (f) is the present invention, (g) is the original image; the first row is the specular highlight rabbit dataset, and the second row is the specular highlight fruit dataset. Detailed Embodiment

[0057] The present invention will be further described below in conjunction with specific embodiments.

[0058] Based on the six - degree - of - freedom robotic arm ABB industrial robot in the actual factory environment, a palletizing robot system platform based on binocular vision as shown in Figure 3 is built. The binocular camera module consists of binocular cameras and is responsible for the acquisition task of cascaded images.

[0059] A method for removing highlights from cascade images based on a conditional generative adversarial network, the steps of which are as follows:

[0060] Step (1), construct a training set

[0061] By obtaining the cascade image I with highlights to be removed H1 and the corresponding cascade image I without highlights H2 to construct a one-to-one cascade surface highlight comparison data set, and perform the same cutting on the images to select the optimal image pair; generate a corresponding clear image I with highlight suppression without any other relevant information assistance S ;

[0062] Step (2), build a cascade image highlight removal model based on a conditional generative adversarial network, and use the training data set after data cutting processing for training;

[0063] The cascade image highlight removal model based on a conditional generative adversarial network includes a main network and a multi-scale gradient cascade module;

[0064] The main network includes a generator and a discriminator;

[0065] The generator adopts the basic architecture of the U-net network, and the U-net network adopts an encoder-decoder structure;

[0066] In the downsampling process of the encoder, the first to fifth spatial context dense layers (Encode-SCFDB) in cascade are adopted;

[0067] The first to fifth spatial context dense layers (Encode-SCFDB) are all used to extract the spatial context background information features between feature pixels, extract and transfer the image pixel background feature information based on multi-path parallel slice-by-slice convolution, and obtain the semantic feature information of the image through a deep dense network; it includes a plurality of spatial context dense blocks SCF in cascade and a first transition layer;

[0068] The spatial context dense block SCF can fully obtain the spatial context background information between the pixels in each row and column of the feature map. Specifically, it first uses the input feature map to apply a 1X1 convolutional bottleneck layer for dimensionality reduction, InstanceNorm instance normalization, and LeakyReLu activation to generate a new feature map Q. Subsequently, the feature map Q is divided into upper and lower branches within the SCF layer, and different directional spatial slice-by-slice convolution strategies of left, right, up, down order and up, down, left, right order are adopted respectively; then, after normalizing and activating the slice information of each path, the two branches are weighted and fused, and then passed through a 3X3 convolution, InstanceNorm instance normalization, and LeakyReLu activation to output the feature map;

[0069] The input of each spatial context dense block is the original feature map and the output feature map of the spatial context dense blocks cascaded in front within the same spatial context dense layer;

[0070] The first transition layer uses 1x1 convolution to compress the feature channels, and then downsamples the feature map through average pooling, reducing the source input feature map to 1 / 2 size;

[0071] The decoding layer includes the sixth to tenth spatial context dense layers (Decode - SCFDB) cascaded in sequence. At each scale, it corresponds one - to - one with the cascaded feature output map of the encoder, outputs the feature map after each initial cascaded enhancement, and sends the feature image into the discriminator of the corresponding scale after a 1x1 convolution; among them, the structure of the tenth spatial context dense layer is the same as that of the first spatial context dense layer, the structure of the ninth spatial context dense layer is the same as that of the second spatial context dense layer, the structure of the eighth spatial context dense layer is the same as that of the third spatial context dense layer, the structure of the seventh spatial context dense layer is the same as that of the fourth spatial context dense layer, and the structure of the sixth spatial context dense layer is the same as that of the fifth spatial context dense layer;

[0072] The first output end of the first spatial context dense layer is connected to the first input end of the first connection layer, and the second output end is connected to the input end of the second spatial context dense layer; the second input end of the first connection layer is connected to the output end of the second 1×1 convolution layer, and the output end is connected to the input end of the first 1×1 convolution layer; the output end of the first 1×1 convolution layer is connected to the input end of the third 3×3 convolution layer; the output end of the third 3×3 convolution layer is connected to the input end of the tenth spatial context dense layer;

[0073] The first output end of the second spatial context dense layer is connected to the first input end of the second connection layer, and the second output end is connected to the input end of the third spatial context dense layer; the second input end of the second connection layer is connected to the output end of the third 1×1 convolution layer, and the output end is connected to the input end of the first upsampling layer; the output end of the first upsampling layer is connected to the input end of the second 1×1 convolution layer and the input end of the second 3×3 convolution layer; the output end of the second 3×3 convolution layer is connected to the input end of the ninth spatial context dense layer;

[0074] The first output end of the third spatial context dense layer is connected to the first input end of the third connection layer, and the second output end is connected to the input end of the fourth spatial context dense layer; the second input end of the third connection layer is connected to the output end of the fourth 1×1 convolution layer, and the output end is connected to the input end of the second upsampling layer; the output end of the second upsampling layer is connected to the input end of the third 1×1 convolution layer and the input end of the first 3×3 convolution layer; the output end of the first 3×3 convolution layer is connected to the input end of the eighth spatial context dense layer;

[0075] The first output end of the fourth spatial context dense layer is connected to the first input end of the fourth connection layer, and the second output end is connected to the input end of the fifth spatial context dense layer and the input end of the seventh spatial context dense block; the second input end of the fourth connection layer is connected to the output end of the fourth upsampling layer, and the output end is connected to the input end of the third upsampling layer; the output end of the third upsampling layer is connected to the input end of the fourth 1×1 convolutional layer;

[0076] The connections between the first to fourth spatial context dense layers and the tenth to seventh spatial context dense blocks respectively adopt the SOS enhancement strategy structure to replace the long-distance skip connection method of the traditional U-net; therefore, the decoder realizes the refinement and enhancement processing of the image. Specifically:

[0077]

[0078] In the formula, ψ k (x) represents the output feature map of the decoder, represents the result after processing by the current spatial context dense block in the decoder, and Up() represents upsampling;

[0079] The discriminator adopts a PatchGAN discriminator, which includes three discriminators D1, D2, and D3 with the same structure, and respectively receives the output images from different scales of the decoder;

[0080] Step (3), using the trained cascade image specular highlight removal model based on the conditional generative adversarial network to realize the specular highlight processing of the cascade surface;

[0081] During the model training process, the total loss function consists of three parts: the adversarial loss function, the diffuse reflection component loss function, and the feature matching loss;

[0082] Adversarial loss function is as follows:

[0083]

[0084] In the formula, G represents the generator, and D k represents the output result of each discriminator, and min G () represents taking the minimum of the loss gap between the generator and the discriminator, represents the conditional loss function:

[0085]

[0086] In the formula, E (x,y) represents the minimum consumption cost of deriving image y from image x, and E (x,z) represents the minimum consumption cost of deriving image z from image x.

[0087] Diffuse reflection component loss function As follows:

[0088]

[0089] where z k and y k are the predicted and true diffuse reflection components of the k-scale image respectively, c k represents the number of channels of the k-scale image (RGB 3 channels), w k and h k are the size of the k-scale image (from 256*256 to 32*32); ∑ represents the accumulation of the loss results of the first k images.

[0090] Feature matching loss is defined as follows:

[0091]

[0092] where N i represents the number of elements in the i-th layer of the discriminator, T is the total number of layers, represents the i-th layer feature of the discriminator D k , x k represents the true image corresponding to the k-scale.

[0093] Therefore, the total loss function is as follows:

[0094]

[0095] where λ 1 , λ 2 are the weights of the feature matching loss function and the diffuse reflection loss respectively.

[0096] Experimental results and analysis

[0097] 1) Experimental settings

[0098] The cascaded image highlight removal model based on conditional generative adversarial network is built based on the Pytorch deep learning framework, the programming language is Python 3.7, the network training server configuration is an eight-core Inter CPU I7, and the graphics processor (GPU) uses NVIDIA GTX 2080Ti with a video memory of 20GB. During training, the adaptive momentum estimation optimization algorithm (Adam) is used as the solver, the momentum parameter β 1 is 0.5, β 2 is the default value, the weights are randomly initialized with a Gaussian distribution, the mean is 0, the standard deviation is 0.02, and a total of 200 epochs are trained. The initial learning rate remains unchanged for the first 170 epochs, and the last 30 epochs are linearly decayed to 0. For the weights of the loss weights, after multiple experiments, λ 1 is 10, λ2 is 0.5. Since most traditional algorithms for removing image highlights are based on color space distribution and matrix operation principles, these algorithms do not require a large number of pictures for verification. Therefore, there is currently no large-scale public database for removing highlight gradient images. For this reason, the present invention takes pictures of real ladders to establish a dataset for training and testing. To simulate the highlight effect, a lighting device is used to irradiate the ladders, and the images of ladder objects under highlight irradiation and without highlight irradiation at the same position are collected respectively. The optimal image pairs are selected by the same cutting of the images, and after being uniformly cropped to a size of 512x512, they are grouped according to whether there is highlight, with a total of 2000 pairs of control images.

[0099] Since the model of the present invention adopts a fully convolutional structure, it is applicable to any picture input. To increase the generalization and universality of the network, the present invention also collects other highlight images and performs dataset expansion processing operations on them until a total of 700 groups are used for generalization training and they are analyzed and compared. The present invention conducts experimental analysis on the images from both objective and subjective aspects. In terms of objective evaluation, the peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) are selected as analysis and evaluation indicators. The larger the PSNR, the smaller the distortion; the larger the SSIM, the closer the picture is to the original image. At the same time, the present invention conducts ablation experiment analysis on the proposed module and loss function.

[0100] 2) Result comparison and analysis

[0101] To evaluate the highlight removal effect of the highlight removal conditional generative adversarial network proposed by the present invention, first, the test ladder highlight images are selected for experiments, and comparisons are made with the methods of Yang [1] , Ramos [2] , Yamamoto [3] , Ye [4] , Yi [5] and the model method based on pix2pixHD respectively. The experimental results are as Figure 6 shown.

[0102] Source of prior art literature:

[0103] [1] Yang Qingxiong, Tang Jinhui, Ahuja N. Efficient and robust specular highlight removal[J]. IEEE Trans on Pattern Analysis and Machine Intelligence, 2015, 37(6): 1304 - 1311.

[0104] [2] Ramos V S, Júnior L G D Q S, Silveira L F D Q. Single image high-light removal for real-time image processing pipelines[J]. IEEE Ac-cess, 2019, 8: 3240-3254.

[0105] [3] Yamamoto T, Nakazawa A. General improvement method of specularcomponent separation using high-emphasis filter andsimilarity function[J]. ITETrans on Media Technology and Applications, 2019, 7(2): 92-102.

[0106] [4] Xin Ye, Jia Zhenhong, Yang Jie, etal. Specular reflection imageenhancement based on a dark channel prior[J]. IEEE Photonics Journal, 2021, 13(1): 1-11.

[0107] [5] Yi Renjiao, Tan Ping, Lin S. Leveraging multi-view image sets forunsupervised intrinsic image decomposition and highlight separation[C] / / Procof AAAI Conference on Artificial Intelligence. 2020: 12685-12692.

[0108] pix2pixHD: Wang Tingchun, Liu Mingyu, Zhu Junyan, et al. High-resolution image synthesis and semantic manipulation with conditional GANs[C] / / Proc of IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE Press, 2018: 8798-8807.

[0109] As Figure 6 shown, it can be seen that the traditional algorithms based on color analysis and optimization have very poor effects in actually processing large-area highlights on the single surface of the ladder. The Ramos method has a poor highlight removal method for such images, with a large abnormal pixel area and the highlight area not being restored, indicating defects in the algorithm. The Yi method suppresses the area near the light point to a certain extent but cannot remove the central highlight. The Ye method cannot detect such highlights well, and the processing result has little difference from the input image. The result image of the Pix2pixHD method has serious color distortion and poor highlight removal effect. The model of the present invention is more delicate in processing the texture details of the ladder, retains the ladder texture to a certain extent, has good color preservation, and the highlight removal result is natural without abnormal pixel problems. Table 1 shows the comparison of the average PSNR and SSMI indicators on the ladder dataset of the present invention. It can be seen that the average PSNR of the model of the present invention on the ladder test set leads the other methods by nearly 11 / dB, and the performance of the traditional algorithm indicators is generally poor, further showing the advantages of the model of the present invention in processing ladder highlight images.

[0110] Table 1 Comparison with Different Highlight Removal Methods in the Ladder Dataset

[0111]

[0112] Note: The bold font is the best result in each column.

[0113] Figure 7It is a comparison chart of the results of stepped images collected at a real workstation for different high - light removal methods and the method of the present invention. The Yamamoto method processed the high - light area in the first - row pictures, but the color restoration was not natural. The Yi method had an insignificant effect on high - light removal. At the same time, it can be seen that when dealing with the high - light area of the background metal (the third row), the traditional algorithm seemed powerless. It could neither suppress and remove the high - light on the stepped edge nor effectively eliminate the strong light in the background, and a large amount of noise was generated in the background. Although the method based on pix2pixHD suppressed the high - light to a certain extent, the color was too strong and not realistic. Since the high - intensity reflection area had completely lost its original features, the algorithm could only perform subsequent elimination and restoration as much as possible through the nearby pixel information. However, the model of the present invention, through the dense feature extraction of the SCFDB module, can not only suppress the high - light on the yellow surface of the steps, but also has a good presentation effect on eliminating the high - intensity metal high - light in the background. It shows that it can effectively extract and utilize the information between the pixels and the background. The overall structure of the image is well - preserved, without serious color distortion, and the restoration effect is good enough to meet the requirements of subsequent processing in actual production.

[0114] To verify the general generalization of the network, in the training of the present invention, the classic high - light images were used to expand the dataset and added to the training. Several representative high - light images were selected for comparative experiments, as Figure 8 shown. Analyzing from the visual aspect, the differences between the methods are very small. Among them, in Figure 8 (a), there was more halo residue above the rabbit's ears and the yellow fruit in the Yang method. In the result of high - light removal of the fruit in the Ye method, there was halo residue in the apple and the restoration was relatively blurred. Table 2 shows the comparison of the average PSNR and SSIM values of each method on the classic high - light dataset. It can be seen that the average performance of the pix2pixHD method is the worst, and the method of the present invention has the best performance, indicating that the model also performs very well on such simple - problem images, with a better overall visual display effect and color restoration, and high - quality generated images.

[0115] Table 2 Comparison with different high - light removal methods in the classic high - light dataset

[0116]

[0117] Note: The bold font is the best result in each column.

[0118] In summary, the model of the present invention has great advantages both in terms of visual effects and index comparison, which shows that the deep encoding-decoding structure adopted by the network of the present invention not only effectively utilizes the advantage of the U-shaped symmetric network structure to extract deep information, but also enhances the field of view through the spatial context dense module and multi-scale gradient cascade, fully extracting the texture and background feature information of the image. Therefore, it can effectively generate images of higher quality, effectively restore the surface features, and make the feature restoration closer to the actual picture.

Claims

1. A method for removing highlights from cascade images based on a conditional generative adversarial network, characterized in that the method comprises the following steps: Step (1), constructing a training set By obtaining the image I of the highlight step to be removed H1 and the corresponding image I of the non-highlight step H2 to construct a one-to-one step surface highlight comparison dataset, and perform the same cutting on the images to select the optimal image pair; generate a corresponding clear image I of highlight suppression without any other relevant information assistance S ; Step (2), building a cascade image highlight removal model based on a conditional generative adversarial network, and training it using the training data set after data cutting and processing; The cascade image highlight removal model based on a conditional generative adversarial network includes a generator and a discriminator; The generator adopts the basic architecture of the U-net network, and the U-net network adopts an encoding-decoding structure; In the downsampling process of the encoder, the first to fifth spatial context dense layers Encode-SCFDB in cascade are used; the first to fifth spatial context dense layers Encode-SCFDB are all used to extract the spatial context background information features between feature pixels, and based on multi-path parallel inter-slice convolution to extract and transmit the image pixel background feature information, and obtain the semantic feature information of the image through a deep dense network; the first to fifth spatial context dense layers Encode-SCFDB all include a plurality of spatial context dense blocks in cascade and a first transition layer; the input of each spatial context dense block is the original feature map and the output feature map of the spatial context dense block cascaded in the previous stage in the same spatial context dense layer; The spatial context dense block obtains the spatial context background information between the pixels in each row and each column of the feature map. Specifically, first, the input feature map is applied with a 1X1 convolutional bottleneck layer for dimensionality reduction, InstanceNorm instance normalization, and LeakyReLu activation to generate a new feature map Q. Subsequently, in the SCF layer, the feature map Q is divided into upper and lower branches, and different direction spatial inter-slice convolution strategies of left, right, up, down order and up, down, left, right order are respectively adopted; then, after normalizing and activating the slice information of each path, the two branches are weighted and fused, and then passed through a 3X3 convolution, InstanceNorm instance normalization, and LeakyReLu activation to output the feature map; The decoding layer includes the sixth to tenth spatial context dense layers Decode-SCFDB in cascade, corresponding to the cascaded feature output maps of the encoder at each scale one by one, outputting the feature map after each initial stage cascade enhancement, and sending the feature image into the discriminator at the corresponding scale after a 1x1 convolution; among them, the structure of the tenth spatial context dense layer is the same as that of the first spatial context dense layer, the structure of the ninth spatial context dense layer is the same as that of the second spatial context dense layer, the structure of the eighth spatial context dense layer is the same as that of the third spatial context dense layer, the structure of the seventh spatial context dense layer is the same as that of the fourth spatial context dense layer, and the structure of the sixth spatial context dense layer is the same as that of the fifth spatial context dense layer; The connection between the first to fourth spatial context dense layers Encode-SCFDB and the tenth to seventh spatial context dense blocks Decode-SCFDB adopts the SOS enhancement strategy structure, specifically: where ψ k (x) represents the decoder output feature map, represents the result after processing by the current spatial context dense layer in the decoder, and Up() represents upsampling; During the model training process, the total loss function consists of three parts: the adversarial loss function, the diffuse component loss function, and the feature matching loss; Adversarial loss function As follows: where G represents the generator, D k represents the output result of each discriminator, and min G () indicates taking the minimum of the loss gap between the generator and the discriminator, represents the conditional loss function: where E (x,y) represents the minimum consumption cost derived from image x to image y, and E (x,z) represents the minimum consumption cost derived from image x to image z; Diffuse reflection component loss function As follows: where z k and y k are the predicted and true diffuse reflection components of the k-th scale image respectively, c k represents the number of channels of the k-th scale image, w k and h k are the size of the k-th scale image; Feature matching loss is defined as follows: where N i represents the number of elements in the i-th layer of the discriminator, T is the total number of layers, represents the i-th layer feature of the discriminator D k , and x k represents the real image corresponding to the k-th scale; Therefore, the total loss function is as follows: where λ 1 and λ 2 are the weights of the feature matching loss function and the diffuse reflection loss, respectively; Step (3): Use the trained cascade image specular highlight removal model based on the conditional generative adversarial network to perform specular highlight processing on the cascade surface.

2. The method according to claim 1, wherein the first transition layer uses 1x1 convolution to compress the feature channels, and then performs downsampling operation on the feature map through average pooling, reducing the source input feature map to 1 / 2 size.

3. The method according to claim 1, wherein The first output end of the first spatial context dense layer is connected to the first input end of the first connection layer, and the second output end is connected to the input end of the second spatial context dense layer; the second input end of the first connection layer is connected to the output end of the second 1×1 convolution layer, and the output end is connected to the input end of the first 1×1 convolution layer; the output end of the first 1×1 convolution layer is connected to the input end of the third 3×3 convolution layer; the output end of the third 3×3 convolution layer is connected to the input end of the tenth spatial context dense layer; The first output end of the second spatial context dense layer is connected to the first input end of the second connection layer, and the second output end is connected to the input end of the third spatial context dense layer; the second input end of the second connection layer is connected to the output end of the third 1×1 convolution layer, and the output end is connected to the input end of the first upsampling layer; the output end of the first upsampling layer is connected to the input ends of the second 1×1 convolution layer and the second 3×3 convolution layer; the output end of the second 3×3 convolution layer is connected to the input end of the ninth spatial context dense layer; The first output end of the third spatial context dense layer is connected to the first input end of the third connection layer, and the second output end is connected to the input end of the fourth spatial context dense layer; the second input end of the third connection layer is connected to the output end of the fourth 1×1 convolution layer, and the output end is connected to the input end of the second upsampling layer; the output end of the second upsampling layer is connected to the input ends of the third 1×1 convolution layer and the first 3×3 convolution layer; the output end of the first 3×3 convolution layer is connected to the input end of the eighth spatial context dense layer; The first output end of the fourth spatial context dense layer is connected to the first input end of the fourth connection layer, and the second output end is connected to the input ends of the fifth spatial context dense layer and the seventh spatial context block; the second input end of the fourth connection layer is connected to the output end of the fourth upsampling layer, and the output end is connected to the input end of the third upsampling layer; the output end of the third upsampling layer is connected to the input end of the fourth 1×1 convolution layer.

4. The method according to claim 1, wherein the discriminator uses a PatchGAN discriminator, which includes three discriminators D1, D2, and D3 with the same structure, and respectively receives the output images from different scales of the decoder.

Citation Information

Patent Citations

  • Magic cube image highlight removing method and device based on generative adversarial network

    CN112686819A