An Image Highlight Removal Method and System Based on an Improved Unet++ Network

Through the improved unet++ network and DDCM-Net network structure, a de-highlight network model is built, which solves the problem of removing highlight areas in the image, and achieves efficient de-highlight and information recovery.

CN114998121BActive Publication Date: 2025-06-03HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210537648.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-17
Publication Date
2025-06-03
Estimated Expiration
2042-05-17

AI Technical Summary

Technical Problem

The prior art is difficult to effectively remove the highlight areas in the image, resulting in loss of image detail information and degradation of image quality.

Method used

The improved unet++ network is adopted and combined with the DDCM-Net network structure, an end-to-end de-highlight network model is built, and the predicted high-light mask, highlight layer map and no-highlight map are obtained through training.

Benefits of technology

It realizes the removal of highlight areas in the image, retaining the original information, and restoring the detailed information of the highlight areas, which is suitable for images with complex colors and textures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114998121B_ABST
    Figure CN114998121B_ABST
Patent Text Reader

Abstract

The present invention discloses an image de - specularization method and system based on an improved Unet++ network. The method comprises the following steps: Step (1), constructing a de - specularization network model, inputting a specular image into the de - specularization network model to obtain a predicted specular mask, a specular layer image, and a specular - free image; Step (2), training the de - specularization network model to obtain network model parameters. The present invention subtracts the specular component from the specular image to obtain the final de - specularized image, achieving the purpose of removing specular highlights and restoring image texture details, and having great adaptability and strong robustness to images with complex colors and textures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of digital image processing, and in particular relates to an image highlight removal method and system based on an improved unet++ network. Background Art

[0002] With the emergence and popularization of computers, the image information that humans can obtain is increasing, and the requirements for information transmission speed, processing speed and processing are also getting higher and higher. Images are an important source of information for people to obtain. The presence of highlights in images will cause a large amount of image detail information to be lost, and the image quality will be greatly reduced. Therefore, how to effectively remove highlights is a technical problem that needs to be solved in this field.

[0003] When an object is exposed to strong light, its smooth surface will form local or global specular reflection, which is called highlight phenomenon. This makes the imaged object surface have saturated pixels, destroying the relevant information of the original area. Therefore, when highlights appear on the surface of an object, the texture and color information will be blocked by the highlights, becoming blurred or even disappearing completely. The emergence of highlights has a great impact on the research in the field of computer vision, mainly in the aspects of object positioning, feature analysis, object recognition, edge detection, image restoration, etc. Most traditional highlight removal algorithms are based on the Lambertian reflection assumption, that is, the surface of the object is completely diffusely reflected. However, in real life, when the smooth surface of an object receives light, specular reflection is prone to occur, and bright spots will be formed on the object, resulting in a highlight area, which leads to the loss of structural information on the surface of the imaged object. Summary of the invention

[0004] In view of the above problems existing in the prior art, the present invention provides an image highlight removal method and system based on an improved unet++ network.

[0005] The present invention adopts the following technical scheme:

[0006] An image highlight removal method based on an improved unet++ network includes a training phase, the specific steps of which are:

[0007] Step (1), constructing a highlight removal network model, inputting the highlight image into the highlight removal network model to obtain a predicted highlight mask, a highlight layer image, and a non-highlight image;

[0008] Step (2), train the de-highlight network model to obtain network model parameters.

[0009] Preferably, step (1) is as follows:

[0010] The specular highlight removal network model is an end-to-end network model that introduces a specular highlight formation model in deep learning and uses domain knowledge to assist and guide deep learning. This model consists of two parts: an improved Unet++ network and a DDCM-Net network structure.

[0011] The formula for the specular highlight formation model (the "specular highlight formation model" can be found in the literature "A multi-task network for joint specular highlight detection and removal", Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2021.) is as follows: Where, I represents the original reflected light intensity image, S represents the specular highlight layer image, M represents the specular highlight mask, and D represents the image without specular highlights. represents element-wise multiplication.

[0012] The improved Unet++ network structure has 5 layers, including a convolutional network, an improved ACBlock network, downsampling, upsampling, and skip connections.

[0013] The convolutional network has a total of 15 sub-convolutional networks. The first, second, third, fourth, and fifth layers have 5, 4, 3, 2, and 1 sub-convolutional networks respectively. Each sub-convolutional network has two convolutional blocks A, and each convolutional block A includes a convolutional layer, a normalization layer, and an activation layer. Among them, the convolutional layer uses a 3×3 convolutional kernel, the sliding step size is 1, and the zero-padding size is 1. The activation layer uses the ReLU function.

[0014] The improved ACBlock network (the "ACBlock network" structure can be found in the literature "Acnet: Strengthening the kernel skeletons for powerful cnn via asymmetric convolution blocks", Proceedings of the IEEE / CVF international conference on computer vision, 2019) uses a convolutional block B1 with a one-dimensional horizontal kernel and a convolutional block B2 with a one-dimensional vertical kernel, and sums the outputs of convolutional block B1 and convolutional block B2. Here, both convolutional block B1 and convolutional block B2 include a convolutional layer and a normalization layer. Among them, the kernel size of convolutional block B1 is 1×3, and the kernel size of convolutional block B2 is 3×1. The output size of the improved ACBlock network is the same as the input size.

[0015] The downsampling uses a max pooling layer with a kernel size of 2×2. The input data of the first layer is 256×256×3, and the output size after passing through the first sub-convolutional network of the first layer is 256×256×32. After downsampling to the second, third, fourth, and fifth layers, the output sizes are 128×128×64, 64×64×128, 32×32×256, and 16×16×512 respectively.

[0016] The upsampling uses a convolutional kernel with a kernel size of 2×2. The output size of the sub-convolutional network of the fifth layer is 16×16×512, and after upsampling to the fourth, third, second, and first layers, the output sizes are 32×32×512, 64×64×256, 128×128×128, and 256×256×64 respectively.

[0017] The skip-layer connection means that the input of each sub-convolutional network in each layer is the concatenation and fusion of the outputs of all previous sub-convolutional networks and the upsampled output features of the next-layer sub-convolutional network, and then the current features are passed to the subsequent network. The input size of the first sub-convolutional network in the first layer is 256×256×3, and the output size is 256×256×32. The input size of the second sub-convolutional network in the first layer is 256×256×96, and the output size is 256×256×32. The input size of the third sub-convolutional network in the first layer is 256×256×128, and the output size is 256×256×32. The input size of the fourth sub-convolutional network in the first layer is 256×256×160, and the output size is 256×256×32. The input size of the fifth sub-convolutional network in the first layer is 256×256×192, and the output size is 256×256×32. The input size of the first sub-convolutional network in the second layer is 128×128×64, and the output size is 128×128×64. The input size of the second sub-convolutional network in the second layer is 128×128×192, and the output size is 128×128×64. The input size of the third sub-convolutional network in the second layer is 128×128×256, and the output size is 128×128×64. The input size of the fourth sub-convolutional network in the second layer is 128×128×320, and the output size is 128×128×64. The input size of the first sub-convolutional network in the third layer is 64×64×128, and the output size is 64×64×128. The input size of the second sub-convolutional network in the third layer is 64×64×384, and the output size is 64×64×128. The input size of the third sub-convolutional network in the third layer is 64×64×512, and the output size is 64×64×128. The input size of the first sub-convolutional network in the fourth layer is 32×32×256, and the output size is 32×32×256. The input size of the second sub-convolutional network in the fourth layer is 32×32×768, and the output size is 32×32×256. Finally, the output of the fifth sub-convolutional network in the first layer passes through a 1×1 convolution and a sigmoid function in sequence to obtain a feature map with a size of 256×256×3.

[0018] The DDCM-Net network has a total of three (see the "DDCM-Net network" structure in the literature "Dense dilatedconvolutions merging network for semantic mapping of remote sensing images", 2019 Joint Urban Remote Sensing Event (JURSE). IEEE, 2019). The input size of the first DDCM-Net is 256×256×3, and the output size is 256×256×3. The input size of the second DDCM-Net is 256×256×6, and the output size is 256×256×3. The input size of the third DDCM-Net is 256×256×9, and the output size is 256×256×3. DDCM-Net integrates local and global context information in the image, which helps to obtain rich features.

[0019] The highlight map is input into the improved unet++ model to obtain the feature map, and then the DDCM-Net is used to obtain the predicted highlight mask, highlight layer map, and non-highlight map in sequence.

[0020] Preferably, step (2) is specifically as follows:

[0021] Step (2-1), obtaining the training dataset: Normalize the training image element values to between 0 and 1.

[0022] Step (2-2), training the network parameters: The initial learning rate is set to 1×10 -8 , and the momentum value is set to 0.95, and the weight decay coefficient is set to 5×10 -4 . The value updated for each network parameter in the model is: the current gradient value × the learning rate + the value after the previous update × the momentum value. Define the loss function as: L = λ s L MSE (S, S′) + λ m L MSE (M, M′) + λ d L MSE (D, D′) + λ k L k (D, D′); where λ s , λ m , λ d , λ k are regularization coefficients, S and S′ are the real and predicted highlight layer maps respectively, M and M′ are the real and predicted highlight masks respectively, D and D′ are the real and predicted non-highlight maps respectively, and L MSE represents the mean square error loss function, and L kIt is expressed as a perceptual loss function ("perceptual loss" can be found in the literature "Perceptual losses for real-time style transfer and super-resolution", European conference on computer vision. Springer, Cham, 2016).

[0023] The training optimizer uses the gradient descent method to end after training and iterating N_1 times on the training dataset. The batch size is B_1, where 3000 ≤ N_1 ≤ 5000 and 4 ≤ B_1 ≤ 20, to obtain the network model parameters.

[0024] Preferably, it further includes step (3), the test phase, and the specific steps of this phase are as follows:

[0025] Step (3-1), normalizing the test image element values to between 0 and 1.

[0026] Step (3-2), inputting the test data into the model to obtain the output data, and then converting it into an RGB format image, which is the predicted image without highlights.

[0027] The present invention also discloses an image dehighlighting system based on an improved unet++ network, which includes the following modules:

[0028] Model construction module: used to construct a dehighlighting network model, input the highlight map into the dehighlighting network model to obtain the predicted highlight mask, highlight layer map, and image without highlights;

[0029] Training module: used to train the dehighlighting network model to obtain the network model parameters.

[0030] Preferably, the model construction module is specifically as follows:

[0031] The dehighlighting network model is an end-to-end network model. A highlight formation model is introduced in deep learning, and domain knowledge is used to assist and guide deep learning; this model includes two parts: an improved unet++ network and a DDCM-Net network structure;

[0032] The formula for the highlight formation model is: Where, I represents the original reflected light intensity image, S represents the highlight layer map, M represents the highlight mask, and D represents the image without highlights, represents element-wise multiplication;

[0033] The improved unet++ network structure has 5 layers, including a convolutional network, an improved ACBlock network, downsampling, upsampling, and skip connections;

[0034] The convolutional network described above has a total of 15 sub-convolutional networks. The first, second, third, fourth, and fifth layers have 5, 4, 3, 2, and 1 sub-convolutional networks respectively. Each sub-convolutional network has two convolutional blocks A, and each convolutional block A includes a convolutional layer, a normalization layer, and an activation layer. Among them, the convolutional layer uses a 3×3 convolutional kernel, a sliding step of 1, and a zero-padding size of 1; the activation layer uses the ReLU function.

[0035] The improved ACBlock network uses a convolutional block B1 with a one-dimensional horizontal kernel and a convolutional block B2 with a one-dimensional vertical kernel, and sums the outputs of the convolutional block B1 and the convolutional block B2. Here, both the convolutional block B1 and the convolutional block B2 include a convolutional layer and a normalization layer. Among them, the kernel size of the convolutional block B1 is 1×3, and the kernel size of the convolutional block B2 is 3×1; the output size of the improved ACBlock network is the same as the input size.

[0036] The downsampling uses a max pooling layer with a kernel size of 2×2; the input data of the first layer is 256×256×3, and the output size after passing through the first sub-convolutional network of the first layer is 256×256×32. The output sizes after downsampling to the second, third, fourth, and fifth layers are 128×128×64, 64×64×128, 32×32×256, and 16×16×512 respectively.

[0037] The upsampling uses a convolutional kernel with a kernel size of 2×2; the output size of the sub-convolutional network of the fifth layer is 16×16×512, and the output sizes after upsampling to the fourth, third, second, and first layers are 32×32×512, 64×64×256, 128×128×128, and 256×256×64 respectively.

[0038] The skip-layer connection means that the input of each sub-convolutional network in each layer is the concatenated fusion of the outputs of all previous sub-convolutional networks and the upsampled output features of the next-layer sub-convolutional network, and then the current features are passed to the subsequent network; the input size of the first sub-convolutional network in the first layer is 256×256×3, and the output size is 256×256×32; the input size of the second sub-convolutional network in the first layer is 256×256×96, and the output size is 256×256×32; the input size of the third sub-convolutional network in the first layer is 256×256×128, and the output size is 256×256×32; the input size of the fourth sub-convolutional network in the first layer is 256×256×160, and the output size is 256×256×32; the input size of the fifth sub-convolutional network in the first layer is 256×256×192, and the output size is 256×256×32; the input size of the first sub-convolutional network in the second layer is 128×128×64, and the output size is 128×128×64; the input size of the second sub-convolutional network in the second layer is 128×128×192, and the output size is 128×128×64; the input size of the third sub-convolutional network in the second layer is 128×128×256, and the output size is 128×128×64; the input size of the fourth sub-convolutional network in the second layer is 128×128×320, and the output size is 128×128×64; the input size of the first sub-convolutional network in the third layer is 64×64×128, and the output size is 64×64×128; the input size of the second sub-convolutional network in the third layer is 64×64×384, and the output size is 64×64×128; the input size of the third sub-convolutional network in the third layer is 64×64×512, and the output size is 64×64×128; the input size of the first sub-convolutional network in the fourth layer is 32×32×256, and the output size is 32×32×256; the input size of the second sub-convolutional network in the fourth layer is 32×32×768, and the output size is 32×32×256; finally, the output of the fifth sub-convolutional network in the first layer passes through a 1×1 convolution and a sigmoid function in sequence to obtain a feature map with a size of 256×256×3;

[0039] There are three DDCM-Net networks in total; the input size of the first DDCM-Net is 256×256×3, and the output size is 256×256×3, the input size of the second DDCM-Net is 256×256×6, and the output size is 256×256×3, the input size of the third DDCM-Net is 256×256×9, and the output size is 256×256×3.

[0040] Preferably, the training model specifically includes:

[0041] A training dataset acquisition sub-module: normalizing the values of training image elements to between 0 and 1;

[0042] A network parameter training sub-module: setting the initial learning rate to 1×10 -8 , setting the momentum value to 0.95, and setting the weight decay coefficient to 5×10 -4 ; the updated value of each network parameter in the highlight removal network model is: the current gradient value × the learning rate + the value after the previous update × the momentum value; defining the loss function as: L = λ s L MSE (S, S′) + λ m L MSE (M, M′) + λ d L MSE (D, D′) + λ k L k (D, D′); where λ s 、λ m 、λ d 、λ k are regularization coefficients, S and S′ are the real and predicted highlight layer maps respectively, M and M′ are the real and predicted highlight masks respectively, D and D′ are the real and predicted non-highlight images respectively, L MSE represents the mean square error loss function, and L k represents the perceptual loss function;

[0043] The training optimizer uses the gradient descent method to train and iterate N_1 times on the training dataset and then ends. The batch size is B_1, 3000 ≤ N_1 ≤ 5000, 4 ≤ B_1 ≤ 20, to obtain the network model parameters.

[0044] Preferably, it further includes a test module. The test module first normalizes the values of test image elements to between 0 and 1; then inputs the test data into the highlight removal network model to obtain output data, and converts it into an RGB format image, which is the predicted non-highlight image.

[0045] The present invention proposes an image highlight removal method and system based on an improved unet++ network, which can remove the highlight area in the image, retain the original information, and restore the highlight area information according to the surrounding information of the highlight. The present invention subtracts the specular component from the highlight map to obtain the final highlight-removed image, achieving the purpose of removing highlights and restoring the texture details of the image, and having strong adaptability and robustness for images with complex colors and textures. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 is the highlight removal network model;

[0047] Figure 2 is the structural diagram of the improved ACBlock network;

[0048] Figure 3 is the system block diagram of an image de - specularization system based on the improved unet++ network of the present invention. Specific implementation manners

[0049] The present invention will be described in detail below in conjunction with the accompanying drawings and embodiments, so that those skilled in the art can better understand the present invention. It should be particularly noted that in the following description, when the detailed description of known functions and designs may obscure the main content of the present invention, these descriptions will be omitted here.

[0050] Embodiment 1

[0051] As Figure 1 - Figure 2 shown, an image de - specularization method based on the improved unet++ network in this embodiment includes a training stage and a testing stage.

[0052] The specific steps of the training stage are as follows:

[0053] Step (1), construct a de - specularization network model:

[0054] The de - specularization network model is an end - to - end network model. A specular highlight formation model is introduced in deep learning, and domain knowledge is used to assist and guide deep learning; this model includes two parts: an improved unet++ network and a DDCM - Net network;

[0055] The formula of the specular highlight formation model is: where I represents the original reflected light intensity image, S represents the specular highlight layer image, M represents the specular highlight mask, D represents the image without specular highlights, represents element - wise multiplication.

[0056] The depth of the improved unet++ network structure is 5 layers, which includes a convolutional network, an improved ACBlock network, downsampling, upsampling, and skip connections. The specific introduction is as follows:

[0057] The convolutional network has a total of 15 sub - convolutional networks. The first, second, third, fourth, and fifth layers have 5, 4, 3, 2, and 1 sub - convolutional networks respectively. Each sub - convolutional network has two convolutional blocks A, and each convolutional block A includes a convolutional layer, a normalization layer, and an activation layer; among them, the convolutional layer uses a 3×3 size convolutional kernel, the sliding step size is 1, and the zero - padding size at the edge is 1; the activation layer uses the ReLU function.

[0058] The improved ACBlock network (for the "ACBlock network" structure, see the literature "Acnet: Strengthening the kernel skeletons for powerful cnn via asymmetric convolution blocks", Proceedings of the IEEE / CVF international conference on computer vision, 2019) uses a convolution block B1 with a one-dimensional horizontal kernel and a convolution block B2 with a one-dimensional vertical kernel, and sums the outputs of the convolution block B1 and the convolution block B2. Here, both the convolution block B1 and the convolution block B2 include a convolutional layer and a normalization layer. Among them, the kernel size of the convolution block B1 is 1×3, and the kernel size of the convolution block B2 is 3×1. The output size of the improved ACBlock network is the same as the input size.

[0059] Max pooling layer with a kernel size of 2×2 is used for downsampling. The input data of the first layer is 256×256×3, and the output size after passing through the first sub-convolutional network of the first layer is 256×256×32. The output sizes after downsampling to the second, third, fourth, and fifth layers are 128×128×64, 64×64×128, 32×32×256, and 16×16×512 respectively.

[0060] Convolution kernel with a kernel size of 2×2 is used for upsampling. The output size of the sub-convolutional network of the fifth layer is 16×16×512, and the output sizes after upsampling to the fourth, third, second, and first layers are 32×32×512, 64×64×256, 128×128×128, and 256×256×64 respectively.

[0061] Skip-layer connection means that the input of each sub-convolutional network in each layer is the concatenation and fusion of the outputs of all previous sub-convolutional networks and the upsampled output features of the next-layer sub-convolutional network, and then the current features are passed to the subsequent network. The input size of the first sub-convolutional network in the first layer is 256×256×3, and the output size is 256×256×32. The input size of the second sub-convolutional network in the first layer is 256×256×96, and the output size is 256×256×32. The input size of the third sub-convolutional network in the first layer is 256×256×128, and the output size is 256×256×32. The input size of the fourth sub-convolutional network in the first layer is 256×256×160, and the output size is 256×256×32. The input size of the fifth sub-convolutional network in the first layer is 256×256×192, and the output size is 256×256×32. The input size of the first sub-convolutional network in the second layer is 128×128×64, and the output size is 128×128×64. The input size of the second sub-convolutional network in the second layer is 128×128×192, and the output size is 128×128×64. The input size of the third sub-convolutional network in the second layer is 128×128×256, and the output size is 128×128×64. The input size of the fourth sub-convolutional network in the second layer is 128×128×320, and the output size is 128×128×64. The input size of the first sub-convolutional network in the third layer is 64×64×128, and the output size is 64×64×128. The input size of the second sub-convolutional network in the third layer is 64×64×384, and the output size is 64×64×128. The input size of the third sub-convolutional network in the third layer is 64×64×512, and the output size is 64×64×128. The input size of the first sub-convolutional network in the fourth layer is 32×32×256, and the output size is 32×32×256. The input size of the second sub-convolutional network in the fourth layer is 32×32×768, and the output size is 32×32×256. Finally, the output of the fifth sub-convolutional network in the first layer passes through a 1×1 convolution and a sigmoid function in sequence to obtain a feature map with a size of 256×256×3.

[0062] The DDCM-Net network has three in total (for the "DDCM-Net network" structure, see the literature "Dense dilatedconvolutions merging network for semantic mapping of remote sensing images", 2019 Joint Urban Remote Sensing Event (JURSE). IEEE, 2019). The input size of the first DDCM-Net is 256×256×3, and the output size is 256×256×3. The input size of the second DDCM-Net is 256×256×6, and the output size is 256×256×3. The input size of the third DDCM-Net is 256×256×9, and the output size is 256×256×3. DDCM-Net integrates local and global context information in the image, which helps to obtain rich features.

[0063] The highlight map is input into the improved unet++ model to obtain the feature map, and then the DDCM-Net is used to obtain the predicted highlight mask, highlight layer map, and non-highlight map in sequence.

[0064] In this embodiment, the improved unet++ network is first used to fuse features at different levels and retain edge detail information to perform preliminary feature extraction on the image to obtain the feature map. Then, the DDCM-Net is used to fuse local and global context information in the image to obtain the predicted highlight mask, highlight layer map, and non-highlight map in sequence.

[0065] Step (2), training the highlight removal network model:

[0066] Step (2-1), obtaining the training data set: Normalize the training image element values to between 0 and 1.

[0067] Step (2-2), training the network parameters: Set the initial learning rate to 1×10 -8 , and set the momentum value to 0.95 and the weight decay coefficient to 5×10 -4 . The value updated for each network parameter in the model is: the current gradient value × the learning rate + the value after the previous update × the momentum value. Define the loss function as: L = λ s L MSE (S, S′) + λ m L MSE (M, M′) + λ d L MSE (D, D′) + λ k L k (D, D′); where λ s 、λ m 、λd , λ k is the regularization coefficient, S and S' are the real and predicted highlight layer maps respectively, M and M' are the real and predicted highlight masks respectively, D and D' are the real and predicted non-highlight maps respectively, and L MSE is expressed as the mean square error loss function, and L k is expressed as the perceptual loss function (for "perceptual loss", see the literature "Perceptual losses for real-time style transfer and super-resolution", European conference on computer vision. Springer, Cham, 2016).

[0068] The training optimizer uses the gradient descent method to train and iterate N_1 times on the training dataset. The batch size is B_1, N_1 = 4000, B_1 = 10, and the network model parameters are obtained.

[0069] Step (3), the testing phase. In this phase, the test image is input into the model to obtain the predicted non-highlight image, specifically as follows:

[0070] Step (3-1), normalize the test image element values to between 0 and 1.

[0071] Step (3-2), input the test data into the model to obtain the output data, and then convert it into an RGB format image, which is the predicted non-highlight image.

[0072] Embodiment 2

[0073] As Figure 3 shown, an image de-highlighting system based on an improved unet++ network in this embodiment includes a model construction module, a training module, and a testing module.

[0074] The model construction module is specifically as follows:

[0075] The de-highlighting network model is an end-to-end network model. A highlight formation model is introduced in deep learning, and domain knowledge is used to assist and guide deep learning; this model includes two parts: an improved unet++ network and a DDCM-Net network;

[0076] The highlight formation model formula is: where I represents the original reflected light intensity image, S represents the highlight layer map, M represents the highlight mask, D represents the non-highlight map, represents element-wise multiplication.

[0077] The improved U-Net++ network structure has 5 layers, including a convolutional network, an improved ACBlock network, downsampling, upsampling, and skip connections, which are introduced as follows:

[0078] The convolutional network has a total of 15 sub-convolutional networks. The first, second, third, fourth, and fifth layers have 5, 4, 3, 2, and 1 sub-convolutional networks respectively. Each sub-convolutional network has two convolutional blocks A, and each convolutional block A includes a convolutional layer, a normalization layer, and an activation layer. Among them, the convolutional layer uses a 3×3 convolutional kernel, the sliding stride is 1, and the zero-padding size is 1. The activation layer uses the ReLU function.

[0079] The improved ACBlock network (for the "ACBlock network" structure, see the literature "Acnet: Strengthening the kernel skeletons for powerful cnn via asymmetric convolution blocks", Proceedings of the IEEE / CVF international conference on computer vision, 2019) uses a convolutional block B1 with a one-dimensional horizontal kernel and a convolutional block B2 with a one-dimensional vertical kernel, and sums the outputs of convolutional block B1 and convolutional block B2. Here, both convolutional block B1 and convolutional block B2 include a convolutional layer and a normalization layer. Among them, the kernel size of convolutional block B1 is 1×3, and the kernel size of convolutional block B2 is 3×1. The output size of the improved ACBlock network is the same as the input size.

[0080] The downsampling uses a max-pooling layer with a kernel size of 2×2. The input data of the first layer is 256×256×3. After passing through the first sub-convolutional network of the first layer, the output size is 256×256×32. After downsampling to the second, third, fourth, and fifth layers, the output sizes are 128×128×64, 64×64×128, 32×32×256, and 16×16×512 respectively.

[0081] The upsampling uses a convolutional kernel with a kernel size of 2×2. The output size of the sub-convolutional network of the fifth layer is 16×16×512. After upsampling to the fourth, third, second, and first layers, the output sizes are 32×32×512, 64×64×256, 128×128×128, and 256×256×64 respectively.

[0082] Skip layer connection means that the input of each sub-convolutional network in each layer is the concatenated fusion of the outputs of all previous sub-convolutional networks and the upsampled output features of the next layer's sub-convolutional network, and then the current features are passed to the subsequent network. The input size of the first sub-convolutional network in the first layer is 256×256×3, and the output size is 256×256×32. The input size of the second sub-convolutional network in the first layer is 256×256×96, and the output size is 256×256×32. The input size of the third sub-convolutional network in the first layer is 256×256×128, and the output size is 256×256×32. The input size of the fourth sub-convolutional network in the first layer is 256×256×160, and the output size is 256×256×32. The input size of the fifth sub-convolutional network in the first layer is 256×256×192, and the output size is 256×256×32. The input size of the first sub-convolutional network in the second layer is 128×128×64, and the output size is 128×128×64. The input size of the second sub-convolutional network in the second layer is 128×128×192, and the output size is 128×128×64. The input size of the third sub-convolutional network in the second layer is 128×128×256, and the output size is 128×128×64. The input size of the fourth sub-convolutional network in the second layer is 128×128×320, and the output size is 128×128×64. The input size of the first sub-convolutional network in the third layer is 64×64×128, and the output size is 64×64×128. The input size of the second sub-convolutional network in the third layer is 64×64×384, and the output size is 64×64×128. The input size of the third sub-convolutional network in the third layer is 64×64×512, and the output size is 64×64×128. The input size of the first sub-convolutional network in the fourth layer is 32×32×256, and the output size is 32×32×256. The input size of the second sub-convolutional network in the fourth layer is 32×32×768, and the output size is 32×32×256. Finally, the output of the fifth sub-convolutional network in the first layer passes through a 1×1 convolution and a sigmoid function in sequence to obtain a feature map with a size of 256×256×3.

[0083] The DDCM-Net network has a total of three (for the "DDCM-Net network" structure, see the literature "Dense dilated convolutions merging network for semantic mapping of remote sensing images", 2019 Joint Urban Remote Sensing Event (JURSE). IEEE, 2019). The input size of the first DDCM-Net is 256×256×3, and the output size is 256×256×3. The input size of the second DDCM-Net is 256×256×6, and the output size is 256×256×3. The input size of the third DDCM-Net is 256×256×9, and the output size is 256×256×3. DDCM-Net fuses local and global context information in the image, which helps to obtain rich features.

[0084] The highlight map is input into the improved unet++ model to obtain the feature map, and then the DDCM-Net is used to obtain the predicted highlight mask, highlight layer map, and non-highlight map in sequence.

[0085] The training module specifically includes:

[0086] Obtain training dataset sub-module: Normalize the training image element values to between 0 and 1.

[0087] Train network parameter sub-module: The initial learning rate is set to 1×10 -8 , and the momentum value is set to 0.95, and the weight decay coefficient is set to 5×10 -4 . The value updated for each network parameter in the model is: the current gradient value × the learning rate + the value after the previous update × the momentum value. Define the loss function as:

[0088] L = λ s L MSE (S, S′) + λ m L MSE (M, M′) + λ d L MSE (D, D′) + λ k L k (D, D′); where λ s , λ m , λ d , λ k are regularization coefficients, S and S′ are the real and predicted highlight layer maps respectively, M and M′ are the real and predicted highlight masks respectively, D and D′ are the real and predicted non-highlight maps respectively, and L MSE represents the mean square error loss function, Lk It is represented as a perceptual loss function (for "perceptual loss", see the literature "Perceptual losses for real-time style transfer and super-resolution", European conference on computer vision. Springer, Cham, 2016).

[0089] The training optimizer uses the gradient descent method to end after training and iterating N_1 times on the training dataset. The batch size is B_1, N_1 = 4000, B_1 = 10, to obtain the network model parameters.

[0090] The test module is specifically as follows:

[0091] Normalize the test image element values to between 0 and 1.

[0092] Input the test data into the model to obtain the output data, and then convert it into an RGB format image, which is the predicted image without highlights.

[0093] The model adopted in the present invention realizes the conversion from a highlighted image to an image without highlights in an end-to-end manner, has a good highlight removal effect, and better retains information such as color and texture.

[0094] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the art of the present invention can make various modifications or supplements to the described specific embodiments or use similar methods to replace them, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.

Claims

1. An image de - highlighting method based on an improved Unet++ network, characterized in that it is carried out according to the following steps: Step (1), construct a de - highlighting network model, input the highlighted image into the de - highlighting network model to obtain the predicted highlight mask, highlight layer map, and non - highlighted image; Step (2), train the de - highlighting network model to obtain network model parameters; Step (1) is specifically as follows: The de - highlighting network model is an end - to - end network model. A highlight formation model is introduced in deep learning, and domain knowledge is used to assist and guide deep learning. The de - highlighting network model includes two parts: an improved Unet++ network and a DDCM - Net network structure; The formula for the highlight formation model is as follows: where I represents the original reflected light intensity image, S represents the highlight layer image, M represents the highlight mask, and D represents the non-highlight image, denotes element-wise multiplication; The improved Unet++ network structure has 5 layers, including a convolutional network, an improved ACBlock network, downsampling, upsampling, and skip connections; The convolutional network has a total of 15 sub - convolutional networks. The first, second, third, fourth, and fifth layers have 5, 4, 3, 2, and 1 sub - convolutional networks respectively. Each sub - convolutional network has two convolutional blocks A, and each convolutional block A includes a convolutional layer, a normalization layer, and an activation layer. Among them, the convolutional layer uses a 3×3 convolutional kernel, the sliding step size is 1, and the zero - padding size at the edge is 1; the activation layer uses the ReLU function; The improved ACBlock network uses a convolutional block B1 with a one - dimensional horizontal kernel and a convolutional block B2 with a one - dimensional vertical kernel, and sums the outputs of convolutional block B1 and convolutional block B2. Here, both convolutional block B1 and convolutional block B2 include a convolutional layer and a normalization layer. Among them, the kernel size of convolutional block B1 is 1×3, and the kernel size of convolutional block B2 is 3×1; the output size of the improved ACBlock network is the same as the input size; The downsampling uses a max - pooling layer with a kernel size of 2×2. The input data of the first layer is 256×256×3. After passing through the first sub - convolutional network of the first layer, the output size is 256×256×32. After downsampling to the second, third, fourth, and fifth layers, the output sizes are 128×128×64, 64×64×128, 32×32×256, and 16×16×512 respectively; The upsampling uses a convolutional kernel with a kernel size of 2×2. The output size of the sub - convolutional network of the fifth layer is 16×16×512. After upsampling to the fourth, third, second, and first layers, the output sizes are 32×32×512, 64×64×256, 128×128×128, and 256×256×64 respectively; The so-called skip-layer connection means that the input of each sub-convolutional network in each layer is the concatenated fusion of the outputs of all previous sub-convolutional networks and the upsampled output features of the next-layer sub-convolutional network, and then the current features are passed to the subsequent network; the input size of the first sub-convolutional network in the first layer is 256×256×3, and the output size is 256×256×32; the input size of the second sub-convolutional network in the first layer is 256×256×96, and the output size is 256×256×32; the input size of the third sub-convolutional network in the first layer is 256×256×128, and the output size is 256×256×32; the input size of the fourth sub-convolutional network in the first layer is 256×256×160, and the output size is 256×256×32; the input size of the fifth sub-convolutional network in the first layer is 256×256×192, and the output size is 256×256×32; the input size of the first sub-convolutional network in the second layer is 128×128×64, and the output size is 128×128×64; the input size of the second sub-convolutional network in the second layer is 128×128×192, and the output size is 128×128×64; the input size of the third sub-convolutional network in the second layer is 128×128×256, and the output size is 128×128×64; the input size of the fourth sub-convolutional network in the second layer is 128×128×320, and the output size is 128×128×64; the input size of the first sub-convolutional network in the third layer is 64×64×128, and the output size is 64×64×128; the input size of the second sub-convolutional network in the third layer is 64×64×384, and the output size is 64×64×128; the input size of the third sub-convolutional network in the third layer is 64×64×512, and the output size is 64×64×128; the input size of the first sub-convolutional network in the fourth layer is 32×32×256, and the output size is 32×32×256; the input size of the second sub-convolutional network in the fourth layer is 32×32×768, and the output size is 32×32×256; finally, the output of the fifth sub-convolutional network in the first layer passes through a 1×1 convolution and a sigmoid function in sequence to obtain a feature map with a size of 256×256×3; There are three DDCM-Net networks in total; the input size of the first DDCM-Net is 256×256×3, and the output size is 256×256×3; the input size of the second DDCM-Net is 256×256×6, and the output size is 256×256×3; the input size of the third DDCM-Net is 256×256×9, and the output size is 256×256×3; Step (2) is specifically as follows: Step (2-1), obtain the training data set: Normalize the training image element values to between 0 and 1; Step (2-2), training network parameters: The initial learning rate is set to 1×10 -8 , the momentum value is set to 0.95, and the weight decay coefficient is set to 5×10 -4 ; The updated value of each network parameter in the highlight removal network model is: the current gradient value × learning rate + the value after the previous update × momentum value; The loss function is defined as: L = λ s L MSE (S, S′) + λ m L MSE (M, M′) + λ d L MSE (D, D′) + λ k L k (D, D′); where λ s , λ m , λ d , λ k are regularization coefficients, S and S′ are the real and predicted highlight layer maps respectively, M and M′ are the real and predicted highlight masks respectively, D and D′ are the real and predicted non-highlight maps respectively, and L MSE represents the mean square error loss function, and L k represents the perceptual loss function; The training optimizer uses the gradient descent method to end after training and iterating N_1 times on the training dataset. The batch size is B_1, where 3000 ≤ N_1 ≤ 5000 and 4 ≤ B_1 ≤ 20, to obtain the network model parameters.

2. A method for removing highlights from images based on an improved unet++ network as claimed in claim 1, characterized in that, it further includes step (3), the testing phase, and the specific steps are as follows: Step (3-1), normalizing the test image element values to between 0 and 1; Step (3-2), inputting the test data into the highlight removal network model to obtain output data, and then converting it into an RGB format image, which is the predicted highlight-free image.

3. An image highlight removal system based on an improved unet++ network, characterized in that, it includes the following modules: Model construction module: used to construct a highlight removal network model, input a highlight map into the highlight removal network model to obtain a predicted highlight mask, a highlight layer map, and a highlight-free map; Training module: used to train the highlight removal network model to obtain network model parameters; The model construction module is specifically as follows: The highlight removal network model is an end-to-end network model. A highlight formation model is introduced in deep learning, and domain knowledge is used to assist and guide deep learning; the highlight removal network model includes two parts: an improved unet++ network and a DDCM-Net network structure; The formula for the highlight formation model is as follows: where I represents the original reflected light intensity image, S represents the highlight layer map, M represents the highlight mask, and D represents the non-highlight image, represents element-wise multiplication; The improved unet++ network structure has 5 layers, including a convolutional network, an improved ACBlock network, downsampling, upsampling, and skip connections; The convolutional network has a total of 15 sub-convolutional networks. The first, second, third, fourth, and fifth layers have 5, 4, 3, 2, and 1 sub-convolutional networks respectively. Each sub-convolutional network has two convolutional blocks A, and each convolutional block A includes a convolutional layer, a normalization layer, and an activation layer; among them, the convolutional layer uses a 3×3 size convolutional kernel, the sliding step size is 1, and the zero-padding size at the edges is 1; the activation layer uses the ReLU function; The improved ACBlock network uses a convolutional block B1 with a one-dimensional horizontal kernel and a convolutional block B2 with a one-dimensional vertical kernel, and sums the outputs of the convolutional block B1 and the convolutional block B2. Here, both the convolutional block B1 and the convolutional block B2 include a convolutional layer and a normalization layer; among them, the kernel size of the convolutional block B1 is 1×3, and the kernel size of the convolutional block B2 is 3×1; the output size of the improved ACBlock network is the same as the input size; The downsampling uses a max pooling layer with a kernel size of 2×2; the input data of the first layer is 256×256×3, and the output size after passing through the first sub-convolutional network of the first layer is 256×256×32. After downsampling to the second, third, fourth, and fifth layers, the output sizes are 128×128×64, 64×64×128, 32×32×256, and 16×16×512 respectively; The upsampling uses a convolutional kernel with a kernel size of 2×2; the output size of the sub-convolutional network in the fifth layer is 16×16×512, and after upsampling, the output sizes to the fourth, third, second, and first layers are 32×32×512, 64×64×256, 128×128×128, and 256×256×64 respectively; The skip connection mentioned refers to that the input of each sub-convolutional network in each layer is the concatenation and fusion of the outputs of all previous sub-convolutional networks and the upsampled output features of the sub-convolutional network in the next layer, and then the current features are passed to the subsequent network; the input size of the first sub-convolutional network in the first layer is 256×256×3, and the output size is 256×256×32; the input size of the second sub-convolutional network in the first layer is 256×256×96, and the output size is 256×256×32; the input size of the third sub-convolutional network in the first layer is 256×256×128, and the output size is 256×256×32; the input size of the fourth sub-convolutional network in the first layer is 256×256×160, and the output size is 256×256×32; the input size of the fifth sub-convolutional network in the first layer is 256×256×192, and the output size is 256×256×32; the input size of the first sub-convolutional network in the second layer is 128×128×64, and the output size is 128×128×64; the input size of the second sub-convolutional network in the second layer is 128×128×192, and the output size is 128×128×64; the input size of the third sub-convolutional network in the second layer is 128×128×256, and the output size is 128×128×64; the input size of the fourth sub-convolutional network in the second layer is 128×128×320, and the output size is 128×128×64; the input size of the first sub-convolutional network in the third layer is 64×64×128, and the output size is 64×64×128; the input size of the second sub-convolutional network in the third layer is 64×64×384, and the output size is 64×64×128; the input size of the third sub-convolutional network in the third layer is 64×64×512, and the output size is 64×64×128; the input size of the first sub-convolutional network in the fourth layer is 32×32×256, and the output size is 32×32×256; the input size of the second sub-convolutional network in the fourth layer is 32×32×768, and the output size is 32×32×256; finally, the output of the fifth sub-convolutional network in the first layer passes through a 1×1 convolution and a sigmoid function in sequence to obtain a feature map with a size of 256×256×3; There are three DDCM-Net networks in total. The input size of the first DDCM-Net is 256×256×3, and the output size is 256×256×3. The input size of the second DDCM-Net is 256×256×6, and the output size is 256×256×3. The input size of the third DDCM-Net is 256×256×9, and the output size is 256×256×3; The training model specifically includes: Training dataset acquisition sub-module: Normalize the training image element values to between 0 and 1; Training network parameter sub-module: The initial learning rate is set to 1×10 -8 , and the impulse value is set to 0.95, and the weight decay coefficient is set to 5×10 -4 ; The updated value of each network parameter in the highlight removal network model is: the current gradient value × learning rate + the value after the previous update × impulse value; The loss function is defined as: L = λ s L MSE (S, S′) + λ m L MSE (M, M′) + λ d L MSE (D, D′) + λ k L k (D, D′); where λ s 、λ m 、λ d 、λ k are regularization coefficients, S and S′ are the true and predicted specular layer maps respectively, M and M′ are the true and predicted specular masks respectively, D and D′ are the true and predicted specular-free maps respectively, and L MSE represents the mean squared error loss function, and L k represents the perceptual loss function; The training optimizer uses the gradient descent method to train and iterate N_1 times on the training dataset and then ends. The batch size is B_1, where 3000 ≤ N_1 ≤ 5000 and 4 ≤ B_1 ≤ 20, to obtain the network model parameters.

4. An image de-highlighting system based on the improved unet++ network as described in claim 3, characterized in that, it further includes a testing module. The testing module first normalizes the test image element values to between 0 and 1; then inputs the test data into the de-highlighting network model to obtain the output data, which is converted into an RGB format image, that is, the predicted non-highlight image.

Citation Information

Patent Citations

  • Image highlight removing method based on U-shaped cavity residual network

    CN111709886A

  • Weak light image enhancement method and device based on U-net + + network

    CN112991227A