A mirror highlight detection and removal method based on a double-flow convolutional neural network

By employing a specular highlight detection and removal method based on a dual-stream convolutional neural network, and utilizing a convolutional block attention module and a highlight extraction module, the problem of image information degradation caused by specular highlights is solved, achieving efficient and accurate highlight removal and improving the accuracy of visual tasks.

CN115311157BActive Publication Date: 2026-06-02ZHEJIANG UNIV OF TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG UNIV OF TECH
Filing Date
2022-07-18
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing specular highlight detection and removal methods are computationally intensive and complex when dealing with large areas of highlights, and have weak highlight detection capabilities in low-saturation areas. The removal results contain visual artifacts and distortions in color, structure, and texture.

Method used

A method based on dual-stream convolutional neural networks is adopted. By constructing a dual-stream feature selection strategy and combining a convolutional block attention module (CBAM) and a specular extraction module, specular features are coarsely extracted and refined in depth. This includes preprocessing, gradient extraction, stepwise downsampling, CBAM processing, and specular refinement removal.

Benefits of technology

It effectively removes specular highlights, reduces the interference of highlights on visual tasks such as target detection, and generates highlight-free images without visual artifacts and color and texture distortion, thus improving the accuracy and efficiency of highlight removal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115311157B_ABST
    Figure CN115311157B_ABST
Patent Text Reader

Abstract

The application discloses a mirror highlight detection and removal method based on a double-flow convolutional neural network, which calculates the gradient of a pixel point in an x direction and a y direction of an input original image with highlights, extracts a first highlight feature mapping, subtracts the first highlight feature mapping after being processed by a first convolution block attention module CBAM from a preprocessed image, and outputs a first-stage highlight-free image; the first highlight feature mapping is progressively down-sampled and reduced in size, each highlight feature mapping after being down-sampled and reduced in size is processed by a convolution block attention module CBAM at each stage, and is subtracted from a highlight-free image output by a previous stage, and finally a rough highlight-free image is obtained. Then, highlight extraction and refinement are performed on the rough highlight-free image by a highlight extraction module to obtain a final highlight-free image. The application can effectively solve the image information degradation problem caused by the mirror highlight, thereby reducing the interference of the highlight on visual tasks such as target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of image restoration, and in particular relates to a method for specular highlight detection and removal based on a two-stream convolutional neural network. Background Technology

[0002] With the development of the automation industry, intelligent robots are finding increasingly diverse applications in smart homes, medical services, military navigation, and the machinery industry. A major focus of research in intelligent robots is solving challenges in visual tasks, and image information degradation caused by specular highlights is a significant one. Specular highlights refer to the numerous bright spots and patches that appear on the surface of highly reflective objects when illuminated. The color, texture, and structure of these areas are degraded or even completely lost. Therefore, image information degradation caused by highlights often has a fatal impact on visual tasks such as target detection and object tracking.

[0003] To eliminate this effect, many researchers have dedicated themselves to detecting and removing highlights in images in recent years. Current methods can be categorized into traditional model-based methods and deep learning-based methods. Model-based highlight detection and removal methods are further divided into single-image and multi-image highlight detection and removal. Single-image highlight detection and removal methods include: methods based on different forms of thresholding, but these methods struggle to locate all highlights due to threshold constraints; methods based on the assumption that only a small portion of the scene contains highlights, but these methods lack the ability to handle large areas of highlights; and methods based on color, spatial, and illumination estimation, or methods based on intrinsic image decomposition, but these fail to produce satisfactory highlight removal results for realistic images with complex backgrounds and lighting. Unlike single-image methods, multi-image methods estimate the location of the light source by assuming the geometry of the known surface, and then superimpose multiple images to separate the highlights. While multi-image methods can achieve more accurate highlight removal results than single-image methods, these methods are computationally intensive and complex, resulting in low practicality.

[0004] While existing model-based single-image speckle detection and removal methods are efficient and easy to implement, they are often unreliable for images with very bright appearances or textures. They frequently misdetect white areas or high-intensity pixels as specks, or fail when large areas of speckle are present in the scene. Deep learning-based approaches overcome these problems. By constructing large-scale real-world datasets and utilizing modern deep learning tools, high-quality speckle removal results can be obtained. Many excellent previous works have designed deep learning-based networks that utilize contextual contrast features to locate speckles and have further proposed multi-task networks for joint speckle detection and removal. Deep learning-based methods break free from the constraints of traditional models, and the generated speckle-free images represent a significant improvement over previous methods. However, shortcomings remain. Existing methods have weak speckle detection capabilities in low-saturation regions, incomplete speckle detection and removal results, and visual artifacts, as well as color, structural, and texture distortions in the compensated pixels, are present in the speckle removal results. Summary of the Invention

[0005] This application proposes a specular highlight detection and removal method based on a dual-stream convolutional neural network to address the problem that specular highlights cause image information degradation, thereby affecting the accuracy of visual tasks such as target recognition.

[0006] To achieve the above objectives, the technical solution of this application is as follows:

[0007] A method for specular highlight detection and removal based on a two-stream convolutional neural network includes:

[0008] The input raw image with highlights is preprocessed, and then the pixel points are calculated. direction and The gradient in the direction is used to extract the first specular feature map;

[0009] After the first specular feature map is processed by the first convolutional block attention module CBAM, it is subtracted from the preprocessed image to output the first-level image without specular highlights.

[0010] The first specular feature map is downsampled and reduced in size step by step. The downsampled and reduced specular feature maps at each level are processed by the convolutional block attention module (CBAM) at each level and then subtracted from the specular-free image output by the previous level to finally obtain a coarse specular-free image.

[0011] The coarse image without highlights is linearly transformed to obtain a linearly transformed coarse image without highlights, which is then input into the highlight extraction module, which includes an encoder and a decoder, to extract refined highlight information.

[0012] The output features of the last-stage convolutional block attention module CBAM are linearly transformed to obtain a linearly transformed specular feature map. Combined with the coarse no-spectrum image after linear transformation and the refined specular information, specular refinement and removal are performed to obtain the final no-spectrum image.

[0013] Furthermore, the first specular feature map is downsampled four times to reduce its size. Each downsampled and reduced-size specular feature map is then processed by a convolutional block attention module (CBAM) at each stage, and subtracted from the specular-free image output from the previous stage to obtain a coarse specular-free image. This process includes:

[0014] The first specular feature map is downsampled once to obtain the second specular feature map, then downsampled once to obtain the third specular feature map, then downsampled once to obtain the fourth specular feature map, and then downsampled once to obtain the fifth specular feature map.

[0015] The second specular feature map is input into the second convolutional block attention module (CBAM). Then, the output of the second convolutional block attention module (CBAM) is subtracted from the first-level no-spectrum image to output the second-level no-spectrum image.

[0016] The third specular feature map is input into the third convolutional block attention module (CBAM). Then, the output of the third convolutional block attention module (CBAM) is subtracted from the second-level no-spectrum image to output the third-level no-spectrum image.

[0017] The fourth specular feature map is input into the fourth convolutional block attention module (CBAM). Then, the output of the fourth convolutional block attention module (CBAM) is subtracted from the third-level no-spectrum image to output the fourth-level no-spectrum image.

[0018] The fifth specular feature map is input into the fifth convolutional block attention module (CBAM). Then, the output of the fifth CBAM is subtracted from the fourth-level no-spectrum image to output the fifth-level no-spectrum image.

[0019] Furthermore, the encoder includes a convolutional layer and a dilated convolutional layer, and the decoder includes a gated convolutional layer and a convolutional layer.

[0020] Furthermore, the output features of the last-stage convolutional block attention module (CBAM) are linearly transformed to obtain a linearly transformed specular feature map. This map is then combined with the coarse, specular-free image obtained after the linear transformation and the refined specular information to perform specular refinement and removal, resulting in the final specular-free image. This is expressed by the following formula:

[0021] σ ;

[0022] in, This represents the final image without highlights, where σ is the Sigmoid activation function and MLP represents a multilayer perceptron. This represents a rough, specular-free image after linear transformation. This represents the specular feature map after linear transformation. This indicates that the highlight information is refined. and They represent and The weight value, This represents the addition of matrix elements. This represents subtracting matrix elements. This indicates residual operations. This indicates a gated convolution operation.

[0023] This application proposes a specular highlight detection and removal method based on a two-stream convolutional neural network. It constructs a coarse extraction process for highlight features and a deep refinement process for highlight features based on a two-stream feature selection strategy, and combines them into a specular highlight detection and removal method based on convolutional block attention and a two-stream convolutional neural network. This method can effectively solve the problem of image information degradation caused by specular highlights, thereby reducing the interference of highlights on visual tasks such as target detection. Attached Figure Description

[0024] Figure 1 This is a flowchart of the specular highlight detection and removal method based on a dual-stream convolutional neural network in this application.

[0025] Figure 2 This is a schematic diagram of the overall structure of the network in an embodiment of this application;

[0026] Figure 3 This is a schematic diagram of the Convolutional Block Attention Module (CBAM). Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0028] In one embodiment, such as Figure 1 As shown, a method for specular highlight detection and removal based on a two-stream convolutional neural network is proposed, including:

[0029] Step S1: Preprocess the input raw image with highlights, and then calculate the pixel points in... direction and The gradient in the direction is used to extract the first specular feature map.

[0030] This embodiment performs image preprocessing by converting the image to grayscale and filtering to reduce image noise. For example, the input original image with highlights is represented as (b, 256, 256, 3), where b represents the batch size of the model training is 4, (256, 256) represents the data size of each image as 256*256, and 3 represents the x, y, and z channels. After grayscale conversion, a grayscale image with the data format (b, 256, 256, 1) is obtained to meet the computational requirements for subsequent highlight feature extraction. The data is then preprocessed by sequentially performing two-dimensional normalization (BatchNorm2D), channel dimension expansion (3, 64), two-dimensional normalization (BatchNorm2D), and ReLU function (ReLU). This application expands the channel dimension from 3 to 64, providing more operational space for subsequent highlight feature extraction.

[0031] This embodiment uses This represents the preprocessed image, specifically the first specular feature map. It is obtained by calculating the gradient of each pixel in image f in the x and y directions. Figure 2 The preliminary feature extraction is used to represent this.

[0032] Step S2: After the first specular feature map is processed by the first convolutional block attention module CBAM, it is subtracted from the preprocessed image to output the first level of no-spectrum image.

[0033] like Figure 2 As shown, this embodiment implements highlight detection and removal in two steps. First, the Convolutional Block Attention (CBAM) module sequentially solves for channel attention and spatial attention, and then a subtraction operation is used to roughly remove highlight features from the original highlighted image. Next, a dual-stream convolutional neural network is used for further deep highlight removal, fully utilizing the spatial information extraction capabilities of the fully convolutional network to complete the deep refinement and removal of highlight features.

[0034] This step and the next step are mainly used to roughly remove highlight features from the highlighted image. The network structure includes five branches, each branch has a convolutional block attention module (CBAM). The output of the CBAM is subtracted from the output of the previous branch to remove highlight features.

[0035] In the first branch, the first specular feature map The input is fed into the first convolutional block attention module (CBAM), and then the output of the first convolutional block attention module (CBAM) is compared with the preprocessed image. Perform a subtraction operation between the two to output the first-level image without highlights. .

[0036] It should be noted that the Convolutional Block Attention Module (CBAM) is as follows: Figure 3 As shown, for The channel attention and spatial attention are solved sequentially to enhance... The features of the Convolutional Block Attention (CBAM) module in this embodiment. The channel attention and spatial attention are solved sequentially. The Convolutional Block Attention (CBAM) module is a relatively mature technology in this field, and will not be described in detail here.

[0037] This embodiment enhances the expressiveness of specular features through the Convolutional Block Attention (CBAM) module, thereby improving the network's ability to detect specular highlights. Finally, specular features are subtracted from the original features, outputting a coarse image without specular highlights and a specular feature map, achieving preliminary specular removal.

[0038] Step S3: The first specular feature map is downsampled and reduced in size step by step. The downsampled and reduced specular feature maps at each level are processed by the convolutional block attention module (CBAM) at each level and then subtracted from the specular-free image output by the previous level to finally obtain a coarse specular-free image.

[0039] This embodiment describes the first specular feature mapping map. Gradually reduce the size by downsampling, such as Figure 2 As shown, for example, the first specular feature map The original size was 256*256. After one downsampling, the second specular feature map was obtained. (128*128), after another downsampling, the third specular feature map is obtained. (64*64), after another downsampling, the fourth specular feature map is obtained. (32*32), after another downsampling, the fifth specular feature map is obtained. (16*16).

[0040] In the second branch, the second specular feature map The input is fed into the second convolutional block attention module (CBAM), and then the output of the second convolutional block attention module (CBAM) is compared with the first-level specular-free image. Perform a subtraction operation between the two to output a second-level image without highlights. ;

[0041] In the third branch, the third specular feature map The input is fed into the third convolutional block attention module (CBAM), and then the output of the third convolutional block attention module (CBAM) is compared with the second-level specular-free image. Perform a subtraction operation between the two to output a third-level image without highlights. ;

[0042] In the fourth branch, the fourth specular feature map The input is fed into the fourth convolutional block attention module (CBAM), and then the output of the fourth convolutional block attention module (CBAM) is compared with the third-level specular-free image. Perform a subtraction operation between the two to output a fourth-level image without highlights. ;

[0043] In the fifth branch, the fifth specular feature map The input is fed into the fifth convolutional block attention module (CBAM), and then the output of the fifth convolutional block attention module (CBAM) is compared with the fourth-level specular-free image. Perform a subtraction operation between the two to output a fifth-level image without highlights. .

[0044] After five branches of the Convolutional Block Attention Module (CBAM) and corresponding subtraction operations, this embodiment obtains a coarsely specular-free image with coarse specular processing, namely the fifth specular-free image. . ㊀ indicates that matrix elements are subtracted one by one.

[0045] This embodiment extracts and effectively utilizes specular features and original features at different levels. Addressing the variability in the size, shape, and position of highlights in an image, iteratively refines the detailed information of low-level features to adapt to different specular situations. High-level features contain rich global environmental information beneficial for highlight localization, while low-level features carry a large amount of detailed information, helping to improve the fine detection and removal of highlights. By utilizing the characteristics of features at different levels, noise in low-level features can be effectively eliminated, filtering specular features from the original features in a progressive manner.

[0046] Step S4: After performing a linear transformation on the coarse image without highlights, the resulting coarse image without highlights is input into the highlight extraction module, which includes an encoder and a decoder to extract refined highlight information.

[0047] This embodiment is for images with minimal highlights. First, through a linear transformation matrix, Mapping from (1,16,16,1024) to (1,64,64,256) yields a coarse, specular-free image after linear transformation. The data is input into the specular extraction module, which leverages the spatial information extraction capabilities of a fully convolutional network for in-depth analysis. Refined specular information .

[0048] The encoder includes a convolutional layer and a dilated convolutional layer, and the decoder includes a gated convolutional layer and a convolutional layer.

[0049] Specifically, such as Figure 2As shown, dim represents the dimension, and depth is extracted through two structures: the encoder and the decoder. Highlights in the image. For example... Figure 1 First of all The original features (1, 64, 64, 256) are expanded to 128 dimensions (1, 64, 128, 128) through convolution. The operation is represented as follows. Then Input into the encoder. First, the specular information is mined using dilated convolutional layers and expanded to 256 dimensions (1, 64, 64, 256), generating features. The operation can be represented as .Will The input is fed into the decoder, where it undergoes gated convolution and upsampling layers to generate 128-dimensional features. The process (1, 64, 128, 128) can be represented as Then, the features are restored to 64 dimensions (1, 128, 128, 64) through upsampling layers and convolutional layers to generate features. The operation can be represented as This is followed by a Sigmoid activation function to generate... .

[0050] The highlight extraction module in this embodiment uses a dual-stream convolutional neural network (BCN), which has excellent spatial information extraction capabilities and can process highlights in images more accurately.

[0051] Step S5: Perform a linear transformation on the output features of the last-stage convolutional block attention module CBAM to obtain a linearly transformed specular feature map. Combine the coarse no-spectrum image after the linear transformation with the refined specular information to perform specular refinement and removal, and obtain the final no-spectrum image.

[0052] In this embodiment, the output features of the fifth convolutional block attention module (CBAM) are linearly transformed from (1,16,16,1024) to (1,256,256,64) to obtain the linearly transformed specular feature map. .

[0053] This embodiment utilizes refined specular information To Further details will be provided. Simultaneously, direct transmission will be performed. Combined with the refined highlight information obtained through in-depth mining exist Based on this, the degraded pixels in the highlight areas are refined to obtain the final highlight removal result. .

[0054] After depth extraction of highlights, they need to be refined and removed. Highlight removal is essentially image compensation; therefore, a highlight feature map is used. and As a guide, in Based on this, gated convolution and residual calculation are combined to assign different weights to different features. The MLP layer with Sigmoid activation function is used for processing, and data filtering is completed by element-wise subtraction to form the final image without highlights.

[0055] The specific operation is expressed by the following formula:

[0056] σ ;

[0057] in, This represents the final image without highlights, σ ​​is the Sigmoid activation function, and MLP represents a multilayer perceptron. The MLP layer is composed of a one-dimensional convolution (256,1,256), ReLU, and another one-dimensional convolution (256,1,256) connected in sequence. This represents a rough, specular-free image after linear transformation. This represents the specular feature map after linear transformation. This indicates that the highlight information is refined. and They represent and The weight value, This represents the addition of matrix elements. This represents subtracting matrix elements. This indicates residual operations. This indicates a gated convolution operation.

[0058] This embodiment uses gated convolution and residual calculation to... Remove the extracted highlight information and output the final result. During information fusion across layers, useful features are selectively preserved to avoid cross-contamination between features from highlight and non-highlight regions. Each convolutional layer is followed by a two-dimensional normalization (BatchNorm2D) and a ReLU function. To remove highlights from the image as much as possible, dilated convolutional layers are used to expand the receptive field, and gated convolutions are used to compensate for image degradation regions, guided by the highlight feature map. The final highlight-free image generated by this method has no visual artifacts or color and texture distortions, outperforming all other existing methods.

[0059] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A method for specular highlight detection and removal based on a two-stream convolutional neural network, characterized in that, The specular highlight detection and removal method based on a two-stream convolutional neural network includes: The input raw image with highlights is preprocessed, and then the pixel points are calculated. direction and The gradient in the direction is used to extract the first specular feature map; After the first specular feature map is processed by the first convolutional block attention module CBAM, it is subtracted from the preprocessed image to output the first-level image without specular highlights. The first specular feature map is downsampled and reduced in size step by step. The downsampled and reduced specular feature maps at each level are processed by the convolutional block attention module (CBAM) at each level and then subtracted from the specular-free image output by the previous level to finally obtain a coarse specular-free image. The coarse image without highlights is linearly transformed to obtain a linearly transformed coarse image without highlights, which is then input into the highlight extraction module. The highlight extraction module includes an encoder and a decoder. The encoder includes a convolutional layer and a dilated convolutional layer, and the decoder includes a gated convolutional layer and a convolutional layer to extract refined highlight information. The output features of the last-stage convolutional block attention module CBAM are linearly transformed to obtain a linearly transformed specular feature map. Combined with the coarse, specular-free image obtained after the linear transformation and the refined specular information, specular thinning and removal are performed to obtain the final specular-free image, expressed by the following formula: s ; in, This represents the final image without highlights, where σ is the Sigmoid activation function and MLP represents a multilayer perceptron. This represents a rough, specular-free image after linear transformation. This represents the specular feature map after linear transformation. This indicates that the highlight information is refined. and They represent and The weight value, This represents the addition of matrix elements. This represents subtracting matrix elements. This indicates residual operations. This indicates a gated convolution operation.

2. The method for specular highlight detection and removal based on a two-stream convolutional neural network according to claim 1, characterized in that, The process of downsampling and reducing the size of the first specular feature map step by step, and then processing each downsampled and reduced-size specular feature map through each level of the Convolutional Block Attention (CBAM) module, subtracting it from the specular-free image output from the previous level, finally yields a coarse specular-free image, including: The first specular feature map is downsampled once to obtain the second specular feature map, then downsampled once to obtain the third specular feature map, then downsampled once to obtain the fourth specular feature map, and then downsampled once to obtain the fifth specular feature map. The second specular feature map is input into the second convolutional block attention module (CBAM). Then, the output of the second convolutional block attention module (CBAM) is subtracted from the first-level no-spectrum image to output the second-level no-spectrum image. The third specular feature map is input into the third convolutional block attention module (CBAM). Then, the output of the third convolutional block attention module (CBAM) is subtracted from the second-level no-spectrum image to output the third-level no-spectrum image. The fourth specular feature map is input into the fourth convolutional block attention module (CBAM). Then, the output of the fourth convolutional block attention module (CBAM) is subtracted from the third-level no-spectrum image to output the fourth-level no-spectrum image. The fifth specular feature map is input into the fifth convolutional block attention module (CBAM). Then, the output of the fifth CBAM is subtracted from the fourth-level no-spectrum image to output the fifth-level no-spectrum image.