Monocular color image normal estimation method based on deep learning

Through deep learning cascade architecture and data set training under multi-illumination conditions, combined with local attention and dynamic scaling mechanism, the robustness and accuracy of monocular normal estimation in complex scenes and lighting changes are solved, and high-precision normal image prediction is achieved.

CN120298475APending Publication Date: 2025-07-11SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510424928.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-11

Smart Images

  • Figure CN120298475A_ABST
    Figure CN120298475A_ABST
Patent Text Reader

Abstract

The invention discloses a monocular color image normal estimation method based on deep learning, and the method comprises the steps: employing a basic data set which comprises a plurality of groups of indoor scene color images and normal images and depth images of the indoor scene color images; training the preliminary normal estimation network, extracting multi-scale features by using a pre-training encoder, performing feature fusion in combination with weighted decoding and a local attention mechanism, and training the network through a composite loss function of cosine similarity loss and L1 loss to output a preliminary normal graph; training a residual optimization network, constructing 9-channel input features including curvature features of a preliminary normal graph and a depth map and a monocular color image, adopting an encoder-decoder architecture, and introducing a dynamic scaling module to adaptively adjust residual amplitude; and taking a target image as input, and adding the output of the initial normal estimation network and the output of the residual optimization network to obtain a refined normal graph. According to the method, the problem of robustness of monocular normal estimation under different illumination is relieved, and meanwhile, the accuracy of monocular normal estimation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision, 3D reconstruction, and graphics rendering, and in particular to a method for monocular color image normal estimation based on deep learning. Background Art

[0002] Monocular color image normal estimation is a technology for inferring the normal direction of an object's surface through a single 2D color image. It plays a core role in the field of 3D visual understanding and is the basis for various computer vision tasks. It has important applications in fields such as 3D reconstruction, lighting modeling, augmented reality, autonomous robot navigation, game development, and industrial inspection. Accurately predicting the normal direction is crucial for understanding the geometric structure of an object.

[0003] Current monocular normal estimation techniques are mainly divided into two categories of methods, namely: 1. Traditional methods: Using geometric information in the image (such as shadows, textures, edges) to infer the surface normal. However, there are many problems with such techniques: severely relying on assumptions such as a single light source and uniform materials, failing when encountering complex scenes and lighting, being sensitive to noise, having poor robustness, and being unable to handle complex geometries; 2. Deep learning-based methods: Constructing an end-to-end model through a neural network, using large-scale datasets, such as paired data of monocular RGB images and normal maps, for supervised or self-supervised training to learn the mapping relationship from a color image to a surface normal map. However, there are also several problems with such methods: relatively low accuracy, unable to obtain high-accuracy and high-fine-grained estimation results in complex scenes, and prone to losing high-frequency geometric details; still having insufficient robustness under conditions of changing environmental lighting, especially in indoor scenes where the light sources are complex and diverse and unevenly distributed. Such a complex lighting environment is extremely likely to produce different shadows and highlights, directly affecting the stability of normal estimation. Summary of the Invention

[0004] The purpose of the present invention is to overcome the shortcomings and deficiencies of the prior art, and propose a method for monocular color image normal estimation based on deep learning. By means of a cascaded architecture of preliminary normal prediction and residual optimization network, the corresponding normal image is estimated from a color image, and a normal estimation result with high accuracy and high robustness is obtained.

[0005] To achieve the above purpose, the technical solution provided by the present invention is: A method for monocular color image normal estimation based on deep learning, comprising the following steps:

[0006] 1) Using a basic dataset, which includes multiple groups of indoor scene color images and their corresponding normal images and depth images. Among them, for each perspective under the same indoor scene, there are multiple differently colored color images and their corresponding normal images and depth images under different lighting conditions. The normal images and depth images are rendered and generated when making this dataset;

[0007] 2) Use the color images in the dataset in step 1) as input to train a preliminary normal estimation network, which is an improved U-Net network. Use the cosine similarity loss and absolute error loss between the normal image predicted by this network and the corresponding ground truth normal image in the dataset as the loss function for network training. Improve the encoder and decoder of the network. The specific improvement is as follows: Use a pre-trained convolutional neural network as the encoder to extract multi-scale high-dimensional features from the input color image. The decoder uses bilinear interpolation for upsampling, combines weighted decoding with features at the corresponding level for fusion and stitching, and introduces a local attention mechanism during the fusion process. Finally, use the Tanh activation function to output the preliminarily estimated normal image;

[0008] 3) Obtain the color images, depth images in the dataset in step 1), and the preliminarily estimated normal images obtained by applying the preliminary normal estimation network in step 2). Extract curvature features from the depth images, and construct 9-channel features with the color images and the preliminarily estimated normal images as input to train a residual optimization network. Introduce a consistency loss on the basis of the loss function in step 2) as the loss function for training the residual optimization network. This consistency loss is used to constrain the consistency of the normal estimation images of color images under the same viewing angle and different lighting conditions. This network adopts an encoder-decoder structure, introduces a dynamic scaling mechanism to dynamically adjust the residual amplitude, and the decoder restores the resolution through upsampling to output the residual map;

[0009] 4) During application, input the monocular target color image into the trained preliminary normal estimation network in step 2) to obtain the preliminarily estimated normal image, then jointly input it with the depth image and color image of the target into the trained residual optimization network in step 3) to obtain the residual map. Finally, add the preliminarily estimated normal image and the residual map and normalize them to output the refined normal image with lighting consistency.

[0010] Furthermore, in step 1), normalize the color images and normal images in the dataset. Normalize the color images to the range [0, 1], and normalize the normal images to the range [-1, 1]. Then perform operations of unified size and data augmentation, including random color jittering and rotation. In the scenario of this dataset, there are a total of 11 different lighting combinations between ambient light, indoor light sources, and outdoor light sources, so as to improve the robustness of the network.

[0011] Furthermore, in step 2), train the preliminary normal estimation network in a supervised learning manner. This network is an improved U-Net network. The input is the color images in the dataset, and the output is the preliminarily estimated normal image. The specific situation of this network is as follows:

[0012] The initial normal estimation network uses the U-Net network as the backbone network. First, a pre-trained convolutional neural network with the global pooling and classification layers removed is used as the encoder. Through five downsamplings, the encoded features at each level are extracted from the color image:

[0013]

[0014] F i = Encoder(F i-1 ), i ∈ {1, 2, 3, 4, 5}

[0015] where I rgb is the input color image, H and W are the height and width of the color image respectively, is the real number field, F0 is the initial encoded feature input from the network to the encoder, F i is the i-th layer encoded feature extracted by the encoder. There are a total of five downsamplings, so i ∈ {1, 2, 3, 4, 5}. The resolution of each layer of encoded features decreases from 256×256 to 16×16 step by step, and the number of channels increases from 32 to 2560. F i-1 is the (i - 1)-th layer encoded feature extracted by the encoder, and Encoder is the encoder;

[0016] Secondly, the decoder of this network contains four levels of upsampling blocks. Each upsampling block gradually restores the resolution of the encoded features from 16×16 to 256×256. First, bilinear interpolation upsampling is used to upsample the encoded features to obtain the upsampled feature U i+1 , and then through skip connections, the encoder feature F i and the upsampled feature U i+1 are concatenated along the channel dimension, and the number of channels is reduced and cross-scale information is fused through a convolutional layer to generate the fused feature X fused :

[0017]

[0018] X fused,j = Conv(Concat(F j , U j+1 ), j ∈ {1, 2, 3, 4}

[0019] where is the low-resolution feature from the (j + 1)-th layer of the decoder, U j+1 is the upsampled feature of the (j + 1)-th layer of the decoder, is the feature map with the lowest resolution in the decoder, that is, the first layer of the decoder. The input is the last layer feature F5 of the encoder, and X fused,j is the fused feature generated by the j-th layer of the decoder, F jis the j-th layer of encoded features extracted by the encoder. Since only the first four layers of the encoder are spliced with the decoder through skip connections to generate fused features, j ∈ {1, 2, 3, 4}, Interpolate is a bilinear interpolation operation, Conv represents a convolution operation, and Concat represents a channel splicing operation; then in each upsampling block, the corresponding fused feature X fused introduces a local attention mechanism for attention enhancement operation. The specific steps are to divide it into local blocks through convolution operation and generate embedding vectors, and then flatten it into a one-dimensional Token sequence t along the spatial dimension:

[0020]

[0021] In the formula, P is the size of the local block, E represents the local block embedding tensor extracted from the fused feature through convolution operation, each of its elements represents an embedding vector of a local block, L is the embedding dimension, is the Token sequence length, Flatten is a function to flatten a continuous range of dimensions, and t represents the flattened one-dimensional original Token sequence; a self-attention mechanism is introduced to extract global attention features from this one-dimensional Token sequence t, which consists of multi-head self-attention and a multi-layer perceptron. The one-dimensional Token sequence is input into a four-layer Transformer encoder. In each layer of the network:

[0022] Q = MLP q (t l-1 )

[0023] K = MLP k (t l-1 )

[0024] V = MLP v (t l-1 )

[0025] In the formula, Q, K, and V are the query matrix, key matrix, and value matrix respectively, MLP q , MLP k , MLP v are the global perceptrons corresponding to Q, K, and V respectively, t l-1 , l ∈ {1, 2, 3, 4} is the input one-dimensional encoded Token sequence or the features of the previous layer; then the weighted self-attention is calculated from the output Q, K, and V:

[0026]

[0027] In the formula, SA is the self-attention function, t l is the attention feature of the current layer, A softmaxis the Softmax activation function, D represents z-score normalization to stabilize the gradient; then the multi-head self-attention module is used to calculate the self-attention of multiple heads simultaneously, and the calculation results are concatenated, and finally a residual connection is added. The formula is as follows:

[0028] MultiHead(t l ) = Concat(head1, head2,..., head h ) + t l-1

[0029] head k = SA(LN(t k , k ∈ {1, 2, 3, 4}

[0030] In the formula, MultiHead is the multi-head self-attention function, head k is the output of the k-th attention head, and there are h = 4 attention heads in total. LN is layer normalization; after each multi-head self-attention operation, a global perceptron and a residual connection are added to obtain the attention feature of the current layer:

[0031] t l = MLP(LN(MultiHead(t l-1 ))) + t l-1

[0032] The one-dimensional encoded Token sequence of each layer of the decoder is re-adjusted to a two-dimensional global attention feature through channel adjustment, and it is bilinearly interpolated and upsampled to the same resolution as the high-dimensional encoded feature to obtain the attention map f A :

[0033] f A,l = Interpolate(View(t l ))

[0034] In the formula, View is the channel adjustment operation, f A,l is the attention map of the current layer; in each layer of the decoder, the fused feature and the attention map are weighted by the Hadamard product to obtain the weighted encoded feature, and a residual connection is included. In the last layer, the convolutional neural network is used to refine the weighted encoded feature with the highest resolution, and it is extended to the size of the input image through bilinear interpolation. Finally, the Tanh activation function is used to output the estimated normal image:

[0035] i normal = A tanh (Conv(X′ fused ⊙ f′ A + Conv(X′ fused )))

[0036] where \(i\) normal is the finally output normal image, \(A\) tanh is the Tanh activation function, \(\odot\) is the Hadamard product operation, Conv(\(X'\) fused ) serves as the residual connection, where \(X'\) fused and \(f'\) A respectively represent the fused features and the attention map obtained in the last upsampling block of the decoder; then the cosine similarity loss function and the absolute error loss function are used to calculate the loss \(L1\) between the output normal image and the corresponding ground truth normal image in the dataset, and it is used to train the network:

[0037] \(L1=\alpha\cdot L\) cos (\(G,\hat{G}\) * )+\(\beta\cdot L\) L1 (\(G,\hat{G}\) * )

[0038] where \(\alpha\) and \(\beta\) are hyperparameters for controlling the loss range, \(G\) is the normal tensor predicted by the preliminary normal estimation network, \(\hat{G}\) * is the ground truth normal tensor in the dataset, \(L\) cos is the cosine similarity loss function, \(L\) L1 is the per-pixel \(L1\) norm loss, i.e., the absolute error loss function, and the calculation methods of the two functions are:

[0039]

[0040] where \(N\) is the number of valid pixels in the image, \(G\) i′ is the normal vector of the \(i'\)-th pixel in the normal map predicted by the preliminary normal estimation network, is the normal vector of the \(i'\)-th pixel in the ground truth normal map.

[0041] Furthermore, in step 3), the residual optimization network is trained in a supervised learning manner. This network is an improved encoder-decoder structure network. The input is a monocular color image, the corresponding depth image, and the preliminarily estimated normal image, and the output is a residual map. The specific situation of this network is as follows:

[0042] First, obtain the color image \(I\) rgb in the dataset in step 1), depth the depth image \(I_d\) pred and the preliminarily estimated normal image \(N\) obtained through the preliminary normal estimation network in step 2) as the network input, where the horizontal gradient \(G_x\) X and the vertical gradient \(G_y\) Y of the depth image are calculated using the Sobel operator:

[0043] \(G_x\) X=Conv2D(Pad horizontal (I depth ),K X )

[0044] G Y =Conv2D(Pad vertical (I depth ),K Y )

[0045] where K X =[[-1,0,1]] / 2 is the horizontal Sobel kernel, and K Y =[[-1],[0],[1]] / 2 is the vertical Sobel kernel. Pad horizontal and Pad vertical are mirror padding in the horizontal and vertical directions respectively. Then, the curvature map C is calculated through discrete second-order differences:

[0046] C=(I depth:,:,:,2: -2I depth:,:,:,1:-1 +I depth:,:,:,:-2 )+(I depth:,:,2:,: -2I depth:,:,1:-1,: +I depth:,:,:-2,: )

[0047] where (I depth:,:,:,2: -2I depth:,:,:,1:-1 +I depth:,:,:,:-2 ) calculates the curvature in the horizontal direction of the depth image, and (I depth:,:,2:,: -2I depth:,:,1:-1,: +I depth:,:,:-2,: ) calculates the curvature in the vertical direction of the depth image. I depth:,:,:,2: represents taking the pixel values of the depth image from the 3rd column to the last column in the width direction. I depth:,:,:,1:-1 represents taking the pixel values of the depth image from the 2nd column to the second-to-last column in the width direction. I depth:,:,:,:-2 represents taking the pixel values of the depth image from the 1st column to the third-to-last column in the width direction. I depth:,:,2:,: represents taking the pixel values of the depth image from the 3rd row to the last row in the height direction. I depth:,:,1:-1,: represents taking the pixel values of the depth image from the 2nd row to the second-to-last row in the height direction. I depth:,:,:-2,: represents taking the pixel values of the depth image from the 1st row to the third-to-last row in the height direction; and they are concatenated along the channel dimension to generate the joint feature map X:

[0048] X=Concat(N pred ,G X (I depth ),G Y (I depth ),C(I depth ),Irgb )

[0049] The encoder of the network consists of four convolutional layers and one downsampling layer. The joint feature map X generated from the network input serves as the input to the encoder. The decoder restores the spatial resolution through bilinear interpolation upsampling and convolutional operations. The last layer uses the Tanh activation function to constrain the value range to [-1, 1] and outputs the basic residual map R base ; The improved part is the introduction of a dynamic scaling mechanism to finely adjust the residuals:

[0050] S = Tanh(Conv(R base )) ⊙ s global

[0051] where s global is a learnable parameter, R base represents the basic residual map output by the decoder, and S is the spatially varying scaling factor map generated by the sub-network; then the Hadamard product operation on the basic residuals can obtain the residual map R res :

[0052] R res = R base ⊙ S

[0053] In the training of the residual optimization network, images with different illuminations under the same perspective are grouped for training. When calculating the loss, a consistency loss constraint is introduced more than in the preliminary normal estimation network to train the network. The loss function L2 of this network is:

[0054] L2 = γ·L′ cos + δ·L′ L1 + ε·L con

[0055]

[0056] where γ, δ, and ε respectively represent the hyperparameters controlling the cosine similarity loss, per-pixel L1 norm loss, and consistency loss weights. L′ cos represents the cosine similarity loss of the residual optimization network, L′ L1 represents the per-pixel L1 norm loss of the residual optimization network, and L con represents the consistency loss of the residual optimization network, represents the predicted normal tensor after adding the residuals, is the ground truth normal tensor, is the normal vector of the i′2-th pixel in the predicted normal map, is the normal vector of the i′2-th pixel in the ground truth normal map, B is the batch dimension, and N Light is the number of illuminations included in each group, The normal mean predicted within the group for the b-th batch, the predicted normals for all illuminations within the same group, i.e., the same viewing angle, are averaged along the N Light dimension. n represents the n-th illumination condition in the group currently, and h′ and w respectively represent the h′-th row and w-th column of the image currently.

[0057] Furthermore, in step 4), the specific steps for obtaining the normal estimation for a monocular color image are as follows: First, obtain a monocular color image to be estimated for normals. Then, use the preliminarily trained normal estimation network in step 2) to obtain a preliminary normal estimation map N pred from the monocular color image. Then, use this normal estimation map, the monocular color image, and its corresponding depth image as inputs to obtain a residual map R using the residual optimization network res . Finally, add the preliminarily estimated normal image and the residual map to obtain the optimized normal estimation map N refined :

[0058] N refined = N pred + R res .

[0059] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0060] 1. The present invention proposes to use a dataset with multiple illumination information under the same viewing angle, enabling the preliminarily trained normal estimation network to improve the robustness and consistency in the face of different illumination conditions.

[0061] 2. The preliminarily trained normal estimation network of the present invention uses a weighted method based on the self-attention mechanism to extract global semantic information, which not only greatly increases the receptive field of the network but also enhances the understanding of the scene illumination information. At the same time, the residual optimization network introduces depth gradient and curvature information, fully integrating global and local features, enabling the predicted normals to have both the overall structure and retain edge details.

[0062] 3. The present invention first performs a rough estimation and then a refinement correction, with a joint training strategy and multi-loss function design, enabling the network to still maintain high-precision prediction in the face of different noises and illumination changes. Compared with other deep learning-based monocular normal estimation methods, it is more suitable for normal reconstruction tasks in complex scenes.

[0063] 4. Compared with other deep learning-based monocular normal estimation methods, the present invention achieves higher accuracy on the currently widely used normal estimation datasets.

[0064] 5. The present invention is of great significance for computer vision tasks and has broad application prospects in the fields of 3D reconstruction, augmented reality, autonomous driving, and industrial inspection. It provides reliable prior information for subsequent tasks such as shape recovery and texture mapping based on normal vectors, and has important theoretical and application values. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 It is a schematic diagram of the logical flow of the method of the present invention.

[0066] Figure 2 It is a schematic diagram of the weighted attention mechanism used in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0067] The present invention will be further described in detail below in conjunction with the embodiments and the drawings, but the embodiments of the present invention are not limited thereto.

[0068] As Figure 1 and Figure 2 shown, this embodiment discloses a method for estimating the normal vector of a monocular color image based on deep learning, and the specific situation is as follows:

[0069] 1) Use a basic dataset, which includes multiple groups of indoor scene color images and their corresponding normal vector images and depth images. Among them, for each perspective in the same indoor scene, there are multiple color images with different illuminations and corresponding normal vector images and depth images. The normal vector images and depth images are rendered and generated when making this dataset;

[0070] Normalize the color images and normal vector images in the dataset. Normalize the color images to the range of [0, 1], and normalize the normal vector images to the range of [-1, 1]. Then perform operations of unified size and data augmentation, including random color jittering and rotation. In the scenes of this dataset, there are a total of 11 different illumination combinations between ambient light, indoor light sources, and outdoor light sources, so as to improve the robustness of the network.

[0071] 2) Use the color images in the dataset in step 1) as the input to train a preliminary normal vector estimation network. This network is an improved U-Net network. Use the cosine similarity loss and absolute error loss between the normal vector image predicted by this network and the corresponding real normal vector image in the dataset as the loss function for network training. Improve the encoder and decoder of the network. The specific improvement is as follows: Use a pre-trained convolutional neural network as the encoder to extract multi-scale high-dimensional features from the input color image. The decoder uses bilinear interpolation for upsampling, combines weighted decoding with features at the corresponding level for fusion and splicing, and introduces a local attention mechanism during the fusion process. Finally, use the Tanh activation function to output the preliminarily estimated normal vector image;

[0072] The preliminary normal estimation network is trained in a supervised learning manner. This network is an improved U-Net network. The input is the color image in the dataset, and the output is the preliminarily estimated normal image. The specific situation of this network is as follows:

[0073] The preliminary normal estimation network uses the U-Net network as the backbone network. First, a pre-trained convolutional neural network with the global pooling and classification layers removed is used as the encoder to extract the encoded features at each level from the color image through five downsamplings:

[0074]

[0075] F i = Encoder(F i-1 ), i ∈ {1, 2, 3, 4, 5}

[0076] In the formula, I rgb is the input color image, H and W are the height and width of the color image respectively, is the real number field, F0 is the initial encoded feature input from the network to the encoder, F i is the i-th layer encoded feature extracted by the encoder. There are a total of five downsamplings, so i ∈ {1, 2, 3, 4, 5}. The resolution of each layer of encoded feature decreases from 256×256 to 16×16 step by step, and the number of channels increases from 32 to 2560. F i-1 is the (i - 1)-th layer encoded feature extracted by the encoder, and Encoder is the encoder;

[0077] Secondly, the decoder of this network contains four levels of upsampling blocks. Each upsampling block gradually restores the resolution of the encoded feature from 16×16 to 256×256. First, bilinear interpolation upsampling is used to upsample the encoded feature to obtain the upsampled feature U i+1 , and then through skip connections, the encoder feature F i and the upsampled feature U i+1 are concatenated along the channel dimension, and the number of channels is reduced and cross-scale information is fused through a convolutional layer to generate the fused feature X fused :

[0078]

[0079] X fused,j = Conv(Concat(F j , U j+1 ), j ∈ {1, 2, 3, 4}

[0080] In the formula, is the low-resolution feature from the (j + 1)-th layer of the decoder, and U j+1 is the feature after upsampling of the (j + 1)-th layer of the decoder. is the feature map with the lowest resolution in the decoder, that is, the first layer of the decoder. The input is the feature F5 of the last layer of the encoder, X fused,j is the fused feature generated by the j-th layer of the decoder, F j is the encoded feature of the j-th layer extracted by the encoder. Since only the first four layers of the encoder are used to splice with the decoder through skip connections to generate the fused feature, so j ∈ {1, 2, 3, 4}. Interpolate is the bilinear interpolation operation, Conv represents the convolution operation, and Concat represents the channel splicing operation; then in each upsampling block, the corresponding fused feature X fused introduces a local attention mechanism for attention enhancement operation. The specific steps are to split into local blocks through convolution operation and generate embedding vectors, and then flatten them into a one-dimensional Token sequence t along the spatial dimension:

[0081]

[0082] In the formula, P is the size of the local block, E represents the local block embedding tensor extracted from the fused feature through convolution operation, each of its elements represents an embedding vector of a local block, L is the embedding dimension, is the Token sequence length, Flatten is a function to flatten a continuous range of dimensions, and t represents the flattened one-dimensional original Token sequence; introduce a self-attention mechanism to extract global attention features from this one-dimensional Token sequence t, which consists of multi-head self-attention and a multi-layer perceptron. Input the one-dimensional Token sequence into a four-layer Transformer encoder. In each layer of the network:

[0083] Q = MLP q (t l-1 )

[0084] K = MLP k (t l-1 )

[0085] V = MLP v (t l-1 )

[0086] In the formula, Q, K, and V are the query matrix, key matrix, and value matrix respectively. MLP q , MLP k , MLP v are the global perceptrons corresponding to Q, K, and V respectively. t l-1 , l ∈ {1, 2, 3, 4} is the input one-dimensional encoded Token sequence or the feature of the previous layer; then calculate the weighted self-attention from the output Q, K, and V:

[0087]

[0088] In the formula, SA is the self-attention function, and t l is the attention feature of the current layer, and A softmax is the Softmax activation function, and D represents z-score normalization to stabilize the gradient; then the multi-head self-attention module is used to calculate the self-attention of multiple heads simultaneously, and the calculation results are concatenated, and finally a residual connection is added. The formula is as follows:

[0089] MultiHead(t l ) = Concat(head1, head2, …, head h ) + t l-1

[0090] head k = SA(LN(t k ), k ∈ {1, 2, 3, 4}

[0091] In the formula, MultiHead is the multi-head self-attention function, and head k is the output of the k-th attention head, and there are a total of h = 4 attention heads, and LN is layer normalization; after each multi-head self-attention operation, a global perceptron and a residual connection are added to obtain the attention feature of the current layer:

[0092] t l = MLP(LN(MultiHead(t l-1 ))) + t l-1

[0093] The one-dimensional encoded Token sequence of each layer of the decoder is readjusted to a two-dimensional global attention feature through channel adjustment, and it is bilinearly interpolated and upsampled to the same resolution as the high-dimensional encoded feature, so as to obtain the attention map f A :

[0094] f A,l = Interpolate(View(t l ))

[0095] In the formula, View is the channel adjustment operation, and f A,l is the attention map of the current layer; in each layer of the decoder, the fused feature and the attention map are weighted by the Hadamard product to obtain the weighted encoded feature, and a residual connection is included. In the last layer, a convolutional neural network is used to refine the weighted encoded feature with the highest resolution, and it is extended to the size of the input image through bilinear interpolation, and finally the estimated normal image is output using the Tanh activation function:

[0096] i normal = A tanh (Conv(X′fused ⊙ f' A + Conv(X' fused )))

[0097] In the formula, i normal is the final output normal image, A tanh is the Tanh activation function, ⊙ is the Hadamard product operation, Conv(X' fused ) is used as the residual connection. Here, X' fused and f' A respectively represent the fused feature and the attention map obtained from the last upsampling block of the decoder; then the cosine similarity loss function and the absolute error loss function are used to calculate the loss L1 between the output normal image and the corresponding ground truth normal image in the dataset, and it is used to train the network:

[0098] L1 = α · L cos (G, G * ) + β · L L1 (G, G * )

[0099] In the formula, α = 0.8 and β = 0.2 are hyperparameters for controlling the loss range, G is the normal tensor predicted by the preliminary normal estimation network, G * is the ground truth normal tensor in the dataset, L cos is the cosine similarity loss function, L L1 is the per-pixel L1 norm loss, i.e., the absolute error loss function. The calculation methods of the two functions are as follows:

[0100]

[0101] In the formula, N is the number of valid pixels in the image, G i′ is the normal vector of the i'-th pixel in the normal map predicted by the preliminary normal estimation network, is the normal vector of the i'-th pixel in the ground truth normal map.

[0102] 3) Obtain the color image, depth image of the dataset in step 1) and the preliminarily estimated normal image obtained by applying the preliminary normal estimation network in step 2). Extract the curvature feature from the depth image, and construct a 9-channel feature with the color image and the preliminarily estimated normal image as the input. Train the residual optimization network. Based on the loss function in step 2), introduce a consistency loss as the loss function for training the residual optimization network. This consistency loss is used to constrain the consistency of the normal estimation images of the color image under the same view and different illumination conditions. This network adopts an encoder-decoder structure, introduces a dynamic scaling mechanism to dynamically adjust the residual amplitude, and the decoder restores the resolution through upsampling to output the residual map;

[0103] Train the residual optimization network in a supervised learning manner. This network is an improved encoder-decoder structure network. The input is a monocular color image, the corresponding depth image, and the initially estimated normal image, and the output is a residual map. The specific situation of this network is as follows:

[0104] First, obtain the color image I of the dataset in step 1) rgb , the depth image I depth and the initially estimated normal image N obtained through the initial normal estimation network in step 2) pred as the network input, where the horizontal gradient G of the depth image is calculated using the Sobel operator X and the vertical gradient G Y :

[0105] G X = Conv2D(Pad horizontal (I depth ), K X )

[0106] G Y = Conv2D(Pad vertical (I depth ), K Y )

[0107] In the formula, K X = [[-1, 0, 1]] / 2 is the horizontal Sobel kernel, K Y = [[-1], [0], [1]] / 2 is the vertical Sobel kernel, Pad horizontal and Pad vertical are mirror padding in the horizontal and vertical directions respectively. Then, calculate the curvature map C through discrete second-order differences:

[0108] C = (I depth:,:,:,2: - 2I depth:,:,:,1:-1 + I depth:,:,:,:-2 ) + (I depth:,:,2:,: - 2I depth:,:,1:-1,: + I depth:,:,:-2,: )

[0109] In the formula, (I depth:,:,:,2: - 2I depth:,:,:,1:-1 + I depth:,:,:,:-2 ) calculates the curvature in the horizontal direction of the depth image, (I depth:,:,2:,: - 2I depth:,:,1:-1,: + I depth:,:,:-2,: ) calculates the curvature in the vertical direction of the depth image, I depth:,:,:,2: represents taking the pixel values of the depth image from the 3rd column to the last column in the width direction, I depth:,:,:,1:-1 represents taking the pixel values of the depth image from the 2nd column to the second-to-last column in the width direction, I depth:,:,:,:-2It represents taking the pixel values of the depth image from the 1st column to the 3rd column from the end in the width direction, I depth:,:,2:,: It represents taking the pixel values of the depth image from the 3rd row to the last row in the height direction, I depth:,:,1:-1,: It represents taking the pixel values of the depth image from the 2nd row to the 2nd row from the end in the height direction, I depth:,:,:-2,: It represents taking the pixel values of the depth image from the 1st row to the 3rd row from the end in the height direction; and concatenating along the channel dimension to generate the joint feature map X:

[0110] X = Concat(N pred , G X (I depth ), G Y (I depth ), C(I depth ), I rgb )

[0111] The encoder of this network consists of four levels of convolutional layers and one level of downsampling layer. The joint feature map X generated from the network input is used as the input of the encoder, while the decoder restores the spatial resolution through bilinear interpolation upsampling and convolutional operations. The last layer uses the Tanh activation function to constrain the value range to [-1, 1] to output the basic residual map R base ; The improved part is to introduce a dynamic scaling mechanism to finely adjust the residuals:

[0112] S = Tanh(Conv(R base )) ⊙ s global

[0113] In the formula, s global is a learnable parameter, initialized to 2.0, R base represents the basic residual map output by the decoder, and S is the spatially varying scaling factor map generated by the sub-network; then performing the Hadamard product operation on the basic residuals can obtain the residual map R res :

[0114] R res = R base ⊙ S

[0115] In the training of the residual optimization network, the pictures with different illuminations under the same perspective are grouped for training. When calculating the loss, a consistency loss constraint is introduced more than the preliminary normal estimation network to train the network. The loss function L2 of this network is:

[0116] L2 = γ · L′ cos + δ · L′ L1 + ε · L con

[0117]

[0118] where γ = 0.5, δ = 0.2, and ε = 0.3 are hyperparameters representing the weights of the control cosine similarity loss, per-pixel L1 norm loss, and consistency loss respectively, and L′ cos represents the cosine similarity loss of the residual optimization network, and L′ L1 represents the per-pixel L1 norm loss of the residual optimization network, and L con represents the consistency loss of the residual optimization network. represents the predicted normal tensor after adding the residual, is the ground truth normal tensor, is the normal vector of the i′2-th pixel in the predicted normal map, is the normal vector of the i′2-th pixel in the ground truth normal map, B is the batch dimension, and N Light is the number of lights included in each group. is the mean of the predicted normals within the group for the b-th batch. The predicted normals for all lights within the same group (same view) are averaged along the N Light dimension. n represents the n-th lighting condition in the current group, and h′ and w represent the h′-th row and w-th column of the current image respectively.

[0119] 4) During application, the monocular target color image is input into the preliminary normal estimation network trained in step 2) to obtain the preliminarily estimated normal image, which is then jointly input into the residual optimization network trained in step 3) together with the depth image and color image of the target to obtain the residual map. Finally, the preliminarily estimated normal image is added to the residual map and normalized to output the refined normal image with lighting consistency.

[0120] The specific steps for obtaining the normal estimation for the monocular color image are as follows: First, a monocular color image to be estimated for the normal is obtained. Then, the preliminary normal estimation network trained in step 2) is used to obtain the preliminary normal estimation map N pred from the monocular color image. Then, this normal estimation map, the monocular color image, and its corresponding depth image are used as inputs to the residual optimization network to obtain the residual map R res . Finally, the preliminarily estimated normal image is added to the residual map to obtain the optimized normal estimation map N refined :

[0121] N refined = N pred + R res .

[0122] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A monocular color image normal estimation method based on deep learning, characterized in that It includes the following steps: 1) Use a basic dataset, which includes multiple groups of indoor scene color images, their normal images, and depth images. For each view of the same indoor scene, there are multiple differently colored color images, their corresponding normal images, and depth images under different lighting conditions. The normal images and depth images are rendered and generated when the dataset is made; 2) Use the color images in the dataset in step 1) as the input to train a preliminary normal estimation network. This network is an improved U-Net network. Use the cosine similarity loss and absolute error loss between the normal image predicted by this network and the corresponding true normal image in the dataset as the loss function for network training, and improve the encoder and decoder of the network. The specific improvement is as follows: Use a pre-trained convolutional neural network as the encoder to extract multi-scale high-dimensional features from the input color image. The decoder uses bilinear interpolation for upsampling, combines weighted decoding with features at the corresponding level for fusion and splicing, and introduces a local attention mechanism during the fusion process. Finally, use the Tanh activation function to output the preliminarily estimated normal image; 3) Obtain the color images, depth images in the dataset in step 1), and the preliminarily estimated normal image obtained by applying the preliminary normal estimation network in step 2). Extract the curvature features from the depth image, and construct a 9-channel feature with the color image and the preliminarily estimated normal image as the input to train a residual optimization network. On the basis of the loss function in step 2), introduce a consistency loss as the loss function for residual optimization network training. This consistency loss is used to constrain the consistency of the normal estimation images of color images under the same view and different lighting conditions. This network adopts an encoder-decoder structure, introduces a dynamic scaling mechanism to dynamically adjust the residual amplitude, and the decoder restores the resolution through upsampling to output a residual map; 4) During application, input the monocular target color image into the preliminarily trained preliminary normal estimation network in step 2) to obtain the preliminarily estimated normal image, and then jointly input it with the depth image and color image of the target into the residual optimization network trained in step 3) to obtain a residual map. Finally, add the preliminarily estimated normal image and the residual map and normalize them to output a refined normal image with lighting consistency; 2. The monocular color image normal estimation method based on deep learning according to claim 1, wherein In step 1), normalize the color images and normal images in the dataset. Normalize the color images to the range of [0, 1], and normalize the normal images to the range of [-1, 1]. Then perform operations of unified size and data augmentation, including random color jitter and rotation. In the scenes of this dataset, there are a total of 11 different lighting combinations between ambient light, indoor light sources, and outdoor light sources, thereby improving the robustness of the network; 3. The monocular color image normal estimation method based on deep learning according to claim 2, wherein In step 2), train the preliminary normal estimation network in a supervised learning manner. This network is an improved U-Net network. The input is the color images in the dataset, and the output is the preliminarily estimated normal image. The specific situation of this network is as follows: The initial normal estimation network uses the U-Net network as the backbone network. First, a pre-trained convolutional neural network with the global pooling layer and classification layer removed is used as the encoder, and encoded features at each level are extracted from the color image through five downsamplings: F i = Encoder(F i-1 ), i ∈ {1, 2, 3, 4, 5} Wherein, I rgb is the input color image, H and W are the height and width of the color image respectively, R is the real number field, F0 is the initial encoded feature input from the network to the encoder, F i is the i-th layer encoded feature extracted by the encoder. There are five downsamplings in total, so i ∈ {1, 2, 3, 4, 5}. The resolution of each layer of encoded feature gradually decreases from 256×256 to 16×16, and the number of channels increases from 32 to 2560. F i-1 is the (i-1)-th layer encoded feature extracted by the encoder, and Encoder is the encoder; Secondly, the decoder of the network contains four levels of upsampling blocks. Each upsampling block gradually restores the resolution of the encoded features from 16×16 to 256×256. First, bilinear interpolation upsampling is used to upsample the encoded features to obtain the upsampled feature U i+1 , and then the encoder feature F i is concatenated with the upsampled feature U i+1 along the channel dimension, and a convolutional layer is used to reduce the number of channels and fuse cross-scale information to generate the fused feature X fused : X fused,j = Conv(Concat(F j , U j+1 ), j ∈ {1, 2, 3, 4} Wherein, is the low-resolution feature from the (j + 1)-th layer of the decoder, and U j+1 is the feature after upsampling in the (j + 1)-th layer of the decoder, is the feature map with the lowest resolution in the decoder, that is, the first layer of the decoder, and the input is the feature F5 of the last layer of the encoder. X fused,j is the fused feature generated in the j-th layer of the decoder, and F j is the encoded feature of the j-th layer extracted by the encoder. Since only the first four layers of the encoder are spliced with the decoder through skip connections to generate the fused feature, so j ∈ {1, 2, 3, 4}. Interpolate is the bilinear interpolation operation, Conv represents the convolution operation, and Concat represents the channel splicing operation; then in each upsampling block, the corresponding fused feature X fused is introduced into the local attention mechanism for attention enhancement operation. The specific steps are to split it into local blocks through convolution operation and generate embedding vectors, and then flatten it into a one-dimensional Token sequence t along the spatial dimension: Wherein, P is the size of the local block, E represents the local block embedding tensor extracted from the fused features through a convolution operation, each element of which represents an embedding vector of a local block, L is the embedding dimension, is the Token sequence length, Flatten is a function to flatten a continuous range of dimensions, and t represents the one-dimensional original Token sequence after flattening; the self-attention mechanism is introduced to extract the global attention features from this one-dimensional Token sequence t, which consists of multi-head self-attention and a multi-layer perceptron, and the one-dimensional Token sequence is input into a four-layer Transformer encoder. In each layer of the network: Q = MLP q (t l-1 ) K = MLP k (t l-1 ) V = MLP v (t l-1 ) Wherein, Q, K, and V are the query matrix, key matrix, and value matrix respectively, and MLP q , MLP k , MLP v are the global perceptrons corresponding to Q, K, and V respectively, and t l-1 , l ∈ {1, 2, 3, 4} is the one-dimensional encoded Token sequence of the input or the feature of the previous layer; subsequently, the weighted self-attention is calculated from the output Q, K, and V: where SA is the self-attention function, t l is the attention feature of the current layer, A softmax is the Softmax activation function, and D represents z-score normalization to stabilize the gradient; then the multi-head self-attention module is used to calculate the self-attention of multiple heads simultaneously, and the calculation results are concatenated, and finally a residual connection is added, and the formula is as follows: MultiHead(t l ) = Concat(head1, head2,..., head h ) + t l-1 head k = SA(LN(t k ), k ∈ {1, 2, 3, 4} where MultiHead is the multi-head self-attention function, and head k is the output of the k-th attention head, and there are a total of h = 4 attention heads. LN is layer normalization; after each multi-head self-attention operation, a global perceptron and a residual connection are connected to obtain the attention feature of the current layer: t l = MLP(LN(MultiHead(t l-1 )))+t l-1 The one-dimensional encoded Token sequences of each layer of the decoder are readjusted into two-dimensional global attention features through channel adjustment, and bilinearly interpolated and upsampled to the same resolution as the high-dimensional encoded features, so as to obtain the attention map f A : f A,l = Interpolate(View(t l )) where View is the channel adjustment operation, and f A,l is the attention map of the current layer; in each layer of the decoder, the fused features and the attention map are subjected to attention weighting through the Hadamard product to obtain weighted encoded features, and residual connections are included. In the last layer, a convolutional neural network is used to refine the weighted encoded features with the highest resolution, and it is extended to the size of the input image through bilinear interpolation. Finally, the estimated normal image is output using the Tanh activation function: i normal = A tanh (Conv(X′ fused ⊙ f′ A + Conv(X′ fused ))) where i normal is the finally output normal image, A tanh is the Tanh activation function, ⊙ is the Hadamard product operation, Conv(X′ fused ) serves as the residual connection, where X′ fused and f A ′ respectively represent the fused feature and the attention map obtained in the last upsampling block of the decoder; then the cosine similarity loss function and the absolute error loss function are used to calculate the loss L1 between the output normal image and the corresponding ground truth normal image in the dataset, and are used to train the network: L1 = α·L cos (G, G * ) + β·L L1 (G, G * ) where α and β are hyperparameters for controlling the loss range, G is the normal tensor predicted by the preliminary normal estimation network, and G * is the ground truth normal tensor in the dataset, and L cos is the cosine similarity loss function, and L L1 is the per-pixel L1 norm loss, i.e., the absolute error loss function. The calculation methods of the two functions are as follows: where N is the number of valid pixels in the image, and G i′ is the normal vector of the i'-th pixel in the normal map predicted by the preliminary normal estimation network, and is the normal vector of the i'-th pixel in the ground truth normal map.

4. The monocular color image normal estimation method based on deep learning according to claim 3, characterized in that In step 3), the residual optimization network is trained in a supervised learning manner. This network is an improved encoder-decoder structure network. The input is a monocular color image, the corresponding depth image, and the initially estimated normal image, and the output is a residual map. The specific situation of this network is as follows: First, obtain the color image I of the dataset in step 1) rgb , the depth image I depth , and the initially estimated normal image N obtained through the initial normal estimation network in step 2) pred as the network input, where the horizontal gradient G of the depth image is calculated using the Sobel operator X and the vertical gradient G Y : G X = Conv2D(Pad horizontal (I depth ), K X ) G Y = Conv2D(Pad vertical (I depth ), K Y ) where K X = [[-1,0,1]] / 2 is the horizontal Sobel kernel, and K Y = [[-1],[0],[1]] / 2 is the vertical Sobel kernel, Pad horizontal and Pad vertical are mirror padding in the horizontal and vertical directions respectively, and then the curvature map C is calculated by discrete second-order differences: C = (I depth:,:,:,2: - 2I depth:,:,:,1:-1 + I depth:,:,:,:-2 ) + (I depth:,:,2:,: - 2I depth:,:,1:-1,: + I depth:,:,:-2,: ) In the formula, (I depth:,:,:,2: - 2I depth:,:,:,1:-1 + I depth:,:,:,:-2 ) calculates the curvature in the horizontal direction of the depth image, (I depth:,:,2:,: - 2I depth:,:,1:-1,: + I depth:,:,:-2,: ) calculates the curvature in the vertical direction of the depth image, I depth:,:,:,2: represents taking the pixel values of the depth image from the 3rd column to the last column in the width direction, I depth:,:,:,1:-1 represents taking the pixel values of the depth image from the 2nd column to the second - last column in the width direction, I depth:,:,:,:-2 represents taking the pixel values of the depth image from the 1st column to the third - last column in the width direction, I depth:,:,2:,: represents taking the pixel values of the depth image from the 3rd row to the last row in the height direction, I depth:,:,1:-1,: represents taking the pixel values of the depth image from the 2nd row to the second - last row in the height direction, I depth:,:,:-2,: represents taking the pixel values of the depth image from the 1st row to the third - last row in the height direction; and stitching along the channel dimension to generate the joint feature map X: X = Concat(N pred , G X (I depth ), G Y (I depth ), C(I depth ), I rgb ) The encoder of this network consists of four convolutional layers and one downsampling layer. The joint feature map X generated from the network input serves as the input to the encoder. The decoder restores the spatial resolution through bilinear interpolation upsampling and convolutional operations, and the last layer uses the Tanh activation function to constrain the value range to [-1, 1] to output the basic residual map R base ; The improved part is the introduction of a dynamic scaling mechanism to finely adjust the residuals: S = Tanh(Conv(R base )) ⊙ s global where s global is a learnable parameter, R base represents the base residual map output by the decoder, and S is the spatially-varying scaling factor map generated by the sub-network; then, performing a Hadamard product operation on the base residual can obtain the residual map R res : R res = R base ⊙ S In the training of the residual optimization network, images with different illuminations under the same perspective are grouped for training. When calculating the loss, a consistency loss constraint is introduced more than in the initial normal estimation network to train the network. The loss function L2 of this network is: L2 = γ·L′ cos + δ·L′ L1 + ε·L con where γ, δ, and ε respectively represent the hyperparameters that control the cosine similarity loss, the per-pixel L1 norm loss, and the consistency loss weight, and L c ′ os represents the cosine similarity loss of the residual optimization network, and L L ′1 represents the per-pixel L1 norm loss of the residual optimization network, and L con represents the consistency loss of the residual optimization network, represents the predicted normal tensor after adding the residual, is the ground truth normal tensor, is the normal vector of the i2′-th pixel in the predicted normal map, is the normal vector of the i2′-th pixel in the ground truth normal map, B is the batch dimension, and N Light is the number of lights included in each group, is the mean of the predicted normals within the group for the b-th batch. The predicted normals for all lights within the same group, i.e., the same viewing angle, are averaged along the N Light dimension. n represents the n-th lighting condition in the current group, and h′ and w respectively represent the h′-th row and the w-th column of the current image.

5. The monocular color image normal estimation method based on deep learning according to claim 4, wherein In step 4), the specific steps for obtaining the normal estimation for a monocular color image are as follows: First, obtain a monocular color image for which the normal is to be estimated. Then, use the preliminary normal estimation network trained in step 2) to obtain a preliminary normal estimation map N from the monocular color image pred , and then use this normal estimation map, the monocular color image, and its corresponding depth image as inputs to obtain a residual map R using the residual optimization network res . Finally, add the preliminarily estimated normal image and the residual map to obtain the optimized normal estimation map N refined : N refined = N pred + R res 。