A photometric stereo reconstruction method based on coaxial light guidance
The coaxial light guidance method in light-field stereoscopy decouples material reflectance effects to accurately estimate surface normals on non-Lambertian surfaces, improving feature representation and spatial structure preservation.
Patent Information
- Application Number
- CN202510631933.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-05-16
AI Technical Summary
When the existing photometric stereoscopic method deals with non-Lambertian surfaces, the normal vector estimation accuracy is not high, and the existing method is difficult to effectively decouple the normal vector information and reflection characteristics, resulting in unsatisfactory estimation effect of complex region.
The photometric stereo reconstruction method with coaxial light-guided is adopted. By constructing the tensor matrix input, the cross-image local and global brightness perception module, shallow texture feature aggregator and deep structure feature extractor are used to separate the light and dark changes, generate global and deep structural features, and finally obtain the surface normal estimation through the normal regressor.
It realizes high-precision normal vector estimation on non-Lambertian surfaces, explicitly decouples normal vector information and reflection characteristics, improves the accuracy and robustness of feature representation, and can capture details and structural information in the image.
Smart Images

Figure CN120147565B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of three-dimensional measurement and computer technology, and particularly relates to a photometric stereo reconstruction method based on coaxial light guidance. Background Art
[0002] Photometric stereo is a non-destructive three-dimensional measurement technology, which is widely used in industrial inspection, endoscopic inspection, biometric identification and other fields. Its basic idea is to use multiple images collected under different illumination conditions to accurately estimate the normal vectors of the three-dimensional object surface. The core of this method lies in inferring the normal vectors of surface points based on the light and dark change clues generated by the illumination and the object geometry. However, most object surfaces in the real world have non-Lambertian reflection characteristics, and non-Lambertian reflection will destroy the light and dark clues generated by the combined action of illumination and object geometry, thereby affecting the accurate estimation of normal vectors. Therefore, achieving high-precision normal vector reconstruction on non-Lambertian surfaces has always been an important challenge in the research of photometric stereo technology.
[0003] Traditional photometric stereo methods deal with non-Lambertian surfaces by establishing a bidirectional reflectance distribution function (BRDF) model of the object surface or removing abnormal regions. However, the applicability of these traditional methods is relatively limited, and they only have certain effects on the surfaces of objects with specific materials. Photometric stereo methods based on deep learning have shown superior performance in dealing with non-Lambertian surfaces. According to the imaging physical model, the BRDF characteristics of the surface and the normal vectors are coupled and interrelated.
[0004] Existing technologies usually combine the physical imaging model with neural networks. These methods usually rely on neural networks to numerically fit the surface BRDF values, but they have not achieved explicit decoupling of the normal vector information and the reflection characteristics, and it is difficult to fundamentally eliminate the interference of spatially varying material characteristics on normal vector estimation. In addition, these methods mostly use a pixel-by-pixel method for normal vector prediction, which destroys the spatial structure information in the image, and the normal vector estimation effect for complex regions such as wrinkles is still not ideal. Summary of the Invention
[0005] To this end, the present invention provides a photometric stereo reconstruction method based on coaxial light guidance to solve the problems raised in the background art.
[0006] To achieve the above object, the present invention provides the following technical solution: A photometric stereo reconstruction method based on coaxial light guidance, comprising:
[0007] Step 1, splicing the illumination direction, the observed image, and the coaxial light image to construct a tensor matrix as the input of the network;
[0008] Step 2: The cross-image local brightness perception module first extracts the local brightness change features of the input tensor under each illumination direction. Subsequently, the cross-image global brightness perception module realizes the information interaction and fusion between local features under different illumination perspectives through three-dimensional convolution, so as to generate a global brightness change feature representation;
[0009] Step 3: The shallow texture feature aggregator first extracts the shallow local spatial information from the brightness change features within the image. Subsequently, the local features are aggregated through a max pooling operation to generate a global shallow spatial feature representation;
[0010] Step 4: The deep structure feature extractor further extracts the deep local structure features from the shallow feature map and realizes feature fusion through a max pooling operation, and finally generates a deep full-structure feature containing rich structure information;
[0011] Step 5: Input the extracted deep full-structure features into the normal regression estimator to obtain the surface normal estimation result.
[0012] Preferably, in Step 1, a coaxial light image is introduced, and the illumination direction, the observation image, and the coaxial light image are stacked as a tensor as the network input ; The coaxial light refers to the light that is collinear with the camera observation direction or the included angle between them is not greater than 10°. Based on the image imaging principle, the coaxial light is used to guide the network to fit the retroreflective model, so as to separate the brightness change features that are not affected by the material. The expression is as follows:
[0013] ;
[0014] where, I is the normalized pixel value of the observation image, I c is the pixel intensity of the coaxial light image, n is the surface normal vector, l is the light source direction, v is the observation direction, l T v represents the inner product of the illumination direction and the observation direction, is the brightness change, is the retroreflective model of the observation image, is the retroreflective model of the coaxial light, is the composite retroreflective model.
[0015] Preferably, in Step 2, the cross-image local brightness perception module under each illumination direction adopts a weight-sharing design, which consists of two serial 1×1 two-dimensional convolutional layers, and a Leaky-ReLU activation function is connected after each convolutional layer to extract the local brightness change features, as shown in the following formula:
[0016] ;
[0017] Among them, is the cross-image local brightness perception module under the i th light direction, θ are the parameters of the cross-image local brightness perception module, and the input is the i th tensor under the light, Φ i , and the output is the local brightness change feature ψ i .
[0018] Preferably, in step 2, a cross-image global brightness perception module is used to fuse the brightness change features under each light direction. This module is based on a three-dimensional convolutional structure, combined with the Leaky-ReLU activation function and the Dropout regularization mechanism to enhance the non-linear expression ability of the features and effectively suppress the overfitting phenomenon. Specifically, all local brightness change features ψ i are stacked , as the input ψ all ; Subsequently, three-dimensional convolutional kernels with sizes of 1×1×1, 3×1×1, and 1×1×1 are sequentially used to process the input features to realize the interaction and fusion of cross-image local brightness change information, and finally extract the global brightness change feature representation;
[0019] ;
[0020] Among them, is the cross-image global brightness perception module, θ are the parameters of the cross-image global brightness perception module, the input is ψ all , and the output global brightness change feature representation is P all .
[0021] Preferably, in step 3, the shallow texture feature aggregator realizes feature extraction by stacking three two-dimensional convolutional modules with Leaky-ReLU activation, and divides the global brightness change feature P all into { P 1, P 2, …, P n}, and splices it with the local brightness change feature ψ i as the input of the shallow texture feature aggregator under each lighting condition;
[0022] ;
[0023] Among them, is a shallow texture feature extractor, θ are the parameters of the shallow texture feature extractor, and the input is concat( ψ i , P i ). The output shallow local texture feature is . Then, all the shallow local texture features are aggregated through a max pooling operation to generate a shallow global texture feature .
[0024] Preferably, in step 4, the deep structure feature extractor includes three modules, and each module is successively composed of a depthwise separable convolution, layer normalization, a linear transformation, and a GELU activation function. Decomposing the standard convolution into a per-channel convolution and a pointwise convolution significantly reduces the number of parameters and the computational cost. The shallow global texture feature is fused with each shallow local texture feature , and the fused feature is used as the input of the deep structure feature extractor to extract deep structure features.
[0025] ;
[0026] Among them, is the deep structure feature extractor, θ are the parameters of the deep structure feature extractor, and the input is . The output deep local structure feature is . Then, all the deep local structure features are fused through a max pooling operation to generate a deep global structure feature .
[0027] Preferably, in step 5, the normal regressor is used to decode the features output by the feature extractor into surface normals. The structure of the normal regressor includes 3 convolutional layers, 1 deconvolutional layer, and 1 L2 normalization layer. The deep global structure feature is input into the normal regressor to obtain a surface normal estimation result:
[0028] ;
[0029] Among them, R is the normal regressor, θ are the parameters of the normal regressor, and the input is the deep global structure feature . The predicted normal vector is .
[0030] The present invention has the following advantages:
[0031] The present invention stitches together the illumination direction, the observed image, and the coaxial light image to construct a tensor matrix as the input of the network. The cross-image local brightness perception module first extracts the local brightness change features of the input tensor under each illumination direction. Subsequently, the cross-image global brightness perception module realizes the information interaction and fusion between local features under different illumination perspectives through three-dimensional convolution, thereby generating a global brightness change feature representation. The shallow texture feature aggregator first extracts the shallow local spatial information from the brightness change features within the image, and then aggregates the local features through a max-pooling operation to generate a global shallow spatial feature representation. The deep structure feature extractor further extracts the deep local structure features from the shallow feature maps and realizes feature fusion through a max-pooling operation, finally generating the deep global structure features. The extracted deep full-structure features are input into the normal regression estimator to obtain the surface normal estimation result. Compared with the prior art, during the image feature extraction process, the shallow texture feature aggregator adopts a hierarchical strategy, operating on local and global features respectively to accurately extract the light and dark change features within the image. This method can comprehensively mine the detailed information and overall information in the image, without relying on the neural network to numerically fit the surface BRDF value, realizing the explicit decoupling of the normal vector information and the reflection characteristics, fundamentally eliminating the interference of spatially varying material characteristics on the normal vector estimation, and greatly improving the accuracy and robustness of the feature representation. Similarly, the deep structure feature extractor also uses the local-global hierarchical extraction method to deeply analyze the deep structure of the image, and through the fine capture of local features, obtain the structural details at the microscopic level in the image, such as the subtle twists and turns of the object edges, the microscopic texture orientation of the material surface, etc., avoiding the destruction of the spatial structure information in the image. Description of the Drawings
[0032] Figure 1 It is a schematic flow chart of the method of the present invention.
[0033] Figure 2 It is the overall architecture diagram of the present invention.
[0034] Figure 3 It is the cross-image brightness perception module of the present invention.
[0035] Figure 4 It is the shallow texture feature extractor and the deep structure feature extractor of the present invention.
[0036] Figure 5 It is the reconstruction diagram of the present invention on the photometric stereo DILIGENT test set. Detailed Embodiments
[0037] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts fall within the scope of protection of the present invention.
[0038] As Figure 1 shown, this embodiment provides a photometric stereo reconstruction method based on coaxial light guidance, including:
[0039] Step 1, splice the illumination direction, the observed image, and the coaxial light image to construct a tensor matrix as the input of the network;
[0040] Step 2, the cross-image brightness perception module first extracts the local brightness change features of the input tensor under each illumination direction, and then realizes the information interaction and fusion between local features under different illumination perspectives through three-dimensional convolution, so as to generate a brightness change feature representation with global perception ability;
[0041] Step 3, the shallow texture feature aggregator first extracts the shallow local spatial information from the brightness change features in the image, and then aggregates the local features through max-pooling operation to generate a shallow spatial feature representation with global perception ability;
[0042] Step 4, the deep structure feature extractor further extracts the deep local spatial features from the shallow feature maps and realizes feature fusion through max-pooling operation, and finally generates a deep global spatial feature containing rich structure information;
[0043] Step 5, input the extracted deep structure features into the normal regression estimator to obtain the surface normal estimation result.
[0044] The present invention trains the network on a coaxial light dataset. These datasets contain complex object surfaces with various non-Lambertian reflection characteristics, and fully consider real imaging phenomena such as shadows, highlights, and multiple reflections. The training dataset contains a total of 114,488 samples, and each sample contains 64 observed images under different illumination directions and 1 coaxial light image. The dataset is divided into a training set (113,343 samples) and a validation set (1,145 samples) according to a ratio of 99:1. In addition, the typical public dataset DILIGENT in the real photometric stereo field is selected as the test set to evaluate the performance of the method of the present invention.
[0045] In an exemplary instance, the observed images under different illumination conditions, Normalize it. At the same time, encode and transform each 1×3 illumination direction vector into an illumination matrix with the same spatial resolution as the observed image. Stack the normalized observed image , the coaxial light image , and the illumination direction , and input them into the network shown in Figure 2 . Based on the image imaging principle, use the coaxial light to guide the network to fit the retroreflective model, so as to separate the light and dark change features that are not affected by the material. The expression is as follows:
[0046] ;
[0047] where I is the normalized pixel value of the observed image, I c is the pixel intensity of the coaxial light image, n is the surface normal vector, l is the light source direction, v is the observation direction, is the light and dark change, is the retroreflective model of the observed image, is the retroreflective model of the coaxial light, is the composite retroreflective model.
[0048] Step 2: The cross-image brightness perception module first extracts the local brightness change features of the input tensor under each illumination direction, and then realizes the information interaction and fusion between local features under different illumination perspectives through three-dimensional convolution, so as to generate a brightness change feature representation with global perception ability.
[0049] In an exemplary example, as Figure 3 shown, the local brightness change feature extractor under each illumination direction adopts a weight-sharing design, which consists of two stacked two-dimensional convolutional layers, and a Leaky-ReLU activation function is connected after each convolutional layer to extract local photometric features, as shown in the following formula:
[0050] ;
[0051] where is the cross-image local brightness perception module under the i th illumination direction, θ is the parameter of the cross-image local brightness perception module, and the input is the tensor i under the Φ i th illumination, and the output local brightness change feature ψ i .
[0052] In an exemplary instance, as Figure 3 shown, a cross-image global brightness perception module is used to fuse the brightness change features under each illumination direction. This module is based on a three-dimensional convolutional structure, combined with a Leaky-ReLU activation function and a Dropout regularization mechanism to enhance the non-linear expression ability of features and effectively suppress the overfitting phenomenon. Specifically, all local brightness change features ψ i are stacked , as the input ψ all ;
[0053]
[0054] Subsequently, three-dimensional convolutional kernels with sizes of 1×1×1, 3×1×1, and 1×1×1 are successively used to process the input features to realize the interaction and fusion of cross-image local brightness change information, and finally, a brightness change feature representation with global perception ability is extracted;
[0055] ;
[0056] Among them, is the cross-image global brightness perception module, θ is the parameter of the cross-image global brightness perception module, the input is ψ all , and the output global brightness change feature is P all .
[0057] Step 3: The shallow texture feature extractor first extracts the shallow image local texture information in the image, and then uses the max pooling operation to fuse the shallow local texture features, thereby generating shallow global texture features;
[0058] In an exemplary instance, as Figure 4 shown, the shallow texture feature extractor realizes feature extraction by stacking three two-dimensional convolutional modules with Leaky-ReLU activation, divides the global photometric feature P all into . And it is concatenated with the local brightness change feature ψ i to form , as the input of the shallow texture feature extractor under each illumination condition;
[0059] ;
[0060] Among them, is the shallow texture feature extractor, θare the parameters of the shallow texture feature extractor. The input is concat( ψ i , P i ), and the output shallow local texture feature is . Then, all the shallow local texture features are fused through a max-pooling operation to generate the shallow global texture feature .
[0061]
[0062] Step 4: The deep structure feature extractor extracts the deep local structure features in the shallow feature map, and then uses the max-pooling operation to fuse the deep local structure features, thereby generating the deep global structure feature;
[0063] In an exemplary instance, as Figure 4 shown, the deep structure feature extractor includes three modules. Each module consists of a depthwise separable convolution, layer normalization, linear transformation, and GELU activation function in sequence. The shallow global texture feature is fused with each shallow local texture feature for fusion , and the fused feature is used as the input of the deep structure feature extractor to extract the deep structure feature.
[0064] ;
[0065] Among them, is the deep structure feature extractor, θ are the parameters of the deep structure feature extractor. The input is , and the output deep local structure feature is . Then, all the deep local structure features are fused through a max-pooling operation to generate the deep global structure feature .
[0066]
[0067] Step 5: Input the extracted deep structure features into the normal regression network to obtain the surface normal estimation result.
[0068] In an exemplary instance, as Figure 2 shown, the normal regression network is used to decode the features output by the feature extractor into surface normals. The structure of the normal regression network includes 3 convolutional layers, 1 deconvolutional layer, and 1 L2 normalization layer. The deep structure feature map is input into the normal regression network to obtain the surface normal estimation result:
[0069] ;
[0070] Among them, R is the normal regressor, θ are the parameters of the normal regressor, and the input is , and the predicted normal vector is .
[0071] In the present invention, by minimizing the error between the normal vector predicted by the normal regressor and the true normal vector, a loss function is constructed to drive the optimization of the network parameters. During the training process, the AdamW optimizer is adopted and its default parameter settings are used. The initial learning rate is set to 0.001 and is decayed at a factor of 0.5 every 5 epochs. The batch size is set to 64 and a total of 35 epochs are trained. After the model training is completed, its weight parameters are saved. Finally, the trained model is used to test the DILIGENT real dataset to verify the normal vector estimation effect of the method of the present invention in the real scenario.
[0072] As Figure 5 shown, it is a comparison diagram of the normal vector reconstruction effect of this embodiment. Among them, the first row is the RGB images of 10 objects in the DILIGENT dataset; the second row is the corresponding true normal vector map; the third row is the result map of the object surface normal vector reconstructed by the method of the present invention; the fourth row is the error distribution map of the reconstructed normal vector and the true normal vector. From Figure 5 the results, it can be seen that the photometric stereo reconstruction method based on coaxial light guidance proposed by the present invention can obtain high normal vector estimation accuracy in different objects and scenarios with complex reflection characteristics, verifying the effectiveness and superiority of this method.
[0073] Although the present invention has been described in detail above with general descriptions and specific embodiments, based on the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection required by the present invention.
Claims
1. A photometric stereo reconstruction method based on coaxial light guiding, characterized in that: Including: Step 1: Stitch the illumination direction, the observed image, and the coaxial light image to construct a tensor matrix as the input of the network. Step 2: The cross-image local brightness perception module first extracts the local brightness change features of the input tensor under each illumination direction. Subsequently, the cross-image global brightness perception module realizes the information interaction and fusion between local features under different illumination perspectives through three-dimensional convolution, thereby generating a global brightness change feature representation. The cross-image local brightness perception module under each illumination direction adopts a weight-sharing design and consists of two serial 1×1 two-dimensional convolutional layers, and a Leaky-ReLU activation function is connected after each convolutional layer to extract local brightness change features, as shown in the following formula: Among them, is the cross-image local brightness perception module under the i-th illumination direction, θ is the parameter of the cross-image local brightness perception module, and the input is the tensor Φ under the i-th illumination i , and the output is the local brightness change feature ψ i ; The cross-image global brightness perception module is used to fuse the brightness change features under each illumination direction; based on the three-dimensional convolution structure, this module combines all local brightness change features ψ i and stacks them {ψ1, ψ2, …, ψ n} as the input ψ all ; subsequently, three-dimensional convolutional kernels with sizes of 1×1×1, 3×1×1, and 1×1×1 are successively used to process the input features, realizing the interaction and fusion of cross-image local brightness change information, and finally extracting the representation with global brightness change features; Among them, is a cross-image global brightness perception module, θ is the parameter of the cross-image global brightness perception module, and the input is ψ all , and the global brightness change feature of the output is represented as P all ; Step 3: The shallow texture feature aggregator first extracts the shallow local spatial information from the brightness change features in the image, and then aggregates the local features through a max pooling operation to generate a global shallow spatial feature representation. The shallow texture feature aggregator realizes feature extraction by stacking three two-dimensional convolutional modules with Leaky-ReLU activation, and represents the global brightness change feature P all is divided into {P1, P2, …, P n}, and is concatenated with the local brightness change feature ψ i as the input of the shallow texture feature aggregator under each illumination condition, and each shallow local texture feature is obtained Among them, is a shallow texture feature extractor, θ is the parameter of the shallow texture feature extractor, and the input is concat(ψ i , P i ), and the output shallow local texture feature is Then, all shallow local texture features are aggregated through a max pooling operation to generate a shallow global texture feature Step 4: The deep structure feature extractor further extracts the deep local structure features from the shallow feature map and realizes feature fusion through a max pooling operation, and finally generates the deep global structure features. Step 5: Input the extracted deep full-structure features into the normal regression estimator to obtain the surface normal estimation result.
2. A photometric stereo reconstruction method based on coaxial light guidance according to claim 1, characterized in that: In step 1, a coaxial light image is introduced, and the illumination direction, the observed image, and the coaxial light image are stacked as a tensor as the network input Φ i ; Based on the image imaging principle, the coaxial light is used to guide the network to fit the retroreflective model, so as to separate the light and dark change features that are not affected by the material. The expression is as follows: where I is the normalized pixel value of the observed image, I c is the pixel intensity of the coaxial light image, n is the surface normal vector, l is the light source direction, v is the observation direction, and l T ·v represents the inner product of the illumination direction and the observation direction, is the light and shade change, f(·) is the retroreflection model of the observed image, F(·) is the retroreflection model of the coaxial light, is the composite retroreflection model.
3. A photometric stereo reconstruction method based on coaxial light guidance according to claim 1, characterized in that: In step 4, the deep structure feature extractor consists of three modules. Each module is successively composed of depthwise separable convolution, layer normalization, linear transformation, and GELU activation function, decomposing the standard convolution into pointwise convolution and depthwise convolution, and fusing the shallow global texture features with each shallow local texture feature to obtain the fused features which are used as the input of the deep structure feature extractor to extract deep local structure features; Among them, is a deep structure feature extractor, θ is the parameter of the deep structure feature extractor, and the input is The output deep local structure features are Then, all deep local structure features are fused through a max pooling operation to generate deep global structure features 4. A photometric stereo reconstruction method based on coaxial light guidance according to claim 1, characterized in that: In step 5, the normal regressor is used to decode the features output by the feature extractor into surface normals, and the structure of the normal regressor inputs the deep global structural features into the normal regressor to obtain the surface normal estimation result: Where R is the normal regressor, θ is the parameter of the normal regressor, and the input is the deep global structural feature The predicted normal vector is
Citation Information
Patent Citations
Face image eye completion method based on self-attention mechanism model generative adversarial network
CN111738940A
Lightweight monocular image depth estimation method and device
CN119941817A