Photometric three-dimensional reconstruction method based on coaxial light guide
By adopting a coaxial light guidance method in photometric stereoscopic technology, extracting features when processing non-Lambertian surfaces and performing normal vector estimation, the problem of insufficient normal vector reconstruction accuracy in the prior art is solved, and higher accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510631933.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-16
AI Technical Summary
The existing photometric stereoscopic technology is difficult to achieve high-precision normal vector reconstruction when processing non-Lambertian surfaces, and the traditional method has limited applicability, and deep learning methods have not yet effectively decoupled normal vector information and reflection characteristics.
The photometric stereo reconstruction method based on coaxial light guidance is adopted. By splicing the light direction, observed images and coaxial light images into tensor matrices as network input, the features are extracted using the cross-image local brightness perception module, the cross-image global brightness perception module, the shallow texture feature aggregator and the deep structure feature extractor, and the normal regressor is finally input to obtain the surface normal estimation result.
The explicit decoupling of normal vector information and reflection characteristics is realized, eliminating the interference of spatially changing material characteristics on normal vector estimation, greatly improving the accuracy and robustness of feature representation.
Smart Images

Figure CN120147565A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of three-dimensional measurement and computer technology, and particularly relates to a photometric stereo reconstruction method based on coaxial light guidance. Background Art
[0002] Photometric stereo is a non-destructive three-dimensional measurement technique widely used in industrial inspection, endoscopic inspection, biometric identification and other fields. Its basic idea is to use multiple images acquired under different lighting conditions to accurately estimate the normal vectors of the three-dimensional object surface. The core of this method lies in inferring the normal vectors of surface points based on the brightness change clues generated by the illumination and the object geometry. However, most object surfaces in the real world have non-Lambertian reflection characteristics, and non-Lambertian reflection will destroy the brightness clues generated by the combined action of illumination and object geometry, thereby affecting the accurate estimation of normal vectors. Therefore, achieving high-precision normal vector reconstruction on non-Lambertian surfaces has always been an important challenge in the research of photometric stereo technology.
[0003] Traditional photometric stereo methods deal with non-Lambertian surfaces by establishing a bidirectional reflectance distribution function (BRDF) model of the object surface or removing abnormal regions. However, the applicability of these traditional methods is relatively limited, and they only have certain effects on the surfaces of objects with specific materials. Photometric stereo methods based on deep learning have shown superior performance in dealing with non-Lambertian surfaces. According to the imaging physical model, the BRDF characteristics of the surface and the normal vectors are coupled and interrelated.
[0004] Existing technologies usually combine the physical imaging model with neural networks. These methods usually rely on neural networks to numerically fit the surface BRDF values, but they have not achieved explicit decoupling of the normal vector information and the reflection characteristics, and it is difficult to fundamentally eliminate the interference of spatially varying material characteristics on the normal vector estimation. In addition, these methods mostly use a pixel-by-pixel method for normal vector prediction, which destroys the spatial structure information in the image, and the normal vector estimation effect for complex regions such as wrinkles is still not ideal. Summary of the Invention
[0005] For this reason, the present invention provides a photometric stereo reconstruction method based on coaxial light guidance to solve the problems raised in the background art.
[0006] To achieve the above object, the present invention provides the following technical solution: A photometric stereo reconstruction method based on coaxial light guidance, comprising: Step 1, splicing the illumination direction, the observed image and the coaxial light image to construct a tensor matrix as the input of the network; Step 2, the cross-image local brightness perception module first extracts the local brightness change features of the input tensor under each illumination direction. Subsequently, the cross-image global brightness perception module realizes the information interaction and fusion between local features under different illumination perspectives through three-dimensional convolution, thereby generating a global brightness change feature representation; Step 3, the shallow texture feature aggregator first extracts the shallow local spatial information from the brightness change features within the image. Subsequently, it aggregates the local features through a max-pooling operation to generate a global shallow spatial feature representation; Step 4, the deep structure feature extractor further extracts the deep local structure features from the shallow feature maps and realizes feature fusion through a max-pooling operation, finally generating a deep full-structure feature containing rich structure information; Step 5, input the extracted deep full-structure features into the normal regression estimator to obtain the surface normal estimation result.
[0007] Preferably, in Step 1, a coaxial light image is introduced, and the illumination direction, the observed image, and the coaxial light image are stacked as a tensor as the network input ; The coaxial light refers to the light that is collinear with the camera observation direction or the included angle is no more than 10°. Based on the image imaging principle, the coaxial light is used to guide the network to fit the retroreflection model, thereby separating the brightness change features that are not affected by the material. The expression is as follows: ; Among them, I is the normalized pixel value of the observed image, I c is the pixel intensity of the coaxial light image, n is the surface normal vector, l is the light source direction, v is the observation direction, l T v represents the inner product of the illumination direction and the observation direction, is the brightness change, is the retroreflection model of the observed image, is the retroreflection model of the coaxial light, is the composite retroreflection model.
[0008] Preferably, in Step 2, the cross-image local brightness perception module under each illumination direction adopts a weight-sharing design, which consists of two serial 1×1 two-dimensional convolutional layers, and a Leaky-ReLU activation function is connected after each convolutional layer to extract the local brightness change features, as shown in the following formula: ; Among them, is the cross-image local brightness perception module under the i th illumination direction,θ are the parameters of the cross-image local brightness perception module. The input is the tensor under the i th illumination Φ i , and the output is the local brightness change feature ψ i .
[0009] Preferably, in step 2, a cross-image global brightness perception module is used to fuse the brightness change features in each illumination direction. This module is based on a three-dimensional convolutional structure, combined with the Leaky-ReLU activation function and the Dropout regularization mechanism to enhance the non-linear expression ability of features and effectively suppress the overfitting phenomenon. Specifically, all local brightness change features ψ i are stacked as the input ψ all ; Subsequently, three-dimensional convolutional kernels with sizes of 1×1×1, 3×1×1, and 1×1×1 are sequentially used to process the input features to achieve the interaction and fusion of cross-image local brightness change information, and finally extract the global brightness change feature representation; ; Among them, is the cross-image global brightness perception module, θ are the parameters of the cross-image global brightness perception module. The input is ψ all , and the output global brightness change feature representation is P all .
[0010] Preferably, in step 3, the shallow texture feature aggregator realizes feature extraction by stacking three two-dimensional convolutional modules with Leaky-ReLU activation. The global brightness change feature P all is divided into { P 1 , P 2 , …, P n}, and concatenated with the local brightness change feature ψ i as the input of the shallow texture feature aggregator under each illumination condition; ; Among them, is the shallow texture feature extractor, θ are the parameters of the shallow texture feature extractor. The input is concat( ψ i , Pi ) and the output shallow local texture features are , and then all the shallow local texture features are aggregated through max-pooling operation to generate shallow global texture features .
[0011] Preferably, in step 4, the deep structure feature extractor includes three modules, and each module is successively composed of depthwise separable convolution, layer normalization, linear transformation, and GELU activation function. The standard convolution is disassembled into pointwise convolution and depthwise convolution, significantly reducing the number of parameters and computational cost. The shallow global texture features are fused with each shallow local texture feature , and the fused features are used as the input of the deep structure feature extractor to extract deep structure features.
[0012] ; wherein, is the deep structure feature extractor, θ are the parameters of the deep structure feature extractor, the input is , and the output deep local structure features are , and then all the deep local structure features are fused through max-pooling operation to generate deep global structure features .
[0013] Preferably, in step 5, the normal regressor is used to decode the features output by the feature extractor into surface normals. The structure of the normal regressor includes 3 convolutional layers, 1 deconvolutional layer, and 1 L2 normalization layer. The deep global structure features are input into the normal regressor to obtain the surface normal estimation result: ; wherein, R is the normal regressor, θ are the parameters of the normal regressor, the input is the deep global structure feature , and the predicted normal vector is .
[0014] The present invention has the following advantages: The present invention splices the illumination direction, the observed image and the coaxial light image to construct a tensor matrix as the input of the network; the cross-image local brightness perception module first extracts the local brightness change characteristics of the input tensor under each illumination direction, and then the cross-image global brightness perception module realizes the information interaction and fusion between local features under different illumination angles through three-dimensional convolution, thereby generating a global brightness change feature representation; the shallow texture feature aggregator first extracts shallow local spatial information from the brightness change characteristics in the image, and then aggregates the local features through the maximum pooling operation to generate a global shallow spatial feature representation; the deep structure feature extractor further extracts deep local structure features from the shallow feature map, and realizes feature fusion through the maximum pooling operation to finally generate deep global structure features; the extracted deep full structure features are input into the normal regressor to obtain the surface normal estimation result. Compared with the prior art, in the image feature extraction process, the shallow texture feature aggregator adopts a hierarchical strategy to operate on local and global features respectively, so as to accurately extract the light and dark change characteristics in the image. This method can fully mine the detailed information and overall information in the image, without relying on the neural network to numerically fit the surface BRDF value, to achieve explicit decoupling of normal vector information and reflection characteristics, fundamentally eliminating the interference of spatially varying material characteristics on normal vector estimation, and greatly improving the accuracy and robustness of feature representation. Similarly, the deep structure feature extractor also uses the local-global hierarchical extraction method to deeply analyze the deep structure of the image, and obtain the structural details at the microscopic level in the image by finely capturing local features, such as the subtle twists and turns of the edge of the object, the microscopic texture direction of the material surface, etc., to avoid destroying the spatial structure information in the image. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a schematic diagram of the method flow of the present invention.
[0016] Figure 2 It is the overall architecture diagram of the present invention.
[0017] Figure 3 It is the cross-image brightness perception module of the present invention.
[0018] Figure 4 It is the shallow texture feature extractor and deep structure feature extractor of the present invention.
[0019] Figure 5 This is the reconstruction image of the present invention on the photometric stereo DILIGENT test set. DETAILED DESCRIPTION
[0020] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0021] As Figure 1 shown, this embodiment provides a photometric stereo reconstruction method based on coaxial light guidance, including: Step 1, splice the illumination direction, the observed image, and the coaxial light image to construct a tensor matrix as the input of the network; Step 2, the cross-image brightness perception module first extracts the local brightness change features of the input tensor under each illumination direction, and then realizes the information interaction and fusion between local features under different illumination perspectives through three-dimensional convolution, so as to generate a brightness change feature representation with global perception ability; Step 3, the shallow texture feature aggregator first extracts the shallow local spatial information from the brightness change features in the image, and then aggregates the local features through max-pooling operation to generate a shallow spatial feature representation with global perception ability; Step 4, the deep structure feature extractor further extracts the deep local spatial features from the shallow feature map and realizes feature fusion through max-pooling operation, and finally generates a deep global spatial feature containing rich structure information; Step 5, input the extracted deep structure features into the normal regression estimator to obtain the surface normal estimation result.
[0022] The present invention trains the network on a coaxial light dataset. These datasets contain complex object surfaces with various non-Lambertian reflection characteristics, fully considering real imaging phenomena such as shadows, highlights, and multiple reflections. The training dataset contains a total of 114,488 samples, and each sample contains 64 observed images under different illumination directions and 1 coaxial light image. The dataset is divided into a training set (113,343 samples) and a validation set (1,145 samples) according to a ratio of 99:1. In addition, the typical public dataset DILIGENT in the real photometric stereo field is selected as the test set to evaluate the performance of the method of the present invention.
[0023] In an exemplary example, the observed images , the coaxial light images under different illumination conditions are normalized. At the same time, each illumination direction vector with a dimension of 1×3 is encoded and transformed into an illumination matrix with the same spatial resolution as the observed image . The normalized observed images , Coaxial light image , Illumination direction Stacked , Input into Figure 2 the network shown in, based on the image imaging principle, using coaxial light to guide the network to fit the retroreflective model, so as to separate the light and dark change features that are not affected by the material, and the expression is as follows: ; Among them, I is the normalized pixel value of the observed image, I c is the pixel intensity of the coaxial light image, n is the surface normal vector, l is the light source direction, v is the observation direction, is the light and dark change, is the retroreflective model of the observed image, is the retroreflective model of the coaxial light, is the composite retroreflective model.
[0024] Step 2, The cross-image brightness perception module first extracts the local brightness change features of the input tensor under each illumination direction, and then realizes the information interaction and fusion between local features under different illumination perspectives through three-dimensional convolution, so as to generate a brightness change feature representation with global perception ability.
[0025] In an exemplary instance, as Figure 3 shown, the local brightness change feature extractor under each illumination direction adopts a weight-sharing design, consists of two stacked two-dimensional convolutional layers, and a Leaky-ReLU activation function is connected after each convolutional layer to extract local photometric features, as follows: ; Among them, is the cross-image local brightness perception module under the i th illumination direction, θ is the parameter of the cross-image local brightness perception module, and the input is the tensor i under the Φ i th illumination, and the output local brightness change feature ψ i .
[0026] In an exemplary instance, as Figure 3As shown in the figure, a cross-image global brightness perception module is used to fuse the brightness change features under each lighting direction. This module is based on a three-dimensional convolutional structure, combined with the Leaky-ReLU activation function and the Dropout regularization mechanism to enhance the non-linear expression ability of features and effectively suppress the overfitting phenomenon. Specifically, all local brightness change features ψ i are stacked , as the input ψ all ;
[0027] Subsequently, three-dimensional convolutional kernels with sizes of 1×1×1, 3×1×1, and 1×1×1 are sequentially used to process the input features to realize the interaction and fusion of cross-image local brightness change information, and finally, a brightness change feature representation with global perception ability is extracted; ; Among them, is the cross-image global brightness perception module, θ are the parameters of the cross-image global brightness perception module, the input is ψ all , and the output global brightness change feature is P all .
[0028] Step 3: The shallow texture feature extractor first extracts the shallow image local texture information in the image, and then uses the max pooling operation to fuse the shallow local texture features, thereby generating the shallow global texture features; In an exemplary example, as Figure 4 shown, the shallow texture feature extractor realizes feature extraction by stacking three two-dimensional convolutional modules with Leaky-ReLU activation, divides the global photometric feature P all into . And concatenates it with the local brightness change feature ψ i to , as the input of the shallow texture feature extractor under each lighting condition; ; Among them, is the shallow texture feature extractor, θ are the parameters of the shallow texture feature extractor, the input is concat( ψ i , P i ), and the output shallow local texture feature is , and then fuse all the shallow local texture features through a max pooling operation to generate shallow global texture features .
[0029]
[0030] Step 4: The deep structure feature extractor extracts the deep local structure features in the shallow feature map, and then uses the max pooling operation to fuse the deep local structure features, thereby generating deep global structure features; In an exemplary instance, as Figure 4 shown, the deep structure feature extractor includes three modules, each module is sequentially composed of a depthwise separable convolution, layer normalization, a linear transformation, and a GELU activation function, and fuses the shallow global texture features with each shallow local texture feature for fusion , and uses the fused features as the input of the deep structure feature extractor to extract deep structure features.
[0031] ; Among them, is the deep structure feature extractor, θ is the parameter of the deep structure feature extractor, the input is , and the output deep local structure feature is , and then fuse all the deep local structure features through a max pooling operation to generate deep global structure features .
[0032]
[0033] Step 5: Input the extracted deep structure features into the normal regression network to obtain the surface normal estimation result.
[0034] In an exemplary instance, as Figure 2 shown, the normal regression network is used to decode the features output by the feature extractor into surface normals. The structure of the normal regression network includes 3 convolutional layers, 1 deconvolutional layer, and 1 L2 normalization layer. Input the deep structure feature map into the normal regression network to obtain the surface normal estimation result: ; Among them, R is the normal regression network, θ is the parameter of the normal regression network, the input is , and the predicted normal vector is .
[0035] In the present invention, by minimizing the error between the normal vector predicted by the normal regressor and the true normal vector, a loss function is constructed to drive the optimization of network parameters. During the training process, the AdamW optimizer is adopted and its default parameter settings are used. The initial learning rate is set to 0.001 and decays at a factor of 0.5 every 5 epochs. The batch size is set to 64 and a total of 35 epochs are trained. After the model training is completed, its weight parameters are saved. Finally, the trained model is used to test the DILIGENT real dataset to verify the normal vector estimation effect of the method of the present invention in real scenarios.
[0036] As Figure 5 shown, it is a comparison diagram of the normal vector reconstruction effect of this embodiment. Among them, the first row is the RGB images of 10 objects in the DILIGENT dataset; the second row is the corresponding true normal vector map; the third row is the result map of the object surface normal vector reconstructed by the method of the present invention; the fourth row is the error distribution map of the reconstructed normal vector and the true normal vector. From Figure 5 the results, it can be seen that the photometric stereo reconstruction method based on coaxial light guidance proposed by the present invention can obtain high normal vector estimation accuracy in different objects and scenarios with complex reflection characteristics, verifying the effectiveness and superiority of this method.
[0037] Although the present invention has been described in detail with general descriptions and specific embodiments above, based on the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection required by the present invention.
Claims
1. A photometric stereo reconstruction method based on coaxial light guidance, characterized in that: include: Step 1: Concatenate the illumination direction, observation image, and coaxial light image to construct a tensor matrix as the input of the network; Step 2: The cross-image local brightness perception module first extracts the local brightness change features of the input tensor under each illumination direction. Then, the cross-image global brightness perception module realizes the information interaction and fusion between local features under different illumination perspectives through three-dimensional convolution, thereby generating a global brightness change feature representation. Step 3: The shallow texture feature aggregator first extracts shallow local spatial information from the brightness change features in the image, and then aggregates the local features through the maximum pooling operation to generate a global shallow spatial feature representation; Step 4: The deep structure feature extractor further extracts deep local structure features from the shallow feature map, and realizes feature fusion through the maximum pooling operation to finally generate deep global structure features; Step 5: Input the extracted deep full-structure features into the normal regressor to obtain the surface normal estimation result.
2. The photometric stereo reconstruction method based on coaxial light guidance according to claim 1, characterized in that: In step 1, the coaxial light image is introduced, and the illumination direction, observation image and coaxial light image are stacked into a tensor as the network input ; Based on the principle of image imaging, the coaxial light guiding network is used to fit the retroreflection model, so as to separate the light and dark change characteristics that are not affected by the material. The expression is as follows: ; in, I is the normalized pixel value of the observed image, I c is the on-axis light image pixel intensity, n is the surface normal vector, l is the light source direction, v is the observation direction, l T v represents the inner product of the illumination direction and the observation direction, It's the change of light and dark. is the inverse reflection model of the observed image, is the retroreflection model of coaxial light, It is a composite retroreflective model.
3. The photometric stereo reconstruction method based on coaxial light guidance according to claim 1, characterized in that: In step 2, the cross-image local brightness perception module under each illumination direction adopts a weight sharing design, which consists of two serial 1×1 two-dimensional convolutional layers, and a Leaky-ReLU activation function is connected after each convolutional layer to extract local brightness change features, as shown in the following formula: ; in, It is i A cross-image local brightness perception module under different illumination directions, θ is the parameter of the local brightness perception module across images, and the input is i A tensor under illumination Φ i , the output is the local brightness change feature ψ i .
4. The photometric stereo reconstruction method based on coaxial light guidance according to claim 3, characterized in that: In step 2, a cross-image global brightness perception module is used to fuse the brightness change features under each illumination direction; this module is based on a three-dimensional convolution structure and combines all local brightness change features. ψ i Stacking , as input ψ all ; Then, three-dimensional convolution kernels with sizes of 1×1×1, 3×1×1 and 1×1×1 are used to process the input features in turn to achieve the interaction and fusion of local brightness change information across images, and finally extract the feature representation with global brightness change; ; in, is a global brightness perception module across images, θ are the parameters of the global brightness perception module across images, and the input is ψ all , the output global brightness change feature is expressed as P all .
5. The photometric stereo reconstruction method based on coaxial light guidance according to claim 3, characterized in that: In step 3, the shallow texture feature aggregator extracts features by stacking three 2D convolutional modules with Leaky-ReLU activations to represent the global brightness change feature. P all Divide into { P 1, P 2,…, P n }, and local brightness change characteristics ψ i Splicing, as the input of the shallow texture feature aggregator under each lighting condition, to obtain each shallow local texture feature ; ; in, is a shallow texture feature extractor, θ is the parameter of the shallow texture feature extractor, and the input is concat( ψ i , P i ), the output shallow local texture feature is , and then all shallow local texture features are pooled by the maximum pooling operation Aggregate to generate shallow global texture features .
6. The photometric stereo reconstruction method based on coaxial light guidance according to claim 1, characterized in that: In step 4, the deep structure feature extractor contains three modules, each of which is composed of depth-wise separable convolution, layer normalization, linear transformation and GELU activation function, which decomposes the standard convolution into channel-by-channel convolution and point-by-point convolution, and extracts the shallow global texture feature. With each shallow local texture feature Fusion, the fused features As the input of the deep structure feature extractor, it is used to extract deep local structure features; ; in, is a deep structure feature extractor, θ are the parameters of the deep structure feature extractor, and the input is , the output deep local structure feature is , and then all deep local structural features are pooled through the maximum pooling operation Fusion to generate deep global structural features .
7. The photometric stereo reconstruction method based on coaxial light guidance according to claim 1, characterized in that: In step 5, the normal regressor is used to decode the features output by the feature extractor into surface normals. The structure of the normal regressor converts the deep global structural features into Input to the normal regressor to obtain the surface normal estimation result: ; in, R is the normal regressor, θ are the parameters of the normal regressor, and the input is the deep global structural features , the predicted normal vector is .
Citation Information
Patent Citations
Multi-graph matching method based on low-rank tensor recovery
CN110443261A
Face image eye completion method based on self-attention mechanism model generative adversarial network
CN111738940A
Image rotation correction method and system, electronic device and storage medium
CN114723639A
Low-light image enhancement method based on parameter estimation
CN116681608A
Lightweight monocular image depth estimation method and device
CN119941817A