Three-dimensional depth map generation method based on L1 weight sparse and information dense complementation

Through the method of L1 weight sparse and multimodal information fusion, combined with the DASnet network optimization depth map generation, the problems of weight redundancy and depth information loss in the three-dimensional reconstruction network are solved, and high-resolution and stable depth map generation are achieved.

CN120339502APending Publication Date: 2025-07-18SHANXI DINGSHENG DIMENSION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510316604.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

There are problems in the existing three-dimensional reconstruction networks such as weight parameters, unstable gradients, low-quality depth map generation, lack of depth information and insufficient resolution, which affect the accuracy and stability of depth maps.

Method used

The method of L1 weight sparse and information intensive completion is adopted, through L1 regularization and multimodal information fusion, combined with DASnet network optimization depth map generation, including feature extraction, cascading cost volume estimation, depth completion and multimodal information assistance, NASNet network is built for depth map recovery.

Benefits of technology

It improves the accuracy and stability of depth map generation, enhances the generalization ability of the model, and can achieve high-resolution depth map reconstruction in different scenarios and lighting conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339502A_ABST
    Figure CN120339502A_ABST
Patent Text Reader

Abstract

According to the three-dimensional depth map generation method based on L1 weight sparse and information dense completion, the sum of absolute values of weights is introduced to serve as a regularization item in L1 regularization, a sparse weight matrix is achieved, irrelevant or redundant features are removed, and therefore the generalization ability and reconstruction precision of a model are improved; information from different data sources is integrated through multi-modal information fusion, so that more comprehensive and accurate data description is provided, and the robustness and stability of reconstruction are improved; the DASNet network can more accurately estimate the depth information of an object in a scene through fine network design and loss function optimization, has relatively strong generalization ability, and can realize stable depth estimation performance under different scenes and illumination conditions. L1 regularization and multi-modal information fusion are added in the three-dimensional reconstruction network, and a DASNet network is adopted to optimize the depth map, so that the accuracy and robustness of three-dimensional reconstruction can be remarkably improved, and better technical support is provided for application in related fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional depth map generation, and particularly to a method for generating a three-dimensional depth map based on L1 weight sparsity and information-intensive completion. Background Art

[0002] There are a large number of weight parameters in the network structure based on three-dimensional reconstruction. These parameters are prone to becoming redundant during the training process. The redundant weights not only increase the computational amount of the network but may also lead to overfitting and a decrease in the generalization ability on data.

[0003] Improper initialization may lead to the disappearance or explosion of gradients during the training process, thus affecting the generation quality of the depth map. Due to the complexity of the network structure, problems such as unstable gradient update and slow convergence speed may be faced during the weight optimization process, which will also affect the accuracy and stability of the depth map.

[0004] During the generation process of the depth map based on the three-dimensional reconstruction network, due to limitations of the input data (such as light changes, occlusion, noise, etc.), the depth information may be missing or noisy. These incomplete depth information is difficult to be effectively utilized during the reconstruction process, thus affecting the clarity and accuracy of the depth map. The resolution of the depth map directly affects its detail representation ability. Existing methods may be limited by computing resources and network structure and are difficult to generate high-resolution depth maps. During the depth map generation process, as the number of network layers increases, the spatial resolution of the feature map will gradually decrease, resulting in the loss of detail information.

[0005] In summary, the existing methods for generating depth maps based on three-dimensional reconstruction networks still have some defects in terms of excessive weight parameters and depth map clarity. In order to overcome these defects, researchers need to continuously explore new network structures, optimization algorithms, and data processing methods to improve the accuracy and clarity of depth map generation. Summary of the Invention

[0006] The present invention provides a method for generating a three-dimensional depth map based on L1 weight sparsity and information-intensive completion to solve the technical problems mentioned in the above background art.

[0007] The present invention provides a method for generating a three-dimensional depth map based on L1 weight sparsity and information-intensive completion, including the following steps:

[0008] S1. Obtain a model, perform multi-view calibration using a depth camera, and write the multi-view source images in sequence as I0, I1, I2, ..., I n ;

[0009] S2. Construct an information feature extraction network. The information feature extraction network includes 8 convolutional calculation layers, and L1 regularization is applied to the convolutional layer;

[0010] S3. Input the source image I i (i = 0, 1, 2,.., n), use a constructed 8 - layer convolutional network to extract corresponding features. Each layer of the convolutional network uses a 3×3 convolutional kernel, a batch normalization layer (i.e., BN layer) and a Relu activation function;

[0011] S4. Calculate the mean μ and variance δ of the features after convolutional processing 2 , and use the mean μ and variance δ 2 to perform normalization processing on each element in the input data to obtain the normalized data,

[0012]

[0013] where, x i represents the feature value at a certain position, m represents the total number of input parameters, μ represents the mean, and δ 2 represents the variance;

[0014] S5. Add L1 regularization to the batch normalization layer (i.e., BN layer), and limit the model complexity by adding a penalty term proportional to the absolute value of the model parameters to the normalized features;

[0015] S6. In the L1 regularization process, introduce a scaling factor γ for each feature channel, multiply it by the output of this channel, then jointly train the network weights and these scaling factors, and impose sparse regularization on the latter. Finally, prune these channels with the scaling factor and fine - tune the pruned network. The calculation process is as follows,

[0016]

[0017] where, γ is used for linear transformation of the normalized data, W represents the trainable weight, the offset β is used to adjust the data offset, μ represents the mean, and δ 2 represents the variance, B is the set of model parameters (x1,.., x m ), and Y represents the channel - level scaling factor;

[0018] S7. Construct a set of channel sparsification scaling factors {Y1, Y2,..., Y m}, set the percentile threshold to 70% to prune 70% of the channels with lower scaling factors, perform scaling factor sorting {0, 0, 0,..., Y n ,..., Y m}, where the proportion of the set with a scaling factor of 0 is 0.3, and the scaling factor is used as the sparsification coefficient to multiply the original input;

[0019] S8. Add a 1×1 convolution in the last feature extraction layer to restore the number of channels. For the source image with an input size of H×W×3, after passing through the feature extraction network, a feature map with a size of H / 4×W / 4×32 is output;

[0020] S9. Perform feature extraction on N input images to obtain n feature maps

[0021] S10. Use a cascaded cost volume to implement depth estimation. First, estimate a low-resolution depth map through a smaller cost volume. Then, based on the depth map output by the previous stage, reduce the depth hypothesis range of the current scale, and use a 3-level cost volume to implement depth map estimation, including two levels of intermediate results and a final depth output;

[0022] S11. First, upsample the n feature depth maps, then use the depth estimation range of the previous layer as a reference to determine the depth estimation range and depth estimation interval of the current layer, and finally output a higher-resolution depth map. A cascaded cost-volume formula is proposed to construct the cost volume in a coarse-to-fine manner; First, upsample the n feature depth maps, then use the depth estimation range of the previous layer as a reference to determine the depth estimation range and depth estimation interval of the current layer, and finally output a higher-resolution depth map. A cascaded cost-volume formula is proposed to construct the cost volume in a coarse-to-fine manner;

[0023] S12. Discretize the depth range of the entire scene into d1, d2, d3,..., d k planes. Taking the depth map estimated at the previous scale as the center, take a fixed depth range R k to determine d min and d max at each pixel position, and a cascaded cost-volume formula is proposed to construct the cost volume in a coarse-to-fine manner.

[0024]

[0025] d k = R k / I k

[0026] R k = d k * I k

[0027] where d k-1 is the depth map after upsampling in the previous stage, R k is the depth hypothesis range at the k-th stage, I k is the hypothesis plane interval at the k-th stage, and d k represents the depth hypothesis plane for k;

[0028] S13. Obtain k planes, each plane corresponding to a homography transformation matrix H. The homography transformation matrix H(d) corresponding to the features at the i-th perspective at depth value d is calculated as follows: i (d) is calculated as shown below:

[0029]

[0030] where: represents the predicted depth of the m-th pixel at the k-th level, is the residual depth of the m-th pixel to be learned at the k+1 stage, K i , R i , are the camera intrinsic matrix and rotation matrix of the i-th reference image respectively. K1 and R1 represent the camera intrinsic matrix and rotation matrix of the source image respectively. represents the normal vector of the source image plane, t i (i = 1, 2, 3.., i) represents the coordinates of the camera center i in the world coordinate system, and t1 - t i represents the translational replacement under different views;

[0031] S14. Use 3D UNet to regularize the rough cost volume C×D×H / 8×W / 8;

[0032] S15. Divide 3D UNet into 3D depthwise convolution and 3D pointwise convolution. These two convolutions are executed sequentially to form a complete convolution;

[0033] S16. Convpoint(Depthwis(V)) has one channel. One convolution kernel only convolves with one channel. 3D depth convolution is independently performed on the cost volume of each channel to obtain an intermediate feature map independent of the channel, as defined by the formula:

[0034]

[0035] where W1 represents the weights of the 3D depth convolution, V ∈ C×D×H×W represents the cost volume, and H, W, C, D represent the length, width, number of channels, and depth value of the cost volume respectively. i, j, u represent position indices, and K, L, M represent the kernel size of the convolution;

[0036] S17. The operation convolution kernel of Pointwise Convolution has a size of 1×1×N, where N is the number of channels of the previous layer. The feature layer from the previous step is weighted and combined in the depth direction to generate a new feature layer. 3D pointwise convolution acts on these channel-independent feature maps to aggregate channel-related information, as defined:

[0037]

[0038] Among them, W2 represents the weights of 3D pointwise convolution, V ∈ C×D×H×W represents the intermediate feature map, and H, W, C, D represent the length, width, number of channels, and depth value of the cost volume respectively. N represents the kernel size of the convolution;

[0039] S18. The 3D depthwise convolution and the 3D pointwise convolution are executed in sequence to form a complete convolution, and its mathematical expression is defined as:

[0040] Conv SepConv (V) = Convpoint(ConvDepth(V))

[0041] Among them, V represents the cost volume, and Conv SepConv (V) represents cost aggregation for the cost volume information;

[0042] S19. Use the softmax operation in the depth direction to regress all values between [0, 1] to form the probability volume P of depth estimation. Finally, multiply the different depth hypothesis plane values by the probability volume P to obtain the low-resolution LR depth map with a size of 1×H / 8×W / 8 Its mathematical expression is defined as:

[0043] P = softmax(V)

[0044]

[0045] Among them, V represents the cost volume, P represents the probability volume, represents the depth map, d min , d max represent the maximum and minimum feature depth values respectively;

[0046] S20. Construct a feature transmission module to infer the missing depth values from the sparse depth map, and integrate additional information into the NasNet structure search framework for depth completion to enhance the model's perception ability. This helps the model better understand the model structure, thereby improving the quality of depth completion;

[0047] S21. To unify the input scale, use the bicubic interpolation algorithm to upsample the low-resolution LR depth map to obtain a depth map with a larger scale

[0048] S22. Adopt the bicubic interpolation algorithm to find 16 neighboring pixel points corresponding to each pixel point in the target image in the original image. These 16 pixel points are respectively located around the target pixel point in the original image, and the weight of each pixel point is calculated according to its distance from the mapping point. The closer the distance, the greater the weight;

[0049] S23. Divide the 16 reference points in S22 into 4 groups in the horizontal direction, perform weighted average calculations separately, and perform another weighted average calculation on the 4 weighted average results obtained in the vertical direction to obtain the final pixel value at each position, that is, generate the depth map.

[0050] S24. Introduce multi-modal RGB image information to assist in depth completion, improve the accuracy and authenticity of the completion result, and use a low-resolution LR depth map with sizes of 3×H×W and 1×H×W as the input;

[0051] S25. NASNet proposes Normal cell and Reduction Cell, and the structures of each cell of the same type are the same and share weights;

[0052] S26. Take the depth map and the RGB image as feature maps respectively and input them into the NASNet Enconder for splicing and fusion at the channel level, followed by a Coustom Enconder layer, and use a convolution with a size of 3×3 and a stride of 2 for convolution calculation. Through multiple operations, gradually restore the densely completed depth map to its original size.

[0053]

[0054] In the formula, is the high-resolution (HD) depth map, is the improved depth map, is the real depth map, Pvalid is the valid point set of the real depth map, and λ is used to balance loss1(p) and loss2(p), which is set to 1.0;

[0055] S27. Set the parameter optimizer to train the 3D reconstruction model;

[0056] Furthermore, in S12, d1 = 48;

[0057] Furthermore, in S27, the initial learning rate is set to 0.0008, the decay weight value for each epoch is set to 0.002. The batch size is set to 16.

[0058] The beneficial effects of the present invention are as follows:

[0059] 1. Adding L1 regularization to the BN layer of the feature information extraction network tends to produce a sparse weight matrix, enabling feature selection, removing unimportant features, and making the model easier to understand and interpret. Setting a pruning rate of 70%, the pruned model has fewer parameters, is more efficient in storage and computation, and maintains the original performance as much as possible.

[0060] 2. RGB images provide rich color, texture, and shape information, while depth maps provide the spatial position and geometric shape information of objects. Using a deep learning model, infer the missing depth values from sparse depth maps. Introduce multi-modal information RGB images to assist in depth completion and improve the accuracy and authenticity of the completion results. The combination of the two can form a more complete and accurate scene description.

[0061] 3. Construct the NASNet network for the optimization of depth maps, leveraging its powerful feature extraction ability and automated network structure design. This network can automatically search for convolutional kernels and network structures suitable for depth map features, thereby improving the accuracy and efficiency of depth map processing. In addition, the NASNet network also has good generalization ability and can achieve stable performance in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0063] Figure 1 is the dense depth map generation process based on 3D reconstruction of the present invention;

[0064] Figure 2 is the schematic diagram of weight pruning with L1 sparse regularization of the present invention;

[0065] Figure 3 is the schematic diagram of sparse-to-dense depth map completion based on DSAnet of the present invention;

[0066] Figure 4 is the schematic diagram of the Normal cell structure in DASnet of the present invention;

[0067] Figure 5 is the schematic diagram of the Reduction Cell structure in DSAnet of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0068] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that for the sake of description, only parts related to the present invention rather than all structures are shown in the accompanying drawings. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0069] Reference herein to "embodiments" means that a particular feature, structure, or characteristic described in connection with the embodiments can be included in at least one embodiment of the invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0070] Please refer to Figures 1 to 5 , a three-dimensional depth map generation method based on L1 weight sparsity and information-intensive completion of the present invention, includes the following steps:

[0071] S1. Obtain a model, perform multi-view calibration using a depth camera, and write the multi-view source images I0, I1, I2,..., I n ;

[0072] S2. Construct an information feature extraction network. The information feature extraction network includes 8 convolutional calculation layers, and L1 regularization is applied in the convolutional layer, which can make the learning model become sparse. The reduction and shrinkage of the effective parameters and scale of the model reduce the generalization error of the learning model and enhance the generalization ability;

[0073] S3. Input the source image I i (i = 0, 1, 2,.., n), and use the constructed 8-layer convolutional network to extract corresponding features. Each layer of the convolutional network uses a 3×3 convolutional kernel, a batch normalization layer (i.e., BN layer) and a Relu activation function;

[0074] S4. Calculate the mean μ and variance δ of the features after convolutional processing 2 , and use the mean μ and variance δ 2 to perform normalization processing on each element in the input data to obtain the normalized data,

[0075]

[0076] where, x i represents the feature value at a certain position, m represents the total number of input parameters, μ represents the mean, and δ 2 represents the variance;

[0077] S5. Add L1 regularization to the batch normalization layer, i.e., the BN layer, to limit the model complexity by adding a penalty term proportional to the absolute value of the model parameters to the normalized features;

[0078] S6. In the process of L1 regularization, introduce a scaling factor γ for each feature channel, multiply it by the output of the channel, then jointly train the network weights and these scaling factors, and impose sparse regularization on the latter. Finally, prune these channels with the scaling factors and fine-tune the pruned network. The calculation process is as follows.

[0079]

[0080] Among them, γ is used for linear transformation of the normalized data, W represents the trainable weights, the offset β is used to adjust the data offset, μ represents the mean, and δ 2 represents the variance, B is the set of model parameters (x1,..,x m ), and Y represents the channel-level scaling factor;

[0081] S7. Construct a collection of channel sparsification scaling factors {Y1, Y2,..., Y m}, set the percentile threshold to 70% to prune 70% of the channels with lower scaling factors, and perform scaling factor sorting {0, 0, 0,..., Y n ,..., Y m}, where the proportion of the collection with a scaling factor of 0 is 0.3, and the scaling factor is used as the sparsification coefficient to multiply the original input;

[0082] S8. Add a 1×1 convolution in the last feature extraction layer to restore the number of channels. For the source image with an input size of H×W×3, after passing through the feature extraction network, a feature map with a size of H / 4×W / 4×32 is output;

[0083] S9. Perform feature extraction on N input images to obtain n feature maps

[0084] S10. Use a cascaded cost volume to implement depth estimation. First, estimate a low-resolution depth map through a smaller cost volume, then reduce the depth hypothesis range of the current scale according to the depth map output by the previous level, and use a 3-level cost volume to implement depth map estimation, including two levels of intermediate results and a final depth output;

[0085] S11. n feature depth maps First, perform upsampling. Then, using the depth estimation range of the previous layer as a reference, determine the modified depth estimation range and depth estimation interval. Finally, output a depth map with a higher resolution. A cascaded cost - volume formula is proposed to construct the cost volume in a coarse - to - fine manner;

[0086] S12. Discretize the depth range of the entire scene into d1, d2, d3,..., d k planes. Taking the depth map estimated at the previous scale as the center, take a fixed depth range R k , and determine d min and d max at each pixel position. A cascaded cost - volume formula is proposed to construct the cost volume in a coarse - to - fine manner.

[0087]

[0088]

[0089] where d k-1 is the depth map after upsampling in the previous stage, R k is the depth hypothesis range at the k - th stage, I k is the hypothesis plane interval at the k - th stage, and d k represents k depth hypothesis planes;

[0090] S13. Obtain k planes, each of which corresponds to a homography transformation matrix H. The homography transformation matrix H i (d) corresponding to the feature of the i - th view at the depth value d is calculated as follows:

[0091]

[0092] where, represents the predicted depth of the m - th pixel at the k - th level, is the residual depth of the m - th pixel to be learned at the k + 1 stage, K i , R i are the camera intrinsic matrix and rotation matrix of the i - th reference image respectively. K1 and R1 represent the camera intrinsic matrix and rotation matrix of the source image respectively. represents the normal vector of the source image plane, and t i (i = 1, 2, 3.., i) represents the coordinates of the camera center i in the world - coordinate system, and t1 - t i represents the translational displacement under different views;

[0093] S14. Use 3D UNet to regularize the rough cost volume C×D×H / 8×W / 8;

[0094] S15. Divide the 3D UNet into 3D depthwise convolution and 3D pointwise convolution. The 3D depthwise convolution performs cost aggregation on the cost volume information in the depth dimension, and the 3D pointwise convolution performs cost aggregation on the cost volume information in the spatial dimension. These two convolutions are executed sequentially to form a complete convolution;

[0095] S16. Convpoint(Depthwis(V)) has one channel. One convolution kernel only convolves with one channel, and 3D depth convolution is independently performed on the cost volume of each channel to obtain an intermediate feature map independent of the channel, as defined by the formula:

[0096]

[0097] where W1 represents the weights of the 3D depth convolution, V ∈ C×D×H×W represents the cost volume, and H, W, C, D represent the length, width, number of channels, and depth value of the cost volume respectively. i, j, u represent position indices, and K, L, M represent the kernel size of the convolution;

[0098] S17. The size of the convolution kernel of the Pointwise Convolution is 1×1×N, where N is the number of channels in the previous layer. The feature layer from the previous step is weighted and combined in the depth direction to generate a new feature layer. The 3D pointwise convolution acts on these channel-independent feature maps to aggregate channel-related information, as defined:

[0099]

[0100] where W2 represents the weights of the 3D pointwise convolution, V ∈ C×D×H×W represents the intermediate feature map, and H, W, C, D represent the length, width, number of channels, and depth value of the cost volume respectively. N represents the kernel size of the convolution;

[0101] S18. The 3D depthwise convolution and the 3D pointwise convolution are executed sequentially to form a complete convolution, and its mathematical expression is defined as:

[0102] Conv SepConv (V) = Convpoint(ConvDepth(V))

[0103] where V represents the cost volume, and Conv SepConv (V) represents cost aggregation of the cost volume information;

[0104] S19. Perform regression on all values between [0, 1] using the softmax operation in the depth direction to form the probability volume P for depth estimation. Finally, multiply the values of different depth hypothesis planes by the probability volume P to obtain a low-resolution LR depth map of size 1×H / 8×W / 8 , and its mathematical expression is defined as:

[0105] P = softmax(V)

[0106]

[0107] where V represents the cost volume and P represents the probability volume, represents the depth map, d min , d max respectively represent the maximum and minimum characteristic depth values;

[0108] S20. Construct a feature transfer module to infer the missing depth values from the sparse depth map and integrate additional information into the depth completion NasNet structure search framework to enhance the model's perception ability. This helps the model better understand the model structure, thereby improving the quality of depth completion;

[0109] S21. Use the bicubic interpolation algorithm to upsample the low-resolution LR depth map to obtain a depth map of a larger scale

[0110] S22. Adopt the bicubic interpolation algorithm to find 16 neighboring pixels corresponding to each pixel point in the target image in the original image. These 16 pixel points are respectively located around the target pixel point in the original image, and the weight of each pixel point is calculated according to its distance from the mapping point. The closer the distance, the greater the weight;

[0111] S23. Divide the 16 reference points in S22 into 4 groups in the horizontal direction and perform weighted average calculations respectively. Then perform another weighted average calculation on the 4 weighted average results obtained in the vertical direction to obtain the final pixel value at each position, that is, generate the depth map

[0112] S24. Introduce multi-modal RGB image information to assist depth completion, improve the accuracy and authenticity of the completion result, and use a low-resolution LR depth map of size 3×H×W and size 1×H×W as the input;

[0113] S25. NASNet proposes Normal cell and Reduction Cell, and the structures of each type of Cell are the same and share weights;

[0114] S26. Input the depth map and the RGB image as feature maps into the NASNet Enconder respectively for splicing and fusion at the channel level, followed by a Coustom Enconder layer, and perform convolution calculations with a convolution size of 3×3 and stride = 2. Through multiple operations, gradually restore the densely completed depth map to its original size.

[0115]

[0116] In the formula, is the high-resolution (HD) depth map, is the improved depth map, is the real depth map, Pvalid is the valid point set of the real depth map, and λ is used to balance loss1(p) and loss2(p), set to 1.0;

[0117] S27. Set the parameter optimizer to train the 3D reconstruction model;

[0118] Specifically, in S12, d k = 48;

[0119] Specifically, in S27, the initial learning rate is set to 0.0008, the decay weight for each epoch is set to 0.002. The batch size is set to 16.

[0120] L1 regularization helps to achieve a sparse weight matrix by introducing the sum of the absolute values of the weights as a regularization term, that is, some weights will tend to 0. L1 regularization can automatically select the features important for the reconstruction task, remove irrelevant or redundant features, thereby improving the generalization ability and reconstruction accuracy of the model. By restricting the size of the weights, L1 regularization helps to prevent the model from overfitting on the training data and improve the performance of the model on the test data.

[0121] Utilize multi-modal information fusion to integrate information from different data sources (such as RGB images, depth maps) to provide a more comprehensive and accurate data description. Data of different modalities contain different information, and after fusion, it can provide a more complete and accurate 3D scene description, improving the robustness and stability of the reconstruction.

[0122] The DASnet network can more accurately estimate the depth information of objects in the scene and has strong generalization ability through fine network design and loss function optimization, and can achieve stable depth estimation performance under different scenes and lighting conditions.

[0123] In summary, adding L1 regularization, multi-modal information fusion to the 3D reconstruction network and optimizing the depth map using the DASnet network can significantly improve the accuracy and robustness of 3D reconstruction, providing better technical support for applications in related fields.

[0124] The above are only embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be included in the patent protection scope of the present invention by the same token.

Claims

1. A method for generating a three-dimensional depth map based on L1 weight sparsity and information-intensive completion, characterized in that It includes the following steps: S1. Obtain a model, perform multi-view calibration using a depth camera, and sequentially write multi-view source images as I0, I1, I2, …, I n ; S2. Construct an information feature extraction network, where the information feature extraction network includes 8 convolutional calculation layers, and L1 regularization is applied in the convolutional layer; S3. Input the source image I i (i = 0, 1, 2,.., n), use a constructed 8-layer convolutional network to extract corresponding features. Each layer of the convolutional network uses a 3×3 convolutional kernel, a batch normalization layer (i.e., BN layer) and a Relu activation function; S4. Calculate the mean μ and variance δ of the features after convolution processing 2 , and use the mean μ and variance δ 2 to perform normalization processing on each element in the input data to obtain the normalized data where x i represents the eigenvalue of a certain position, m represents the total number of input parameters, μ represents the mean value, and δ 2 represents the variance; S5. Add L1 regularization to the batch normalization layer, i.e., the BN layer, and limit the model complexity by adding a penalty term proportional to the absolute value of the model parameters to the normalized features; S6. In the processing of L1 regularization, a scaling factor γ is introduced for each feature channel, multiplied by the output of this channel, then the network weights and these scaling factors are jointly trained, and sparse regularization is imposed on the latter. Finally, these channels are pruned with the scaling factor, and the pruned network is fine-tuned. The calculation process is as follows. Among them, γ is used for linear transformation of the normalized data, W represents the trainable weights, the offset β is used to adjust the data offset, μ represents the mean, and δ 2 represents the variance, B is the set of model parameters (x1,..,x m ), and Y represents the channel-level scale factor; S7. Construct a set of channel sparsification ratio factors {Y1, Y2,..., Y m}, set the percentile threshold to 70% to prune 70% of the channels with lower ratio factors, and perform ratio factor sorting {0, 0, 0,..., Y n ,..., Y m}, where the proportion of the set with a ratio factor of 0 is 0.3, and the ratio factor is used as the sparsification coefficient to multiply the original input; S8. Add a 1×1 convolution in the last feature extraction layer to restore the number of channels. For the source image with an input size of H×W×3, after passing through the feature extraction network, a feature map with an output size of H / 4×W / 4×32 is obtained; S9. On N input images Feature extraction is performed to obtain n feature maps S10. Use a cascaded cost volume to implement depth estimation. First, estimate a low-resolution depth map through a smaller cost volume, then reduce the depth hypothesis range of the current scale according to the depth map output by the previous level, and use a 3-level cost volume to implement depth map estimation, including two levels of intermediate results and a final depth output; S11, n feature depth maps First, perform upsampling. Then, using the depth estimation range of the previous layer as a reference, determine the modified depth estimation range and depth estimation interval. Finally, output a depth map with a higher resolution, and propose a cascaded cost - volume formula to construct the cost volume in a coarse - to - fine manner; S12. Discretize the depth range of the entire scene into d1, d2, d3,..., d k planes. Taking the depth map estimated at the previous scale as the center, take a fixed depth range R k , and determine d min and d max at each pixel position, and propose a cascaded cost - volume formula to construct the cost volume in a coarse - to - fine manner. d k = R k / I k R k = d k * I k where d k-1 is the depth map after sample loading in the previous stage, R k is the depth hypothesis range in the k-th stage, I k is the assumed plane interval in the k-th stage, d k represents k depth hypothesis planes; S13. Obtain k planes, each of which corresponds to a homography transformation matrix H. The homography transformation matrix H(d) corresponding to the feature of the i-th perspective at the depth value d is calculated as follows: i (d) is calculated as shown below, Among them, represents the predicted depth of the m-th pixel at the k-th level, is the residual depth of the m-th pixel to be learned at the k+1 stage, K i , R i , are the camera intrinsic matrix and rotation matrix of the i-th reference image respectively. K1 and R1 represent the camera intrinsic matrix and rotation matrix of the source image respectively. represents the normal vector of the source image plane, t i (i = 1, 2, 3.., i) represents the coordinates of the camera optical center i in the world coordinate system, t1 - t i represents the translational substitution under different views; S14. Use 3D UNet to regularize the rough cost volume C×D×H / 8×W / 8; S15. Divide 3D UNet into 3D depthwise convolution and 3D pointwise convolution, and these two convolutions are executed sequentially to form a complete convolution; S16. Convpoint(Depthwis(V)) for one channel, where one convolution kernel only convolves with one channel, and 3D depth convolution is independently performed on the cost volume of each channel to obtain an intermediate feature map independent of the channel, as defined by the formula where W1 represents the weight of the three-dimensional depth convolution, V∈C×D×H×W represents the cost volume, and H, W, C, D represent the length, width, number of channels, and depth value of the cost volume respectively. i, j, u represent position indices, and K, L, M represent the kernel size of the convolution; S17. The operation convolution kernel size of Pointwise Convolution is 1×1×N, where N is the number of channels in the previous layer. The feature layer of the previous step is weighted and combined in the depth direction to generate a new feature layer. 3D pointwise convolution acts on these channel-independent feature maps to aggregate channel-related information, as defined: where W2 represents the weight of the three-dimensional pointwise convolution, V∈C×D×H×W represents the intermediate feature map, and H, W, C, D represent the length, width, number of channels, and depth value of the cost volume respectively. N represents the kernel size of the convolution; S18. The 3D depthwise convolution and the 3D pointwise convolution are executed sequentially to form a complete convolution, and its mathematical expression is defined as: Conv SepConv (V) = Convpoint(ConvDepth(V)) Among them, V represents the cost volume, and Conv SepConv (V) represents cost aggregation for the cost volume information; S19. Use the softmax operation in the depth direction to regress all values between [0, 1] to form the probability volume P for depth estimation. Finally, multiply the values of different depth hypothesis planes by the probability volume P to obtain a low-resolution LR depth map of size 1×H / 8×W / 8. Its mathematical expression is defined as: P = softmax(V) Among them, V represents the cost volume, and P represents the probability volume. represents the depth map, and d min , d max respectively represent the maximum and minimum characteristic depth values. S20. Construct a feature transmission module to infer the missing depth values from the sparse depth map and integrate additional information into the NasNet structure search framework for depth completion; S21. Use the bicubic interpolation algorithm to upsample the low-resolution LR depth map to obtain a depth map at a larger scale S22. Use the bicubic interpolation algorithm to find 16 neighboring pixels corresponding to each pixel in the target image in the original image. These 16 pixels are located around the target pixel in the original image, and the weight of each pixel is calculated according to its distance from the mapping point. The closer the distance, the greater the weight. S23. Divide the 16 reference points in S22 into 4 groups in the horizontal direction, perform weighted average calculations respectively, and then perform a weighted average calculation on the 4 weighted average results obtained in the vertical direction to obtain the final pixel value at each position, that is, generate a depth map. S24. Introduce multi-modal RGB image information to assist depth completion, improve the accuracy and authenticity of the completion result, and use a low-resolution (LR) depth map with a size of 3×H×W and a size of 1×H×W as input; S25. NASNet proposed Normal cell and Reduction Cell, and the structures of each cell of the same type are the same and share weights. S26. Input the depth map and the RGB image into the NASNet Enconder respectively as feature maps, perform splicing and fusion at the channel level, followed by a Coustom Enconder layer. Use convolution calculation with a convolution size of 3×3 and stride = 2. Through multiple operations, gradually restore the depth map after dense completion to its original size. where is the high-resolution (HD) depth map, is the improved depth map, is the true depth map, Pvalid is the valid point set of the true depth map, and λ is used to balance loss1(p) and loss2(p), which is set to 1.0; S27. Set the parameter optimizer to train the three-dimensional reconstruction model.

2. The three-dimensional depth map generation method based on L1 weight sparsity and information-intensive completion according to claim 1, characterized in that, In S12, d1 = 48.

3. The three-dimensional depth map generation method based on L1 weight sparsity and information-intensive completion according to claim 1, wherein In S27, the initial learning rate is set to 0.0008, and the decay weight for each epoch is set to 0.

002. The batch size is set to 16.

Citation Information

Cited By

  • Rapid damage-free digital reconstruction method and device for cultural relics and ancient buildings

    CN122289571A