Inspection scene multi-view three-dimensional reconstruction method based on adaptive feature enhancement
By using adaptive feature enhancement network, multi-scale depth estimation and depth map fusion methods in the inspection scenario, the problems of feature extraction difficulties and insufficient network generalization in the inspection scenario are solved, and high-quality three-dimensional reconstruction effect is achieved.
Patent Information
- Application Number
- CN202510130271.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to effectively extract the features of weak texture areas and non-Luber body surfaces in inspection scenarios, resulting in problems such as large voids and missing edge areas in the three-dimensional reconstruction results, and the network generalization is insufficient, so the effect is not ideal when used directly in three-dimensional reconstruction of inspection scenarios.
A multi-view three-dimensional reconstruction method for patrol scenes based on adaptive feature enhancement is adopted, including adaptive feature enhancement network, multi-scale depth estimation and depth map fusion. Through adaptive feature enhancement network extraction and enhance feature information, multi-scale depth estimation performs depth estimation from coarse to fine, depth map fusion generates dense three-dimensional color point clouds, and improves the generalization ability of the network through transfer learning fine-tuning.
It improves the network's feature perception ability of the inspection scenario, reduces memory consumption, improves the integrity and detail accuracy of three-dimensional reconstruction, improves the network's generalization ability in inspection scenarios, and achieves ideal reconstruction results.
Smart Images

Figure CN119991963A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multi-view inspection scene based on adaptive feature enhancement Figure 3 The invention discloses a three-dimensional reconstruction method, and belongs to the technical field of three-dimensional reconstruction. Background Art
[0002] Intelligent robots can perform patrol, monitoring, early warning and other security tasks in different scenarios by integrating multiple advanced technologies such as navigation and positioning, artificial intelligence and wireless communication. As the core task of digital security construction, autonomous inspection by intelligent robots relies on accurate perception and understanding of the surrounding environment of the inspection scene. By reconstructing a detailed three-dimensional model of the surrounding environment, the robot can perceive the three-dimensional information of obstacles, plan obstacle avoidance paths in advance, and improve the safety and efficiency of autonomous navigation. The robot can also accumulate and learn the environmental information and inspection task execution experience obtained from each three-dimensional reconstruction, continuously optimize behavior patterns and decision-making strategies, and improve adaptability and perception and understanding capabilities in different inspection environments.
[0003] From the perspective of sensors, 3D reconstruction is divided into monocular camera multi-view reconstruction and LiDAR point cloud reconstruction. Considering the high cost of LiDAR equipment, the large amount of collected point cloud data makes storage and transmission requirements high. The robot's inspection endurance will be shortened during use, and the generated model lacks color and texture information. Therefore, a monocular camera carried by an intelligent robot is used to perform multi-view reconstruction of the inspection scene. Figure 3 3D color point cloud reconstruction. Due to the presence of weeds, bushes, trees and other influencing factors in the inspection scene, it is difficult to extract the features of weak texture areas and non-Lambertian surfaces in the scene image. The characteristics of the scene are not felt enough, resulting in prominent problems such as large holes and missing edge areas in the point cloud of complex inspection scenes (such as electrical equipment, garbage collection boxes, etc.). In addition, when processing large-scale scenes or high-resolution images, memory consumption increases significantly. Existing multi-view Figure 3 Most 3D reconstruction methods are trained on public datasets of small indoor scenes (such as DTU datasets). The network generalization is insufficient, and the reconstruction effect is not ideal when directly used for 3D reconstruction of inspection scenes. Summary of the invention
[0004] In order to overcome the problems existing in the prior art, the present invention aims to provide a multi-view inspection scene based on adaptive feature enhancement. Figure 3The network includes an adaptive feature enhancement network, multi-scale depth estimation and depth map fusion. Aiming at the problem of difficulty in extracting features of weak texture areas and non-Lambertian surface in inspection scenes and the problem of low perception, an adaptive feature enhancement network for inspection scenes is proposed to improve the network learning ability and the perception of detail features such as boundary information and texture information. Aiming at the problem of excessive video memory consumption, a three-level cascade architecture is established to perform multi-scale depth estimation of inspection scenes from coarse to fine. Depth map fusion is used to filter and denoise the depth map, and the filtered depth map is fused with the inspection scene image to generate a dense three-dimensional color point cloud of the inspection scene. Aiming at the problem of insufficient generalization of the network, the network is first pre-trained on a public data set, and then transfer learning and fine-tuning are performed through the inspection scene data set to improve the reconstruction effect of the inspection scene.
[0005] In order to achieve the above-mentioned invention object, the technical solution adopted by the present invention is: a multi-view inspection scene based on adaptive feature enhancement Figure 3 The reconstruction method comprises the following steps:
[0006] Step 1: Build a depth map fusion framework based on LiDAR and monocular camera self-calibration to obtain color images and depth map data of the inspection scene, which includes the following sub-steps:
[0007] (a) Self-calibration of LiDAR and monocular camera. The color image and 3D point cloud data of the inspection scene are obtained through the monocular camera of the intelligent robot and the LiDAR installed above the monocular camera. The black and white checkerboard calibration method and a self-calibration method of LiDAR and camera are introduced to obtain the camera intrinsic parameters and the conversion matrix from the LiDAR coordinate system to the camera coordinate system. ;
[0008] (b) Inspection scene depth map completion, which consists of two parts: inspection scene sparse depth map generation and inspection scene sparse depth map completion:
[0009] 1) Generation of sparse depth map of inspection scene. Suppose the point cloud of the inspection scene is ,in It is the inspection scene point cloud data. Indicates Points, is the number of point clouds, and the points pass Convert to the camera coordinate system, then project it onto the color image of the inspection scene through the camera internal parameters, remove invalid points outside the image range, and obtain the point cloud Sparse depth map projected on the image ,in Yes By transforming the matrix and the camera intrinsic parameters projected to the pixel coordinates of the image, for Corresponding depth value, is the number of valid point clouds projected into the image range;
[0010] 2) Completion of sparse depth maps of inspection scenes. Since sparse depth maps are obtained Some pixels in the image have no depth value, resulting in holes of varying sizes. A depth map completion method is introduced to complete the holes, using a greedy strategy: effective depth pixels Surrounding empty values may have similar values, so fill the original sparse depth map first Medium effective depth pixels The surrounding small holes are filled in the order of large holes, and finally a 5×5 median blur is used to remove noise, remove outliers on the local edges, and then a 5×5 Gaussian blur is used to smooth the local plane to obtain the final inspection scene depth map. ,in , is the number of valid pixels within the completed image range;
[0011] Step 2: Build a multi-view inspection scene based on adaptive feature enhancement Figure 3 The 3D reconstruction network includes an adaptive feature enhancement network, multi-scale depth estimation, and depth map fusion. The adaptive feature enhancement network is used to extract and enhance the feature information of the input inspection scene image. The multi-scale depth estimation combines the feature map extracted by the adaptive feature enhancement network and the camera pose information obtained by the incremental SFM technology. Through a three-level cascade structure, the scene multi-view depth estimation and depth map generation are performed from coarse to fine. The depth map fusion is to fuse all color images and corresponding depth maps into a 3D color point cloud of the inspection scene, which specifically includes the following sub-steps:
[0012] (a) Construct an adaptive feature enhancement network based on the feature pyramid network, with an input size of Multi-view inspection scene ,in, For reference view, is the source view, is the number of multiple views, extract multiple views The multi-scale detail features such as scene boundary information and texture information are output as three-scale feature maps with global and detail semantic information. ,in, is the reference view feature map, is the source view feature map, The size and number of channels are , , , the adaptive feature enhancement network consists of three parts:
[0013] 1) Bottom-up multi-scale feature extraction. First, use a 3×3 convolution to extract the size of Multi-view inspection scene The size and number of channels are extracted. Then, a 3×3 convolution and an attention module SimAM are used to complete the initial feature extraction. At this time, the feature map is , Secondly, use a 5×5 convolution with a step size of 2 to downsample the feature map, and then use two 3×3 convolutions and an attention module SimAM to complete the secondary feature extraction. At this time, the feature map size and the number of channels are The third feature extraction is the same as the second feature extraction operation. The feature map size and number of channels are ;
[0014] 2) Top-down multi-scale feature aggregation,First, the third extraction feature map of the bottom-up multi-scale feature extraction part remains unchanged, and the size and number of channels are , the first and second output feature maps are converted to 32 channels through a 1×1 convolution. Secondly, the feature maps after the third and second 1×1 convolutions are upsampled by 2 times using the nearest neighbor interpolation method, and are added to the feature maps after the second and first 1×1 convolutions to complete the feature aggregation operation. At this time, the size and number of channels of the three scale feature maps are , , ;
[0015] 3) Construct an adaptive perception module for feature (APMF) consisting of deformable convolution, Bn layer, and Relu layer. The three scale feature maps output by the top-down multi-scale feature aggregation part are output through APMF to finally extract the feature map. The size and number of channels are , , , corresponding to the input feature maps of stage 1, stage 2, and stage 3 in multi-scale depth estimation;
[0016] (b) Multi-scale depth estimation. A three-level cascade structure is constructed. Each level represents a stage. Depth estimation is performed in three stages. Each stage of the cascade structure includes two parts: cost volume construction and depth prediction. The input of each stage is the camera pose information and the feature map extracted by the adaptive feature enhancement network. The camera pose information is obtained by the incremental Sfm technology. The feature map is the three-scale feature map output by the adaptive feature enhancement network. The size and number of channels are respectively , , , respectively as the input of stage 1, stage 2, and stage 3 in multi-scale depth estimation. All three stages perform cost volume construction and depth prediction to output depth maps. Stage 1 outputs a depth map that roughly estimates the depth range. , the size is , stage 2 and stage 3 output depth map And the final predicted depth map , the sizes are , ,The difference between stage 2 and stage 3 and stage 1 is that stage 2 and stage 3 respectively use the estimated depth value of the previous stage as the median of the depth range, and achieve the purpose of cascade depth map refinement by reducing the assumed depth range, depth interval and the number of assumed planes;
[0017] Cost volume construction is to measure the feature similarity between multiple views of the inspection scene. It requires three steps. First, multiple discrete depth hypothesis planes are determined. Second, the source view extracted by the adaptive feature enhancement network is Corresponding Warp to the reference view via homography The feature volume is constructed on multiple hypothetical planes. Finally, the feature volumes are fused together to construct a three-dimensional cost volume. Depth prediction is to regularize the cost volume into a probability volume, and then find the expectation of the probability volume along the depth direction to obtain the estimated depth value. The specific process of cost volume construction and depth prediction is as follows:
[0018] 1) Determine the depth hypothesis plane. First, under the camera frustum corresponding to the reference view, establish a plane that is consistent with the principal optical axis of the reference camera. The depth hypothetical plane is perpendicular to each other, and the depth range of stage 1 is set to ,from arrive , evenly selected within this range depth hypothesis plane, then The sampling depth value corresponding to the plane , described by formula (1),
[0019]
[0020] in, , Indicates depth range The minimum and maximum values of is the depth interval size of stage 1, The value range is 1 to ,In the cascade structure, stage 2 and stage 3 use the estimated depth value of the previous stage as the median of the depth range, and refine the cascade depth map by reducing the assumed depth range, depth interval and the number of assumed planes. Stage 2 and stage 3 reduce the assumed depth range, which is described by formula (2).
[0021]
[0022] in, Indicates the stage, with values of 1 and 2. , It is a stage and the depth hypothesis range of the previous stage, It is a stage Assume the reduction factor of the range, Representation stage Depth range The minimum value of Representation stage Depth range The maximum value of For stage Reference view pixels The predicted depth of the image is reduced, and the depth interval is described by formula (3).
[0023]
[0024] in, , The stages and the depth interval size of the previous stage, It is a stage The reduction factor of the depth interval, given the stage Assumption Range and depth interval size , then the corresponding number of assumed planes is described by formula (4):
[0025]
[0026] At this point, stage In the The sampling depth value corresponding to the plane , described by formula (5),
[0027]
[0028] in, Indicates the stage, with values of 1 and 2. Representation stage Depth range The minimum value of The value range is 1 to ,Cascaded depth map refinement uses adaptive depth sampling of the above formula to use computing and memory resources in a more appropriate estimation range, which can significantly reduce computing time and video memory consumption;
[0029] 2) Construct a feature body and enhance the feature map extracted by the adaptive feature enhancement network Warp to the reference view by the homography transformation matrix The feature volume is constructed on the hypothetical plane of the camera cone, and the homography transformation is described by formula (6):
[0030]
[0031] in, is the unit vector referring to the main axis of the camera, Represents the posture information, for The corresponding camera intrinsic parameters, is the rotation matrix and Translation matrix, , , For reference view Corresponding camera pose information, , , For the Zhang Yuan view corresponds to the camera pose information, is the identity matrix, represents the projection equation, for The pixels on express Project to Pixels, Indicates the sampling depth value At correspond The eigenvalues are projected onto The homography transformation matrix of is , because the three-stage cascade is used for coarse-to-fine depth estimation, stage 1 is to roughly estimate the depth range, and stage 2 and stage 3 are respectively combined with the depth range of the previous stage. By reducing the assumed depth range, depth interval and the number of assumed planes, the accuracy of the feature map volume is improved. According to the cascade depth map refinement method, formula (6) is modified to obtain stage The homography matrix of is described by formula (7):
[0032]
[0033] in, Indicates the stage, with values of 1 and 2. For stage Reference view pixels The predicted depth, For the stage To learn The residual depth of Construct more and more detailed depth planes and transform them through formula (7) to obtain the feature volume: , with a resolution of ,in, , is the feature map size, is the number of assumed planes, is the number of channels of the feature body, and the resolution of the feature body in stage 1 is , stage 2 is , stage 3 is ;
[0034] 3) Feature fusion, in order to measure The feature differences between views are calculated by using the variance-based cost metric. Aggregated into a unified matching cost body, described by formula (8):
[0035]
[0036] in, For the price body, is a variance-based cost metric, is the mean of all features, is the number of views, the cost volume Resolution is ,in, , is the feature map size, is the number of assumed planes, is the number of channels of the feature body, and the cost body resolution after feature body fusion stage 1 is , stage 2 is , stage 3 is ;
[0037] 4) Depth prediction: In order to filter noise and prevent overfitting, the cost volume is transformed into Encoding and decoding are regularized, and the softmax function is used to calculate the different depth sampling values of each pixel along the depth dimension The probability distribution at the depth is obtained , with a resolution of ,in, , is the feature map size, Assuming the number of planes, the resolution of the probability volume in stage 1 is , the resolution of stage 2 is , stage 3 is , the depth value is estimated using the softargmin function along the depth direction, which is described by formula (9):
[0038]
[0039] in, , is the minimum and maximum depth corresponding to the reference view pixel, is the depth sampling value, Different sampling depth values for pixels The probability value under the condition is estimated for all pixels to obtain the depth map. Stage 1 generates a depth map that roughly estimates the depth range. , the size is , stage 2 generates depth map , the size is , Final depth map of stage 3 , the size is ;
[0040] (c) Depth map fusion, including filtering and fusion. Filtering includes geometric constraints and photometric constraints. The geometric constraint is that there is a point in the reference view. , the estimated depth is ,Will Projected to the source view point, The depth of ,Will The corresponding point is then reprojected to a point in the reference image point, The estimated depth of the point correspondence is , if the constraints are met, it can be described by formula (10):
[0041]
[0042] but Satisfy geometric constraint consistency, photometric constraint is that the network passes through the probability body Get the depth map At the same time, for each point on the H and W planes, Calculate the sum of the probabilities of the four neighbors along the direction The direction takes the maximum probability and gets a probability map, depth map The more the depths of the pixels are concentrated near a certain depth, the higher the probability of accurate depth judgment of the point is. is 0.8, filtering out probabilities and less than Pixels of
[0043] Fusion is to combine multiple inspection scenes Figure 3 The 3D reconstruction network generates predicted depth maps corresponding to multiple perspectives. After filtering, a specific visualization fusion algorithm that minimizes occlusion and conflict is adopted to generate dense 3D color point cloud data of the inspection scene. ,in, 3D color point cloud data for inspection scenes The colored dots in For color points The coordinates of For color points Color information, is the number of points in the three-dimensional color point cloud;
[0044] Step 3: Multi-view inspection scenes Figure 3 The 3D reconstruction network is fine-tuned through transfer learning. To improve the inspection scene reconstruction capability, the training process is divided into two stages: pre-training and fine-tuning. The pre-training stage is based on the network structure of step 2 above, with the DTU dataset as the basic dataset, and the pre-training model weights are obtained through pre-training. The fine-tuning stage is based on the pre-training model weights, with the inspection scene data produced in step 1 as the fine-tuning dataset, and all layers are fine-tuned and trained. Although fine-tuning increases the training time of the network, all layers are fine-tuned and optimized, making the proposed network model more suitable for the inspection scene 3D reconstruction task.
[0045] Step 4: Divide the inspection scene dataset and configure the experimental environment to train and test the network model. The DTU public dataset is used in the pre-training stage, and the inspection scene dataset is used in the fine-tuning stage. Its structure is consistent with the multi-view stereo matching dataset BlendedMVS, which includes the inspection scene color image, camera pose and depth map data. The loss function uses the focal loss function. The loss weights of each stage are 1.0, 1.0, and 1.0. The total loss is the sum of the losses of each stage. The evaluation indicators of the pre-training test stage use the indicators of the DTU public dataset, including accuracy (Accuracy, Acc), completeness (Completeness, Comp), and overall index (Overall, OA). OA is described by formula (11).
[0046]
[0047] The overall indicator OA avoids the limitations of a single indicator by combining accuracy and completeness, and is a reliable basis for evaluating the quality of reconstruction. The evaluation indicators of the BlendedMVS dataset are used in the fine-tuning stage, namely, the end point error (EPE) representing the average difference between the predicted depth and the true depth value of the inspection scene dataset, and the percentage of pixels with an error greater than 1 depth pixel. , the percentage of pixels with a depth error greater than 3 pixels , the depth pixel is the resolution of the depth map in the depth direction, and the overall indicators OA, EPE, , The lower it is, the smaller the error is and the better the reconstruction effect is.
[0048] The beneficial effects of the present invention are: a multi-view inspection scene based on adaptive feature enhancement Figure 3 The method comprises the following steps: (1) constructing a depth map fusion framework based on laser radar and monocular camera self-calibration to obtain color images and depth map data of the inspection scene; (2) constructing a multi-view inspection scene based on adaptive feature enhancement. Figure 3 Dimensional reconstruction network, (3) multi-view inspection scene Figure 3 Dimensional reconstruction network for transfer learning fine-tuning,
[0049] (4) Divide the inspection scene data set and configure the experimental environment to train and test the network model. Figure 3 dimensional reconstruction network, including an inspection scene adaptive feature enhancement network, multi-scale depth estimation and depth fusion. The inspection scene adaptive feature enhancement network is based on a feature pyramid network, which introduces an attention module into each layer of bottom-up multi-scale feature extraction. By assigning attention weights to scene feature information in three-dimensional space, feature information is enriched, and a feature adaptive perception module based on deformable convolution is constructed at the network output. By expanding the receptive field of the feature extraction network, the network's sampling capability for feature maps of different scales is enhanced, and the perception capability for detail features such as boundary information and texture information is improved. Multi-scale depth estimation combines the feature map extracted by the adaptive feature enhancement network with the camera pose information, and performs multi-view depth estimation and depth map generation from coarse to fine through a three-level cascade structure. Depth map fusion is to fuse all color images and corresponding depth maps into a three-dimensional color point cloud of the inspection scene. The multi-view proposed in the present invention Figure 3 In the pre-training stage, the overall index of the network on the public data set is better than that of other network models. In the transfer learning fine-tuning stage, the evaluation indexes on the inspection scene data set are lower than the values before the baseline network and fine-tuning. Fine-tuning makes the network more suitable for the reconstruction of inspection scenes and achieves ideal results. The point cloud reconstructed by the present invention in the weak texture and non-Lambertian surface areas of the public data set and inspection scene is the most complete, and more refined surface details are restored. This is because in the process of feature extraction by the adaptive feature enhancement network, on the one hand, the network extracts multi-scale features from the bottom up and aggregates multi-scale features from the top down to fuse features and semantic information of different scales; on the other hand, the attention module is introduced and the feature adaptive perception module is constructed to expand the receptive field, enhance the network feature extraction capability, enrich the feature information, and improve the perception of detail features such as boundary information and texture information. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 It is a flow chart of the steps of the method of the present invention.
[0051] Figure 2 It is a schematic diagram of the depth map fusion framework of the present invention.
[0052] Figure 3 It is a schematic diagram of the overall architecture of the network model of the present invention.
[0053] Figure 4 It is a schematic diagram of the adaptive feature enhancement network of the present invention.
[0054] Figure 5 Schematic diagram of the attention module of the present invention.
[0055] Figure 6 It is a schematic diagram of the feature adaptive perception module of the present invention.
[0056] Figure 7 Schematic diagram of the camera viewing cone of the present invention.
[0057] Figure 8 It is a schematic diagram of cascade depth map refinement of the present invention.
[0058] Fig. 9 It is a schematic diagram of the transfer learning fine-tuning process of the present invention. DETAILED DESCRIPTION
[0059] like Figure 1 As shown in the figure, a multi-view inspection scene based on adaptive feature enhancement Figure 3 The reconstruction method comprises the following steps:
[0060] Step 1: Build a depth map fusion framework based on LiDAR and monocular camera self-calibration to obtain color images and depth map data of the inspection scene. The schematic diagram of the depth map fusion framework is as follows: Figure 2 As shown, it specifically includes the following sub-steps:
[0061] (a) Self-calibration of LiDAR and monocular camera. The color image and 3D point cloud data of the inspection scene are obtained through the monocular camera of the intelligent robot and the LiDAR installed above the monocular camera. The black and white checkerboard calibration method and a self-calibration method of LiDAR and camera are introduced to obtain the camera intrinsic parameters and the conversion matrix from the LiDAR coordinate system to the camera coordinate system. ;
[0062] (b) Inspection scene depth map completion, which consists of two parts: inspection scene sparse depth map generation and inspection scene sparse depth map completion:
[0063] 1) Generation of sparse depth map of inspection scene. Suppose the point cloud of the inspection scene is ,in It is the inspection scene point cloud data. Indicates Points, is the number of point clouds, and the points pass Convert to the camera coordinate system, then project it onto the color image of the inspection scene through the camera internal parameters, remove invalid points outside the image range, and obtain the point cloud Sparse depth map projected on the image ,in Yes By transforming the matrix and the camera intrinsic parameters projected to the pixel coordinates of the image, for Corresponding depth value, is the number of valid point clouds projected into the image range;
[0064] 2) Completion of sparse depth maps of inspection scenes. Since sparse depth maps are obtained Some pixels in the image have no depth value, resulting in holes of varying sizes. A depth map completion method is introduced to complete the holes, using a greedy strategy: effective depth pixels Surrounding empty values may have similar values, so fill the original sparse depth map first Medium effective depth pixels The surrounding small holes are filled in the order of large holes, and finally a 5×5 median blur is used to remove noise, remove outliers on the local edges, and then a 5×5 Gaussian blur is used to smooth the local plane to obtain the final inspection scene depth map. ,in
[0065] , is the number of valid pixels within the completed image range;
[0066] Step 2: Build a multi-view inspection scene based on adaptive feature enhancement Figure 3 The overall architecture of the network model is as follows: Figure 3 As shown, it includes an adaptive feature enhancement network, multi-scale depth estimation, and depth map fusion. The adaptive feature enhancement network is used to extract and enhance the feature information of the input inspection scene image. The multi-scale depth estimation combines the feature map extracted by the adaptive feature enhancement network and the camera pose information obtained by the incremental SFM technology. Through a three-level cascade structure, the scene multi-view depth estimation and depth map generation are performed from coarse to fine. The depth map fusion is to fuse all color images and corresponding depth maps into a three-dimensional color point cloud of the inspection scene, which specifically includes the following sub-steps:
[0067] (a) Building an adaptive feature enhancement network based on the feature pyramid network, such as Figure 4 As shown, the input size is Multi-view inspection scene ,in, For reference view, is the source view, is the number of multiple views, extract multiple views The multi-scale detail features such as scene boundary information and texture information are output as three-scale feature maps with global and detail semantic information. ,in, is the reference view feature map, is the source view feature map, The size and number of channels are , , , the adaptive feature enhancement network consists of three parts:
[0068] 1) Bottom-up multi-scale feature extraction. First, use a 3×3 convolution to extract the size of Multi-view inspection scene The size and number of channels are extracted. The feature map is then extracted using a 3×3 convolution and an attention module SimAM. The attention module SimAM is as follows: Figure 5 As shown, the feature map is , Secondly, use a 5×5 convolution with a step size of 2 to downsample the feature map, and then use two 3×3 convolutions and an attention module SimAM to complete the secondary feature extraction. At this time, the feature map size and the number of channels are The third feature extraction is the same as the second feature extraction operation. The feature map size and number of channels are ;
[0069] 2) Top-down multi-scale feature aggregation,First, the third extraction feature map of the bottom-up multi-scale feature extraction part remains unchanged, and the size and number of channels are , the first and second output feature maps are converted to 32 channels through a 1×1 convolution. Secondly, the feature maps after the third and second 1×1 convolutions are upsampled by 2 times using the nearest neighbor interpolation method, and are added to the feature maps after the second and first 1×1 convolutions to complete the feature aggregation operation. At this time, the size and number of channels of the three scale feature maps are , , ;
[0070] 3) Construct a feature adaptive perception module (APMF) consisting of deformable convolution, Bn layer, and Relu layer. The feature adaptive perception module APMF is as follows: Figure 6 As shown in the figure, the three scale feature maps output by the top-down multi-scale feature aggregation part are output through APMF to finally extract the feature map, and the size and number of channels are respectively , , , corresponding to the input feature maps of stage 1, stage 2, and stage 3 in multi-scale depth estimation;
[0071] (b) Multi-scale depth estimation, constructed as Figure 3 The three-level cascade structure shown in the figure, each level represents a stage, and depth estimation is performed in three stages. Each stage of the cascade structure includes two parts: cost volume construction and depth prediction. The input of each stage is the camera pose information and the feature map extracted by the adaptive feature enhancement network. Among them, the camera pose information is obtained by the incremental Sfm technology, and the feature map is the three-scale feature map output by the adaptive feature enhancement network. The size and number of channels are respectively , , , respectively as the input of stage 1, stage 2, and stage 3 in multi-scale depth estimation. All three stages perform cost volume construction and depth prediction to output depth maps. Stage 1 outputs a depth map that roughly estimates the depth range. , the size is , stage 2 and stage 3 output depth map And the final predicted depth map , the sizes are , ,The difference between stage 2 and stage 3 and stage 1 is that stage 2 and stage 3 respectively use the estimated depth value of the previous stage as the median of the depth range, and achieve the purpose of cascade depth map refinement by reducing the assumed depth range, depth interval and the number of assumed planes;
[0072] Cost volume construction is to measure the feature similarity between multiple views of the inspection scene. It requires three steps. First, multiple discrete depth hypothesis planes are determined. Second, the source view extracted by the adaptive feature enhancement network is Corresponding Warp to the reference view via homography The feature volume is constructed on multiple hypothetical planes. Finally, the feature volumes are fused together to construct a three-dimensional cost volume. Depth prediction is to regularize the cost volume into a probability volume, and then find the expectation of the probability volume along the depth direction to obtain the estimated depth value. The specific process of cost volume construction and depth prediction is as follows:
[0073] 1) Determine the depth hypothesis plane. First, under the camera frustum corresponding to the reference view, the camera frustum is as follows: Figure 7 As shown, the principal optical axis of the reference camera is established The depth hypothetical plane is perpendicular to each other, and the depth range of stage 1 is set to ,from arrive , evenly selected within this range depth hypothesis plane, then The sampling depth value corresponding to the plane , described by formula (1),
[0074]
[0075] in, , Indicates depth range The minimum and maximum values of is the depth interval size of stage 1, The value range is 1 to ,In the cascade structure, stage 2 and stage 3 use the estimated depth value of the previous stage as the median of the depth range, and ,refine the cascade depth map by narrowing the assumed depth range, depth interval and reducing the number of ,assumed planes.,The cascade depth map refinement of stage 1 and 2 is shown in the example. Figure 8 As shown, stages 2 and 3 reduce the assumed depth range, which is described by formula (2):
[0076]
[0077] in, Indicates the stage, with values of 1 and 2. , It is a stage and the depth hypothesis range of the previous stage, It is a stage Assume the reduction factor of the range, Representation stage Depth range The minimum value of Representation stage Depth range The maximum value of For stage Reference view pixels The predicted depth of the image is reduced, and the depth interval is described by formula (3).
[0078]
[0079] in, , The stages and the depth interval size of the previous stage, It is a stage The reduction factor of the depth interval, given the stage Assumption Range and depth interval size , then the corresponding number of assumed planes is described by formula (4):
[0080]
[0081] At this point, stage In the The sampling depth value corresponding to the plane , described by formula (5),
[0082]
[0083] in, Indicates the stage, with values of 1 and 2. Representation stage Depth range The minimum value of The value range is 1 to ,Cascaded depth map refinement uses adaptive depth sampling of the above formula to use computing and memory resources in a more appropriate estimation range, which can significantly reduce computing time and video memory consumption;
[0084] 2) Construct a feature body and enhance the feature map extracted by the adaptive feature enhancement network Warp to the reference view by the homography transformation matrix The feature volume is constructed on the hypothetical plane of the camera cone, and the homography transformation is described by formula (6):
[0085]
[0086] in, is the unit vector referring to the main axis of the camera, Represents the posture information, for The corresponding camera intrinsic parameters, is the rotation matrix and Translation matrix, , , For reference view Corresponding camera pose information, , , For the Zhang Yuan view corresponds to the camera pose information, is the identity matrix, represents the projection equation, for The pixels on express Project to Pixels, Indicates the sampling depth value At correspond The eigenvalues are projected onto The homography transformation matrix of is , because the three-stage cascade is used for coarse-to-fine depth estimation, stage 1 is to roughly estimate the depth range, and stage 2 and stage 3 are respectively combined with the depth range of the previous stage. By reducing the assumed depth range, depth interval and the number of assumed planes, the accuracy of the feature map volume is improved. According to the cascade depth map refinement method, formula (6) is modified to obtain stage The homography matrix of is described by formula (7):
[0087]
[0088] in, Indicates the stage, with values of 1 and 2. For stage Reference view pixels The predicted depth, For the stage To learn The residual depth of Construct more and more detailed depth planes and transform them through formula (7) to obtain the feature volume: , with a resolution of ,in, , is the feature map size, is the number of assumed planes, is the number of channels of the feature body, and the resolution of the feature body in stage 1 is , stage 2 is , stage 3 is ;
[0089] 3) Feature fusion, in order to measure The feature differences between views are calculated by using the variance-based cost metric. Aggregated into a unified matching cost body, described by formula (8):
[0090]
[0091] in, For the price body, is a variance-based cost metric, is the mean of all features, is the number of views, the cost volume Resolution is ,in, , is the feature map size, is the number of assumed planes, is the number of channels of the feature body, and the cost body resolution after feature body fusion stage 1 is , stage 2 is , stage 3 is ;
[0092] 4) Depth prediction: In order to filter noise and prevent overfitting, the cost volume is transformed into Encoding and decoding are regularized, and the softmax function is used to calculate the different depth sampling values of each pixel along the depth dimension The probability distribution at the depth is obtained , with a resolution of ,in, , is the feature map size, Assuming the number of planes, the resolution of the probability volume in stage 1 is , the resolution of stage 2 is , stage 3 is , the depth value is estimated using the softargmin function along the depth direction, which is described by formula (9):
[0093]
[0094] in, , is the minimum and maximum depth corresponding to the reference view pixel, is the depth sampling value, Different sampling depth values for pixels The probability value under the condition is estimated for all pixels to obtain the depth map. Stage 1 generates a depth map that roughly estimates the depth range. , the size is , stage 2 generates depth map , the size is , Final depth map of stage 3 , the size is ;
[0095] (c) Depth map fusion, including filtering and fusion. Filtering includes geometric constraints and photometric constraints. The geometric constraint is that there is a point in the reference view. , the estimated depth is ,Will Projected to the source view point, The depth of ,Will The corresponding point is then reprojected to a point in the reference image point, The estimated depth of the point correspondence is , if the constraints are met, it can be described by formula (10):
[0096]
[0097] but Satisfy geometric constraint consistency, photometric constraint is that the network passes through the probability body Get the depth map At the same time, for each point on the H and W planes, Calculate the sum of the probabilities of the four neighbors along the direction The direction takes the maximum probability and gets a probability map, depth map The more the depths of the pixels are concentrated near a certain depth, the higher the probability of accurate depth judgment of the point is. is 0.8, filtering out probabilities and less than Pixels of
[0098] Fusion is to combine multiple inspection scenes Figure 3 The 3D reconstruction network generates predicted depth maps corresponding to multiple perspectives. After filtering, a specific visualization fusion algorithm that minimizes occlusion and conflict is adopted to generate dense 3D color point cloud data of the inspection scene. ,in, 3D color point cloud data for inspection scenes The colored dots in For color points The coordinates of For color points Color information, is the number of points in the three-dimensional color point cloud;
[0099] Step 3: Multi-view inspection scenes Figure 3 The network is reconstructed to perform transfer learning fine-tuning. The transfer learning fine-tuning process is as follows: Fig. 9 As shown in the figure, in order to improve the inspection scene reconstruction capability, the training process is divided into two stages, the pre-training stage and the fine-tuning stage. The pre-training stage is based on the network structure of step 2 above, with the DTU dataset as the basic dataset, and the pre-training model weights are obtained by pre-training. The fine-tuning stage is based on the pre-training model weights, with the inspection scene data produced in step 1 as the fine-tuning dataset, and all layers are fine-tuned and trained. Although fine-tuning increases the training time of the network, all layers are fine-tuned and optimized, making the proposed network model more suitable for the inspection scene 3D reconstruction task;
[0100] Step 4: Divide the inspection scene dataset and configure the experimental environment to train and test the network model. The DTU public dataset is used in the pre-training stage, and the inspection scene dataset is used in the fine-tuning stage. Its structure is consistent with the multi-view stereo matching dataset BlendedMVS, which includes the inspection scene color image, camera pose and depth map data. The loss function uses the focal loss function. The loss weights of each stage are 1.0, 1.0, and 1.0. The total loss is the sum of the losses of each stage. The evaluation indicators of the pre-training test stage use the indicators of the DTU public dataset, including accuracy (Accuracy, Acc), completeness (Completeness, Comp), and overall index (Overall, OA). OA is described by formula (11):
[0101]
[0102] The overall indicator OA avoids the limitations of a single indicator by combining accuracy and completeness, and is a reliable basis for evaluating the quality of reconstruction. The evaluation indicators of the BlendedMVS dataset are used in the fine-tuning stage, namely, the end point error (EPE) representing the average difference between the predicted depth and the true depth value of the inspection scene dataset, and the percentage of pixels with an error greater than 1 depth pixel. , the percentage of pixels with a depth error greater than 3 pixels , the depth pixel is the resolution of the depth map in the depth direction, and the overall indicators OA, EPE, , The lower it is, the smaller the error is and the better the reconstruction effect is.
[0103] Multi-view inspection scene based on adaptive feature enhancement Figure 3 The results of the 3D reconstruction network on the DTU public dataset are shown in Table 1, and the reconstruction results in the inspection scenario are shown in Table 2.
[0104] Table 1 Reconstruction results on the DTU public dataset
[0105]
[0106] Table 2 Reconstruction results in inspection scenario
[0107]
Claims
1. A method for multi-view 3D reconstruction of inspection scenes based on adaptive feature enhancement, characterized in that: The following steps are involved: Step 1: Build a depth map fusion framework based on LiDAR and monocular camera self-calibration to obtain color images and depth map data of the inspection scene, which includes the following sub-steps: (a) Self-calibration of LiDAR and monocular camera. The color image and 3D point cloud data of the inspection scene are obtained through the monocular camera of the intelligent robot and the LiDAR installed above the monocular camera. The black and white checkerboard calibration method and a self-calibration method of LiDAR and camera are introduced to obtain the camera intrinsic parameters and the conversion matrix from the LiDAR coordinate system to the camera coordinate system. ; (b) Inspection scene depth map completion, which consists of two parts: inspection scene sparse depth map generation and inspection scene sparse depth map completion: 1) Generation of sparse depth map of inspection scene. Suppose the point cloud of the inspection scene is ,in It is the inspection scene point cloud data. Indicates Points, is the number of point clouds, and the points pass Convert to the camera coordinate system, then project it onto the color image of the inspection scene through the camera internal parameters, remove invalid points outside the image range, and obtain the point cloud Sparse depth map projected on the image ,in Yes By transforming the matrix and the camera intrinsic parameters projected to the pixel coordinates of the image, for Corresponding depth value, is the number of valid point clouds projected into the image range; 2) Completion of sparse depth maps of inspection scenes. Since sparse depth maps are obtained Some pixels in the image have no depth value, resulting in holes of varying sizes. A depth map completion method is introduced to complete the holes, using a greedy strategy: effective depth pixels Surrounding empty values may have similar values, so fill the original sparse depth map first Medium effective depth pixels The surrounding small holes are filled in the order of large holes, and finally a 5×5 median blur is used to remove noise, remove outliers on the local edges, and then a 5×5 Gaussian blur is used to smooth the local plane to obtain the final inspection scene depth map. ,in , is the number of valid pixels within the completed image range; Step 2: Build a multi-view 3D reconstruction network for inspection scenes based on adaptive feature enhancement, including adaptive feature enhancement network, multi-scale depth estimation, and depth map fusion. The adaptive feature enhancement network is used to extract and enhance the feature information of the input inspection scene image. The multi-scale depth estimation combines the feature map extracted by the adaptive feature enhancement network and the camera pose information obtained by the incremental SFM technology. Through the three-level cascade structure, the scene multi-view depth estimation and depth map generation are performed from coarse to fine. The depth map fusion is to fuse all color images and corresponding depth maps into a three-dimensional color point cloud of the inspection scene, which specifically includes the following sub-steps: (a) Construct an adaptive feature enhancement network based on the feature pyramid network, with an input size of Multi-view inspection scene ,in, For reference view, is the source view, is the number of multiple views, extract multiple views The multi-scale detail features such as scene boundary information and texture information are output as three-scale feature maps with global and detail semantic information. ,in, is the reference view feature map, is the source view feature map, The size and number of channels are , , , the adaptive feature enhancement network consists of three parts: 1) Bottom-up multi-scale feature extraction. First, use a 3×3 convolution to extract the size of Multi-view inspection scene The size and number of channels are extracted. Then, a 3×3 convolution and an attention module SimAM are used to complete the initial feature extraction. At this time, the feature map is , Secondly, use a 5×5 convolution with a step size of 2 to downsample the feature map, and then use two 3×3 convolutions and an attention module SimAM to complete the secondary feature extraction. At this time, the feature map size and the number of channels are The third feature extraction is the same as the second feature extraction operation. The feature map size and number of channels are ; 2) Top-down multi-scale feature aggregation,First, the third extraction feature map of the bottom-up multi-scale feature extraction part remains unchanged, and the size and number of channels are , the first and second output feature maps are converted to 32 channels through a 1×1 convolution. Secondly, the feature maps after the third and second 1×1 convolutions are upsampled by 2 times using the nearest neighbor interpolation method, and are added to the feature maps after the second and first 1×1 convolutions to complete the feature aggregation operation. At this time, the size and number of channels of the three scale feature maps are , , ; 3) Construct an adaptive perception module for feature (APMF) consisting of deformable convolution, Bn layer, and Relu layer. The three scale feature maps output by the top-down multi-scale feature aggregation part are output through APMF to finally extract the feature map. The size and number of channels are , , , corresponding to the input feature maps of stage 1, stage 2, and stage 3 in multi-scale depth estimation; (b) Multi-scale depth estimation. A three-level cascade structure is constructed. Each level represents a stage. Depth estimation is performed in three stages. Each stage of the cascade structure includes two parts: cost volume construction and depth prediction. The input of each stage is the camera pose information and the feature map extracted by the adaptive feature enhancement network. The camera pose information is obtained by the incremental Sfm technology. The feature map is the three-scale feature map output by the adaptive feature enhancement network. The size and number of channels are respectively , , , respectively as the input of stage 1, stage 2, and stage 3 in multi-scale depth estimation. All three stages perform cost volume construction and depth prediction to output depth maps. Stage 1 outputs a depth map that roughly estimates the depth range. , the size is , stage 2 and stage 3 output depth map And the final predicted depth map , the sizes are , ,The difference between stage 2 and stage 3 and stage 1 is that stage 2 and stage 3 respectively use the estimated depth value of the previous stage as the median of the depth range, and achieve the purpose of cascade depth map refinement by reducing the assumed depth range, depth interval and the number of assumed planes; Cost volume construction is to measure the feature similarity between multiple views of the inspection scene. It requires three steps. First, multiple discrete depth hypothesis planes are determined. Second, the source view extracted by the adaptive feature enhancement network is Corresponding Warp to the reference view via homography The feature volume is constructed on multiple hypothetical planes. Finally, the feature volumes are fused together to construct a three-dimensional cost volume. Depth prediction is to regularize the cost volume into a probability volume, and then find the expectation of the probability volume along the depth direction to obtain the estimated depth value. The specific process of cost volume construction and depth prediction is as follows: 1) Determine the depth hypothesis plane. First, under the camera frustum corresponding to the reference view, establish a plane that is consistent with the principal optical axis of the reference camera. The depth hypothetical plane is perpendicular to each other, and the depth range of stage 1 is set to ,from arrive , evenly selected within this range depth hypothesis plane, then The sampling depth value corresponding to the plane , described by formula (1), ; in, , Indicates depth range The minimum and maximum values of is the depth interval size of stage 1, The value range is 1 to ,In the cascade structure, stage 2 and stage 3 use the estimated depth value of the previous stage as the median of the depth range, and refine the cascade depth map by reducing the assumed depth range, depth interval and the number of assumed planes. Stage 2 and stage 3 reduce the assumed depth range, which is described by formula (2). ; in, Indicates the stage, with values of 1 and 2. , It is a stage and the depth hypothesis range of the previous stage, It is a stage Assume the reduction factor of the range, Representation stage Depth range The minimum value of Representation stage Depth range The maximum value of For stage Reference view pixels The predicted depth of the , narrowing the depth interval, is described by formula (3), ; in, , The stages and the depth interval size of the previous stage, It is a stage The reduction factor of the depth interval, given the stage Assumption Range and depth interval size , then the corresponding number of assumed planes is described by formula (4): ; At this point, stage In the The sampling depth value corresponding to the plane , described by formula (5), ; in, Indicates the stage, with values of 1 and 2. Representation stage Depth range The minimum value of The value range is 1 to ,Cascaded depth map refinement uses adaptive depth sampling of the above formula to use computing and memory resources in a more appropriate estimation range, which can significantly reduce computing time and video memory consumption; 2) Construct a feature body and enhance the feature map extracted by the adaptive feature enhancement network Warp to the reference view by the homography transformation matrix The feature volume is constructed on the hypothetical plane of the camera cone, and the homography transformation is described by formula (6): ; in, is the unit vector referring to the main axis of the camera, Represents the posture information, for The corresponding camera intrinsic parameters, is the rotation matrix and Translation matrix, , , For reference view Corresponding camera pose information, , , For the Zhang Yuan view corresponds to the camera pose information, is the identity matrix, represents the projection equation, for The pixels on express Project to Pixels, Indicates the sampling depth value At correspond The eigenvalues are projected onto The homography transformation matrix of is , because the three-stage cascade is used for coarse-to-fine depth estimation, stage 1 is to roughly estimate the depth range, and stage 2 and stage 3 are respectively combined with the depth range of the previous stage. By reducing the assumed depth range, depth interval and the number of assumed planes, the accuracy of the feature map volume is improved. According to the cascade depth map refinement method, formula (6) is modified to obtain stage The homography matrix of is described by formula (7): ; in, Indicates the stage, with values of 1 and 2. For stage Reference view pixels The predicted depth, For the stage To learn The residual depth of Construct more and more detailed depth planes and transform them through formula (7) to obtain the feature volume: , with a resolution of ,in, , is the feature map size, is the number of assumed planes, is the number of channels of the feature body, and the resolution of the feature body in stage 1 is , stage 2 is , stage 3 is ; 3) Feature fusion, in order to measure The feature differences between views are calculated by using the variance-based cost metric. Aggregated into a unified matching cost body, described by formula (8): ; in, For the price body, is a variance-based cost metric, is the mean of all features, is the number of views, the cost volume Resolution is ,in, , is the feature map size, is the number of assumed planes, is the number of channels of the feature body, and the cost body resolution after feature body fusion stage 1 is , stage 2 is , stage 3 is ; 4) Depth prediction: In order to filter noise and prevent overfitting, the cost volume is transformed into Encoding and decoding are regularized, and the softmax function is used to calculate the different depth sampling values of each pixel along the depth dimension The probability distribution at the depth is obtained , with a resolution of ,in, , is the feature map size, Assuming the number of planes, the resolution of the probability volume in stage 1 is , the resolution of stage 2 is , stage 3 is , the depth value is estimated using the softargmin function along the depth direction, which is described by formula (9): ; in, , is the minimum and maximum depth corresponding to the reference view pixel, is the depth sampling value, Different sampling depth values for pixels The probability value under the condition is estimated for all pixels to obtain the depth map. Stage 1 generates a depth map that roughly estimates the depth range. , the size is , stage 2 generates depth map , the size is , Final depth map of stage 3 , the size is ; (c) Depth map fusion, including filtering and fusion. Filtering includes geometric constraints and photometric constraints. The geometric constraint is that there is a point in the reference view. , the estimated depth is ,Will Projected to the source view point, The depth estimate is ,Will The corresponding point is then reprojected to a point in the reference image point, The estimated depth of the point correspondence is , if the constraints are met, it can be described by formula (10): ; but Satisfy geometric constraint consistency, photometric constraint is that the network passes through the probability body Get the depth map At the same time, for each point on the H and W planes, Calculate the sum of the probabilities of the four neighbors along the direction The direction takes the maximum probability and gets a probability map, depth map The more the depths of the pixels are concentrated near a certain depth, the higher the probability of accurate depth judgment of the point is. is 0.8, filtering out probabilities and less than Pixels of Fusion is to generate dense 3D color point cloud data of the inspection scene by filtering the predicted depth maps corresponding to multiple perspectives generated by the multi-view 3D reconstruction network of the inspection scene, and then adopt a specific visualization fusion algorithm that minimizes occlusion and conflict. ,in, 3D color point cloud data for inspection scenes The colored dots in For color points The coordinates of For color points Color information, is the number of points in the three-dimensional color point cloud; Step 3: Perform transfer learning fine-tuning on the multi-view 3D reconstruction network of the inspection scene. To improve the ability to reconstruct the inspection scene, the training process is divided into two stages: pre-training and fine-tuning. The pre-training stage is based on the network structure of step 2 above, with the DTU dataset as the basic dataset, and the pre-training model weights are obtained through pre-training. The fine-tuning stage is based on the pre-training model weights, with the inspection scene data produced in step 1 as the fine-tuning dataset, and all layers are fine-tuned and trained. Although fine-tuning increases the training time of the network, all layers are fine-tuned and optimized, making the proposed network model more suitable for the task of 3D reconstruction of inspection scenes. Step 4: Divide the inspection scene dataset and configure the experimental environment to train and test the network model. The DTU public dataset is used in the pre-training stage, and the inspection scene dataset is used in the fine-tuning stage. Its structure is consistent with the multi-view stereo matching dataset BlendedMVS, which includes the inspection scene color image, camera pose and depth map data. The loss function uses the focal loss function. The loss weights of each stage are 1.0, 1.0, and 1.
0. The total loss is the sum of the losses of each stage. The evaluation indicators of the pre-training test stage use the indicators of the DTU public dataset, including accuracy (Accuracy, Acc), completeness (Completeness, Comp), and overall index (Overall, OA). OA is described by formula (11): ; The overall indicator OA avoids the limitations of a single indicator by combining accuracy and completeness, and is a reliable basis for evaluating the quality of reconstruction. The evaluation indicators of the BlendedMVS dataset are used in the fine-tuning stage, namely, the end point error (EPE) representing the average difference between the predicted depth and the true depth value of the inspection scene dataset, and the percentage of pixels with an error greater than 1 depth pixel. , the percentage of pixels with a depth error greater than 3 pixels , the depth pixel is the resolution of the depth map in the depth direction, and the overall indicators OA, EPE, , The lower it is, the smaller the error is and the better the reconstruction effect is.
Citation Information
Cited By
Digital monitoring platform for gas drainer based on image recognition
CN120259141A
Outdoor boundless scene three-dimensional reconstruction method based on multi-scale features and deep supervision neural radiation field
CN120431275A
Homography assisted unmanned aerial vehicle landing method based on computer vision and deep learning
CN120780007A
Equal-scale image projection method and equipment based on Leiyu fusion, and storage medium
CN121235961A
Unmanned aerial vehicle oblique photography three-dimensional reconstruction method and system based on machine vision
CN122244333A