Pipeline inner wall defect detection method, system and product based on point cloud and image depth fusion
By using a point cloud and image deep fusion method, three-dimensional reconstruction of the inner wall of the pipeline and quantitative assessment of defects were achieved, solving the problem of low utilization of three-dimensional information in existing technologies and realizing high-precision defect detection and quantitative calculation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI JINXING MUNICIPAL DESIGN CONSULTING CO LTD
- Filing Date
- 2026-01-20
- Publication Date
- 2026-05-08
AI Technical Summary
Existing drainage pipeline inspection methods suffer from low utilization of three-dimensional information, severe information loss, difficulty in accurately identifying defects under complex pipeline conditions, and lack of precise measurement capabilities for defect dimensions, especially quantitative calculation of three-dimensional physical quantities such as crack depth and depression volume.
A point cloud and image deep fusion method is adopted, which achieves three-dimensional reconstruction of the inner wall of the pipeline and quantitative assessment of defects by using two-dimensional image feature upscaling and three-dimensional voxel-level feature fusion, combined with BEV and perspective feature map dual-view detection. Sparse three-dimensional convolutional neural network is used for feature interaction and fusion.
It significantly improves the accuracy of defect identification, can accurately quantify the size of internal defects in pipelines with an accuracy of ±5mm, improves the system's detection performance in complex environments, and reduces the false alarm rate.
Smart Images

Figure CN121998930A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of drainage pipeline defect detection, specifically to a method, system, and product for detecting pipeline inner wall defects based on point cloud and image deep fusion. Background Technology
[0002] Urban water supply and drainage pipelines are an important component of urban municipal infrastructure, responsible for transporting sewage, rainwater, and domestic water. As the pipeline network ages, defects such as cracks, corrosion, disconnections, and misalignments frequently appear in the pipeline system. These defects severely affect the structural integrity and function of the pipelines. If not properly addressed, they can lead to pipeline leaks, and in severe cases, can cause serious consequences such as environmental pollution and ground subsidence.
[0003] Existing methods for inspecting drainage pipes include: 1) Closed-circuit television (CCTV) inspection: This method acquires two-dimensional video through cameras moving inside the pipes, relying on experienced inspectors for subjective interpretation. However, it suffers from several drawbacks: it only provides two-dimensional image information, lacking depth and geometric dimension information; the interpretation results are qualitative rather than quantitative; it is greatly affected by environmental factors such as lighting and water quality; its work efficiency is relatively low; and it involves high labor costs and significant errors in subjective interpretation.
[0004] 2) LiDAR (LiDAR) Inspection: This method uses lidar to acquire 3D point clouds of the pipe cross-section to assess the geometric distortion of the pipe's cross-section. However, it has limitations such as the point cloud itself lacking texture and color information, making it difficult to accurately distinguish between surface dirt, water stains, and actual structural defects; limited resolution for minute defects (such as fine cracks); and inability to output quantitative indicators of defect depth.
[0005] 3) Two-dimensional projection fusion scheme: Some studies project laser point clouds into pseudo-color depth maps and fuse them with grayscale or color images in a two-dimensional pixel plane to achieve defect identification. However, since the projection process essentially compresses three-dimensional information into two dimensions, the rich information of the normal distance (depth dimension) is severely compressed. The fusion is still carried out in a two-dimensional plane, which cannot fully utilize the geometric constraints of three-dimensional space. The final quantitative detection results of defects (such as crack width) are still based on two-dimensional pixel estimation, which limits the accuracy; it cannot accurately handle complex three-dimensional defect morphologies.
[0006] In summary, the main problems with existing technologies are low utilization of three-dimensional information and significant information loss, particularly along the normal (depth) direction. This leads to difficulties in distinguishing surface contaminants from actual defects in two-dimensional images under complex pipeline conditions (e.g., with deposits, water accumulation, or bubbles inside the pipe); insufficient ability to identify complex pipeline deformations (e.g., ellipticization, collapse) and deep defects (e.g., internal corrosion, stress corrosion); and insufficient accuracy in defect size estimation based on two-dimensional pixel images, typically ±15–30 mm, lacking the ability to accurately measure defect size and quantify three-dimensional physical quantities such as crack depth and indentation volume. These indicators are crucial for structural bearing capacity verification based on limit states and for prioritizing repairs. Summary of the Invention
[0007] The purpose of this application is to overcome the shortcomings of the prior art and provide a method, system and product for detecting defects in the inner wall of pipelines based on the deep fusion of point cloud and image. It integrates the three-dimensional point cloud of LiDAR with high-resolution digital image information to realize the three-dimensional precise reconstruction of the inner wall of the pipeline and the quantitative assessment of defects. It is applicable to the fine inspection, structural condition assessment and maintenance decision-making of municipal basic pipeline facilities such as drainage pipes, combined sewer pipes, pressure pipes and box culverts.
[0008] Firstly, this application provides a method for detecting defects in the inner wall of a pipe based on point cloud and image depth fusion, the technical solution of which includes the following steps: S100, acquire the two-dimensional image sequence and point cloud sequence of the inner wall of the pipe respectively; S200: Feature extraction is performed on the two-dimensional image sequence to obtain two-dimensional image feature vectors. The two-dimensional image feature vectors are then up-dimensionalized, and back-projected onto a pre-constructed three-dimensional voxel grid space to obtain the two-dimensional image voxel feature volume V. img ; S300, the point cloud sequence is mapped to the three-dimensional voxel grid space, feature extraction is performed to obtain point cloud feature vectors, and then the point cloud voxel feature volume V is obtained through the relationship between the point cloud and the three-dimensional voxels. pts ; S400, within the three-dimensional voxel grid space, perform the two-dimensional image voxel feature volume V corresponding to the three-dimensional voxels. img and point cloud voxel feature V pts Feature fusion yields fused voxel feature body V. fused ; S500, based on fusion voxel feature V fused The BEV feature map and perspective feature map of the inner wall of the pipe are obtained respectively. Based on the BEV feature map and perspective feature map, a three-dimensional voxel mask for identifying three-dimensional defects in the inner wall of the pipe is obtained, and the defect point cloud corresponding to the three-dimensional voxel mask is obtained. S600 calculates three-dimensional defects on the inner wall of a pipeline based on defect point clouds.
[0009] By adopting the above technical solution, the lossless fusion of two-dimensional image information and three-dimensional point cloud information is achieved, avoiding the loss of depth information in traditional two-dimensional projection. At the same time, defect detection and segmentation are performed simultaneously from two perspectives: BEV feature map and perspective feature map. This enables accurate identification and quantitative calculation of three-dimensional defects on the inner wall of the pipeline.
[0010] As a preferred embodiment, the specific method for obtaining the two-dimensional image voxel feature volume in S200 is as follows: S201, For any pixel (u,v) on a two-dimensional image, calculate its back-projection frustum ray in the three-dimensional voxel grid space; Let the homogeneous coordinates of the pixel be... Its normalized coordinates are, The back-projected cone ray is obtained through the cone ray equation. ,have: ; in, To obtain the position of the camera optical center in the global coordinate system of a 2D image, Let K be the ray direction, K be the camera intrinsic calibration matrix, T be the extrinsic calibration matrix, and t∈[0,∞) be the ray parameters; S202, in the three-dimensional voxel mesh space V, the back-projected view cone ray is calculated using the ray traversal algorithm. The set of all voxels passed through ; S203, extract the feature vector corresponding to pixel (u, v) According to the weighted rule, the rays are scattered into the voxels through which they pass, where C represents the set of real numbers, and C is the number of feature channels in a two-dimensional image. S204, for the same voxel v projected by multiple rays i Weighted average pooling is used to obtain the two-dimensional image voxel feature volume V. img ,have: ; Wherein, S(v) i ) is projected onto voxel v i The set of all pixels, w k f is the weight corresponding to the k-th pixel. k The feature vector corresponding to the k-th pixel is the image voxel feature volume. , where (L,W,H) are the voxel grid size and C is the number of feature channels in the two-dimensional image.
[0011] By adopting the above technical solution, the back-projection cone rays of two-dimensional image pixels are calculated, the set of voxels through which the rays pass is determined by the ray traversal algorithm, and the pixel feature vectors are scattered into each voxel according to the weight configuration, so as to preserve complete three-dimensional depth information in the voxel feature volume of the two-dimensional image.
[0012] As a preferred embodiment, the specific method for obtaining the point cloud voxel feature volume in S300 is as follows: S301 maps the point cloud in the global coordinate system to the voxel mesh V; S302, for each voxel v i Calculate multiple statistical features of its internal point cloud subsets; S303 combines multiple statistical features into a point cloud feature vector, and performs MLP encoding on the point cloud feature vector using a multilayer perceptron to obtain the pixel volume v. i Point cloud voxel features Where (L, W, H) are the voxel mesh sizes, and D pts Let v be the number of feature channels in the point cloud. For voxels that do not contain point clouds... i V pts (v i )=0.
[0013] Preferably, in S301, the point cloud in the global coordinate system is... Each point in a point cloud has its own voxel v. i index i, voxel v i The internal point cloud subset is P i ; In S302, each voxel v i Several statistical features include the centroid, covariance matrix, eigenvalues, and point density for each voxel v. i Calculate its internal point cloud subset The specific methods for obtaining multiple statistical characteristics are as follows: For the center of mass, we have: ; For the covariance matrix, we have: ; For eigenvalues, we have: ; For point density, we have: ; In the formula, Represents the set of real numbers, T is the extrinsic calibration matrix, and V cell The volume of a single voxel.
[0014] By adopting the above technical solution, it is possible to extract point cloud features in a unified pre-constructed three-dimensional voxel mesh space, statistically encode the point cloud, and thus generate point cloud voxel feature volumes.
[0015] As a preferred embodiment, the specific method for feature fusion of two-dimensional image voxel feature volume and point cloud voxel feature volume in S400 is as follows: S401, for voxel v i The two-dimensional image voxel feature volume and the point cloud voxel feature volume are concatenated in the feature channel dimension to obtain the concatenated feature volume; S402 uses a sparse three-dimensional convolutional neural network to perform multi-layer convolution calculations on the spliced feature volume; S403 uses pooling operations at different depths of the network to generate multi-scale features and capture defects at different scales. S404, Output fused voxel feature Where (L, W, H) are the voxel mesh sizes, and D fused This represents the number of feature channels for the fusion voxel.
[0016] By adopting the above technical solution, point cloud voxel feature volume and image voxel feature volume are concatenated through channels, and then feature interaction and fusion are performed through a sparse three-dimensional convolutional neural network to generate a fused voxel feature volume containing both geometric and texture information.
[0017] As a preferred option, the S500 specifically includes: S501, Max pooling is performed on the fused voxel feature to obtain the BEV feature map; S502 utilizes a differentiable projection layer to reproject the fused voxel features onto the camera image plane based on the projection matrix from voxels to camera pixels, obtains the projection function, and uses average pooling or aggregate pooling to obtain the perspective feature map. S503, Construct an alignment loss function to constrain the feature alignment of the BEV feature map and perspective feature map; S504 configures two-dimensional convolutional segmentation heads on the BEV feature map and perspective feature map respectively, calculates segmentation probability maps separately, and then sums them by weight to obtain the final segmentation probability map; S505 performs binarization thresholding on the final segmentation probability map, obtains the three-dimensional voxel mask where the final segmentation probability map calculation result exceeds the threshold, and converts the three-dimensional voxel mask into a defect point cloud through inverse mapping.
[0018] By employing the above technical solution, a dual-view feature map is obtained by generating a BEV feature map (bird's-eye view) through pooling and a perspective feature map through voxel projection. By constructing a feature alignment network, defect detection and segmentation are performed simultaneously under both BEV and perspective views, outputting a 3D defect mask. This approach comprehensively utilizes the advantages of both perspectives, improving the stability of defect detection.
[0019] As a preferred embodiment, in S600, the calculation of three-dimensional defects on the inner wall of the pipe includes clustering the defect point cloud, extracting the three-dimensional skeleton of the clustered defect point cloud, performing spline fitting on the skeleton, and calculating the defect length, defect width, and defect depth respectively.
[0020] By adopting the above technical solution, it is possible to output three-dimensional quantitative indicators of the length, width, and depth of defects based on the defect point cloud.
[0021] Secondly, this application discloses a pipeline inner wall defect detection system based on point cloud and image deep fusion. The technical solution adopted includes: a data acquisition module, a two-dimensional image feature upscaling module, a point cloud voxelization feature encoding module, a voxel feature fusion module, a defect detection module, and a defect calculation module, to implement the above-mentioned defect detection method.
[0022] Preferably, the data acquisition module includes a coordinate system definition submodule, which defines the coordinate system of the global coordinates according to the shape of the inner wall of the pipe.
[0023] Thirdly, this application discloses a computer program product, the technical solution of which includes a computer program or instructions, enabling the computer program or instructions to implement the steps in the above-mentioned method for detecting defects in the inner wall of a pipeline based on point cloud and image depth fusion.
[0024] In summary, this application includes at least one of the following beneficial technical effects: 1. This application preserves the geometric and topological information of the three-dimensional space of the inner wall of the tested pipeline by feature upscaling of two-dimensional image features and voxel-level fusion with three-dimensional point cloud features, thus avoiding the loss of depth information caused by traditional projection schemes and significantly improving the accuracy of defect identification under complex pipeline conditions.
[0025] 2. This application directly calculates the physical parameters of the length, width, and depth of defects based on the real three-dimensional point cloud coordinates of the inner wall of the pipe being tested. It can accurately quantify the size of internal defects in the pipe with an accuracy of ±5mm, which is far superior to the traditional ±15–30mm based on pixel estimation, and meets the needs of engineering applications for structural safety assessment.
[0026] 3. This application combines the advantages of LiDAR's insensitivity to light and camera's rich texture, which greatly improves the system's detection performance in complex environments such as water accumulation, stains, and uneven lighting, and has strong anti-interference ability and low false alarm rate.
[0027] 4. Based on the obtained fusion voxel feature, this application employs a multi-view collaborative mechanism to enhance the robustness of defect detection. The limitation of a single viewpoint is eliminated through a feature alignment mechanism using both BEV and perspective feature maps. The BEV viewpoint allows for global observation of defect distribution along the path, while the perspective viewpoint captures subtle texture details. The combination of these two approaches significantly improves the detection rate of minute cracks and hidden defects. Attached Figure Description
[0028] Figure 1 This is a flowchart illustrating a method for detecting defects in the inner wall of a pipe based on point cloud and image depth fusion, according to an embodiment of this application. Figure 2 This is a schematic diagram illustrating the principle of obtaining a two-dimensional image voxel feature body by performing feature dimensionality upscaling in S200 of this application embodiment; Figure 3 This is a schematic diagram of the deep fusion network structure of two-dimensional image voxel features and point cloud voxel features in S400 of this application embodiment; Figure 4 This is a schematic diagram of the network structure for aligning BEV and perspective feature map features from two perspectives and calculating three-dimensional voxel mask in S500 of this application embodiment. Figure 5 This is a schematic diagram of the geometric principle of three-dimensional physical quantization calculation of cracks in S600 of an embodiment of this application; Figure 6 This is a schematic diagram of the architecture of a pipeline inner wall defect detection system based on point cloud and image depth fusion according to an embodiment of this application. Detailed Implementation
[0029] This specific embodiment is merely an explanation of this application and is not intended to limit it. After reading this specification, those skilled in the art can make modifications to this embodiment without contributing any inventive step, but such modifications are protected by patent law as long as they are within the scope of this application.
[0030] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. It should be noted that in the optional embodiments of this application, the object information and other related data involved require the permission or consent of the object when the embodiments of this application are applied to specific products or technologies, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. That is to say, if the embodiments of this application involve data related to the object, it needs to be obtained with the authorization and consent of the object, the authorization and consent of the relevant departments, and in compliance with the relevant laws, regulations, and standards of the country and region. If personal information is involved in the embodiments, the acquisition of all personal information requires the consent of the individual. If sensitive information is involved, the separate consent of the information subject is required, and the embodiments also need to be implemented with the authorization and consent of the object.
[0031] The embodiments of this application will now be described in further detail with reference to the accompanying drawings. Example
[0032] Taking a municipal drainage circular pipe as an example, this application's method for defect detection and defect quantification is explained. Please refer to... Figure 1 Specifically, it includes the following steps.
[0033] S100, acquire the two-dimensional image sequence and point cloud sequence of the inner wall of the pipe respectively.
[0034] Specifically, a two-dimensional image sequence is acquired using an industrial imaging camera, and a point cloud sequence is acquired using a LiDAR (Light Detection and Ranging) radar. Hardware triggers are used for time synchronization, ensuring that the camera and LiDAR acquisition times are consistent, with a time error controlled within 10ms. The camera and LiDAR together form a camera-LiDAR imaging platform.
[0035] Before conducting formal defect detection, sensor calibration and coordinate system unification are required. Multiple frames of point clouds and images, including a checkerboard calibration board, are acquired within the calibrated pipeline scene for subsequent extrinsic parameter matrix calibration.
[0036] The specific steps to obtain the extrinsic calibration matrix of the camera-LiDAR imaging platform are as follows: S101 uses a checkerboard calibration board to simultaneously acquire LiDAR point clouds and camera images (no fewer than 15 sets).
[0037] S102 detects chessboard corner points in a point cloud, approximately 50–100 points; in an image, it detects chessboard corner points, typically using an 8×8 grid with a total of 64 points.
[0038] S103, Establish the correspondence between the 3D corner points detected in the point cloud and the 2D corner points detected in the image.
[0039] S104, Solving the extrinsic parameter matrix T using the RANSAC algorithm. ext To minimize the projection error, we have: ; Where π represents the camera projection model. For the i-th 3D corner point, It corresponds to the 2D image points.
[0040] S105 verifies the accuracy of external parameter calibration. The reprojection error should be less than 1 pixel, corresponding to a spatial accuracy of approximately 5 mm.
[0041] Then, trajectory calculation is performed on the camera-LiDAR imaging platform. The specific steps are as follows: S106 uses LiDAR's Iterative Closest Point Algorithm (ICP algorithm) to calculate the relative pose transformation between two adjacent point clouds.
[0042] S107, fusing path length data and inertial measurement unit (IMU) data from the camera-LiDAR imaging platform, uses an extended Kalman filter to calculate the time-by-time pose trajectory of the camera-LiDAR imaging platform, resulting in: ; ; In the formula, Let u be the state estimate at time k. k To control the input, Z k For the observed value, w k v k These are process noise and observation noise, respectively.
[0043] S108, summing the accumulated state estimates yields the global trajectory sequence, i.e., the pose transformation matrices T0, T1, T2, ... T for each frame. N .
[0044] Then, coordinate unification is performed, with the following steps: S109, In this embodiment, the observed object is a circular pipe. A global cylindrical coordinate system for the pipe is defined, with the pipe axis as the z-axis and the pipe inlet as the origin. The cylindrical coordinate system is (s, θ, r), where s is the axial distance, θ is the circumferential angle, and r is the radial distance.
[0045] S110, transform the point cloud and 2D image data of all frames to a unified global pipeline cylindrical coordinate system. The pose transformation matrix of the i-th frame is T. i .
[0046] Thus, the coordinate system of the two-dimensional image sequence and the point cloud sequence is realized.
[0047] S200: Feature extraction is performed on the two-dimensional image sequence to obtain two-dimensional image feature vectors. The two-dimensional image feature vectors are then up-dimensionalized, and back-projected onto a pre-constructed three-dimensional voxel grid space to obtain the two-dimensional image voxel feature volume V. img .
[0048] In the embodiments of this application, a pre-trained two-dimensional convolutional neural network (such as ResNet-50) is used to extract features from the image sequence to obtain multi-scale feature maps. It should be noted that those skilled in the art will understand that other feature engineering and algorithm models can also be used to extract features from two-dimensional image sequences.
[0049] In this application, a two-dimensional feature vector upscaling technique is adopted, which involves back-projecting the two-dimensional image feature vectors into a three-dimensional voxel space, instead of the traditional method of projecting point clouds onto a planar image. This reverse operation method avoids the compression loss of three-dimensional information.
[0050] To achieve the above objectives, a three-dimensional voxel mesh space needs to be pre-constructed. Based on the global pipeline cylindrical coordinate system (s, θ, r), in this embodiment, the axial resolution Δs = 2 mm, totaling 5000 units, corresponding to a 10 m pipeline; the circumferential resolution: Δθ = 1°, totaling 360 units; and the radial resolution: Δr = 5 mm, totaling 40 units, corresponding to a pipeline with a radius of 200 mm; voxel mesh dimensions... Where C=128 is the number of feature channels in a two-dimensional image.
[0051] After completing the construction of the three-dimensional voxel mesh space, please refer to Figure 2 The two-dimensional image voxel feature volume V is obtained. img The specific method includes the following steps.
[0052] S201, For any pixel (u,v) on a two-dimensional image, calculate its back-projected view cone ray in the three-dimensional voxel grid space.
[0053] Let the homogeneous coordinates of the pixel be... Its normalized coordinates are, The back-projected cone ray is obtained through the cone ray equation. ,have: ; in, To obtain the position of the camera optical center in the global coordinate system of a 2D image, Let K be the ray direction, K be the camera intrinsic calibration matrix, T be the extrinsic calibration matrix, and t∈[0,∞) be the ray parameter, reflecting the distance of the pixel from the origin (i.e., the camera optical center).
[0054] S202, in the three-dimensional voxel mesh space V, the back-projected view cone ray is calculated using the ray traversal algorithm. The set of all voxels passed through The ray traversal algorithm can be either the Amanatides-Woo algorithm or the 3D DDA algorithm.
[0055] S203, extract the feature vector corresponding to pixel (u, v) According to the weighted rule, the rays are scattered into the voxels through which they pass, where Let C represent the set of real numbers, and C be the number of feature channels in a two-dimensional image.
[0056] More specifically, the weighting rules can be calculated using different methods.
[0057] When there is no prior knowledge of point cloud depth, a uniformly distributed weight can be used, as follows: ; w i is the weight corresponding to the i-th pixel, and n is the total number of voxels passing through each voxel.
[0058] If there is a priori point cloud depth d pc We can use the depth prior weight calculation formula, which is: ;in, d represents the distance from the voxel center to the camera's optical signal; pc The distance from the nearest point in the laser point cloud along the direction of the view cone ray to the optical center is denoted by α. α is an adjustment parameter, set based on experience with the point cloud and camera operating environment, typically α ∈ [0.5, 2.0]. This weighting formula ensures a strong correlation between the feature vector projected onto the voxel and the point cloud depth information, improving the geometric consistency of the fused features. This weighting calculation formula is used in the embodiments of this application.
[0059] Similarly, the distance from the pixel to the optical center can be used as a weight for distance attenuation calculation, resulting in: ; Among them, t i Let β be the ray parameter of the i-th pixel, and β be the attenuation coefficient.
[0060] S204, after completing the weight calculation, for the same voxel v projected by multiple rays... iWeighted average pooling is used to obtain the two-dimensional image voxel feature volume V. img ,have: ; Wherein, S(v) i ) is projected onto voxel v i The set of all pixels, w k f is the weight corresponding to the k-th pixel. k The feature vector corresponding to the k-th pixel is the two-dimensional image voxel feature volume. .
[0061] Compared with traditional two-dimensional image or point cloud two-dimensional projection schemes, the above scheme preserves complete three-dimensional spatial information, avoids information loss, and can utilize the prior information of point cloud depth for weight calculation, thereby improving the accuracy of feature allocation. It can also handle complex three-dimensional geometric relationships, such as pipe deformation and depressions.
[0062] S300, the point cloud sequence is mapped to the three-dimensional voxel grid space, feature extraction is performed to obtain point cloud feature vectors, and then the point cloud voxel feature volume V is obtained through the relationship between point cloud and voxels. pts The specific steps include the following.
[0063] S301, the point cloud in the global coordinate system Mapped to the voxel grid space V. Each point p i voxel index (i s i θ i r ) by its cylindrical coordinates (s i ,θ i ,r i ) Determined, there are: i s =s i / Δs;i θ =θ i / Δθ;i r =r i / Δr.
[0064] S302, for each voxel v i Calculate multiple statistical features of its internal point cloud subsets.
[0065] Each voxel v i Several statistical features include the centroid, covariance matrix, eigenvalues, and point density for each voxel v. i Calculate its internal point cloud subset The specific methods for obtaining multiple statistical characteristics are as follows.
[0066] For the center of mass, we have: ; For the covariance matrix, we have: ; For eigenvalues, we have: ; For point density, we have: ; In the formula, Represents the set of real numbers, T is the extrinsic calibration matrix, and V cell The volume of a single voxel.
[0067] S303 combines multiple statistical features into a point cloud feature vector, and then uses a multilayer perceptron (MLP) to encode the point cloud feature vector: ; ; Where W1 and W2 are weight matrices, and b1 and b2 are bias vectors. The ReLU activation function is used, and Dpts=128 represents the number of feature channels in the point cloud.
[0068] Thus, the pixel volume v is obtained. i Point cloud voxel features For voxels that do not contain point clouds i V pts (v i )=0.
[0069] S400, within the three-dimensional voxel grid space, perform the two-dimensional image voxel feature volume V corresponding to the three-dimensional voxels. img and point cloud voxel feature V pts Feature fusion yields fused voxel feature body V. fused The specific steps include the following.
[0070] S401, for voxel v i V2D image voxel feature volume V img V with point cloud voxel feature pts By concatenating the features along the feature channel dimension, we obtain the concatenated feature body, which includes: ; V concat For each stitched feature, the voxel features of each stitched feature contain both the geometric information of the point cloud and the texture information of the two-dimensional image.
[0071] S402 uses a sparse 3D convolutional neural network (Sparse 3D CNN) to stitch together the feature volume V. concat Perform multi-layer convolution calculations. Sparse convolution only calculates non-empty voxels, significantly reducing computational cost.
[0072] For more details, please see Figure 3 The network structure of a sparse 3D convolutional neural network includes multiple sparse 3D convolutional layers with a kernel size of 1 and a stride of 1. Each convolutional layer is followed by a ReLU activation function and batch normalization, and multiple residual blocks are added to accelerate convergence.
[0073] Input layer: Receives tensors concatenated from the channels.
[0074] Sparse coding layer: Converts sparse voxel features into COO (coordinate format) or CSR (compressed sparse line) format, and only performs calculations on non-empty voxels.
[0075] Convolutional layer group: contains 4–8 sparse 3D convolutional layers, all with a kernel size of 3×3×3, a stride of 1, and padding of 1.
[0076] Non-linear activation: A ReLU activation function is applied after each convolutional layer.
[0077] Batch Normalization: Batch normalization is performed before the activation function.
[0078] Residual connections: Residual blocks are designed to speed up convergence, and the residual blocks span 2–3 convolutional layers.
[0079] Output layer: Outputs fused voxel features.
[0080] S403 uses pooling operations at different depths in the network to generate multi-scale features and capture defects at different scales.
[0081] S404, Output fused voxel feature D fused =256 represents the number of feature channels for fusion voxels.
[0082] S500, please refer to Figure 4 Based on the fusion voxel feature V fused The BEV feature map and perspective feature map of the inner wall of the pipe are obtained respectively. Based on the BEV feature map and perspective feature map, a three-dimensional voxel mask for identifying three-dimensional defects in the inner wall of the pipe is obtained, and the defect point cloud corresponding to the three-dimensional voxel mask is obtained.
[0083] It should be noted that, although the fusion voxel feature V fusedWhile it incorporates geometric information from point clouds and texture information from two-dimensional images, relying solely on a single viewpoint makes it susceptible to interference from that single dimension, leading to a high false alarm rate. For example, uneven lighting, water accumulation, air bubbles, and sediment inside pipes can cause high false alarm rates for texture-based detection. Geometric distortion, especially in pipes with irregular cross-sections or severe deformations, distorts the shape of gaps, making it difficult to accurately capture depth changes along the pipe wall normal. This application utilizes a feature alignment network combining BEV and perspective views, leveraging the advantages of both perspectives to improve the stability of defect detection. Specifically, it includes the following steps.
[0084] S501, Max pooling is performed on the fused voxel features along the radial direction r of the pipeline: ; Two-dimensional BEV feature map obtained BEV feature maps are similar to unfolded views (bird's-eye view) of the inner wall of a pipe, used for macroscopic observation of the distribution of defects along the axial and circumferential directions of the pipe.
[0085] S502 utilizes a differentiable projection layer, based on the projection matrix from voxels to camera pixels. Reproject the fused voxel features back onto the camera image plane: ; The projection function is:
[0086] The aggregation function can be either average pooling or max pooling, to obtain the perspective feature map M. persp The perspective feature map preserves texture details and lighting information from the original camera viewpoint.
[0087] S503, Construct an alignment loss function to impose feature alignment constraints on the BEV feature map and perspective feature map: ; Warp is a geometric transformation function that constrains the consistency of features in the defect region from both perspectives. This represents the mapping from BEV coordinates to perspective coordinates. Alignment parameters are optimized using gradient backpropagation.
[0088] S504 configures two-dimensional convolutional segmentation heads on both the BEV feature map and the perspective feature map: ; ; The final segmentation probability map S is obtained through weighted fusion. final: ; Where α is the balance coefficient, which is usually set to 0.5.
[0089] S505 performs binarization thresholding on the final segmentation probability map to generate a 3D voxel mask representing the defect code. The threshold is typically set to 0.5. ; By inverse mapping, the 3D voxel mask is converted into a defect point cloud P. defect .
[0090] In the above scheme, the BEV perspective provides a global view, making it easy to detect the distribution of defects along the pipe axis and circumference. The perspective view preserves texture details, making it easy to detect subtle defects and complex textures. Feature alignment constraints eliminate the limitations of a single perspective and improve detection stability.
[0091] S600, please refer to Figure 5 The calculation of three-dimensional defects on the inner wall of a pipeline is based on a defect point cloud. The specific steps include the following.
[0092] S601, Crack length calculation.
[0093] Density-based clustering algorithm (DBSCAN) is used to cluster the defect point set P. defect Clustering was performed to separate multiple independent crack point sets C. k : ; in, is the neighborhood radius, and min_pts is the minimum number of points required to form a cluster.
[0094] For each crack point set C k Perform 3D skeletonization processing and extract the skeleton point set Skel representing the direction of the crack center. k Skeletonization methods employ a 3D generalization of the Voronoi diagram or minimize the distance to the boundary. The skeleton point set is typically 5-10% the size of the original point set.
[0095] For skeleton point set Perform B-spline curve fitting (centerline fitting) to obtain the parameterized curve equation: ; Where, d i N is the control point. i,k (u) represents the k-th order B-spline basis function, typically a 3rd order (cubic) B-spline. The fitting process uses the least squares method to minimize the sum of squared residuals. ; Where uj =j / m is the parameter value.
[0096] The geodesic length is calculated by numerical integration of the parameterized curve. Simpson's rule or the trapezoidal rule is used. ; Where Δu = 1 / M, and M is usually 1000 to ensure integration accuracy. The final output is the crack length L. k , unit mm.
[0097] S602, Crack width calculation.
[0098] Based on skeleton spline curves (centerline) The cross-section position is set at a fixed interval of ΔL = 10mm: ; Where L is the total length of the crack.
[0099] At each cross-section location Calculate the tangent vector, normal vector, and binormal vector of the centerline to construct a local Frenet coordinate system. : ; ; ; The normal section includes the center point. And perpendicular to the tangent vector The plane.
[0100] Crack point set C k Projecting points within 5mm of the mid-distance section location onto the plane of that section: ; Extract the set of edge points of the crack within the cross section using convex hull or boundary detection algorithms. .
[0101] For the j-th section, calculate the maximum Euclidean distance between the edge points as the local width at that section: ; The width of all cross sections is counted to obtain the width distribution. Take the maximum and average values: ; ; Output maximum crack width W max and average width W avg , unit mm.
[0102] S603, Defect depth calculation.
[0103] Select a healthy pipe wall point cloud P within a radius R = 50 mm around the defect area. healthy : ; Fit the local reference surface S using the least squares method or RANSAC algorithm. base .
[0104] In this embodiment, for a circular pipe, the reference surface is a cylindrical surface, and the equation is: ; in, Let r0 be the tube axis projection and r0 be the fitted radius. Optimization is performed using least squares: ; For the defect point set P defect Each point p in i Calculate the shortest distance (normal distance) from the cylindrical surface to the fitted reference plane. For a cylindrical surface, this distance is: ; For general curved surfaces, polynomial surface fitting is used: z = f(x, y), where f is a second- or third-order polynomial. The nearest point is solved using projection iteration or the Gauss-Newton method, and the normal distance is calculated.
[0105] The statistical measure of normal distance is used, with the maximum value taken as the maximum depth and the average value calculated as the average depth. ; ; Output the maximum depth of the defect D max and average depth D avg , unit mm.
[0106] This completes the calculation of the physical quantities of three-dimensional defects in the inner wall of the pipeline. It should be noted that the above defect calculation method does not constitute a limitation on the technical solution of this application.
[0107] S700 outputs the physical quantities of three-dimensional defects in the inner wall of the pipe, including detailed information on all defects and a visualized three-dimensional defect model.
[0108] Example 2 For rectangular pipes or box culverts, simply adjust the cylindrical coordinate system to a Cartesian coordinate system; the above algorithm framework remains exactly the same.
[0109] The voxel mesh in a Cartesian coordinate system is defined as follows: .
[0110] The BEV feature map in Cartesian coordinate system is an unfolded top view.
[0111] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0112] Using the above method, the calculated three-dimensional physical quantities of the pipe inner wall cracks have the following accuracy: the relative error for length is controlled within ±5%, and the absolute error within ±5mm; the absolute error for width is controlled within ±3mm; and the absolute error for depth is controlled within ±5mm. These accuracy parameters were verified through comparative calibration with a high-precision three-dimensional laser scanner (accuracy ±1 mm).
[0113] Example 3 Please see Figure 6 This application provides a pipeline inner wall defect detection system based on point cloud and image deep fusion, including a data acquisition module 1, a two-dimensional image feature upscaling module 2, a point cloud voxelization feature encoding module 3, a voxel feature fusion module 4, a defect detection module 5, and a defect calculation output module 6, to implement the aforementioned defect detection method. The data acquisition module 1 includes a coordinate system definition submodule, which defines the global coordinate system based on the shape of the pipeline inner wall.
[0114] More specifically, in the embodiments of this application, the hardware of the data acquisition module 1 is configured as follows: The detection mobile platform is a wheeled crawler adapted to pipes with a diameter of 400–600 mm; the lidar is a 16-line rotating LiDAR with a scanning frequency of 10 Hz and a single scan of 100,000 points; the imaging camera is an industrial monochrome camera with a resolution of 1280×960 and a frame rate of 30 fps.
[0115] The computer device used for data computation in this application is an embedded industrial PC equipped with a GPU (NVIDIA RTX 3070) (or equivalent or higher) and 16GB of memory. The communication interface uses Gigabit Ethernet for real-time transmission of raw data and detection results. The computer device includes a non-volatile storage medium for executing a computer program product. The computer program product contains an executable Python script, pre-trained deep learning model weight files (PyTorch or TensorFlow format), and configuration files. When executed, the computer program product can implement the steps in the aforementioned method for detecting defects in the inner wall of pipes based on point cloud and image deep fusion.
[0116] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the pipeline inner wall defect detection system based on point cloud and image deep fusion described above can be referred to the corresponding process in the aforementioned method embodiments, and will not be repeated here.
[0117] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0118] This application achieves deep fusion of 2D images and 3D point cloud data by acquiring a large amount of high-frequency laser point cloud data (typically 10 Hz × 100,000 points / frame) and multi-view high-resolution images. By fully utilizing complementary information, the accuracy of pipeline defect detection is significantly improved. This application solves the problem of 3D information loss caused by traditional 2D projection fusion, achieving 3D physical quantification of pipeline defects with millimeter-level accuracy, directly supporting structural safety assessment and precise maintenance decisions. The advantages compared with traditional detection methods are shown in the table below.
[0119]
[0120] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for detecting defects in the inner wall of a pipe based on point cloud and image depth fusion, characterized in that, Includes the following steps: S100, acquire the two-dimensional image sequence and point cloud sequence of the inner wall of the pipe respectively; S200: Feature extraction is performed on the two-dimensional image sequence to obtain two-dimensional image feature vectors. The two-dimensional image feature vectors are then up-dimensionalized, and back-projected onto a pre-constructed three-dimensional voxel grid space to obtain the two-dimensional image voxel feature volume V. img ; S300, the point cloud sequence is mapped to the three-dimensional voxel grid space, feature extraction is performed to obtain point cloud feature vectors, and then the point cloud voxel feature volume V is obtained through the relationship between the point cloud and the three-dimensional voxels. pts ; S400, within the three-dimensional voxel grid space, perform the two-dimensional image voxel feature volume V corresponding to the three-dimensional voxels. img and point cloud voxel feature V pts Feature fusion yields fused voxel feature body V. fused ; S500, based on fusion voxel feature V fused The BEV feature map and perspective feature map of the inner wall of the pipe are obtained respectively. Based on the BEV feature map and perspective feature map, a three-dimensional voxel mask for identifying three-dimensional defects in the inner wall of the pipe is obtained, and the defect point cloud corresponding to the three-dimensional voxel mask is obtained. S600 calculates three-dimensional defects on the inner wall of a pipeline based on defect point clouds.
2. The method for detecting defects in the inner wall of a pipe based on point cloud and image depth fusion according to claim 1, characterized in that, In S200, the specific method for obtaining the two-dimensional image voxel feature volume is as follows: S201, For any pixel (u,v) on a two-dimensional image, calculate its back-projection frustum ray in the three-dimensional voxel grid space; Let the homogeneous coordinates of the pixel be... Its normalized coordinates are, The back-projected cone ray is obtained through the cone ray equation. ,have: ; in, To obtain the position of the camera optical center in the global coordinate system of a 2D image, Let K be the ray direction, K be the camera intrinsic calibration matrix, T be the extrinsic calibration matrix, and t∈[0,∞) be the ray parameters; S202, in the three-dimensional voxel mesh space V, the back-projected view cone ray is calculated using the ray traversal algorithm. The set of all voxels passed through ; S203, extract the feature vector corresponding to pixel (u, v) According to the weighted rule, the rays are scattered into the voxels through which they pass, where C represents the set of real numbers, and C is the number of feature channels in a two-dimensional image. S204, for the same voxel v projected by multiple rays i Weighted average pooling is used to obtain the two-dimensional image voxel feature volume V. img ,have: ; Wherein, S(v) i ) is projected onto voxel v i The set of all pixels, w k f is the weight corresponding to the k-th pixel. k The feature vector corresponding to the k-th pixel is the image voxel feature volume. , where (L,W,H) are the voxel grid size and C is the number of feature channels in the two-dimensional image.
3. The method for detecting defects in the inner wall of a pipe based on point cloud and image depth fusion according to claim 1, characterized in that, In S300, the specific method for obtaining point cloud voxel feature volumes is as follows: S301 maps the point cloud in the global coordinate system to the voxel mesh V; S302, for each voxel v i Calculate multiple statistical features of its internal point cloud subsets; S303 combines multiple statistical features into a point cloud feature vector, and performs MLP encoding on the point cloud feature vector using a multilayer perceptron to obtain the pixel volume v. i Point cloud voxel features Where (L, W, H) are the voxel mesh sizes, and D pts Let v be the number of feature channels in the point cloud. For voxels that do not contain point clouds... i V pts (v i )=0.
4. The method for detecting defects in the inner wall of a pipe based on point cloud and image depth fusion according to claim 3, characterized in that, In S301, the point cloud in the global coordinate system is Each point in a point cloud has its own voxel v. i index i, voxel v i The internal point cloud subset is P i ; In S302, each voxel v i Several statistical features include the centroid, covariance matrix, eigenvalues, and point density for each voxel v. i Calculate the subset of its internal point cloud The specific methods for obtaining multiple statistical characteristics are as follows: For the center of mass, we have: ; For the covariance matrix, we have: ; For eigenvalues, we have: ; For point density, we have: ; In the formula, Represents the set of real numbers, T is the extrinsic calibration matrix, and V cell The volume of a single voxel.
5. The method for detecting defects in the inner wall of a pipe based on point cloud and image depth fusion according to claim 1, characterized in that, In S400, the specific method for feature fusion of two-dimensional image voxel feature volumes and point cloud voxel feature volumes is as follows: S401, for voxel v i The two-dimensional image voxel feature volume and the point cloud voxel feature volume are concatenated in the feature channel dimension to obtain the concatenated feature volume; S402 uses a sparse three-dimensional convolutional neural network to perform multi-layer convolution calculations on the spliced feature volume; S403 uses pooling operations at different depths of the network to generate multi-scale features and capture defects at different scales. S404, Output fused voxel feature Where (L, W, H) are the voxel mesh sizes, and D fused This represents the number of feature channels for the fusion voxel.
6. The method for detecting defects in the inner wall of a pipe based on point cloud and image depth fusion according to claim 1, characterized in that, The S500 specifically includes: S501, Max pooling is performed on the fused voxel feature to obtain the BEV feature map; S502 utilizes a differentiable projection layer to reproject the fused voxel features onto the camera image plane based on the projection matrix from voxels to camera pixels, obtains the projection function, and uses average pooling or aggregate pooling to obtain the perspective feature map. S503, Construct an alignment loss function to constrain the feature alignment of the BEV feature map and perspective feature map; S504 configures two-dimensional convolutional segmentation heads on the BEV feature map and perspective feature map respectively, calculates segmentation probability maps separately, and then sums them by weight to obtain the final segmentation probability map; S505 performs binarization thresholding on the final segmentation probability map, obtains the three-dimensional voxel mask where the final segmentation probability map calculation result exceeds the threshold, and converts the three-dimensional voxel mask into a defect point cloud through inverse mapping.
7. The method for detecting defects in the inner wall of a pipe based on point cloud and image depth fusion according to claim 6, characterized in that, In S600, the calculation of three-dimensional defects on the inner wall of the pipeline includes clustering the defect point cloud, extracting the three-dimensional skeleton of the clustered defect point cloud, performing spline fitting on the skeleton, and then calculating the defect length, defect width and defect depth respectively.
8. A pipeline inner wall defect detection system based on point cloud and image deep fusion, characterized in that, It includes a data acquisition module, a two-dimensional image feature upscaling module, a point cloud voxelization feature encoding module, a voxel feature fusion module, a defect detection module, and a defect calculation module, to implement the defect detection method according to any one of claims 1 to 7.
9. A pipeline inner wall defect detection system based on point cloud and image depth fusion according to claim 8, characterized in that, The data acquisition module includes a coordinate system definition submodule, which defines the coordinate system of the global coordinates based on the shape of the inner wall of the pipe.
10. A computer program product, characterized in that, The computer program product includes a computer program or instructions that enable the computer program or instructions to perform the steps in the pipe inner wall defect detection method based on point cloud and image depth fusion as described in any one of claims 1 to 7.