Unstructured road pit extraction method based on unmanned aerial vehicle point cloud data and optical image fusion
Through drone remote sensing technology and image point cloud fusion algorithm, the high-precision and high-efficiency problems of unstructured road pothole detection are solved, and the accurate extraction of road boundaries and potholes is achieved, which improves the reliability and efficiency of detection.
Patent Information
- Application Number
- CN202510440584.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-11
AI Technical Summary
The existing unstructured road pothole detection technology mainly relies on vehicle-mounted sensors, which have problems such as poor road conditions adaptability, limited sensor field of view, and data quality is susceptible to vehicle speed and load interference, making it difficult to achieve high-precision and high-efficiency pothole extraction.
The road data is obtained by using drone remote sensing technology, combining image semantic segmentation and point cloud scattered contour algorithm to extract road boundaries, and high-precision and high-efficiency extraction of potholes is achieved through image two-dimensional and point cloud three-dimensional fusion algorithms.
It realizes high-precision and efficient extraction of potholes in unstructured roads, reduces false extraction of non-road areas, and improves the reliability and efficiency of detection.
Smart Images

Figure CN120298899A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of unstructured road pothole extraction, and particularly to a method for extracting unstructured road potholes based on the fusion of UAV point cloud data and optical images. Background Art
[0002] Unstructured roads refer to roads without lane lines and clear boundaries. The shapes of the roads are diverse, the road surface grades are relatively low, and the surrounding environments of the roads are complex. It is very difficult to distinguish between road areas and non-road areas. In recent years, autonomous driving technologies have been rapidly implemented in fields such as logistics distribution, sanitation operations, and intelligent mines. The pothole detection of unstructured roads has become a core issue restricting the safety of in-vehicle autonomous driving technologies. Concave and uneven potholes are one of the most serious diseases of unstructured road surfaces, causing violent vibrations of vehicles, damaging components, and also causing the vehicle body to tilt and bump, thus triggering accidents. Therefore, studying accurate detection methods for unstructured road potholes is of great significance for the safe production and transportation of vehicles.
[0003] In the existing technologies, the research on structured road pothole extraction has become mature, mainly relying on vehicle-mounted sensors (such as vibration sensors, vehicle-mounted cameras, and vehicle-mounted radars), and realizing detection by analyzing vehicle driving trajectories, two-dimensional image textures, or local point cloud features. It has disadvantages such as poor road condition adaptability, limited sensor vision, and the data quality being easily interfered by vehicle speed and load. Its detection algorithms are mainly divided into four categories: the first category is to extract two-dimensional texture features of potholes based on image data; the second category is to obtain three-dimensional information of potholes by using laser point cloud data; the third category is to fuse data from two sensors to extract pothole features; the fourth category is to obtain relevant data such as vehicle trajectories through vibration sensors.
[0004] The detection of potholes on unstructured roads faces many challenges. It started relatively late and inherited the method of obtaining data through vehicle-mounted sensors. There are mainly two existing detection algorithms: one is the deep learning method based on images, and the other is the local feature method based on point clouds. Among them, the deep learning method based on images mainly utilizes the difference between road potholes and the background to detect targets by extracting pothole defect features. Hu et al. proposed to extract pothole defect feature information from the collected image data based on road texture features and verify it through damaged road defects, which can detect pothole defects in specific scenarios. Ruan Shunling et al. proposed an FCOS per-pixel regression model to detect potholes in mining areas, realizing the recognition of negative obstacles in front of lightweight and fast unmanned vehicles. However, these methods rely on the low-altitude shooting perspective of vehicle-mounted cameras and are easily affected by external factors such as the shape, size, and background of pothole defects, resulting in low detection accuracy. The drone platform can obtain global road images from a high-altitude perspective and effectively reduce background interference by combining multi-angle shooting. The method based on point cloud geometric features mainly utilizes the local features of point cloud data to detect targets by extracting features such as the jump in the front and rear edge distances and the height sink of potholes. Li Jiawei proposed a ground segmentation model based on fan surface fitting and a positive obstacle detection model based on density clustering, which improved the detection accuracy of potholes on unstructured roads in open-pit mines through methods such as piecewise polynomial fitting and adaptive clustering radius. However, the road conditions of unstructured roads are complex, and the point cloud density of vehicle-mounted lidar drops sharply when the longitudinal / transverse slope of the road is large, resulting in poor feature extraction effects and easy omission and misdetection of some potholes. The drone-mounted lidar can vertically scan the road surface and ensure data uniformity through strip-shaped point cloud acquisition, significantly improving the feature stability under complex terrains.
[0005] In summary, the existing research mainly detects potholes on unstructured roads based on single data, and the application scenarios are all concentrated in the vehicle-mounted environment. The extraction of potholes on unstructured roads with a single data source fails to utilize the complementary characteristics of point clouds and images, restricting the improvement of detection effects and making it difficult to meet the high-precision extraction requirements of potholes under complex road conditions at the same time. At the same time, vehicle-mounted technology is limited by spatial coverage ability and sensor configuration, with high data acquisition costs and low efficiency, making it difficult to meet the detection requirements of unstructured roads. In addition, the boundary of unstructured roads is blurred, making it difficult to accurately distinguish between road surfaces and non-road areas, easily causing mis-extraction of potholes in non-road areas and affecting the accuracy of detection. Therefore, this invention uses drone remote sensing to obtain road data in the target area and proposes a method for extracting potholes on unstructured roads based on the fusion of drone point cloud data and optical images, realizing the high-precision and high-efficiency extraction of potholes on unstructured roads. Summary of the Invention
[0006] Aiming at the above technical problems, the object of the present invention is to overcome the deficiencies of the prior art and propose an unstructured road pothole extraction method based on the fusion of UAV point cloud data and optical images. The road data of the target area is obtained by UAV remote sensing. From the perspective of multi-source data fusion processing, image semantic segmentation (DeepLab Version3Plus, DeepLabv3+) and point cloud scatter contour algorithm (Alpha-Shapes) are used to extract the road range respectively. After weighted average fusion, the road boundary range is determined. Combining the DeepLabv3+ model and the Canny algorithm for rough extraction of two-dimensional road potholes, using the tensor voting algorithm to extract the geometric features of the point cloud for rough extraction of three-dimensional potholes, and using the comprehensive hierarchical fusion algorithm to fuse the pothole extraction results to achieve accurate pothole extraction. The specific process includes: experimental data collection and preprocessing, road boundary fusion extraction, two-dimensional pothole rough extraction of images, three-dimensional pothole rough extraction of point clouds, hierarchical fusion of rough extraction potholes, accuracy verification and other steps. The advantages of this method are: quickly obtaining the road data of the target area through UAV remote sensing, making full use of the texture and semantic information of the image and the three-dimensional geometric structure of the point cloud, and realizing the high-precision and high-efficiency extraction of unstructured road potholes through the designed multi-layer fusion algorithm.
[0007] To achieve the above functions, the present invention provides an unstructured road pothole extraction method based on the fusion of UAV point cloud data and optical images, including the following steps:
[0008] S1: Experimental data collection and preprocessing: Obtain complete and high-quality optical images and lidar point cloud data of the unstructured road in the target area through UAV oblique photography and airborne LiDAR technology. The preprocessing steps are as follows:
[0009] (1) Image preprocessing mainly includes image cropping, pixel de-mean and dataset construction; crop the size of the image to adapt to the computer processing performance, and pixel de-mean standardizes the data features, focusing the image on texture features, and then construct a road boundary and road pothole dataset for model construction;
[0010] (2) Point cloud preprocessing mainly includes non-ground point filtering, coordinate transformation and data correspondence; non-ground point filtering focuses the point cloud on the road surface information, effectively reducing noise interference, coordinate transformation ensures the correct alignment of the point cloud and image data, and data correspondence integrates the point cloud and image data into a unified road surface model;
[0011] S2: Road boundary fusion extraction: Based on the processed point cloud and image data, use the DeepLabv3+ model to roughly extract the road boundary of the image, use the Alpha-Shapes algorithm to roughly extract the road boundary of the point cloud, perform weighted averaging on the results of both to determine the final road boundary range, and finally perform buffer expansion on the road boundary range to ensure the integrity of the road range;
[0012] S3: Rough extraction of potholes in image semantic segmentation: Crop the road range in S2 from the image data to obtain the image road range, extract the semantic label potholes through the trained DeepLabv3+ model, and use the edge detection (Canny) algorithm to enhance the accuracy and continuity of the pothole edges to obtain the candidate potholes in the image plane coordinate system and their two-dimensional pothole contours;
[0013] S4: Rough extraction of potholes by point cloud tensor voting: Crop the road range in S2 from the point cloud data to obtain the point cloud road range, and generate the corresponding road intensity map. By analyzing the tensor voting results of the point cloud, identify the regions with significant geometric changes and determine whether the potholes in the candidate potholes are regarded as potholes;
[0014] S5: Hierarchical fusion of roughly extracted potholes: Based on the pothole results roughly extracted from the image and point cloud, first screen the high-confidence regions based on the majority decision, determine the regions where both are potholes as pothole regions, and mark the high-confidence regions of single data as candidate regions; then perform morphological dilation on the marked final pothole regions. If the candidate regions of single data around are morphologically continuous and geometrically similar to the dilated main pothole regions, merge them; for the candidate regions of single data that are not around, use the MRF theory to construct a probability graph model, dynamically adjust the MRF parameters according to the image entropy value and point cloud roughness to obtain the Markov potential function, and then use the ICM algorithm to solve the potential function to obtain the distribution result with the maximum probability and determine the road pothole regions;
[0015] S6: Accuracy verification: Compare with the true road pothole segmentation results of the division, calculate PA (Pixel Accuracy), MPA (Mean Pixel Accuracy of Classes), and MIoU (Mean Intersection over Union) to evaluate the accuracy of road pothole extraction.
[0016] Furthermore, the data acquisition and preprocessing in step S1 mainly include the following steps:
[0017] (1) Data acquisition
[0018] Use the drone oblique photography and airborne LiDAR technology to traverse the unstructured road area and collect complete and high-precision optical images and lidar point cloud data;
[0019] (2) Image preprocessing
[0020] Image preprocessing mainly includes image cropping and pixel mean subtraction; the size of the cropped image is adjusted to adapt to the processing performance of the computer, and pixel mean subtraction standardizes the data features, focusing the attention of the image on texture features; the specific processing steps of image preprocessing are as follows:
[0021] 1) Image cropping
[0022] Based on the collected image data, an orthophoto map is generated and cropped to a suitable size;
[0023] 2) Pixel mean subtraction
[0024] When training on natural images, it doesn't make much sense to estimate the mean and variance for each pixel individually because theoretically the statistical properties of any part of the image should be the same as those of other parts, which is called the stationarity of the image. Subtracting the image mean can focus the attention of the image on texture features rather than illumination intensity;
[0025] The pixel mean is obtained by summing and averaging the R, G, and B channels of all images respectively, resulting in three values; the calculation formula for the pixel mean is as follows:
[0026]
[0027] In the formula, I n is the nth image in the dataset, i, j, and k are the row index, column index, and number of channels of the image matrix respectively, i = 1, 2,..., r, j = 1, 2,..., c, Rm is the pixel average of the R channel, Gm is the pixel average of the G channel, and Bm is the pixel average of the B channel;
[0028] (3) Point cloud preprocessing
[0029] Point cloud preprocessing mainly includes non-ground point filtering, coordinate transformation, and data correspondence; the specific processing steps of point cloud preprocessing are as follows:
[0030] 1) Non-ground point filtering
[0031] Based on the elevation information of the point cloud, an elevation threshold is set, and the height change of each point relative to its neighborhood is calculated. If the height change of a certain point is greater than the set threshold, then this point is classified as a non-ground point, and the points of non-ground objects such as vegetation and guardrails in the point cloud data are removed;
[0032] 2) Coordinate transformation
[0033] Calculate the projection matrix, which can map the point cloud into the image space; the calculation formula for the projection matrix is as follows:
[0034] T = P rect ×R rect ×T rv2c (2)
[0035] Wherein, P rect —— The projection matrix for making different cameras coplanar; R rect —— The projection matrix for making images coplanar; T rv2c —— The projection matrix from the radar space to the image space, including rotation transformation and translation transformation;
[0036] 3) Data correspondence
[0037] In order to make the point cloud and the image coplanar, it is necessary to divide the x, y, and z of each point of the point cloud by the z value of the point simultaneously to complete the projection process, and obtain the projection of the point cloud on the image plane; the calculation formula for data correspondence is as follows:
[0038] P OUT = T × P_in (3)
[0039] Wherein, P OUT is the coordinate matrix of the point cloud projected into the image space, and P_in is the transpose of the point cloud data.
[0040] Furthermore, step S2 for road boundary fusion extraction mainly includes the following steps:
[0041] (1) Coarse extraction of the road boundary in the image
[0042] Adopt DeepLabV3+ as the road boundary semantic segmentation model, train the model, input the preprocessed road image into the trained road boundary detection model for detection, and obtain the detection result;
[0043] (2) Coarse extraction of the road boundary in the point cloud
[0044] The Alpha-Shapes algorithm can reconstruct the set shape from the unordered point set and construct the polygonal boundary of the point cloud. Adopt the Alpha-Shapes algorithm to extract the envelope surface, and then judge whether the pixel point is within the point cloud area according to the point cloud range. If it is, it is considered as the road area; the specific steps of the Alpha-Shapes algorithm are as follows:
[0045] 1) Start from any point in the point set S, denote this point as P1, find the point set whose distance from P1 is less than 2 × α, denote it as S′, and take any point in the set S′ and denote it as P2; the calculation formula for the center point P3 is as follows:
[0046]
[0047] Wherein, X1, Y1, X2, Y2, X3, Y3 are the coordinates of points P1, P2, P3, S 2 P1P2 =(x1 - x2)2 +(y1 - y2) 2 , where α is the radius;
[0048] 2) Calculate the distance L from all points except P2 to the center P3 in the set S′, and perform the following judgments;
[0049] a. If the L of all points is greater than α and there are no other points in the upper two circles, it means there are no other points in the circle, and P1 and P2 are boundary point pairs. Connect P1P2 and execute 4;
[0050] b. If there exists a point whose L is less than or equal to L and there are other points in the lower two circles, it means there are other points in the circle, then the point pair P1 and P2 does not form a boundary, and execute 3;
[0051] 3) Perform the two-step judgments of 1 and 2 on the next point in S′ until all points in S′ have been judged; if all do not meet the conditions described in a, then the P1 point does not belong to the boundary point;
[0052] 4) Take the next point in S as P1 and execute 1, 2, and 3 until all points in S have been traversed once;
[0053] (3) Fusion of the rough extraction results of the road boundary
[0054] In the preprocessing, the data of the two have been corresponding. First, estimate the boundary confidence according to the image quality and boundary clarity. Areas with higher image quality, clear and continuous boundaries can be set with higher confidence; then determine the confidence according to the point cloud density, aggregation degree of boundary points, etc. In places where the road area boundary is clear and the point cloud density is high, a higher confidence can be given. Then perform weighted fusion on the results of the two, and only extract the xy boundary information without considering the point cloud height information z; the calculation formula for weighted average fusion is as follows:
[0055] I f I(x,y) = ω1I1(x,y) + ω2I2(x,y) (5)
[0056] In the formula, I f (x,y) is the image after weighted average fusion, ω1 and ω2 are the weight coefficients of the image I1(x,y) and the point cloud I2(x,y) respectively, and satisfy ω 1+ ω2 = 1;
[0057] Based on the rough extraction results of the image road boundary and the point cloud road boundary, according to the above algorithm, the final road range is obtained;
[0058] (4) Expansion of the road boundary buffer zone
[0059] In order to prevent the incomplete extraction of road boundaries, which may lead to the omission of subsequent road potholes, the extracted road range is buffered according to the data size, and 1 / 10 of the average road width is taken for expansion to obtain the expanded road boundary range.
[0060] Furthermore, step S3 of image two-dimensional pothole rough extraction mainly includes the following steps:
[0061] (1) Image Semantic Label Pothole Extraction
[0062] The road image is cropped using the road range extracted from S2 to obtain a sub-image containing the road area to reduce unnecessary background interference. The processed image is input into the trained DeepLabv3+ model to output a semantic segmentation map in which each pixel is labeled as its category. The output semantic segmentation map includes: 0-background; 1-road; 2-pothole; all pixel areas marked as "pothole" are screened in the segmentation result to generate preliminary two-dimensional pothole candidate areas;
[0063] (2) 2D pothole edge refinement extraction
[0064] The Sobel operator is used to calculate the gradient of the pothole area of the semantic segmentation result to generate a preliminary edge map. The Gaussian filter, gradient amplitude calculation, non-maximum suppression and double threshold processing of the Canny algorithm are combined to enhance the accuracy and continuity of the edge. Morphological operations are used to repair broken edges and remove redundant noise. Finally, the contour detection algorithm is used to extract and smooth the edge to generate an accurate two-dimensional pothole contour. The formula for calculating the gradient map of the boundary by the Sobel operator is as follows:
[0065]
[0066] Where: G x , G y are the horizontal and vertical gradients of the image.
[0067] Furthermore, step S4 of roughly extracting three-dimensional potholes from point clouds mainly includes the following steps:
[0068] (1) Road Strength Map Generation
[0069] The present invention adopts the Inverse Distance Weighted (IDW) algorithm, takes the intensity of the road point cloud and the distance between the point clouds as weights, cuts out the point cloud road range data according to the road boundary range extracted by S2, and interpolates it into a two-dimensional road intensity image, which provides a basis for the geometric analysis of pothole areas;
[0070] On a two-dimensional plane, assuming the center point of each grid is (x, y), it is necessary to calculate its intensity value I(x, y). For each interpolation point, search for N known points (x i , y i , I i ) within its neighborhood range, and select neighborhood points by a fixed radius;
[0071] Neighborhood point distance d i Calculation formula:
[0072]
[0073] Neighborhood point weight w i Calculation formula:
[0074]
[0075] Calculation of the intensity value I(x, y) of the interpolation point:
[0076]
[0077] According to the above formula, map the interpolated intensity value to a two-dimensional grid to form a two-dimensional road intensity map;
[0078] (2) Coarse extraction of potholes by tensor voting
[0079] Based on the generated road intensity map, initialize the tensor information of each point, calculate the local geometric structure of the point cloud through neighborhood propagation, identify regions with significant changes in curvature and normal vector, and determine whether these regions conform to the characteristics of potholes, and output the result of coarse extraction of potholes in the point cloud;
[0080] 1) Tensor voting initialization: According to the normal vector and curvature information of each point in the point cloud, construct an initial tensor; The tensor initialization is expressed as:
[0081]
[0082] In the formula: T is the given tensor, λ1, λ2, λ3 represent the principal curvatures, and e1, e2, e3 represent the corresponding principal directions;
[0083] 2) Tensor propagation: Propagate the tensor information within the neighborhood point set N(j) of each point, and use the kernel function to weight and accumulate the tensors of the neighborhood points; The tensor propagation formula is:
[0084] T j =∑ i∈N(j) T j ·ωi j (11)
[0085] In the formula: ωi jis a weight function based on the distance and direction between points, used to control the propagation influence range; ωi j The calculation formula is as follows:
[0086]
[0087] In the formula: ||i - j|| 2 is the Euclidean distance between point i and point j, the angle between the normal vectors, σ d and σ θ are parameters to control the influence of distance and direction;
[0088] 3) Geometric feature extraction: Perform eigenvalue decomposition on the propagated tensor, analyze the principal curvature distribution of each point, judge the linear feature by λ1 >> λ2 ≈ λ2, judge the surface feature by λ1 ≈ λ2 >> λ2, and judge the point - like feature by λ1 ≈ λ2 ≈ λ2, so as to identify the significant geometric change points in the pothole area, screen the points with significant eigenvalue changes, especially the points with higher curvature (the points where λ1 is significantly greater than other eigenvalues), and mark them as potential pothole points;
[0089] 4) Geometric verification: According to the distribution of significant points, perform clustering based on the DBSCAN algorithm to generate three - dimensional candidate regions, and by analyzing the shape features such as the aspect ratio and normal vector consistency of the candidate regions, eliminate the regions that do not conform to the characteristics of road potholes, and output the three - dimensional point cloud regions that conform to the pothole characteristics. The geometric features such as depth, shape, and normal vector of these regions are consistent with the expected features of road potholes.
[0090] Furthermore, the rough extraction of pothole hierarchical fusion in step S5 mainly includes the following steps:
[0091] (1) Majority decision to screen high - confidence potholes
[0092] Perform a threshold judgment operation based on the rough extraction results of potholes from images and point clouds, mark the regions where the confidence levels C of both are higher than the threshold as the determined pothole regions D (i,p) , mark the regions where the confidence level of a single data is higher than the threshold as candidate regions D s , and mark the regions where the confidence level of a single data is lower than the veto threshold as non - pothole regions D'; The majority - decision pothole screening formula is as follows:
[0093]
[0094] In the formula: the image confidence level threshold T i = 0.7, the point cloud confidence level threshold T p = 0.65; the image veto threshold T i ' = 0.3, the point cloud veto threshold T p ' = 0.25;
[0095] (2) Pothole expansion in the peripheral area driven by morphological features
[0096] Potholes usually appear as irregular circles or polygons. The edges of real potholes may cause missed detection in some parts of a single sensor due to occlusion or data noise. For the pothole area D (i,p) The candidate area D around it s Determine whether it is merged into the pothole area through morphological features. The non-D (i,p) The peripheral D s The area is reserved for MRF refinement; for the above D (i,p) The area is subjected to morphological dilation. If the peripheral D s The area and the dilated D (i,p) The areas are morphologically continuous and geometrically similar, then they are merged; the specific calculation steps are as follows:
[0097] 1) Morphological preprocessing
[0098] Perform a dilation operation on the high-confidence pothole area D (i,p) to generate an expanded area D ( ′ i,p) , and the dilation radius r i is dynamically set according to the size of D (i,p) ; r i The calculation formula is as follows:
[0099]
[0100] In the formula: A i is the area of the i-th high-confidence pothole area, r i is the equivalent radius of each pothole, and k is an empirical coefficient, k = 0.5;
[0101] 2) Candidate area screening
[0102] Extract the intersection range of the expanded area D ( ′ i,p) and the candidate area D s The intersecting candidate area D c is expressed as:
[0103] D c = D s ∩D d (15)
[0104] 3) Morphological similarity determination
[0105] Calculate the morphological similarity between the intersecting candidate area D c and the pothole area. If the threshold is met, they are merged; the contact length between the candidate area and the boundary of the main pothole needs to be greater than 20% of its perimeter, and the area does not exceed 50% of the main pothole;
[0106] a. Extract the pothole area D through the Canny algorithm (i,p) With the candidate area D c 's boundary, count D c The number of pixels where the area boundary overlaps with D (i,p) area boundary, and determine whether its contact length is greater than 20% of the perimeter of D (i,p) ;
[0107] b. Calculate the area ratio with the pixel points of the pothole area D (i,p) and the intersecting candidate area D c as the basic unit, count the number of pixel points in the two areas, and determine whether the number of pixel points in the D c area is less than 50% of the D (i,p) area;
[0108] (3) Construction of the MRF fusion model
[0109] Based on the candidate areas around the non-high-confidence areas, construct the MRF fusion model; to construct a total energy function with physical significance, dynamically adjust the MRF parameters according to the characteristics of the image and the point cloud, optimize the pothole determination, and the dynamic weights include the image entropy value η i and the point cloud roughness θ i ;
[0110] The image entropy value η i Quantify the image texture features through the Gray-level Co-occurrence Matrix (GLCM), calculate its entropy value, and the high entropy value corresponds to the granular noise inside the pothole, expressed as:
[0111] η i = ∑ i,j P(i,j)·logP(i,j) (16)
[0112] The roughness θ i is defined as follows: Taking the i-th point of the point cloud as the center, use the points within its neighborhood range R i to fit the reference plane S, and the directed vertical distance from point i to the reference plane S is recorded as the roughness of point i; the plane equation and the calculation formula of the roughness are respectively:
[0113] A×x + B×y + C×z + D = 0 (17)
[0114]
[0115] In the formula: A, B, C, D are the parameters of the plane model S, θ i is the roughness of point P, x i , y i , z i are the coordinates of point P;
[0116] The optimized MRF energy function is expressed as:
[0117] E(x, y1, y2) = h × ∑ i x i -β × ∑ ij x i × x j -η i × ∑ i x i × y 1i -θ i × ∑ i x i × y 2i (19)
[0118] In the formula: it includes a unary potential and three binary potentials. The unary potential is the label itself, indicating that the classification result is related to its own assumed label; the binary potentials include the product of the current pixel label and the labels of surrounding pixels, the product of the score of the current pixel with the image entropy value and the current pixel label, and the product of the current pixel label and the current point cloud label, which respectively represent that the classification result of the current pixel is related to the reliability of the label, the categories of surrounding pixels, and the radar observation value;
[0119] (4) Solving by the ICM algorithm
[0120] The purpose of the present invention is to obtain the segmentation result, so the specific result of the joint distribution is not solved; the ICM algorithm is used to solve the optimized MRF energy function. First, initialize, mark the pixels in the uncertain area as the initial guess (all as potholes), and then traverse each pixel to calculate the energy values of two states (pothole / non-pothole):
[0121] E 坑洼 = h - β∑ j∈N(i) x j -η i y 1i -θ i y 2i , E 非坑洼 = 0 (20)
[0122] Furthermore, the accuracy verification in step S6 mainly includes the following steps:
[0123] By statistically analyzing the pixel classification results, the pothole area extracted by the algorithm is compared with the real annotation data, and the performance of the model is quantitatively evaluated from both the overall and category levels; comparing the extraction result of the algorithm with the real annotation data, by statistically analyzing the pixel classification results, the three indexes of PA, MPA, and MIoU are directly calculated:
[0124] The calculation formula of PA is as follows:
[0125]
[0126] The calculation formula of MPA is as follows:
[0127]
[0128] The calculation formula of MIoU is as follows:
[0129]
[0130] Where: TP is the true positive, FP is the false positive, FN is the false negative, TN is the true negative, and P i is the pixel accuracy of each category.
[0131] Beneficial effects: Compared with the prior art, a method for extracting unstructured road potholes based on the fusion of UAV point cloud data and optical images has the following advantages:
[0132] (1) The present invention uses UAV remote sensing technology to collect unstructured road data, without relying on ground vehicles or other auxiliary devices, and can quickly obtain road data containing multi-dimensional information such as images and point clouds;
[0133] (2) Before extracting unstructured road potholes, the present invention extracts the road range through the fusion of images and point clouds, reduces the mis-extraction of potholes in non-road areas, and improves the reliability of pothole extraction;
[0134] (3) In the extraction of unstructured road potholes, the present invention combines the advantages of images and laser point clouds, combines image semantic segmentation and point cloud tensor voting technology to roughly extract potholes, and through the designed rough extraction pothole hierarchical fusion algorithm, makes the advantages of image point clouds complementary to each other to obtain the accurate pothole range. Description of the Drawings
[0135] The description of the content of the present invention will become obvious and easy to understand in combination with the following drawings, where:
[0136] Figure 1 is the flow chart of a method for extracting unstructured road potholes based on the fusion of UAV point cloud data and optical images of the present invention; Detailed Embodiments
[0137] According to Figure 1 the steps shown, a method for extracting unstructured road potholes based on the fusion of UAV point cloud data and optical images of the present invention will be described in detail.
[0138] Step 1: Data collection and preprocessing. It includes the following specific steps:
[0139] (1) Data acquisition
[0140] Using drone oblique photography and airborne LiDAR technology, traverse the unstructured road area to collect complete and high-precision optical images and LiDAR point cloud data;
[0141] (2) Image preprocessing
[0142] Image preprocessing mainly includes image cropping and pixel mean removal; crop the size of the image to adapt to the computer processing performance, and pixel mean removal standardizes the data features, focusing the attention of the image on the texture features; the specific processing steps of image preprocessing are as follows:
[0143] 1) Image cropping
[0144] Based on the collected image data, generate an orthophoto map and crop it to a suitable size;
[0145] 2) Pixel mean removal
[0146] When training on natural images, it doesn't make much sense to estimate the mean and variance for each pixel individually because theoretically the statistical properties of any part of the image should be the same as other parts, which is called the stationarity of the image. Subtracting the image mean can focus the attention of the image on the texture features rather than the illumination intensity;
[0147] The pixel mean is obtained by summing and averaging the R, G, and B channels of all images respectively, resulting in three values; the pixel mean calculation formula is as follows:
[0148]
[0149] In the formula, I n is the nth image in the dataset, i, j, and k are the row index, column index, and number of channels of the image matrix respectively, i = 1, 2,..., r, j = 1, 2,..., c, Rm is the pixel average value of the R channel, Gm is the pixel average value of the G channel, and Bm is the pixel average value of the B channel;
[0150] (3) Point cloud preprocessing
[0151] Point cloud preprocessing mainly includes non-ground point filtering, coordinate transformation, and data correspondence; the specific processing steps of point cloud preprocessing are as follows:
[0152] 1) Non-ground point filtering
[0153] Based on the elevation information of the point cloud, set the elevation threshold, calculate the height change of each point relative to its neighborhood. If the height change of a certain point is greater than the set threshold, then this point is classified as a non-ground point, and remove the points of non-ground objects such as vegetation and guardrails from the point cloud data;
[0154] 2) Coordinate transformation
[0155] Calculate the projection matrix, which can map the point cloud into the image space. The calculation formula of the projection matrix is as follows:
[0156] T = P rect ×R rect ×T rv2c (25)
[0157] In the formula, P rect ——The projection matrix that makes different cameras coplanar; R rect ——The projection matrix that makes the images coplanar; T rv2c ——The projection matrix from the radar space to the image space, including rotation transformation and translation transformation;
[0158] 3) Data correspondence
[0159] To make the point cloud and the image coplanar, it is necessary to divide the x, y, and z of each point of the point cloud by the z value of the point to complete the projection process and obtain the projection of the point cloud on the image plane. The calculation formula for data correspondence is as follows:
[0160] P OUT = T × P_in (26)
[0161] In the formula, P OUT is the coordinate matrix of the point cloud projected into the image space, and P_in is the transpose of the point cloud data.
[0162] Step 2: Road boundary fusion extraction. It includes the following specific steps:
[0163] (1) Coarse extraction of the road boundary in the image
[0164] Use DeepLabV3+ as the road boundary semantic segmentation model, train the model, and input the preprocessed road image into the trained road boundary detection model for detection to obtain the detection result;
[0165] (2) Coarse extraction of the road boundary in the point cloud
[0166] The Alpha-Shapes algorithm can reconstruct the set shape from the unordered point set and construct the polygonal boundary of the point cloud. Use the Alpha-Shapes algorithm to extract the envelope surface, and then judge whether the pixel point is in the point cloud area according to the point cloud range. If it is, it is considered to be the road area. The specific steps of the Alpha-Shapes algorithm are as follows:
[0167] 1) Start from any point in the point set S, which is denoted as P1, find the point set whose distance from P1 is less than 2×α, denoted as S′, and randomly select a point in the set S′ and denote it as P2; The calculation formula for the center point P3 is as follows:
[0168]
[0169] Wherein, X1, Y1, X2, Y2, X3, and Y3 are the coordinates of points P1, P2, and P3, S 2 P1P2 =(x1 - x2) 2 +(y1 - y2) 2 , and α is the radius;
[0170] 2) Calculate the distance L from all points except P2 to the center P3 in the set S', and perform the following judgments;
[0171] a. If the L of all points is greater than α, there are no other points in the upper two circles, indicating that there are no other points in the circle, and P1 and P2 are boundary point pairs. Connect P1P2 and execute 4;
[0172] b. If there exists a point whose L is less than or equal to L, there are other points in the lower two circles, indicating that there are other points in the circle. Then the point pair P1 and P2 does not form a boundary, and execute 3;
[0173] 3) Perform the two-step judgments of 1 and 2 on the next point in S' until all points in S' are judged; if none of them meet the conditions described in a, then point P1 does not belong to the boundary point;
[0174] 4) Take the next point in S as P1 and execute 1, 2, and 3 until all points in S are traversed once;
[0175] (3) Fusion of the rough extraction results of the road boundary
[0176] In the preprocessing, the data of the two have been corresponding. First, estimate the boundary confidence according to the image quality and boundary clarity. Areas with higher image quality, clear and continuous boundaries can be set with higher confidence; then determine the confidence according to the point cloud density, aggregation degree of boundary points, etc. In places where the boundary of the road area is clear and the point cloud density is high, a higher confidence can be given. Then perform weighted fusion on the results of the two, and only extract the xy boundary information without considering the point cloud height information z; the calculation formula for weighted average fusion is as follows:
[0177] I f I(x,y)=ω1I1(x,y)+ω2I2(x,y) (28)
[0178] Wherein, I f (x,y) is the image after weighted average fusion, ω1 and ω2 are the weight coefficients of the image I1(x,y) and the point cloud I2(x,y) respectively, and satisfy ω 1+ ω2 = 1;
[0179] Based on the results of the rough extraction of the road boundary from the image and the rough extraction of the road boundary from the point cloud, the final road range is obtained according to the above algorithm;
[0180] (4) Expansion of the road boundary buffer
[0181] To prevent incomplete extraction of the road boundary from resulting in missed detection of subsequent road potholes, the extracted road range is expanded with a buffer according to the data size. The average width of the road is taken as 1 / 10 for expansion to obtain the expanded road boundary range.
[0182] Step 3: Rough extraction of two-dimensional potholes from the image. It includes the following specific steps:
[0183] (1) Extraction of potholes from the semantic labels of the image
[0184] Use the road range extracted from S2 to crop the road image to obtain a sub-image containing the road area, so as to reduce unnecessary background interference. Input the processed image into the trained DeepLabv3+ model to output a semantic segmentation map, where each pixel is labeled as the category it belongs to. The output semantic segmentation map includes: 0 - background; 1 - road; 2 - pothole; In the segmentation result, filter all pixel regions labeled as "pothole" to generate a preliminary two-dimensional pothole candidate region;
[0185] (2) Refined extraction of the edges of two-dimensional potholes
[0186] Use the Sobel operator to calculate the gradient of the pothole region in the semantic segmentation result to generate a preliminary edge map. Combine the Gaussian filtering, gradient magnitude calculation, non-maximum suppression, and double-threshold processing of the Canny algorithm to enhance the accuracy and continuity of the edges, and use morphological operations to repair broken edges and remove redundant noise. Finally, extract and smooth the edges through the contour detection algorithm to generate an accurate two-dimensional pothole contour; The formula for calculating the gradient map of the boundary by the Sobel operator is as follows:
[0187]
[0188] In the formula: G x 、G y are the horizontal and vertical gradients of the image.
[0189] Step 4: Rough extraction of point cloud potholes. It includes the following specific steps:
[0190] (1) Generation of the road intensity map
[0191] The present invention adopts the Inverse Distance Weighted (IDW) algorithm, uses the intensity of the road point cloud and the distance between point clouds as weights, cuts out the point cloud road range data according to the road boundary range extracted in S2, interpolates it into a two-dimensional road intensity image, and provides a basis for the geometric analysis of pothole areas;
[0192] On the two-dimensional plane, assume that the center point of each grid is (x, y), and its intensity value I(x, y) needs to be calculated. For each interpolation point, search for N known points (x i , y i , I i ) within its neighborhood range, and select neighborhood points through a fixed radius;
[0193] The distance d of the neighborhood point i Calculation formula:
[0194]
[0195] The weight w of the neighborhood point i Calculation formula:
[0196]
[0197] Calculation of the intensity value I(x, y) of the interpolation point:
[0198]
[0199] According to the above formula, map the interpolated intensity value to a two-dimensional grid to form a two-dimensional road intensity map;
[0200] (2) Coarse extraction of potholes by tensor voting
[0201] Based on the generated road intensity map, initialize the tensor information of each point, calculate the local geometric structure of the point cloud through neighborhood propagation, identify the regions where the curvature and normal vector change significantly, determine whether these regions conform to the characteristics of potholes, and output the result of the coarse extraction of potholes in the point cloud;
[0202] 1) Tensor voting initialization: Construct an initial tensor according to the normal vector and curvature information of each point in the point cloud; The tensor initialization is expressed as:
[0203]
[0204] In the formula: T is the given tensor, λ1, λ2, λ3 represent the principal curvatures, and e1, e2, e3 represent the corresponding principal directions;
[0205] 2) Tensor propagation: Propagate the tensor information within the neighborhood point set N(j) of each point, and use the kernel function to weight and accumulate the tensors of the neighborhood points; The tensor propagation formula is:
[0206] T j = ∑ i∈N(j) T j · ωi j (34)
[0207] Where: ωi j is a weight function based on the distance and direction between points, used to control the propagation influence range; ωi j The calculation formula is as follows:
[0208]
[0209] Where: ||i - j|| 2 is the Euclidean distance between point i and point j, the angle between the normal vectors, σ d and σ θ are parameters for controlling the influence of distance and direction;
[0210] 3) Geometric feature extraction: Perform eigenvalue decomposition on the propagated tensor, analyze the principal curvature distribution of each point, judge the linear feature by λ1 >> λ2 ≈ λ2, judge the surface feature by λ1 ≈ λ2 >> λ2, and judge the dot feature by λ1 ≈ λ2 ≈ λ2, so as to identify the significant geometric change points in the pitted area, screen the points with significant eigenvalue changes, especially the points with higher curvature (the points where λ1 is significantly greater than other eigenvalues), and mark them as potential pitted points;
[0211] 4) Geometric verification: Based on the distribution of significant points, generate three-dimensional candidate regions by clustering using the DBSCAN algorithm, and eliminate the regions that do not conform to the characteristics of road potholes by analyzing the shape characteristics such as the aspect ratio and normal vector consistency of the candidate regions, and output the three-dimensional point cloud regions that conform to the pothole characteristics. The geometric characteristics such as depth, shape, and normal vector of these regions are consistent with the expected characteristics of road potholes.
[0212] Step 5: Coarse extraction of pothole hierarchical fusion. The following specific steps are included:
[0213] (1) Majority decision to screen high-confidence potholes
[0214] Perform a threshold judgment operation on the rough extraction results of potholes based on images and point clouds, mark the regions where the confidence levels C of both are higher than the threshold as the determined pothole regions D (i,p) , mark the regions where the confidence level of a single data is higher than the threshold as candidate regions D s , and mark the regions where the confidence level of a single data is lower than the veto threshold as non-pothole regions D'; The majority decision pothole screening formula is as follows:
[0215]
[0216] Where: the image confidence threshold T i = 0.7, the point cloud confidence threshold T p = 0.65; the image rejection threshold T i ' = 0.3, the point cloud rejection threshold T p ' = 0.25;
[0217] (2) Pothole expansion in the peripheral area driven by morphological features
[0218] Potholes usually appear as irregular circles or polygons. The true pothole edges may cause missed detections in some parts of a single sensor due to occlusion or data noise. For the pothole area D (i,p) in the peripheral candidate area D s judge whether it is merged into the pothole area through morphological features. For the D (i,p) in the non-peripheral D s area, leave it for MRF refinement; perform morphological dilation on the above D (i,p) area. If the D s area in the periphery and the dilated D (i,p) area are morphologically continuous and geometrically similar, then merge them; the specific calculation steps are as follows:
[0219] 1) Morphological preprocessing
[0220] Perform a dilation operation on the high-confidence pothole area D (i,p) to generate the expanded area D ( ', and the dilation radius r i,p) is dynamically set according to the D i size; the calculation formula of r (i,p) is as follows: i In the formula: A i is the area of the i-th high-confidence pothole area, r i is the equivalent radius of each pothole, and k is an empirical coefficient, k = 0.5;
[0221]
[0222]
[0223] 2) Candidate area screening
[0224] Extract the intersection range of the expanded area D ( ' i,p) and the candidate area D s . The intersecting candidate area D c is expressed as:
[0225] D c = D s ∩D d (38)
[0226] 3) Morphological similarity determination
[0227] Calculate the intersection candidate region D c The morphological similarity with the pothole region. If the threshold is met, merge them; the contact length between the candidate region and the main pothole boundary should be greater than 20% of its perimeter, and the area should not exceed 50% of the main pothole;
[0228] a. Extract the pothole region D through the Canny algorithm (i,p) With the candidate region D c 's boundary, count the number of overlapping pixels between the boundary of region D c and the boundary of region D (i,p) to determine whether its contact length is greater than 20% of the perimeter of D (i,p) ;
[0229] b. Calculate the area ratio with the pothole region D (i,p) and the intersection candidate region D c using the pixel points as the basic unit. Count the number of pixel points in the two regions and determine whether the number of pixel points in region D c is less than 50% of the number of pixel points in region D (i,p) ;
[0230] (3) Construction of the MRF fusion model
[0231] Based on the candidate regions around the non-high-confidence area, construct the MRF fusion model; to construct a total energy function with physical significance, dynamically adjust the MRF parameters according to the characteristics of the image and the point cloud to optimize the pothole determination. The dynamic weights include the image entropy value η i and the point cloud roughness θ i ;
[0232] The image entropy value η i Quantify the image texture features through the Gray-level Co-occurrence Matrix (GLCM) and calculate its entropy value. A high entropy value corresponds to the granular noise inside the pothole, which is expressed as:
[0233] η i =∑ i,j P(i,j)·logP(i,j) (39)
[0234] The roughness θ i is defined as follows: Taking the point i in the point cloud as the center, use the points within its neighborhood range R i to fit the reference plane S. The directed vertical distance from point i to the reference plane S is recorded as the roughness of point i. The plane equation and the calculation formula for roughness are respectively:
[0235] A×x + B×y + C×z + D=0 (40)
[0236]
[0237] Where: A, B, C, and D are parameters of the planar model S, and θ i is the roughness of point P, and x i , y i , z i are the coordinates of point P;
[0238] The optimized MRF energy function is expressed as:
[0239] E(x, y1, y2) = h × ∑ i x i - β × ∑ ij x i × x j - η i × ∑ i x i × y 1i - θ i × ∑ i x i × y 2i (42)
[0240] Where: It includes a unary potential and three binary potentials. The unary potential is its own label, indicating that the classification result is related to its own assumed label; the binary potentials include the product of the current pixel label and the surrounding pixel labels, the product of the score of the current pixel added with the image entropy value and the current pixel label, and the product of the current pixel label added with the point cloud quality index and the current point cloud label, which respectively represent that the classification result of the current pixel is related to the reliability of the label, the category of the surrounding pixels, and the radar observation value;
[0241] (4) Solving with the ICM algorithm
[0242] The purpose of the present invention is to obtain the segmentation result, so the specific result of the joint distribution is not solved; the ICM algorithm is used to solve the optimized MRF energy function. First, initialize, mark the pixels in the uncertain area as the initial guess (all as potholes), and then traverse each pixel to calculate the energy values of two states (pothole / non-pothole):
[0243] E 坑洼 = h - β∑ j∈N(i) x j - η i y 1i - θ i y 2i , E 非坑洼 = 0 (43)
[0244] Step 6: Accuracy verification. It includes the following specific steps:
[0245] By counting the pixel classification results, the pothole areas extracted by the algorithm are compared and analyzed with the real annotation data, and the performance of the model is quantitatively evaluated from both the overall and category levels; by comparing the extraction results of the algorithm with the real annotation data and counting the pixel classification results, three indicators, namely PA, MPA, and MIoU, are directly calculated:
[0246] The calculation formula of PA is as follows:
[0247]
[0248] The calculation formula of MPA is as follows:
[0249]
[0250] The calculation formula of MIoU is as follows:
[0251]
[0252] In the formula: TP is the true positive, FP is the false positive, FN is the false negative, TN is the true negative, and P i is the pixel accuracy of each category.
[0253] A method for extracting unstructured road potholes based on the fusion of UAV point cloud data and optical images according to the present invention aims at the requirements of safe production and transportation of vehicles on unstructured roads. Based on image and point cloud data, deep learning and tensor voting algorithms are adopted, and through a multi-source data hierarchical fusion algorithm, accurate extraction of unstructured road potholes and specific parameter information is realized.
[0254] The above are only the best embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. An unstructured road pothole extraction method based on the fusion of drone point cloud data and optical images, characterized in that, It includes the following steps: S1: Experimental data acquisition and preprocessing: Obtain complete and high-quality optical images and LiDAR point cloud data of the unstructured road in the target area through UAV oblique photography and airborne LiDAR technology. The preprocessing steps are as follows: (1) Image preprocessing mainly includes image cropping, pixel de-mean, and dataset construction; crop the image size to adapt to the computer processing performance, pixel de-mean standardizes the data features, focuses the image on texture features, and then constructs the road boundary and road pothole datasets for model construction; (2) Point cloud preprocessing mainly includes non-ground point filtering, coordinate transformation, and data correspondence; non-ground point filtering focuses the point cloud on road surface information, effectively reducing noise interference, coordinate transformation ensures the correct alignment of the point cloud and image data, and data correspondence integrates the point cloud and image data into a unified road surface model; S2: Road boundary fusion extraction: Based on the processed point cloud and image data, use the DeepLabv3+ model to roughly extract the road boundary of the image, roughly extract the road boundary of the point cloud through the Alpha-Shapes algorithm, perform weighted averaging on the results of both to determine the final road boundary range, and finally perform buffer expansion on the road boundary range to ensure the integrity of the road range; S3: Coarse extraction of potholes by image semantic segmentation: Cut out the road range in S2 from the image data to obtain the image road range, extract the semantic label potholes through the trained DeepLabv3+ model, and use the edge detection (Canny) algorithm to enhance the accuracy and continuity of the pothole edges to obtain the candidate potholes and their two-dimensional pothole contours in the image plane coordinate system; S4: Coarse extraction of potholes by point cloud tensor voting: Cut out the road range in S2 from the point cloud data to obtain the point cloud road range, and generate the corresponding road intensity map. By analyzing the tensor voting results of the point cloud, identify the regions with significant geometric changes and determine whether the potholes in the candidate potholes are regarded as potholes; S5: Hierarchical fusion of coarsely extracted potholes: Based on the pothole results of image and point cloud coarsely extracted potholes, first screen the high-confidence regions based on the majority decision, determine the regions where both are potholes as pothole regions, and mark the high-confidence regions of single data as candidate regions; then perform morphological dilation on the marked final pothole regions. If the candidate regions of single data around are morphologically continuous and geometrically similar to the dilated main pothole regions, merge them; For the candidate regions of single data that are not around, construct a probability graph model using the MRF theory, dynamically adjust the MRF parameters according to the image entropy value and point cloud roughness to obtain the Markov potential function, and then use the ICM algorithm to solve the potential function to obtain the distribution result with the maximum probability and determine the road pothole regions; S6: Accuracy verification: Compare with the true road pothole segmentation results of the division, calculate PA (Pixel Accuracy), MPA (Mean Pixel Accuracy of Classes), and MIoU (Mean Intersection over Union) to evaluate the road pothole extraction accuracy.
2. The method for extracting unstructured road potholes based on the fusion of UAV point cloud data and optical images according to claim 1, wherein, The step S1 includes the following steps: (1) Data acquisition Using drone oblique photography and airborne LiDAR technology, traverse the unstructured road area to collect complete and high-precision optical images and LiDAR point cloud data; (2) Image preprocessing Image preprocessing mainly includes image cropping and pixel de-mean. Crop the size of the image to adapt to the computer processing performance, and pixel de-mean standardizes the data features, focusing the image's attention on texture features. The specific processing steps of image preprocessing are as follows: 1) Image cropping Based on the collected image data, generate an orthophoto map and crop it to a suitable size; 2) Pixel de-mean When training on natural images, it doesn't make much sense to estimate the mean and variance for each pixel individually because theoretically, the statistical properties of any part of the image should be the same as other parts, which is called the stationarity of the image. Subtracting the image mean can focus the image's attention on texture features rather than light intensity; The pixel mean is obtained by summing and averaging the R, G, and B channels of all images respectively, resulting in three values. The pixel mean calculation formula is as follows: Where, I n is the nth picture in the dataset, i, j, and k are the row index, column index, and number of channels of the image matrix respectively, i = 1, 2,..., r, j = 1, 2,..., c, Rm is the pixel average value of the R channel, Gm is the pixel average value of the G channel, and Bm is the pixel average value of the B channel; (3) Point cloud preprocessing Point cloud preprocessing mainly includes non-ground point filtering, coordinate transformation, and data correspondence. The specific processing steps of point cloud preprocessing are as follows: 1) Non-ground point filtering Based on the elevation information of the point cloud, set an elevation threshold, calculate the height change of each point relative to its neighborhood. If the height change of a certain point is greater than the set threshold, then this point is classified as a non-ground point, and remove the points of non-ground objects such as vegetation and guardrails from the point cloud data; 2) Coordinate transformation Calculate the projection matrix, which can map the point cloud into the image space. The calculation formula of the projection matrix is as follows: T = P rect × R rect × T rv2c (2) Where, P rect —— The projection matrix for making different cameras coplanar; R rect —— The projection matrix for making images coplanar; T rv2c —— The projection matrix from radar space to image space, including rotation transformation and translation transformation; 3) Data correspondence To make the point cloud and the image coplanar, it is necessary to divide the x, y, and z of each point of the point cloud by the z value of that point to complete the projection process, obtaining the projection of the point cloud on the image plane. The data correspondence calculation formula is as follows: P OUT = T × P_in (3) Wherein, P OUT is the coordinate matrix of the point cloud projected into the image space, and P_in is the transpose of the point cloud data.
3. A method for extracting unstructured road potholes based on the fusion of UAV point cloud data and optical images according to claim 1, characterized in that, The S2 includes the following steps: (1) Coarse extraction of image road boundaries Adopt DeepLabV3+ as the road boundary semantic segmentation model, train the model, and input the preprocessed road image into the trained road boundary detection model for detection to obtain the detection result; (2) Coarse extraction of point cloud road boundaries The Alpha-Shapes algorithm can reconstruct the set shape from an unordered point set and construct the polygonal boundary of the point cloud. Using the Alpha-Shapes algorithm, extract the envelope surface, and then judge whether a pixel point is within the point cloud area according to the point cloud range. If it is, it is considered to be the road area. The specific steps of the Alpha-Shapes algorithm are as follows: 1) Start from any point in the point set S, denote this point as P1, find the point set whose distance from P1 is less than 2×α, denoted as S′, and randomly select a point in the set S′ and denote it as P2. The calculation formula for the center point P3 is as follows: wherein, X1, Y1, X2, Y2, X3, and Y3 are the coordinates of points P1, P2, and P3 S 2 P1P2 =(x1 - x2) 2 +(y1 - y2) 2 , and α is the radius 2) Calculate the distance L from all points except P2 in the set S′ to the center point P3, and perform the following judgment; a. If the L of all points is greater than α and there are no other points in the upper two circles, indicating that there are no other points in the circle, P1 and P2 are boundary point pairs, connect P1P2, and execute 4; b. If there is a point whose L is less than or equal to L, and there are other points in the two circles below, it means that there are other points in the circles. Then this point does not constitute a boundary for P1 and P2, and execute 3; 3) Perform step 1 and 2 on the next point in S′ until all points in S′ are judged; if none of them meet the requirements in step a, then point P1 does not belong to the boundary point; 4) Take the next point in S as P1, and execute 1, 2, and 3 until all points in S have been traversed; (3) Fusion of road boundary rough extraction results The data of the two have been matched in the preprocessing. First, the boundary confidence is estimated according to the image quality and boundary clarity. A higher confidence can be set for areas with high image quality, clear and continuous boundaries. The confidence is then determined according to the point cloud density, the concentration of boundary points, etc. In places where the road area has clear boundaries and high point cloud density, a higher confidence can be assigned. Then, the results of the two are weighted fused, and only the xy boundary information is extracted, without considering the point cloud height information z. The calculation formula for weighted average fusion is as follows: I f (x,y) = ω1I1(x,y) + ω2I2(x,y) (5) where I f (x, y) is the image after weighted average fusion, ω1 and ω2 are the weight coefficients of the image I1(x, y) and the point cloud I2(x, y) respectively, and satisfy ω 1+ ω2 = 1; Based on the rough extraction results of image road boundaries and point cloud road boundaries, the final road range is obtained according to the above algorithm; (4) Road boundary buffer zone expansion In order to prevent the incomplete extraction of road boundaries, which may lead to the omission of subsequent road potholes, the extracted road range is buffered according to the data size, and 1 / 10 of the average road width is taken for expansion to obtain the expanded road boundary range.
4. A method for extracting unstructured road potholes based on the fusion of UAV point cloud data and optical images according to claim 1, characterized in that, The S3 comprises the following steps: (1) Image Semantic Label Pothole Extraction The road image is cropped using the road range extracted from S2 to obtain a sub-image containing the road area to reduce unnecessary background interference. The processed image is input into the trained DeepLabv3+ model to output a semantic segmentation map in which each pixel is labeled as its category. The output semantic segmentation map includes: 0-background; 1-road; 2-pothole; all pixel areas marked as "pothole" are screened in the segmentation result to generate preliminary two-dimensional pothole candidate areas; (2) 2D pothole edge refinement extraction The Sobel operator is used to calculate the gradient of the pothole area of the semantic segmentation result to generate a preliminary edge map. The Gaussian filter, gradient amplitude calculation, non-maximum suppression and double threshold processing of the Canny algorithm are combined to enhance the accuracy and continuity of the edge. Morphological operations are used to repair broken edges and remove redundant noise. Finally, the contour detection algorithm is used to extract and smooth the edge to generate an accurate two-dimensional pothole contour. The formula for calculating the gradient map of the boundary by the Sobel operator is as follows: Where: G x and G y are the horizontal and vertical gradients of the image.
5. A method for extracting unstructured road potholes based on the fusion of UAV point cloud data and optical images according to claim 1, characterized in that, The S4 comprises the following steps: (1) Road Strength Map Generation The present invention adopts the Inverse Distance Weighted (IDW) algorithm, takes the intensity of the road point cloud and the distance between the point clouds as weights, cuts out the point cloud road range data according to the road boundary range extracted by S2, and interpolates it into a two-dimensional road intensity image, which provides a basis for the geometric analysis of pothole areas; On a two-dimensional plane, assuming the center point of each grid is (x, y), it is necessary to calculate its intensity value I(x, y). For each interpolation point, search for N known points (x i , y i , I i ) within its neighborhood range, and select neighborhood points by fixing the radius; Neighbor point distance d i Calculation formula: Neighborhood point weight w i Calculation formula: The intensity value I(x,y) of the interpolation point is calculated as follows: According to the above formula, map the interpolated intensity values to a two-dimensional grid to form a two-dimensional road intensity map; (2) Coarse extraction of potholes by tensor voting Based on the generated road intensity map, initialize the tensor information of each point, calculate the local geometric structure of the point cloud through neighborhood propagation, identify the regions with significant changes in curvature and normal vector, determine whether these regions conform to the characteristics of potholes, and output the coarse extraction result of potholes in the point cloud; 1) Tensor voting initialization: According to the normal vector and curvature information of each point in the point cloud, construct an initial tensor; The tensor initialization is expressed as: In the formula: T is the given tensor, λ1, λ2, and λ3 represent the principal curvatures, and e1, e2, and e3 represent the corresponding principal directions; 2) Tensor propagation: Propagate the tensor information within the neighborhood point set N(j) of each point, and use the kernel function to weighted accumulate the tensors of neighborhood points; The tensor propagation formula is: T j = ∑ i∈N(j) T j · ω ij (11) where: ωi j is a weight function based on the distance and direction between points, used to control the propagation influence range; ωi j The calculation formula is as follows: where: ||i - j|| 2 is the Euclidean distance between point i and point j, the angle between the normal vectors, σ d and σ θ are parameters that control the influence of distance and direction; 3) Geometric feature extraction: Perform eigenvalue decomposition on the propagated tensor, analyze the principal curvature distribution of each point, judge the linear feature by λ1 >> λ2 ≈ λ2, judge the surface feature by λ1 ≈ λ2 >> λ2, and judge the point feature by λ1 ≈ λ2 ≈ λ2, so as to identify the significant geometric change points in the pothole area, screen the points with significant eigenvalue changes, especially the points with higher curvature (the points where λ1 is significantly greater than other eigenvalues), and mark them as potential pothole points; 4) Geometric verification: According to the distribution of significant points, perform clustering based on the DBSCAN algorithm to generate three-dimensional candidate regions, and by analyzing the shape features such as the aspect ratio and normal vector consistency of the candidate regions, eliminate the regions that do not conform to the characteristics of road potholes, and output the three-dimensional point cloud regions that conform to the characteristics of potholes. The geometric features such as the depth, shape, and normal vector of these regions are consistent with the expected features of road potholes.
6. The method for extracting unstructured road potholes based on the fusion of UAV point cloud data and optical images according to claim 1, wherein, The S5 includes the following steps: (1) Majority decision to screen high-confidence potholes Based on the pothole rough extraction results of images and point clouds, perform a threshold judgment operation, and mark the area where the confidence levels C of both are higher than the threshold as the determined pothole area D (i,p) , mark the area where the confidence level of a single data is higher than the threshold as the candidate area D s , mark the area where the confidence level of a single data is lower than the veto threshold as the non-pothole area D'; the majority decision pothole screening formula is as follows: Where: the image confidence threshold T i = 0.7, the point cloud confidence threshold T p = 0.65; the image rejection threshold T i ' = 0.3, the point cloud rejection threshold T p ' = 0.25; (2) Expansion of potholes in the surrounding area driven by morphological features Potholes usually appear as irregular circles or polygons. The edges of real potholes may cause missed detections in some parts of a single sensor due to occlusion or data noise; for the pothole area D (i,p) The candidate area D s around it is judged whether it is combined into the pothole area, not D (i,p) The D s around it is left for MRF refinement processing; for the above D (i,p) area, morphological dilation is performed. If the D s area around it and the dilated D (i,p) area are morphologically continuous and geometrically similar, they are combined; the specific calculation steps are as follows: 1) Morphological preprocessing Perform a dilation operation on the high-confidence pothole area D (i,p) to generate an expanded area D ( ′ i,p) , where the dilation radius r i is dynamically set according to the size of D (i,p) ; the calculation formula for r i is as follows: Where: A i is the area of the i-th high-confidence pothole region, r i is the equivalent radius of each pothole, k is an empirical coefficient, and k = 0.5; 2) Screening of candidate regions Extract the extended region D ( ′ i,p) The range that intersects with the candidate region D s The intersecting candidate region D c Is expressed as: D c = D s ∩ D d (15) 3) Judgment of morphological similarity Calculate the intersection candidate region D c Calculate the morphological similarity with the pothole region, and merge if the threshold is met; the contact length between the candidate region and the main pothole boundary needs to be greater than 20% of its perimeter, and the area does not exceed 50% of the main pothole; a. Extract the pothole area D through the Canny algorithm (i,p) and the candidate area D c boundary, and count the D c number of pixels where the area boundary overlaps with the D (i,p) area boundary, and determine whether its contact length is greater than D (i,p) 20% of the perimeter; b. Using the pitted area D (i,p) and the intersecting candidate area D c as the basic unit to calculate the area ratio, count the number of pixel points in the two areas, and judge whether the number of pixel points in area D c is less than 50% of that in area D (i,p) ; (3) Construction of MRF fusion model Construct an MRF fusion model based on candidate regions around non-high-confidence areas; to construct a total energy function with physical significance, dynamically adjust the MRF parameters according to the characteristics of the image and the point cloud, optimize the pothole determination, and the dynamic weights include the image entropy value η i and the point cloud roughness θ i ; Image entropy value η i Quantify the image texture features through the Gray-level Co-occurrence Matrix (GLCM), calculate its entropy value, and the high entropy value corresponds to the granular noise inside the pothole, which is expressed as: η i = ∑ i,j P(i, j)·log P(i, j) (16) Roughness θ i is defined as follows: taking point i of the point cloud as the center, using the points within its neighborhood range R i to fit the reference plane S. The directed vertical distance from point i to the reference plane S is denoted as the roughness of point i. The calculation formulas for the plane equation and roughness are as follows: A×x + B×y + C×z + D = 0 (17) Where: A, B, C, and D are parameters of the planar model S, and θ i is the roughness of point P, and x i , y i , z i are the coordinates of point P; The optimized MRF energy function is expressed as: E(x, y1, y2) = h × ∑ i x i -β × ∑ ij x i × x j -η i × ∑ i x i × y 1i -θ i × ∑ i x i × y 2i (19) In the formula: it includes a unary potential and three binary potentials. The unary potential is the label of itself, indicating that the classification result is related to its own assumed label; The binary potentials include the product of the current pixel label and the labels of surrounding pixels, the product of the score of the current pixel with the image entropy value and the current pixel label, and the product of the current pixel label and the current point cloud label, which respectively represent that the classification result of the current pixel is related to the reliability of the label, the category of surrounding pixels, and the radar observation value; (4) Solving by ICM algorithm The purpose of the present invention is to obtain the segmentation result, so do not solve the specific result of the joint distribution; Use the ICM algorithm to solve the optimized MRF energy function. First, perform initialization, mark the pixels in the uncertain area as the initial guess (all are potholes), and then traverse each pixel to calculate the energy values of two states (pothole / non-pothole): E 坑洼 = h - β∑ j∈N(i) x j - η i y 1i - θ i y 2i ,E 非坑洼 = 0 (20).
7. A method for extracting unstructured road potholes based on the fusion of UAV point cloud data and optical images according to claim 1, characterized in that The S6 includes the following steps: By counting the pixel classification results, the pothole areas extracted by the algorithm are compared and analyzed with the real annotation data, and the performance of the model is quantitatively evaluated from both the overall and category levels; by comparing the algorithm extraction results with the real annotation data, and by counting the pixel classification results, three indicators of PA, MPA, and MIoU are directly calculated: The calculation formula of PA is as follows: The calculation formula of MPA is as follows: The calculation formula of MIoU is as follows: Where: TP is the true positive, FP is the false positive, FN is the false negative, TN is the true negative, and P i is the pixel accuracy for each category.
Citation Information
Cited By
Method and device for measuring volume of material pile in construction scene
CN121259077A
A method and device for measuring the volume of material piles in a construction scenario
CN121259077B