Road surface pit detection method and device
By performing plane fitting, rotation and semantic eigenvalue screening on pavement pothole point cloud data, the problem of difficult to guarantee the accuracy of pavement pothole detection in the prior art is solved, and higher detection accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510227015.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-20
AI Technical Summary
It is difficult for the prior art to accurately detect pavement holes, especially under the influence of factors such as lighting conditions, shooting angles and surface materials of objects, the accuracy of reconstructing three-dimensional shapes in two-dimensional images is difficult to ensure.
By plane fitting the acquired pothole point cloud data, the target fitting plane and normal vector are determined, the fitting plane is rotated to the horizontal plane, the geometric features and elevation difference between the target point and the neighbor point cloud are obtained, weighted aggregation is performed to determine the semantic eigenvalue, the pothole point cloud collection is selected, and the depth and volume parameters of the pothole are determined based on this.
The accuracy and reliability of pavement pit detection are improved. Through the screening and weighted aggregation of semantic eigenvalues, noise interference is effectively reduced and the overall accuracy of point cloud data is improved.
Smart Images

Figure CN120178265A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to a method and device for detecting road potholes. Background Art
[0002] In recent years, with the continuous leap of sensor technology, portable laser scanners have achieved efficient and high-precision data acquisition capabilities. This technological innovation has even penetrated into smartphone devices such as the iPhone, marking a new height of technology popularization. Such devices enable relevant personnel to quickly and comprehensively scan and detect road potholes at relatively low costs in areas such as sidewalks, low-grade roads, or tree-lined roads where traditional detection methods are difficult to reach or have limited efficiency. This technology not only demonstrates high popularity but also greatly broadens the application scenarios with its portability advantage.
[0003] However, images are essentially two-dimensional. Even if the images contain clues such as perspective and shadows, these clues are often affected by various factors such as lighting conditions, shooting angles, and object surface materials, so it is difficult to accurately interpret them. And reconstructing the three-dimensional shape of an object from a two-dimensional image is a complex process that requires precise algorithms and a large amount of computing resources. Even if the three-dimensional shape of the pothole can be reconstructed, due to the limitations of image information, the accuracy of the reconstruction result is difficult to guarantee. Summary of the Invention
[0004] In view of this, the present invention proposes a method and device for detecting road potholes.
[0005] The technical solution of the present invention is implemented as follows: In the first aspect of the present invention, a method for detecting road potholes is provided, including:
[0006] Performing plane fitting on the acquired pothole point cloud data to determine a target fitting plane and a target normal vector corresponding to the target fitting plane;
[0007] Rotating the target fitting plane to the original space horizontal plane according to the included angle between the target normal vector and the z-axis direction of the original space coordinate system to obtain a corrected point cloud set;
[0008] Obtaining geometric features between a target point and neighbor point clouds, determining the weight corresponding to each neighbor point cloud by using the elevation difference between the target point and the neighbor point clouds, obtaining the semantic feature value of the target point through weighted aggregation, and screening out a pothole point cloud set from the corrected point cloud set based on the semantic feature value; the target point is any point in the corrected point cloud set;
[0009] Determining the depth parameter and volume parameter of the pothole based on the pothole point cloud set.
[0010] Based on the above technical solutions, preferably, the method of performing plane fitting on the obtained pothole point cloud data to determine the target fitting plane and the target normal vector corresponding to the target fitting plane includes:
[0011] Select at least three sets of point data from the pothole point cloud data, determine at least two sets of point cloud vectors corresponding to the at least three sets of point data, and calculate the cross product of the point cloud vectors to determine the target normal vector;
[0012] Obtain the distance parameter of each point in the pothole point cloud data to the candidate plane, and determine the inliers as the points whose distance parameters do not exceed the preset threshold; the candidate plane is any plane between the two farthest point clouds in the direction of the target normal vector;
[0013] Select the candidate plane with the largest number of inliers from the multiple candidate planes as the target fitting plane.
[0014] Based on the above technical solutions, preferably, the method of rotating the target fitting plane to the original space horizontal plane according to the angle between the target normal vector and the z-axis direction of the original space coordinate system to obtain the corrected point cloud set includes:
[0015] Determine the rotation matrix R according to the angle between the target normal vector and the z-axis direction of the original space coordinate system:
[0016]
[0017] where I is a 3×3 identity matrix, θ is the angle between the target normal vector and the z-axis direction of the original space coordinate system, and [v] × is the rotation axis, v = (v x , v y , v z ), and the rotation axis is the cross product of the target normal vector and the z-axis of the original space coordinate system;
[0018] Rotate the target fitting plane to the original space horizontal plane based on the rotation matrix to obtain the corrected point cloud set.
[0019] Based on the above technical solutions, preferably, the method of obtaining the geometric features between the target point and the neighbor point cloud includes obtaining the geometric features by constructing an umbrella surface; the principle of the umbrella surface is:
[0020]
[0021] where α is the direction of the umbrella surface, m is the number of planes in the umbrella surface, and n i represents the normal vector of the i-th plane.
[0022] Based on the above technical solutions, preferably, determining the weight corresponding to each of the neighboring point clouds by using the elevation difference between the target point and the neighboring point clouds, and obtaining the semantic feature value of the target point through weighted aggregation includes:
[0023] Obtaining the elevation difference between the target point and each of the neighboring point clouds;
[0024] Based on the magnitude of the elevation difference, matching a corresponding weight to each of the neighboring point clouds; the weight corresponding to the neighboring point cloud with a large elevation difference is greater than the weight corresponding to the neighboring point cloud with a small elevation difference;
[0025] Performing weighted aggregation based on the geometric features between the target point and the neighboring point clouds and the weights to obtain the semantic feature value of the target point.
[0026] Based on the above technical solutions, preferably, performing weighted aggregation based on the geometric features between the target point and the neighboring point clouds and the weights to obtain the semantic feature value of the target point includes using the following formula to perform weighted aggregation on the geometric features based on the weights to obtain the semantic feature value:
[0027] h θ (x i , x j ) = h θ (α ij (x j - x i )) + h φ (x i );
[0028] Wherein, x i is the target point; x j is the neighboring point; α ij is the weight corresponding to the neighboring point; h is a non - linear function, and θ and φ are trainable parameters.
[0029] Based on the above technical solutions, preferably, determining the depth parameter and volume parameter of the pothole based on the pothole point cloud set includes:
[0030] Determining the difference between the maximum point cloud elevation and the minimum point cloud elevation in the pothole point cloud set as the depth parameter of the pothole;
[0031] Simulating the pothole using the voxel method to determine the volume parameter of the pothole.
[0032] Even more preferably, a second aspect of the present invention provides a road surface pothole detection device, including: a plane fitting module, a rotation correction module, a semantic segmentation module, and a determination module; wherein,
[0033] The plane fitting module is configured to perform plane fitting on the acquired pothole point cloud data to determine a target fitting plane and a target normal vector corresponding to the target fitting plane;
[0034] The rotation correction module is configured to rotate the target fitting plane to the original space horizontal plane according to the angle between the target normal vector and the z-axis direction of the original space coordinate system, so as to obtain a corrected point cloud set;
[0035] The semantic segmentation module is configured to obtain geometric features between a target point and neighbor point clouds, and determine weights corresponding to each neighbor point cloud by using the elevation difference between the target point and the neighbor point clouds, obtain a semantic feature value of the target point through weighted aggregation, and screen out a pothole point cloud set from the corrected point cloud set based on the semantic feature value; the target point is any point in the corrected point cloud set;
[0036] The determination module is configured to determine the depth parameter and volume parameter of the pothole based on the pothole point cloud set.
[0037] More preferably, a third aspect of the present invention provides an electronic device, including a processor and a memory; the memory stores a computer program, wherein the computer program, when executed by the processor, implements the road surface pothole detection method described in the first aspect.
[0038] More preferably, a fourth aspect of the present invention provides a computer storage medium, on which a computer program is stored, wherein the computer program, when executed by a processor, implements the road surface pothole detection method described in the first aspect.
[0039] A road surface pothole detection method of the present invention has the following beneficial effects compared with the prior art:
[0040] 1. By obtaining geometric features between a target point and neighbor point clouds, combining the elevation difference between the target point and the neighbor point clouds to determine weights corresponding to each neighbor point cloud, and performing weighted aggregation, a semantic feature value of the target point is obtained through weighted aggregation, and based on this, the semantic attributes of pothole points, ground points, and noise points are differentiated, improving the accuracy of road surface pothole detection.
[0041] 2. Plane fitting is performed on the pothole point cloud data to obtain a target normal vector corresponding to the target fitting plane, and then the target fitting plane is rotated according to the angle between the target normal vector and the z-axis direction of the original space coordinate system, so as to realize the overall rotation of the pothole point cloud data and complete the correction transformation of the pothole point cloud data, improving the accuracy and reliability of the point cloud set.
[0042] 3. By screening the neighboring points based on the elevation difference between the target point and the neighboring point clouds, the interference of noise points is excluded, the difficulty of weighted aggregation is reduced, and the finally selected neighboring nodes are more representative, thereby improving the reliability of the aggregation result. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0044] Figure 1 It is a schematic flowchart of a road pothole detection method provided by an embodiment of the present invention;
[0045] Figure 2 It is a schematic structural diagram of the YOLOv8 network model provided by an embodiment of the present invention;
[0046] Figure 3 It is a schematic structural diagram of the geometric features provided by an embodiment of the present invention;
[0047] Figure 4 It is a comparison diagram of different neighbor search strategies provided by an embodiment of the present invention;
[0048] Figure 5 It is a schematic flowchart of the working process of the neighbor explorer provided by an embodiment of the present invention;
[0049] Figure 6 It is a schematic diagram of the principle of the voxel method provided by an embodiment of the present invention;
[0050] Figure 7 It is a schematic structural diagram of a road pothole detection device provided by an embodiment of the present invention;
[0051] Figure 8 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0053] In some embodiments, as Figure 1 shown, Figure 1Schematic flowchart of a road pothole detection method provided by an embodiment of the present invention; A road pothole detection method provided by the present invention includes:
[0054] S110. Perform plane fitting on the acquired pothole point cloud data to determine a target fitting plane and a target normal vector corresponding to the target fitting plane.
[0055] Here, the pothole to be detected can be continuously scanned around for about 10 seconds based on the lidar mode of the Polycam software installed in the system, and the point cloud data is downsampled to 5000 points. Each point in the point cloud data p contains its three-dimensional information in space, that is, p i =(x i , y i , z i ). Reasonable downsampling of the acquired pothole point cloud data can reduce the data volume while retaining key information, which helps to maintain the accuracy of the data while reducing the data complexity.
[0056] Plane fitting based on the pothole point cloud data mainly involves finding an optimal fitting plane, that is, the target fitting plane, from the discrete three-dimensional point cloud data, so that the sum of the squares of the distances from these points to the target fitting plane is the smallest. The target fitting plane can be represented by a plane equation, such as ax + by + cz + d, where a, b, c are the components of the normal vector of the plane, and d is the constant term. The plane normal vector can be determined by the coefficients a, b, c, that is, the normal vector is (a, b, c).
[0057] In some embodiments, S110. Perform plane fitting on the acquired pothole point cloud data to determine a target fitting plane and a target normal vector corresponding to the target fitting plane, including:
[0058] Select at least three sets of point data from the pothole point cloud data, determine the corresponding at least two sets of point cloud vectors based on the at least three sets of point data, and calculate the cross product of the point cloud vectors to determine the target normal vector;
[0059] Obtain the distance parameter of each point in the pothole point cloud data to the candidate plane, and determine the points with the distance parameter not exceeding the preset threshold as inliers; the candidate plane is any plane between the two point clouds farthest in the direction of the target normal vector;
[0060] Screen out the candidate plane with the largest number of inliers from multiple candidate planes as the target fitting plane.
[0061] In this embodiment, the RANSAC (Random Sample Consensus) algorithm can be used for plane fitting. The RANSAC algorithm is an iterative method for estimating a parametric model from a dataset containing a large amount of noise and outliers. By randomly selecting a set of data points to assume a model, and calculating the distances from other points to this model to determine the inliers that meet the model requirements, and then re - estimating the model using all the inliers, and repeating this process until the best model is found. For the pothole point cloud data, using the RANSAC algorithm can effectively eliminate the influence of noise and outliers.
[0062] In one example, randomly select 3 points p1=(x1, y1, z1), p2=(x2, y2, z2), p3=(x3, y3, z3) from the pothole point cloud data. These 3 points can be used to represent the fitting plane. The expression of the fitting plane is ax + by + cz + d = 0. Construct two vectors: from point p1=(x1, y1, z1) to point p2=(x2, y2, z2) and from point p1=(x1, y1, z1) to point p3=(x3, y3, z3), two vectors can be obtained: v1=(x2 - x1, y2 - y1, z2 - z1), v2=(x3 - x1, y3 - y1, z3 - z1). Calculate the cross - product of the two vectors, and the normal vector n=(a, b, c) of the plane can be obtained. The normal vector of the plane can be expressed as: n = v1×v2. To determine the parameter d, substitute any known point, such as p1=(x1, y1, z1). According to a·x1 + b·y1 + c·z1 + d = 0, it can be obtained that: d=-(a·x1 + b·y1 + c·z1). For each point p i =(x i , y i , z i ) in the point cloud data, calculate the distance to this fitting plane Select a preset threshold τ. For each point, if D i >τ, it is called an outlier, otherwise it is an inlier, and count the number of outliers and inliers respectively. Repeat the above steps M times, and select the result with the most inliers as the final target fitting plane. The number of repetitions M is calculated by the following formula: In the formula, e is the outlier rate, that is, the probability that a point in the data is an outlier, s is the minimum number of points required to determine the model, and p is the confidence level of obtaining a good model. In this example, the values of each parameter can be τ = 0.03, e = 0.7, p = 0.99, s = 3. While ensuring sufficient fitting accuracy requirements, the amount of calculation is reduced.
[0063] S120, according to the angle between the target normal vector and the z - axis direction of the original space coordinate system, rotate the target fitting plane to the original space horizontal plane to obtain the corrected point cloud set.
[0064] In this embodiment, the rotation axis can be determined first through the target normal vector and the z-axis direction of the original space coordinate system, and thus the included angle can be determined. Based on this, the target fitting plane is rotated to the original space horizontal plane. Correspondingly, the whole of the pothole point cloud data is rotated, so as to obtain the corrected point cloud set.
[0065] In some embodiments, S120, rotating the target fitting plane to the original space horizontal plane according to the included angle between the target normal vector and the z-axis direction of the original space coordinate system to obtain a corrected point cloud set, includes:
[0066] Determine the rotation matrix R according to the included angle between the target normal vector and the z-axis direction of the original space coordinate system:
[0067]
[0068] where I is a 3×3 identity matrix, θ is the included angle between the target normal vector and the z-axis direction of the original space coordinate system, and [v] × is the skew-symmetric matrix of the rotation axis v = (v x , v y , v z ), and the rotation axis is the cross product of the target normal vector and the z-axis of the original space coordinate system;
[0069] Rotate the target fitting plane to the original space horizontal plane based on the rotation matrix to obtain a corrected point cloud set.
[0070] In this embodiment, the rotation axis of the rotation operation can be calculated in advance. The rotation axis is the cross product of the normal vector n and the z-axis, that is:
[0071]
[0072] This indicates that the rotation axis is in the xy plane. It should be noted that the vector v must be normalized.
[0073] The rotation angle θ can be calculated through the included angle between the normal vector and the z-axis. The cosine value of the included angle is the included angle between the z component of the normal vector and the unit vector:
[0074]
[0075] On this basis, the rotation matrix R and the skew-symmetric matrix [v] of the rotation axis v can be determined × :
[0076]
[0077] Rotating the original pothole point cloud data p = (x, y, z) based on the rotation matrix R can obtain the corrected point cloud set, greatly simplifying the registration process and improving the registration accuracy.
[0078] S130. Obtain the geometric features between the target point and the neighboring point clouds, determine the weight corresponding to each neighboring point cloud using the elevation difference between the target point and the neighboring point clouds, obtain the semantic feature value of the target point through weighted aggregation, and screen out the pothole point cloud set from the corrected point cloud set based on the semantic feature value; the target point is any point in the corrected point cloud set.
[0079] The geometric feature is to describe the geometric shape formed by a point and its surrounding neighboring points. The geometric features between the same point and different neighboring point clouds often vary in shape and direction. The elevation difference refers to the vertical distance or height difference between two points. For the pothole scenario, the elevation feature carries more information: the elevation difference between the ground point and its neighbors is small, while the elevation difference between the pothole point and its neighbors is large. The semantic feature value can be a feature value obtained by numericalizing the comprehensive elevation difference. According to the range interval where the semantic feature value is located, the type of the corresponding point can be determined.
[0080] In some embodiments, for S130, obtaining the geometric features between the target point and the neighboring point clouds includes obtaining the geometric features by constructing an umbrella surface; the principle of the umbrella surface is:
[0081]
[0082] where α is the direction of the umbrella surface, m is the number of planes in the umbrella surface, and n i represents the normal vector of the i-th plane.
[0083] The umbrella curvature is obtained by multiplying and summing the direction vectors of the surrounding neighboring points and the target point with the normal vector of the target point. However, the normal vector information is not included in the point cloud data, and the umbrella curvature cannot be directly used. Therefore, the normal vector of each plane in the umbrella surface is used to represent the overall direction. During the calculation process, the neighboring points are projected onto the xy plane, and triangles are constructed clockwise from 0° to 359°. At the same time, to maintain the local consistency of the normal direction, the target point is projected onto the line where each normal vector is located. If the projected point is in the reverse region of the normal vector direction, the normal vector direction is reversed. By observing and analyzing the shape of the umbrella surface, geometric features such as symmetry, smoothness, and curvature change can be extracted.
[0084] In some embodiments, for S130, determining the weight corresponding to each neighboring point cloud using the elevation difference between the target point and the neighboring point clouds and obtaining the semantic feature value of the target point through weighted aggregation includes:
[0085] Obtain the elevation difference between the target point and each neighboring point cloud;
[0086] Match corresponding weights to each neighbor point cloud based on the magnitude of the elevation difference; the weight corresponding to the neighbor point cloud with a large elevation difference is greater than the weight corresponding to the neighbor point cloud with a small elevation difference;
[0087] Perform weighted aggregation based on the geometric features and weights between the target point and the neighbor point clouds to obtain the semantic feature value of the target point.
[0088] In this embodiment, the point cloud space where the target point is located can be pre-divided into a series of voxels, and only one point, such as the mass point of the voxel, is retained in each voxel, so as to reduce the computational amount while ensuring the geometric features of the point cloud. On this basis, assign corresponding weights to the neighbor points according to the magnitude of the elevation difference between the target point and each neighbor point cloud. The greater the elevation difference, the greater the weight. By giving a higher weight to the points with a larger elevation difference, the significant features in the terrain can be captured more sensitively.
[0089] In some embodiments, S130, perform weighted aggregation based on the geometric features and weights between the target point and the neighbor point clouds to obtain the semantic feature value of the target point, including using the following formula to perform weighted aggregation on the geometric features based on the weights to obtain the semantic feature value:
[0090] h θ (x i , x j ) = h θ (α ij (x j - x i )) + h φ (x i );
[0091] Among them, x i is the target point; x j is the neighbor point; α ij is the weight corresponding to the neighbor point; h is a non-linear function, and θ and φ are trainable parameters.
[0092] In an alternative embodiment, semantic segmentation can be performed on the corrected point cloud set through PointPSSN to screen out the pothole point cloud set. Please refer to Figure 2 , Figure 2 which is the structural schematic diagram of the PointPSSN model provided by the embodiment of the present invention; referring to the classic PointNet++ model, a hierarchical structure including an encoder and a decoder is used.
[0093] Before the encoder, a geometric feature extractor is added to extract the geometric features between the target point and the neighbor points. For the semantic segmentation task of potholes, please refer to Figure 3 , Figure 3 which is the structural schematic diagram of the geometric features provided by the embodiment of the present invention; the geometric shape formed by the ground point and its neighbor points is vertically upward, such asFigure 3 as shown in (b); while the pit edge points and pit points are obliquely upward, as Figure 3 shown in (a). By fitting the surface information around the target points, the geometric features carried by them are extracted, and the feature information of each point is strengthened, so that the feature dimension of each point can be expanded from 3D to 6D.
[0094] For the encoder part (SetAbstraction), it mainly includes five small parts: downsampling layer, neighbor finder, feature aggregation layer, MLPs, and Reduction. The functions of each part are as follows:
[0095] (1) Downsampling layer, which is configured to downsample the overall point cloud to select the target points for aggregating neighbor information. This process uses FarthestPoint Sampling (FPS). Specifically, its working steps are as follows: First, randomly select a point in the original point cloud as the first target point; repeat the following two steps until the specified number of target points have been found; calculate the distance between all points and the nearest target point in the point cloud data; select the point with the largest distance as the new target point.
[0096] (2) Neighbor finder, which is configured to find more suitable neighbors for the target points and provide weighted information according to importance. For the pit scenario, the elevation feature carries more information: the elevation difference between the ground points and their neighbors is small, while the elevation difference between the pit points and their neighbor points is large. Please refer to Figure 4 , Figure 4 which is a comparison chart of different neighbor finding strategies provided by the embodiments of the present invention; for dense point clouds, there will be too many neighbors that are too close to the point itself, resulting in difficulty in aggregating effective information, as Figure 4 shown in (b), while the neighbor selection scheme as shown in Figure 4 (c) can effectively aggregate more feature information. Therefore, it is particularly important to select and focus on aggregating neighbor points carrying important features from the surrounding points.
[0097] In an exemplary example, please refer to Figure 5 , Figure 5 which is a schematic diagram of the working process of the neighbor finder provided by the embodiments of the present invention. The neighbor point exploration process can be as follows: First, find 2×K neighbors for the target point as its receptive field, and then subtract the elevation coordinate value of the neighbor point from the coordinate value of the target point and take the absolute value to obtain the elevation difference γ of the neighbor. Sort γ in descending order and select the first K neighbor points as the final neighbors of the target point.
[0098] To calculate the neighbor weight of each target point, the elevation difference γ is used as the initial weight c ij , and after normalization, the final weight α is obtained ij, and finally perform weighted aggregation on neighbor features. Specifically, first set x i is the target point, γ ij is the jth neighbor x of target point i ij The elevation difference with the target point. In order to align the comparison of the attention coefficients between the neighbors of different points, the SoftMax function is used to normalize the coefficients of all neighbors to each point, that is, the weight value of each neighbor point, calculated as follows:
[0099]
[0100] Where: K is the final number of neighbors.
[0101] (3) Feature aggregation layer: This layer is configured to aggregate the neighbor features found by the neighbor finder for the target point and perform weighted aggregation according to the weight information provided by the neighbor finder. The aggregation formula is: θ (x i , x j )=h θ (α ij (x j -x i ))+h φ (x i ).
[0102] In the formula, x i is the target point; x j For x i Neighbor point of ij Find the target point x for the neighbor i Neighbor x j weight; h is a nonlinear function, which is a single-layer fully connected layer in this article; θ and φ are trainable parameters.
[0103] (4) MLPs, this layer is configured to perform feature mining on the features aggregated by the target points through a multi-layer perceptron. Specifically, the formula of this layer is as follows: y = w x + b; where x is the input target point feature information matrix, whose size is n × c1, n is the number of input points, and c1 is the feature dimension carried by the input point; w is a learnable weight matrix, whose size is c1 × c2; b is a learnable bias vector, whose length is c1; y is the output information, whose size is n × c2.
[0104] (5) Reduction: This layer is configured to output the target point.
[0105] There are a total of 4 SetAbstraction layers and corresponding FeaturePropagation (feature propagation) layers in the PointPSSN model. The downsampling layers reduce the target points to N / 2, N / 8, N / 32, and N / 128 respectively, and the MLPs layers expand the feature information to 64, 128, 256, and 512 dimensions respectively.
[0106] Here, the FeaturePropagation layer is used to decode the results obtained by the SetAbstraction layer. Each layer mainly includes three parts:
[0107] (1) Inverse interpolation. In order to restore the point cloud p2 obtained by downsampling in the SetAbstraction layer to the number of points of p1 before downsampling, for each point x in p2, find the k points in p1 that are closest to it in the original point cloud coordinate space, and perform inverse interpolation through the following formula to obtain new point features:
[0108]
[0109] Among them,
[0110] (2) Concat. This layer uses the idea of skip connection to splice the features obtained by inverse interpolation with the feature information in the corresponding SetAbstraction, so as to retain more local feature information and improve the learning and representation ability of the overall model.
[0111] (3) MLPs. This layer mines the features obtained by inverse interpolation through a multi-layer perceptron. It passes through 4 Feature Propagation layers, corresponding to the SetAbstraction layer. It should be noted that the point features input to the first FeaturePropagation layer are the results spliced from the features of the penultimate layer and the second-to-last layer of Set Abstraction, which can effectively retain multi-scale features and thus improve the performance of the model. The overall model passes through four FeaturePropagation layers. Each layer gradually restores the number of points to N / 32, N / 8, N / 2, and N, and the feature dimensions gradually become 256, 256, 128, and 128. Finally, the output of the fully connected layer is an N×2 matrix, representing the semantic segmentation result of each point finally. For each point, if the value of the first dimension is larger, it is classified as "other points", otherwise it is "pothole points".
[0112] S140. Determine the depth parameter and volume parameter of the pothole based on the pothole point cloud set.
[0113] In some embodiments, in S140, determining the depth parameter and volume parameter of the pothole based on the pothole point cloud set includes:
[0114] Determining the difference between the maximum point cloud elevation and the minimum point cloud elevation in the pothole point cloud set as the depth parameter of the pothole;
[0115] Using the voxel method to simulate the pothole to determine the volume parameter of the pothole.
[0116] In this embodiment, the depth of the pothole can be obtained by subtracting the minimum elevation from the maximum elevation of the pothole points obtained by semantic segmentation using the PointPSSN model, that is:
[0117] D i = maxz i - minz i
[0118] In the formula, D i is the depth of the i-th pothole; z i is the elevation value of all points in the i-th pothole.
[0119] Use the voxel method to evaluate the volume of the pothole. Specifically, please refer to Figure 6 , Figure 6 which is a schematic diagram of the principle of the voxel method provided by the embodiment of the present invention; the space is divided into a large number of cubes. The cubes with point clouds falling in them are solid, otherwise they are empty voxels. Calculate the number of empty voxel grids between each voxel grid containing the pothole edge points and the ground, and finally sum the volumes of all the empty voxel grids to obtain the final pothole volume. The calculation formula is as follows:
[0120]
[0121] In the formula: V i is the volume of the i-th pothole; n is the number of voxel grids containing the pothole edge points; b j is the number of empty voxel grids between the j-th voxel grid containing the pothole edge points and the ground; v is the volume of each voxel grid.
[0122] In an example, comparing the road surface pothole detection method with the current state-of-the-art PointMLP model, the accuracy, F1-score, and intersection over union are improved by 0.229%, 1.314%, and 1.971% respectively. In addition, 14 potholes were scanned outside the dataset and tested. Compared with the true measurement values, the average error of the calculated values of the pothole depth by this model is 9.08%, and the average error of the calculated values of the volume is 9.04%, which can effectively quantify the three-dimensional information of the pothole.
[0123] In some embodiments, please refer to Figure 7 ,Figure 7 The figure is a schematic structural diagram of a road surface pothole detection device provided by an embodiment of the present invention. The present invention provides a road surface pothole detection device 700, including: a plane fitting module 710, a rotation correction module 720, a semantic segmentation module 730, and a determination module 740; wherein,
[0124] The plane fitting module 710 is configured to perform plane fitting on the acquired pothole point cloud data to determine a target fitting plane and a target normal vector corresponding to the target fitting plane;
[0125] The rotation correction module 720 is configured to rotate the target fitting plane to the original space horizontal plane according to the included angle between the target normal vector and the z-axis direction of the original space coordinate system to obtain a corrected point cloud set;
[0126] The semantic segmentation module 730 is configured to acquire geometric features between a target point and neighboring point clouds, determine weights corresponding to each neighboring point cloud by using the elevation difference between the target point and the neighboring point clouds, obtain a semantic feature value of the target point through weighted aggregation, and screen out a pothole point cloud set from the corrected point cloud set based on the semantic feature value; the target point is any point in the corrected point cloud set;
[0127] The determination module 740 is configured to determine the depth parameter and volume parameter of the pothole based on the pothole point cloud set.
[0128] In some embodiments, the plane fitting module 710 is specifically configured to:
[0129] Select at least three sets of point data from the pothole point cloud data, determine at least two sets of point cloud vectors corresponding to the at least three sets of point data, calculate the cross product of the point cloud vectors, and determine the target normal vector;
[0130] Obtain the distance parameter from each point in the pothole point cloud data to a candidate plane, and determine the points with the distance parameter not exceeding a preset threshold as inliers; the candidate plane is any plane between the two point clouds farthest in the direction of the target normal vector;
[0131] Screen out the candidate plane with the largest number of inliers from multiple candidate planes as the target fitting plane.
[0132] In some embodiments, the rotation correction module 720 is specifically configured to:
[0133] Determine a rotation matrix R according to the included angle between the target normal vector and the z-axis direction of the original space coordinate system:
[0134]
[0135] wherein, I is a 3×3 identity matrix, θ is the included angle between the target normal vector and the z-axis direction of the original space coordinate system, [v] ×is the skew-symmetric matrix of the rotation axis v=(v x , v y , v z ), and the rotation axis is the cross product of the target normal vector and the z-axis of the original space coordinate system;
[0136] Rotate the target fitting plane to the original space horizontal plane based on the rotation matrix to obtain the corrected point cloud set.
[0137] In some embodiments, the semantic segmentation module 730 is specifically configured to obtain geometric features by constructing an umbrella surface; the principle of the umbrella surface is:
[0138]
[0139] i where α is the direction of the umbrella surface, m is the number of planes in the umbrella surface, and n
[0140]
[0141]
[0142]
[0143]
[0144]
[0145]
[0146]
[0147]
[0148] h θ (x i , x j ) = h θ (α ij (x j - x i )) + h φ (x i );
[0146] where x i is the target point; x j is the neighbor point; α ij is the weight corresponding to the neighbor point; h is a non-linear function, and θ and φ are trainable parameters.
[0147] In some embodiments, the determination module 740 is specifically configured to:
[0148] Determine the depth parameter of the pothole by taking the difference between the maximum point cloud elevation and the minimum point cloud elevation in the pothole point cloud set;
[0149] Simulate the pothole using the voxel method to determine the volume parameter of the pothole.
[0150] It should be noted that the road surface pothole detection device provided in the embodiments of the present application and the road surface pothole detection method provided in the embodiments of the present application are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the foregoing road surface pothole detection method, and the repeated parts will not be elaborated.
[0151] In some embodiments, please refer to Figure 8 , Figure 8 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. An electronic device 800 provided in an embodiment of the present application includes a processor 810 and a memory 820; the memory 820 stores a computer program, wherein the computer program, when executed by the processor, implements the above-mentioned road surface pothole detection method.
[0152] Specifically, the processor 810 may include, for example, a general microprocessor, an instruction set processor, and / or a related chipset and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), and so on. The processor 810 may also include on-board memory for caching purposes. The processor 810 may be a single processing unit or multiple processing units for performing different actions of the method flow according to the embodiments of the present application.
[0153] The memory 820 may be, for example, any medium capable of containing, storing, transmitting, propagating, or transferring instructions. For example, the memory 820 may include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, components, or propagation media. Specific examples of the memory 820 include: magnetic storage devices, such as magnetic tapes or hard disk drives (HDDs); optical storage devices, such as compact discs (CD-ROMs); may also be, for example, random access memory (RAM) or flash memory; and / or wired / wireless communication links.
[0154] The present application also provides a computer-readable medium, on which a computer program is stored, and when the program is executed by the processor, it implements the above-mentioned road surface pothole detection method. The computer-readable medium may be included in the device / device / system described in the above embodiments; or it may exist separately without being assembled into the device / device / system. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiments of the present application is implemented.
[0155] According to an embodiment of the present application, a computer-readable medium may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wired, optical fiber cable, radio frequency signal, etc., or any suitable combination of the above.
[0156] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly recited in the present application. In particular, without departing from the spirit and teachings of the present application, the features recited in the various embodiments and / or claims of the present application can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present application. Therefore, the scope of the present application should not be limited to the above embodiments, but should be determined not only by the appended claims, but also by the equivalents of the appended claims. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A road pothole detection method, characterized in that: include: Performing plane fitting on the acquired pothole point cloud data to determine a target fitting plane and a target normal vector corresponding to the target fitting plane; According to the angle between the target normal vector and the z-axis direction of the original space coordinate system, the target fitting plane is rotated to the horizontal plane of the original space to obtain a corrected point cloud set; Acquire geometric features between the target point and the neighbor point cloud, and determine the weight corresponding to each of the neighbor point clouds using the elevation difference between the target point and the neighbor point cloud, obtain the semantic feature value of the target point by weighted aggregation, and filter out the pothole point cloud set from the corrected point cloud set based on the semantic feature value; The target point is any point in the corrected point cloud set; Determine the depth parameter and volume parameter of the pothole based on the pothole point cloud set.
2. The road pothole detection method according to claim 1, characterized in that: The performing plane fitting on the acquired pothole point cloud data to determine a target fitting plane and a target normal vector corresponding to the target fitting plane includes: Selecting at least three examples of point data from the pothole point cloud data, determining at least two corresponding groups of point cloud vectors based on the at least three examples of point data, and performing cross product on the point cloud vectors to determine the target normal vector; Obtaining a distance parameter from each point in the pothole point cloud data to a candidate plane, and determining a point whose distance parameter does not exceed a preset threshold as an interior point; the candidate plane is any plane between the two point clouds farthest in the direction of the target normal vector; The candidate plane with the largest number of inliers is selected from the plurality of candidate planes as the target fitting plane.
3. The road pothole detection method according to claim 1, characterized in that: The step of rotating the target fitting plane to the horizontal plane of the original space according to the angle between the target normal vector and the z-axis direction of the original space coordinate system to obtain a corrected point cloud set includes: According to the angle between the target normal vector and the z-axis direction of the original space coordinate system, the rotation matrix R is determined: Where I is a 3×3 unit matrix, θ is the angle between the target normal vector and the z-axis of the original space coordinate system, [v] × The rotation axis v = (v x , v y , v z ), the rotation axis is the cross product of the target normal vector and the z-axis of the original space coordinate system; The target fitting plane is rotated to a horizontal plane of the original space based on the rotation matrix to obtain a corrected point cloud set.
4. The road pothole detection method according to claim 1, characterized in that: The step of obtaining the geometric features between the target point and the neighbor point cloud includes obtaining the geometric features by constructing an umbrella surface; the umbrella surface principle is: Among them, α is the direction of the umbrella surface, m is the number of planes in the umbrella surface, and n i Represents the normal vector of the i-th plane.
5. The road pothole detection method according to claim 1, characterized in that: The step of determining the weight corresponding to each of the neighbor point clouds by using the elevation difference between the target point and the neighbor point clouds, and obtaining the semantic feature value of the target point by weighted aggregation includes: Obtaining the elevation difference between the target point and each of the neighbor point clouds; Based on the size of the elevation difference, a corresponding weight is matched for each of the neighbor point clouds; the weight corresponding to the neighbor point cloud with a large elevation difference is greater than the weight corresponding to the neighbor point cloud with a small elevation difference; Weighted aggregation is performed based on the geometric features between the target point and the neighbor point cloud and the weight to obtain a semantic feature value of the target point.
6. The road pothole detection method according to claim 5, characterized in that: The weighted aggregation based on the geometric features between the target point and the neighbor point cloud and the weight to obtain the semantic feature value of the target point includes: performing weighted aggregation on the geometric features based on the weight using the following formula to obtain the semantic feature value: h θ (x i ,x j )=h θ (α ij (x j -x i ))+h φ (x i ); Among them, x i is the target point; x j is the neighbor point; α ij is the weight corresponding to the neighbor point; h is a nonlinear function, and θ and φ are trainable parameters.
7. The road pothole detection method according to claim 1, characterized in that: The determining of the depth parameter and volume parameter of the pothole based on the pothole point cloud set includes: Determine the difference between the maximum point cloud elevation and the minimum point cloud elevation in the pothole point cloud set as the depth parameter of the pothole; The pothole is simulated by using a voxel method to determine the volume parameters of the pothole.
8. A road pothole detection device, characterized in that: include: Plane fitting module, rotation correction module, semantic segmentation module and determination module; among them, The plane fitting module is configured to perform plane fitting on the acquired pothole point cloud data, and determine a target fitting plane and a target normal vector corresponding to the target fitting plane; The rotation correction module is configured to rotate the target fitting plane to the horizontal plane of the original space according to the angle between the target normal vector and the z-axis direction of the original space coordinate system to obtain a corrected point cloud set; The semantic segmentation module is configured to obtain geometric features between the target point and the neighbor point cloud, and determine the weight corresponding to each of the neighbor point clouds using the elevation difference between the target point and the neighbor point cloud, obtain the semantic feature value of the target point by weighted aggregation, and filter out the pothole point cloud set from the modified point cloud set based on the semantic feature value; the target point is any point in the modified point cloud set; The determination module is configured to determine the depth parameter and volume parameter of the pothole based on the pothole point cloud set.
9. An electronic device comprising a processor and a memory; the memory stores a computer program, wherein: When the computer program is executed by the processor, the computer program implements the road pothole detection method according to any one of claims 1 to 7.
10. A computer storage medium, characterized in that: A computer program is stored thereon, wherein when the computer program is executed by a processor, the road pothole detection method as claimed in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Porous asphalt concrete gap blockage identification method based on pavement noise signals
CN121049135A
Method for identifying void blockage of porous asphalt concrete based on road noise signal
CN121049135B