Asphalt pavement track detection method and related equipment
By using an improved RAFT-Stereo network and an onboard multi-camera system, the accuracy and cost issues of rutting detection on asphalt pavements in urban highways with weak texture have been resolved, achieving high-precision and low-cost rutting detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CENT SOUTH UNIV
- Filing Date
- 2026-04-01
- Publication Date
- 2026-05-01
AI Technical Summary
Existing methods for detecting rutting on asphalt pavements are unable to achieve millimeter-level reconstruction accuracy in urban highways with weak textures. Traditional equipment is costly, complex to maintain, and cannot meet the needs of high-frequency, low-cost testing.
An improved RAFT-Stereo network is used for stereo matching, combined with an on-board multi-camera data acquisition platform and a lightweight image restoration network. Through disparity estimation and dense point cloud reconstruction, high-precision road rut detection is achieved.
Achieving millimeter-level reconstruction accuracy under weak texture and complex lighting conditions improves measurement reliability, reduces inspection costs, and is suitable for high-frequency inspection of urban highways.
Smart Images

Figure CN121962147A_ABST
Abstract
Description
A method and related equipment for detecting rutting in asphalt pavement Technical Field
[0001] This invention relates to the field of road rutting detection technology, and in particular to a method and related equipment for detecting rutting on asphalt pavement. Background Technology
[0002] Highways are crucial infrastructure supporting the national economy, and their structural performance directly impacts traffic safety and transportation efficiency. Under the combined effects of long-term traffic loads and complex environmental factors, rutting in asphalt pavements has become one of the most common and significantly detrimental forms of structural damage. Studies show that for every 1cm increase in rutting depth, vehicle fuel consumption rises by approximately 2-3%, and the risk of driving in rainy weather is significantly increased, potentially leading to hydroplaning accidents. Therefore, developing high-precision, high-efficiency, and low-cost automated detection methods for rutting is of great significance for ensuring safe road operation and promoting the intelligent development of transportation infrastructure.
[0003] To achieve this goal, road surface inspection technology has evolved from manual to automated inspection. Traditional manual inspection methods rely on professionals using simple tools such as rulers and levels for on-site measurements, which suffers from inherent drawbacks such as low efficiency, strong subjectivity, and poor safety. With technological advancements, automated inspection systems based on active sensing technologies such as laser scanning and infrared imaging have gradually been applied. Among them, the three-dimensional line laser profiler, with its millimeter-level accuracy in single-section measurements, has become the mainstream equipment for regular inspections of high-grade highways. However, these high-precision devices generally suffer from high cost, complex maintenance, and stringent environmental requirements. These factors collectively limit the feasibility of achieving high-frequency, low-cost, and comprehensive inspections in wide-area road networks, making it difficult to meet the daily maintenance and management needs of ultra-large-scale road networks.
[0004] In recent years, the rapid development of computer vision and deep learning technologies has significantly promoted the progress of 3D perception methods. Among them, depth prediction methods based on stereo vision have shown broad application prospects in the field of road surface defect detection due to their advantages such as low hardware cost, high system integration, and dense information acquisition. Compared with active sensing devices such as LiDAR, the cost of binocular camera systems is typically only 1 / 10 to 1 / 20 of theirs, making them more suitable for vehicle-mounted deployment and large-scale applications. Furthermore, they can acquire high-density 3D structural information while meeting detection accuracy requirements. Driven by the demand for low-cost, high-precision road surface inspection equipment, vision-based road surface defect detection has become an important research direction.
[0005] With the development of deep learning, the performance of stereo matching methods in weakly textured scenes has been significantly improved. The end-to-end Pyramid Stereo Matching Network (PSMNet) effectively enhances global context awareness by introducing spatial pyramid pooling and 3D convolutional aggregation; the Group-wise Correlation Stereo Network (GWCNet) further improves the discriminativeness and robustness of feature matching by constructing group cost volumes; the end-to-end hierarchical NAS framework LEAStereo introduces neural architecture search into stereo matching, achieving synergistic optimization of efficiency and accuracy; and ACVNet improves the model's adaptability under complex lighting and weakly textured conditions through attention cost volumes and multi-scale feature fusion. However, in urban highway scenes with common lighting abrupt changes, reflections, and watermark interference, the cross-scene generalization ability and accuracy stability of these methods still need further verification. Summary of the Invention
[0006] This invention provides a method and related equipment for detecting rutting on asphalt pavement, the purpose of which is to achieve millimeter-level reconstruction accuracy in the weak texture environment of real urban roads to improve measurement reliability.
[0007] To achieve the above objectives, this invention provides a method for detecting rutting on asphalt pavement, comprising: Step 1, acquiring a set of original asphalt pavement images of the target road section, and using a portion of the original asphalt pavement image data from the set as a training set; Step 2, training an improved RAFT-Stereo network using the training set to obtain a stereo matching model, and inputting the original asphalt pavement images other than those in the training set into the stereo matching model for disparity estimation, then converting them into a high-precision dense pavement point cloud; Step 3, performing spatial registration and fusion on the high-precision dense pavement point cloud to obtain a fused pavement point cloud; Step 4: Extract cross-sectional lines from the stitched road point cloud to obtain the cross-sectional lines of the target road segment, and use the cross-sectional lines of the target road segment to calculate the rut depth of the target road segment. The improved RAFT-Stereo network includes a feature extraction module for extracting matching features and contextual features, a feature enhancement module for enhancing the global consistency of features, a correlation volume construction module for calculating the matching relationship between features, a recursive update module for predicting disparity increments and correcting pixel correspondences, and an upsampling module for restoring low resolution to the original image resolution.
[0008] Furthermore, the vehicle-mounted multi-camera data acquisition platform collects a set of original asphalt pavement images on the target road section.
[0009] Furthermore, the vehicle-mounted multi-camera data acquisition platform includes: multiple industrial-grade binocular structured light cameras equipped with infrared filters, a multi-device synchronous data acquisition unit, and an aluminum bracket; the aluminum bracket is located at the rear of the inspection vehicle; all industrial-grade binocular structured light cameras are mounted on the aluminum bracket, with the lenses of the industrial-grade binocular structured light cameras pointing vertically downwards toward the road surface; the output terminals of all industrial-grade binocular structured light cameras are electrically connected to the input terminals of the multi-device synchronous data acquisition unit.
[0010] Furthermore, before step 2, the method includes: using a lightweight general image restoration network to perform super-resolution enhancement on the original asphalt pavement images other than the training set, to obtain the enhanced asphalt pavement images.
[0011] Furthermore, the improved RAFT-Stereo network is trained using a training set to obtain a stereo matching model. This process includes: loading the weight parameters of the improved RAFT-Stereo network pre-trained on a large-scale synthetic dataset as initial weight parameters; freezing the parameters of the feature extraction module and the related volume construction module; and using a progressive learning approach to initially train the parameters of the feature enhancement module and the recursive update module to obtain the initially trained RAFT-Stereo network; and fine-tuning the parameters of all modules in the initially trained RAFT-Stereo network on a self-built real-world scene dataset to obtain the stereo matching model.
[0012] Furthermore, step 3 includes: extracting local features from the original asphalt pavement images other than the training set to obtain local feature points; performing feature matching on the local feature points using a feature matching algorithm to obtain a set of matching point pairs; normalizing the color vectors of each matching point pair in the set and calculating the Euclidean distance to obtain the color distance, and selecting valid matching point pairs from the set based on the color distance; combining the internal parameters, external parameters, and depth information of the industrial-grade binocular structured light camera, transforming the valid matching pairs using a back-projection operator to obtain the initial mapping result; estimating the rigid body transformation parameters of the high-precision dense pavement point cloud based on the initial mapping result to obtain the rigid body transformation parameter estimation result; and optimizing and fusing the high-precision dense pavement point cloud based on the rigid body transformation parameter estimation result to obtain the fused pavement point cloud.
[0013] This invention also provides an asphalt pavement rutting detection device, comprising: an acquisition module for acquiring a set of original asphalt pavement images of a target road section, and using a portion of the original asphalt pavement image data from the original asphalt pavement image set as a training set; a disparity estimation module for training an improved RAFT-Stereo network using the training set to obtain a stereo matching model, and inputting original asphalt pavement images other than those in the training set into the stereo matching model for disparity estimation, which is then converted into a high-precision dense pavement point cloud; and a fusion module for spatial registration and fusion of the high-precision dense pavement point cloud to obtain a fused pavement point cloud; and a calculation module. The calculation module is used to extract the cross-sectional lines of the stitched road point cloud to obtain the cross-sectional lines of the target road segment, and to calculate the rut depth of the target road segment using the cross-sectional lines of the target road segment. The improved RAFT-Stereo network includes a feature extraction module for extracting matching features and contextual features, a feature enhancement module for enhancing the global consistency of features, a correlation volume construction module for calculating the matching relationship between features, a recursive update module for predicting disparity increments and correcting pixel correspondences, and an upsampling module for restoring low resolution to the original image resolution.
[0014] The present invention also provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a method for detecting rutting on asphalt pavement.
[0015] The present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements a method for detecting rutting on asphalt pavement.
[0016] The above-mentioned solution of the present invention has the following beneficial effects: The present invention uses a portion of the original asphalt pavement image data of the original asphalt pavement image set of the target road section as a training set to train the improved RAFT-Stereo network to obtain a stereo matching model. The original asphalt pavement images other than those in the training set are input into the stereo matching model for disparity estimation and then converted into a high-precision dense road point cloud. Spatial registration and fusion are performed on the high-precision dense road point cloud to obtain a fused road point cloud. Cross-sectional lines are extracted from the stitched road point cloud to obtain the cross-sectional lines of the target road section. The rutting depth of the target road section is calculated using the cross-sectional lines of the target road section. Compared with the prior art, the present invention uses disparity estimation to complete the three-dimensional reconstruction of the road point cloud, which can not only accurately restore the road geometry, but also maintain high matching stability under weak texture and complex lighting conditions, thereby achieving millimeter-level reconstruction accuracy in the weak texture environment of real urban highways to improve measurement reliability.
[0017] Other beneficial effects of the present invention will be described in detail in the following detailed description section. Attached Figure Description
[0018] Figure 1 is a flowchart of an embodiment of the present invention; Figure 2 is a structural diagram of the improved RAFT-Stereo network of an embodiment of the present invention; Figure 3 is a result diagram of rut depth distribution in an embodiment of the present invention; Figure 4 is a structural diagram of the asphalt pavement rut detection device in an embodiment of the present invention; Figure 5 is a structural diagram of the terminal equipment in an embodiment of the present invention. Detailed Implementation
[0019] To make the technical problems, solutions, and advantages of this invention clearer, a detailed description will be provided below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0020] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0021] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a locking connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0022] Furthermore, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0023] This invention addresses existing problems by providing a method and related equipment for detecting rutting on asphalt pavements.
[0024] As shown in Figure 1, an embodiment of the present invention provides a method for detecting rutting on asphalt pavement, comprising: Step 1, acquiring a set of original asphalt pavement images of the target road segment, and using a portion of the original asphalt pavement image data from the original asphalt pavement image set as a training set; Step 2, training an improved RAFT-Stereo network using the training set to obtain a stereo matching model, and inputting the original asphalt pavement images other than those in the training set into the stereo matching model for disparity estimation, and then converting them into a high-precision dense pavement point cloud; Step 3, performing spatial registration and fusion on the high-precision dense pavement point cloud to obtain a fused pavement point cloud; Step 4, extracting cross-sectional lines from the stitched pavement point cloud to obtain the cross-sectional lines of the target road segment, and using the cross-sectional lines of the target road segment to calculate the rutting depth of the target road segment.
[0025] In order to achieve efficient three-dimensional detection of road ruts, this embodiment of the invention uses an on-board multi-camera data acquisition platform to collect a set of original asphalt pavement images on the target road section.
[0026] Specifically, the vehicle-mounted multi-camera data acquisition platform includes: multiple industrial-grade binocular structured light cameras equipped with infrared filters, a multi-device synchronous data acquisition unit, and an aluminum bracket; the aluminum bracket is located at the rear of the inspection vehicle; all industrial-grade binocular structured light cameras are mounted on the aluminum bracket, with the lenses of the industrial-grade binocular structured light cameras pointing vertically downwards toward the road surface; the output terminals of all industrial-grade binocular structured light cameras are electrically connected to the input terminals of the multi-device synchronous data acquisition unit.
[0027] It should be noted that the industrial-grade binocular structured light camera consists of four Gemini 335L industrial-grade binocular structured light cameras. It adopts a 95mm long baseline design and combines active infrared speckle projection with passive binocular vision. It can achieve a spatial relative accuracy of better than ≤0.3% within a working range of 0.25-6m. It is also equipped with an infrared filter, which significantly improves the imaging stability and depth quality under outdoor high dynamic lighting conditions.
[0028] To ensure spatiotemporal consistency among multi-view images, this embodiment of the invention utilizes a multi-device synchronous data acquisition unit as the master clock source. Through the dedicated 8-pin synchronization interface of the Gemini 335L industrial-grade binocular structured light camera, a unified hardware trigger pulse and a timestamp clearing signal are sequentially cascaded. This hardware-level triggering and signal cascading mechanism achieves high-precision data synchronization with an exposure center time deviation of ≤5ms for all cameras, and adds a unified high-precision hardware timestamp to all output raw asphalt pavement image data.
[0029] In this embodiment of the invention, the aluminum bracket is designed as a triangular truss structure, which utilizes the high rigidity of the triangular structure to suppress the impact of vehicle vibration on imaging quality during travel. The industrial-grade binocular structured light camera is installed as follows: four industrial-grade binocular structured light cameras are fixed vertically downward on the aluminum bracket behind the detection vehicle, with the lens 1 meter above the road surface. It can simultaneously acquire original asphalt pavement image data at a resolution of 848×480 and a frame rate of 60fps at a driving speed of 40-60km / h.
[0030] Because asphalt pavement images generally have limited resolution, weak texture features, low contrast, and strong surface structure repetition, performing stereo matching directly on the original resolution image often faces problems such as insufficient matching cost discrimination and unstable disparity estimation. Therefore, in order to improve the local texture discernibility of the input image, this embodiment of the invention uses a lightweight general image restoration network to perform super-resolution enhancement on the original asphalt pavement images other than the training set before performing step 2, resulting in an enhanced asphalt pavement image. The spatial resolution of the enhanced asphalt pavement image is increased from 848×480 to 1696×960. Its core purpose is to restore and enhance the high-frequency details and local gradient information in the image, providing the improved RAFT-Stereo network with clearer texture and more discernible features, which helps to improve the matching robustness and disparity estimation smoothness of weak texture areas and asphalt pavement areas with continuous deformation features.
[0031] In the most preferred embodiment of the present invention, before inputting the enhanced asphalt pavement image into the improved RAFT-Stereo network, the input image data also needs to be normalized to linearly map its pixel values to the range of [-1,1], thereby ensuring the numerical stability of the network during training and inference.
[0032] Specifically, as shown in Figure 2, the improved RAFT-Stereo network includes a feature extraction module for extracting matching features and contextual features, a feature enhancement module for enhancing the global consistency of features, a correlation volume construction module for calculating the matching relationship between features, a recursive update module for predicting disparity increments and correcting pixel correspondences, and an upsampling module for restoring low resolution to the original image resolution.
[0033] In this embodiment of the invention, the feature extraction module employs a dual-branch feature extraction strategy, constructing matching features and context features separately. The matching features are extracted by the feature extraction network and used for similarity calculation between candidate left and right views. The context features are used to initialize the hidden state of the recursive update module, providing prior information for subsequent iterative optimization. The feature enhancement module is a spatial attention module, embedded in the medium-resolution feature extraction layer of the feature extraction module. It explicitly models the relationships between a wide range of pixels through a self-attention mechanism, thereby enhancing the global consistency of features. Simultaneously, an adaptive window attention layer is set in the feature enhancement module to dynamically select between global attention and window attention modes based on the feature map size. When the feature size is large, the feature map is divided into multiple local windows, and attention calculations are performed within the windows, thus limiting the computational complexity to a controllable range. When the feature size is small, global attention is directly used to fully model long-distance dependencies. The correlation volume construction module is used to construct a one-dimensional correlation volume to characterize the matching relationship between left and right features in the horizontal direction. This correlation volume is grouped at multiple scales. The model is designed to simultaneously capture local fine-grained matching information and large-scale displacement relationships. In this embodiment, the relevant volume construction module performs relevant calculations on features within the disparity candidate range. This avoids the high computational and storage overhead of traditional cost volumes in 3D space, significantly improving computational efficiency. The recursive update module, based on a multi-scale GRU structure, decomposes context features into multiple gated branches. These branches correspond to update gates, reset gates, and candidate states, respectively. In each iteration, the disparity increment is predicted based on the current disparity, relevant volume features, and context state, and the pixel correspondence is gradually corrected, enabling the model to achieve high-precision convergence with fewer iterations. After each recursive update, the model generates an intermediate disparity prediction result and restores the low-resolution disparity map to the original image resolution using an upsampling module. This upsampling process is not a simple interpolation operation but uses a learned weight mask to perform convex combination of disparities in the local neighborhood, effectively suppressing noise while preserving edge details. This significantly improves the reconstruction quality of high-resolution disparity maps in edge and low-texture regions.
[0034] Specifically, the improved RAFT-Stereo network is trained using a training set to obtain a stereo matching model. This includes: loading the weight parameters of the improved RAFT-Stereo network pre-trained on a large-scale synthetic dataset as initial weight parameters; freezing the parameters of the feature extraction module and the related volume construction module; using a progressive learning approach to initially train the parameters of the feature enhancement module and the recursive update module to obtain the initially trained RAFT-Stereo network; and fine-tuning the parameters of all modules in the initially trained RAFT-Stereo network on a self-built real-world scene dataset to obtain the stereo matching model.
[0035] This invention employs a progressive learning approach. By appropriately setting batch size, training steps, and data augmentation methods, the model gradually adapts to the changes in feature distribution brought about by the addition of new structures. During training, a variable batch size is used to reduce memory usage and enhance the stability of gradient updates. Simultaneously, prior data augmentation techniques such as multi-scale spatial transformation and color enhancement are employed to improve the model's generalization ability to complex real-world scenarios. To further improve training efficiency and resource utilization, the system enables a mixed-precision training mode, significantly reducing memory consumption and accelerating training speed while ensuring numerical stability.
[0036] Since a single camera cannot effectively cover the entire lane area while ensuring measurement accuracy, this embodiment of the invention adopts a multi-view binocular camera collaborative acquisition method. For the image overlap rate of about 40% between adjacent binocular cameras, a registration method based on image feature matching and then solving the three-dimensional rigid body transformation is used to spatially register and fuse the high-precision dense road point cloud.
[0037] Specifically, step 3 includes: extracting local features from the original asphalt pavement images other than the training set to obtain local feature points; performing feature matching on the local feature points using a feature matching algorithm to obtain a set of matching point pairs; normalizing the color vectors of each matching point pair in the set and calculating the Euclidean distance to obtain the color distance, and selecting valid matching point pairs from the set based on the color distance; combining the internal parameters, external parameters, and depth information of the industrial-grade binocular structured light camera, transforming the valid matching pairs using a back-projection operator to obtain the initial mapping result; estimating the rigid body transformation parameters of the high-precision dense pavement point cloud based on the initial mapping result to obtain the rigid body transformation parameter estimation result; and optimizing and fusing the high-precision dense pavement point cloud based on the rigid body transformation parameter estimation result to obtain the fused pavement point cloud.
[0038] In this embodiment of the invention, the expression for filtering valid matching point pairs from the set of matching point pairs based on color distance is: ;in, , Both represent the normalized color vectors of the matching point cloud. This represents the distance threshold.
[0039] Specifically, the rigid body transformation parameter estimation for high-precision dense road point clouds based on the initial mapping results is implemented as follows: First, stable feature point pairs are extracted from the overlapping areas of the pre-registered point cloud pairs. Then, the spatial coordinates corresponding to these feature point pairs are used as input, and the least squares optimization method is applied to solve for the rotation matrix and translation vector, i.e., the rigid body transformation parameters. During the solution process, the iterative nearest point algorithm is used to reduce the spatial error between point pairs through iterative optimization, and finally, high-precision rigid body transformation parameter estimation results are obtained, which are used to unify the point clouds to be registered into the reference coordinate system to achieve accurate alignment of the point clouds.
[0040] Specifically, based on the estimation results of rigid body transformation parameters, the high-precision dense road point cloud is optimized and fused to obtain the fused road point cloud. This includes: using the KD tree data structure to efficiently establish the spatial correspondence between the two point clouds; and using the KD tree to quickly retrieve the nearest neighbor point of the first group of point clouds for each point in the second group of point clouds, thereby achieving a one-to-one correspondence between the point clouds.
[0041] In this embodiment of the invention, after completing the spatial registration and fusion of high-precision dense road point clouds, road cross sections are extracted at fixed intervals of 1m along the vehicle travel direction for rut depth analysis. Due to factors such as drainage cross slope design, local defects, and changes in vehicle posture, the original cross section often exhibits overall tilt and local anomalies. In this embodiment of the invention, relatively stable areas at both ends of the cross section are selected as the benchmark. The posture and tilt of the cross section are corrected by rotation transformation. Abnormal cross sections caused by vehicle bumps, lateral offset, accompanying defects, or road debris are removed during the sample construction stage to ensure data reliability. Then, the envelope method is used to automatically analyze each cross section, calculate the rut depth of the left and right wheel tracks respectively, and take the maximum value as the rut index of the cross section, forming the rut depth distribution along the road mileage direction, as shown in Figure 3. As can be seen from Figure 3, the rut depth of the overall road section fluctuates within the range of 3mm to 10mm, with significantly increased rut values in local locations, with the maximum depth exceeding 20mm, indicating obvious local cumulative deformation in this section. The difference between the left and right wheel track depth curves reflects the uneven distribution of lateral load and the asymmetry of the stress on the road structure.
[0042] Specifically, the envelope method is used to automatically analyze each cross-section, and the calculation expressions for the rut depths of the left and right wheel tracks are as follows: ;in, Indicates the depth of the ruts. Represents discrete sampling points. This represents the estimated elevation value. , Represents the cross-sectional curve of the road surface. Indicates elevation error. This represents the upper envelope.
[0043] In order to evaluate the overall performance of different binocular stereo matching methods, this invention selected six representative algorithms covering traditional optimization methods and deep learning models for comparative analysis. All experiments were conducted on a unified hardware and software platform to ensure the fairness and reproducibility of the evaluation results.
[0044] The dataset used in the experiment was a self-built binocular stereo vision dataset, which included synchronously acquired left and right view images and their corresponding ground truth (GT) disparities. The GT disparities were stored in .pfm format, possessing pixel-level accuracy and capable of being directly used for quantitative error calculation. During the evaluation process, error metrics were calculated only within the effective GT region. For invalid disparity regions (such as occlusion or measurement failure regions), they were removed by pruning to ensure fair comparison of different methods within the same effective area. Six indicators were selected from three dimensions: matching accuracy, error distribution, and computational efficiency. Endpoint error (EPE) is the most critical accuracy indicator, used to measure the mean absolute error between predicted and true disparity. Bad3 error rate reflects the proportion of pixels with prediction errors greater than 3 pixels, used to evaluate the robustness of the algorithm when generating large errors. D1 error rate further examines the proportion of pixels with prediction errors exceeding a set threshold, used to measure the algorithm's mismatch situation in complex regions. Root mean square error (RMSE) and mean absolute error (MAE) characterize the stability of the algorithm from the perspectives of overall error distribution and average deviation, respectively. In addition, frames per second (FPS) records the inference speed of the algorithm on a unified test platform, used to evaluate its computational efficiency and real-time potential for actual deployment. The specific results are shown in Table 1 below: Table 1 Quantitative Comparison Results of Different Binocular Stereo Matching Methods
[0045] The overall performance comparison reveals a clear trade-off between accuracy and computational efficiency among different methods. The traditional stereo matching method, SGM, still boasts a significant computational efficiency advantage on CPU platforms (7.57 FPS), but its matching accuracy in real-world complex scenarios is clearly insufficient, with EPE and Bad3 accuracy reaching 10.758 px and 8.17% respectively, failing to meet the application requirements of high-precision 3D reconstruction and detailed road surface morphology analysis.
[0046] Both PSM Net and Gwc Net, deep learning methods based on cost volume aggregation, achieved significant improvements in accuracy compared to SGM. PSM Net demonstrated superior overall error stability, with the lowest RMSE (2.236 px), indicating smaller overall fluctuations in its prediction results. Gwc Net, on the other hand, showed advantages in suppressing significant mismatches, with the lowest D1 score (3.52%), while maintaining relatively high inference speed (1.01 FPS) within the deep learning model, demonstrating a good balance between accuracy and efficiency.
[0047] Furthermore, improved deep learning methods such as LEA Stereo and ACV Net exhibit similar performance levels across multiple metrics, with overall accuracy improvements over PSM Net and Gwc Net. However, the model provided in this embodiment demonstrates the most outstanding performance in cross-domain generalization accuracy, achieving optimal or near-optimal results in key metrics such as EPE and Bad3 (EPE: 4.018px, Bad3: 2.05%), significantly outperforming other comparative methods. This is primarily due to its recursive multi-scale update mechanism based on a relevance pyramid, enabling the network to more effectively address issues such as weak texture regions, occlusion, and noise interference in complex real-world scenes. Although the inference speed of the model provided in this embodiment is relatively low (0.31 FPS), its significant advantages in accuracy and robustness make it more suitable for precision measurement and 3D reconstruction tasks with high reconstruction quality requirements. In summary, the method provided in this embodiment is the most advantageous choice for application scenarios with extremely high accuracy requirements and relatively low real-time requirements.
[0048] This invention uses a portion of the original asphalt pavement image data from the original asphalt pavement image set of the target road segment as a training set to train an improved RAFT-Stereo network, obtaining a stereo matching model. The original asphalt pavement images other than those in the training set are then input into the stereo matching model for disparity estimation and converted into a high-precision dense pavement point cloud. Spatial registration and fusion are performed on the high-precision dense pavement point cloud to obtain a fused pavement point cloud. Cross-sectional lines are extracted from the stitched pavement point cloud to obtain the cross-sectional lines of the target road segment, and the rutting depth of the target road segment is calculated using these cross-sectional lines. Compared with existing technologies, this invention utilizes disparity estimation to complete the 3D reconstruction of the pavement point cloud, which not only accurately restores the pavement geometry but also maintains high matching stability under weak texture and complex lighting conditions. This achieves millimeter-level reconstruction accuracy in the weak texture environment of real urban highways, thereby improving measurement reliability.
[0049] Corresponding to the asphalt pavement rutting detection method described in the above embodiments, as shown in Figure 4, this embodiment of the invention also provides an asphalt pavement rutting detection device 100, which includes: an acquisition module 101, used to acquire a set of original asphalt pavement images of the target road section, and use a portion of the original asphalt pavement image data in the original asphalt pavement image set as a training set; a disparity estimation module 102, used to train an improved RAFT-Stereo network using the training set to obtain a stereo matching model, and input the original asphalt pavement images other than the training set into the stereo matching model for disparity estimation and then convert them into a high-precision dense pavement point cloud; and a fusion module 103, used to perform high-precision dense pavement point cloud fusion. The dense road point cloud is spatially registered and fused to obtain the fused road point cloud; the calculation module 104 is used to extract the cross-sectional lines of the stitched road point cloud to obtain the cross-sectional lines of the target road segment, and to calculate the rut depth of the target road segment using the cross-sectional lines of the target road segment; the improved RAFT-Stereo network includes a feature extraction module for extracting matching features and contextual features, a feature enhancement module for enhancing the global consistency of features, a correlation volume construction module for calculating the matching relationship between features, a recursive update module for predicting disparity increment and correcting pixel correspondence, and an upsampling module for restoring low resolution to the original image resolution.
[0050] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0051] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0052] This invention also provides a terminal device, as shown in FIG5. The terminal device D10 of this embodiment includes: at least one processor D100 (only one processor is shown in FIG5), a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100. When the processor D100 executes the computer program D102, it implements the above-mentioned asphalt pavement rutting detection method.
[0053] The terminal device D10 can be a desktop computer, laptop, handheld computer, server, server cluster, or cloud server, etc. This terminal device may include, but is not limited to, a processor D100 and a memory D101. Those skilled in the art will understand that Figure 5 is merely an example of the terminal device D10 and does not constitute a limitation on the terminal device D10. It may include more or fewer components than shown, or combine certain components, or different components; for example, it may also include input / output devices, network access devices, etc.
[0054] The processor D100 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0055] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as a hard disk or memory of the terminal device D10. In other embodiments, the memory D101 may be an external storage device of the terminal device D10, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the terminal device D10. Furthermore, the memory D101 may include both internal and external storage units of the terminal device D10. The memory D101 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory D101 can also be used to temporarily store data that has been output or will be output.
[0056] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0057] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0058] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements a method for detecting ruts on asphalt pavements.
[0059] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to a building device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks.
[0060] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for detecting rutting in asphalt pavement, characterized in that, include: Step 1: Collect a set of original asphalt pavement images of the target road section, and use a portion of the original asphalt pavement image data from the set as a training set. Step 2: Train the improved RAFT-Stereo network using the training set to obtain a stereo matching model. Input the original asphalt pavement image (excluding the training set) into the stereo matching model for disparity estimation and convert it into a high-precision dense pavement point cloud. Step 3: Perform spatial registration and fusion on the high-precision dense pavement point cloud to obtain a fused pavement point cloud. Step 4: Extract cross-sectional lines from the stitched pavement point cloud to obtain the cross-sectional lines of the target road segment. Calculate the rut depth of the target road segment using the cross-sectional lines of the target road segment. The improved RAFT-Stereo network includes a feature extraction module for extracting matching features and contextual features, a feature enhancement module for enhancing global consistency of features, a correlation volume construction module for calculating the matching relationship between features, a recursive update module for predicting disparity increments and correcting pixel correspondences, and an upsampling module for restoring low resolution to the original image resolution.
2. The method for detecting rutting in asphalt pavement according to claim 1, characterized in that, The original asphalt pavement image set was collected on the target road section using a vehicle-mounted multi-camera data acquisition platform.
3. The method for detecting rutting in asphalt pavement according to claim 2, characterized in that, The vehicle-mounted multi-camera data acquisition platform includes: multiple industrial-grade binocular structured light cameras equipped with infrared filters, a multi-device synchronous data acquisition unit, and an aluminum bracket; the aluminum bracket is located at the rear of the inspection vehicle; all industrial-grade binocular structured light cameras are mounted on the aluminum bracket, and the lenses of the industrial-grade binocular structured light cameras are vertically downward facing the road surface; the output terminals of all industrial-grade binocular structured light cameras are electrically connected to the input terminals of the multi-device synchronous data acquisition unit.
4. The method for detecting rutting in asphalt pavement according to claim 1, characterized in that, Before step 2, the method further includes: using a lightweight general image restoration network to perform super-resolution enhancement on the original asphalt pavement images other than those in the training set, to obtain enhanced asphalt pavement images.
5. The method for detecting rutting in asphalt pavement according to claim 1, characterized in that, The improved RAFT-Stereo network is trained using the training set to obtain a stereo matching model, including: loading the weight parameters of the improved RAFT-Stereo network pre-trained on a large-scale synthetic dataset as initial weight parameters; freezing the parameters of the feature extraction module and the related volume construction module, and performing preliminary training on the parameters of the feature enhancement module and the recursive update module using a progressive learning approach to obtain a pre-trained RAFT-Stereo network; and fine-tuning the parameters of all modules in the pre-trained RAFT-Stereo network on a self-built real-world scene dataset to obtain the stereo matching model.
6. The method for detecting rutting in asphalt pavement according to claim 1, characterized in that, Step 3 includes: extracting local features from the original asphalt pavement images other than the training set to obtain local feature points; performing feature matching on the local feature points using a feature matching algorithm to obtain a set of matching point pairs; normalizing the color vectors of each matching point pair in the set and calculating the Euclidean distance to obtain the color distance, and selecting valid matching point pairs from the set based on the color distance; combining the internal parameters, external parameters, and depth information of the industrial-grade binocular structured light camera, transforming the valid matching pairs using a back-projection operator to obtain an initial mapping result; estimating the rigid body transformation parameters of the high-precision dense pavement point cloud based on the initial mapping result to obtain the rigid body transformation parameter estimation result; and optimizing and fusing the high-precision dense pavement point cloud based on the rigid body transformation parameter estimation result to obtain the fused pavement point cloud.
7. A device for detecting ruts in asphalt pavement, characterized in that, include: The acquisition module is used to acquire a set of original asphalt pavement images of the target road section, and to use a portion of the original asphalt pavement image data in the set of original asphalt pavement images as a training set. The disparity estimation module is used to train the improved RAFT-Stereo network using the training set to obtain a stereo matching model, and input the original asphalt pavement image other than the training set into the stereo matching model for disparity estimation and then convert it into a high-precision dense pavement point cloud; the fusion module is used to perform spatial registration and fusion on the high-precision dense pavement point cloud to obtain a fused pavement point cloud. The calculation module is used to extract the cross-sectional lines of the stitched road point cloud to obtain the cross-sectional lines of the target road segment, and to calculate the rut depth of the target road segment using the cross-sectional lines of the target road segment. The improved RAFT-Stereo network includes a feature extraction module for extracting matching features and contextual features, a feature enhancement module for enhancing the global consistency of features, a correlation volume construction module for calculating the matching relationship between features, a recursive update module for predicting disparity increments and correcting pixel correspondences, and an upsampling module for restoring low resolution to the original image resolution.
8. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the asphalt pavement rut detection method as described in any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the asphalt pavement rutting detection method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Asphalt pavement rut three-dimensional shape automatic generation method
CN116091714A
Highway pavement disease intelligent detection method and system based on multi-source data fusion and YOLO optimization algorithm
CN121190981A
Ice and snow airport runway flatness measuring system and method based on point cloud
CN121498530A
Asphalt pavement rut depth image detection and cross section reconstruction method and system thereof
CN121599979A