Road surface crack detection device and method based on machine vision

By combining multimodal data acquisition and deep learning models, the accuracy and adaptability issues of pavement crack detection in existing technologies have been solved, enabling efficient and accurate crack identification and quantitative measurement in complex environments, and supporting urban road maintenance and management.

CN121095229BActive Publication Date: 2026-07-21WUHAN MUNICIPAL ENG DESIGN & RES INST
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WUHAN MUNICIPAL ENG DESIGN & RES INST
Filing Date
2025-09-24
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing machine vision-based road crack detection devices have poor recognition accuracy under varying lighting conditions and complex environments, with high false positive and false negative rates. They also cannot accurately quantify crack geometry, resulting in insufficient adaptability and difficulty in meeting the needs of urban road maintenance.

Method used

A multimodal data acquisition module is adopted, including a high-definition RGB camera, a binocular stereo vision module, and an infrared thermal imaging module. Combined with a deep learning model, it can simultaneously acquire image details, spatial depth information, and temperature anomalies. By fusing multi-source information, the detection accuracy is improved, and quantitative information such as the width and length of the crack is output.

Benefits of technology

It significantly improves the accuracy and robustness of crack detection, enables real-time and efficient detection in complex environments, outputs rich quantitative information, adapts to different road materials and weather conditions, and supports real-time maintenance decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121095229B_ABST
    Figure CN121095229B_ABST
Patent Text Reader

Abstract

The present application relates to the field of road crack detection, and particularly relates to a road crack detection device and method based on machine vision, which comprises a multi-modal data acquisition module, an illumination unit, a central processing unit, a display module and a storage module, a positioning and communication unit, a power supply module and a support damping structure, the multi-modal data acquisition module comprising a high-definition RGB camera module, a binocular stereo vision module and an infrared thermal imaging module; the detection method comprises the following steps: S01 data acquisition and preprocessing, S02 multi-modal fusion and input preparation, S03 crack identification processing, S04 post-processing and parameter extraction, and S05 result output and storage. The present application solves the problems of poor adaptability, insufficient real-time performance, high false detection and missed detection rates, and imperfect quantitative measurement in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of pavement crack detection, and specifically to a pavement crack detection device and method based on machine vision. Background Technology

[0002] The maintenance and management of urban roads requires the timely detection of early-stage defects such as pavement cracks. Traditional methods relying on manual inspections are inefficient, subjective, and pose safety hazards. In recent years, automated detection equipment using vehicle-mounted imaging systems and computer vision algorithms has improved inspection efficiency, but many shortcomings still remain.

[0003] Currently, existing technologies disclose road surface crack detection devices and methods based on machine vision, including: an image acquisition module, an image preprocessing module, an image feature extraction module, and a crack detection module. The image acquisition module acquires road surface crack images and divides them into training and testing sets. The image preprocessing module preprocesses the training set. The image feature extraction module constructs a feature extraction model based on the training set to extract image features from the road surface crack images. The crack detection module detects road surface cracks based on the feature extraction model and the testing set. A series of image preprocessing algorithms are employed to improve the efficiency of image feature extraction; multi-class recognition is performed, and all extracted image features are converted into feature data, improving the recognition accuracy of the recognition model.

[0004] For example, existing devices mostly rely on a single visible light camera to acquire information, making them prone to missed or false detections of cracks when encountering changes in lighting or complex road surface textures. Especially when there are obstructions such as road markings, patches, or shadows, deep learning models may incorrectly identify these non-crack features as cracks, causing false alarms. Furthermore, due to differences in road materials and environmental conditions, current models have poor adaptability when applied across different scenarios, requiring significant manual recalibration and training. In addition, most publicly available solutions can only output the location or outline of cracks, failing to accurately quantify key indicators such as crack width and depth; this makes it difficult for detection results to directly guide maintenance decisions.

[0005] In summary, there is an urgent need for an intelligent detection device and method that can integrate multiple sensor information, improve recognition accuracy and environmental adaptability, and output crack geometry dimensions, in order to overcome the limitations of existing technologies. Summary of the Invention

[0006] Based on the above description, this invention provides a deep learning-based urban road crack detection device and method to address the problems of poor adaptability, insufficient real-time performance, high false positive and false negative rates, and imperfect quantitative measurement in existing technologies. Specifically, this solution aims to: improve the versatility of the device under different road materials and environments, achieve real-time high-precision crack identification while vehicles are in motion, significantly reduce misjudgments caused by interference factors, and output quantitative information such as crack width and length.

[0007] On the one hand, the technical solution of the present invention to solve the above-mentioned technical problems is as follows: a road surface crack detection device based on machine vision.

[0008] It includes a multimodal data acquisition module, an illumination unit, a central processing unit, a display module and a storage module, a positioning and communication unit, a power supply module, and a support and vibration reduction structure;

[0009] The multimodal data acquisition module includes a high-definition RGB camera module, a binocular stereo vision module, and an infrared thermal imaging module. The multimodal data acquisition module and the lighting unit are uniformly installed on the support and vibration reduction structure, which is fixedly connected to the top or front structure of the vehicle.

[0010] The central processing unit is electrically connected to the multimodal data acquisition module, the positioning and communication unit, the display module, and the storage module;

[0011] The power supply module provides multiple stable power sources for the aforementioned modules.

[0012] It is worth noting that by integrating a high-definition RGB camera, a binocular stereo vision module, and an infrared thermal imaging module into a multimodal data acquisition approach, the system simultaneously acquires image details, spatial depth information, and temperature anomalies of road surface cracks. This improves the accuracy and robustness of crack detection, especially in complex environments (such as low light, temperature differences, or surface contamination), enabling stable identification of damage characteristics. The lighting unit and multimodal sensors are uniformly mounted on a support vibration-damping structure, which is fixedly connected to the vehicle. This effectively resists vibration interference generated during vehicle movement, ensuring stable and continuous image data acquisition and enhancing the practicality and imaging quality of the equipment in dynamic inspections. The central processing unit integrates data processing, model inference, and control functions. Through efficient connections with data acquisition, positioning communication, display, and storage modules, it achieves real-time crack identification, location, and result visualization output. It also facilitates remote transmission and subsequent management, improving operational efficiency and information closure. The power supply module provides multiple regulated power supplies to ensure continuous and stable operation of each submodule in a vehicle environment, improving the overall reliability of the system and making it suitable for long-term, high-frequency road inspection tasks.

[0013] Based on the above technical solution, the present invention can be further improved as follows.

[0014] Furthermore, the high-definition RGB camera module includes at least one 2-megapixel progressive scan industrial camera, which supports global shutter, is equipped with an 8mm wide-angle lens, has automatic aperture and shutter control functions, and has an exposure time range of 5μs–100ms.

[0015] Its output image is transmitted to the central processing unit via a GigE Vision or USB3 Vision interface.

[0016] Furthermore, the binocular stereo vision module includes two grayscale industrial cameras, which are set at both ends of the slide rail to form a stereo vision system with a baseline of 0.5 meters.

[0017] The exposure is kept consistent by a synchronous trigger, and the installation angle is deflected by about 30° to generate a road surface parallax map and convert it into a depth map. The depth map has a spatial resolution of no more than 5 mm at an effective viewing distance of less than 5 meters.

[0018] Secondly, the technical solution of the present invention to solve the above-mentioned technical problems is as follows: a road surface crack detection method based on machine vision, comprising the following steps:

[0019] Data acquisition and preprocessing; control the high-definition RGB camera module to continuously capture road surface images at a frame rate of approximately 20 Hz, and simultaneously acquire images from the binocular stereo vision module and the infrared thermal imaging module; perform distortion correction and perspective transformation on the RGB images to obtain a top-view orthophoto; calculate the disparity map of the stereo image and convert it into a depth map; resample the infrared thermal image to a coordinate system that matches the RGB image;

[0020] Multimodal fusion and input preparation: The RGB image, depth image, and temperature image are concatenated in the channel dimension to form a six-channel tensor. The depth image and temperature image are normalized and then subjected to position encoding before being used as the input to the neural network.

[0021] Crack identification processing: An improved encoder-decoder deep learning model is adopted. The encoder includes a residual network structure and an attention mechanism, and a Transformer structure is introduced for global context modeling. The decoder consists of a semantic segmentation branch and an object detection branch, and outputs a crack probability mask and a set of bounding boxes.

[0022] Post-processing and parameter extraction: Morphological processing of the mask image is performed, and the bounding box is fused to achieve instance-level segmentation. Further extraction of crack length, width, depth and location information is then performed.

[0023] Results output and storage: The crack outline and attribute information are overlaid on the original image or map, saved to local storage in time series and uploaded to the cloud platform.

[0024] Through the aforementioned technical solution, during the data acquisition phase, the synchronous collaboration of an RGB camera, binocular stereo vision, and infrared thermal imaging module enables the fusion of multimodal information on image texture, spatial structure, and thermal anomalies. Preprocessing techniques such as perspective transformation, distortion correction, and image alignment yield standardized top-view images and depth and temperature maps in a unified coordinate system, significantly improving subsequent recognition accuracy and algorithm robustness. Furthermore, channel-level stitching forms a six-channel fusion tensor, and non-image channels are normalized and positionally encoded, ensuring spatial and numerical scale coordination between depth and thermal imaging data and visual data, providing richer crack characterization features for the neural network.

[0025] The core recognition algorithm introduces an improved encoder-decoder structure. The encoder utilizes a residual network and attention mechanism to extract multi-scale features, and combines a Transformer structure to enhance global perception capabilities, effectively identifying minute cracks and occluded areas in complex backgrounds. The decoder outputs crack probability maps and bounding boxes through dual branches, achieving fine semantic segmentation and target detection in synergy. In the post-processing stage, instance-level crack segmentation is completed by fusing masks and bounding boxes, and geometric parameters such as length, width, and depth, as well as geographical location, are extracted to meet the needs of quantitative analysis of road defects and engineering decision-making. Furthermore, the detection results are output as a superimposed graphic and structured data format, and data closure is achieved through local and cloud storage, providing traceable and quantifiable technical support for road maintenance. This method outperforms traditional image recognition algorithms and single vision models in terms of detection accuracy, spatial geometric recognition, global semantic understanding, and information output completeness, demonstrating high engineering application value.

[0026] Furthermore, the distortion correction and perspective transformation of the RGB image in the data acquisition and preprocessing are performed by projecting the original image onto a plane coordinate system with the vehicle center as the reference, using a pre-calibrated homography matrix, so that the image size is consistent and facilitates subsequent scale analysis.

[0027] Furthermore, the six-channel tensor fused in the multimodal fusion and input preparation includes: RGB three channels, grayscale depth... Figure 1 Channel, grayscale temperature Figure 1 Channels, and location-encoded channels; where depth maps and temperature maps are processed through mean-variance normalization or linear mapping.

[0028] Furthermore, the encoder adopts a ResNet-34 backbone structure with channel attention modules and spatial attention modules, and connects a Transformer module after each layer of the backbone to perform global context modeling of the feature maps.

[0029] The decoder includes a semantic segmentation branch and a detection branch:

[0030] The semantic segmentation branch uses a U-Net structure for stepwise upsampling and employs an attention gate mechanism to fuse multi-scale features.

[0031] The detection branch constructs multi-scale feature maps based on the FPN network and outputs crack bounding boxes using a YOLO-style detection head.

[0032] Furthermore, the crack length calculation method in the post-processing and parameter extraction is as follows: extract the skeleton pixels of the crack mask, project them onto the geographic coordinate system, and then accumulate the skeleton curve distance; the crack width is obtained by transforming the distance from the skeleton to the edge to obtain the average width and the maximum width; the crack depth is calculated by the depth difference between the mask area and the surrounding area.

[0033] The crack location is calculated by matching the image frame positioning information with the mask center coordinates, and mapping it to global positioning coordinates or road station numbers.

[0034] The formula for calculating road station numbers is as follows:

[0035] ;

[0036] in Given the location of the reference station, Let its coordinates be given.

[0037] Furthermore, the results are displayed in the following way: crack outlines, detection boxes and attribute text are superimposed on the original image or road map, and low-confidence targets are marked with different colors or icons to prompt manual review; the detection data is packaged and stored in the form of images, masks and attribute JSON / XML.

[0038] Furthermore, the deep learning model training employs a multi-task loss function, including: pixel-level binary cross-entropy loss for segmentation masks, L1 regression loss for detecting bounding boxes, and cross-entropy loss for class confidence, which are jointly used to update model parameters via backpropagation.

[0039] In addition to real labeled images, the training data also includes synthetic images and augmented samples. A transfer learning strategy is used for the infrared and depth channels, which first pre-trains on the synthetic data and then fine-tunes on the real data to improve the robustness and accuracy of crack detection under different environments.

[0040] Compared with the prior art, the technical solution of this application has the following beneficial technical effects:

[0041] 1. Multi-source information fusion for accuracy and reliability: The device simultaneously acquires visible light, depth, and infrared data, perceiving road conditions from three perspectives: texture, morphology, and temperature. Depth information ensures the capture of the three-dimensional geometric features of cracks, while infrared information highlights internal material defects, thus significantly reducing false detections caused by changes in lighting and road coatings. Even against shadows or complex textures, the model can identify real cracks based on abrupt changes in depth, demonstrating strong anti-interference capabilities. Experiments show that compared to systems using only RGB cameras, this approach improves the detection rate of minute cracks while significantly reducing the false alarm rate.

[0042] 2. Real-time and efficient, suitable for vehicle use: This solution employs an optimized deep learning architecture and hierarchical processing strategy to ensure high frame rate operation on embedded GPUs. Through a two-stage algorithm of rapid detection followed by fine segmentation, the average processing time per frame is significantly reduced, enabling real-time detection of vehicles traveling at normal speeds. Furthermore, the combination of YOLO-like single-shot detection and Transformer global feature extraction allows the model to efficiently identify cracks of different sizes. After actual road testing, the system can stably achieve a processing capacity of over 20 frames per second, meeting the requirements for non-stop inspection of urban roads.

[0043] 3. Strong adaptability and good generalization: The deep model incorporates multimodal data augmentation and domain adaptation strategies during training and continuously learns from new samples in different regions. Therefore, it maintains good detection performance for various road types (asphalt, cement, paved bricks, etc.) and weather conditions. The sensor parameters of the device (such as exposure and gain) can be automatically adjusted according to ambient light to ensure consistent image quality under different lighting conditions. In addition, the model can be quickly adapted to new urban roads with a small amount of additional training, enabling cross-regional deployment without recalibration, greatly improving the versatility of the solution.

[0044] 4. Abundant output information, facilitating decision-making: This device not only identifies the location of cracks but also directly measures their dimensional parameters, providing estimates of the crack's length, width range, and depth. These quantitative indicators provide maintenance departments with a scientific basis to assess the severity and development trend of cracks and accurately formulate repair and reinforcement plans. Simultaneously, the system automatically numbers and maps cracks, facilitating on-site location by subsequent maintenance personnel and significantly improving maintenance management efficiency.

[0045] 5. Modular Structure and Controllable Cost: The device adopts a modular design, with sensors and computing units connected via standard interfaces, facilitating installation and debugging. Mature and reliable components are prioritized in hardware selection, with core sensors based on industrial-grade cameras and rangefinders, significantly reducing costs compared to imported integrated inspection vehicles. For small and medium-sized cities, simplified configurations (such as omitting thermal imagers or using a monocular + structured light solution) can be chosen to reduce costs; for high-requirement scenarios, the number of high-definition cameras or higher-precision lasers can be added. This invention possesses excellent scalability and cost-effectiveness, and is expected to be widely applied in practical road maintenance. Attached Figure Description

[0046] Figure 1 This is a side view of the overall structure of the machine vision-based road crack detection device provided in Embodiment 1 of the present invention.

[0047] Figure 2 This is a schematic diagram of the overall structure of the machine vision-based road crack detection device provided in Embodiment 1 of the present invention.

[0048] Figure 3 This is a schematic diagram of the support and vibration reduction structure of the machine vision-based road crack detection device provided in Embodiment 1 of the present invention.

[0049] Figure 4 A reference diagram showing the usage status of the machine vision-based road crack detection device provided in Embodiment 1 of the present invention;

[0050] Figure 5 This is a system framework diagram of the machine vision-based road surface crack detection device provided in Embodiment 1 of the present invention;

[0051] Figure 6 This is a flowchart of the machine vision-based road surface crack detection method provided in Embodiment 2 of the present invention;

[0052] Figure 7 This is a visual diagram of the crack recognition results of the machine vision-based road crack detection method provided in Embodiment 2 of the present invention.

[0053] Figure 8 This is a schematic diagram of the crack recognition model architecture of the machine vision-based road crack detection method provided in Embodiment 2 of the present invention.

[0054] Figure 9 This is a data fusion and input preparation structure diagram for the machine vision-based road crack detection method provided in Embodiment 2 of the present invention.

[0055] Figure 10 This is a flowchart of the crack post-processing and parameter extraction process of the machine vision-based road crack detection method provided in Embodiment 2 of the present invention.

[0056] Figure reference numerals: 1. Multimodal data acquisition module; 11. High-definition RGB camera module; 12. Binocular stereo vision module; 13. Infrared thermal imaging module;

[0057] 21. Lighting unit; 22. Central processing unit; 23. Display module; 24. Storage module; 25. Positioning and communication unit; 26. Power supply module;

[0058] 3. Supporting vibration damping structure; 31. Mounting bracket; 32. Vibration damping structure; 33. Connecting frame; 34. Connecting triangular plate; 35. Inclined frame;

[0059] 41. Connecting block; 42. Bearing block; 43. Connecting shaft; 44. Connecting hoop; 45. Clamping plate; 46. Upper clamping part; 47. Lower clamping part; 48. Guide shaft; 49. Buffer component. Detailed Implementation

[0060] To facilitate understanding of this application, a more complete description will be provided below with reference to the accompanying drawings, which illustrate embodiments of the present application. However, the present application can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that the disclosure of this application will be thorough and complete.

[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application.

[0062] Example 1:

[0063] refer to Figure 1 and Figure 4 A machine vision-based road surface crack detection device is installed on a detection vehicle. The device includes a multimodal data acquisition module 1, an illumination unit 21, a central processing unit 22, a display module 23, a storage module 24, a positioning and communication unit 25, a power supply module 26, and a support and vibration damping structure 3. The modules are electrically and structurally connected to form a complete system. The multimodal data acquisition module 1 includes a high-definition RGB camera module 11, a binocular stereo vision module 12, and an infrared thermal imaging module 13. The multimodal data acquisition module 1 is detachably and fixedly connected to the vehicle via the support and vibration damping structure 3. The specific structure is as follows:

[0064] The high-definition RGB camera module 11 is preferably an RGB camera, ideally a 2-megapixel progressive scan industrial-grade imaging unit, supporting a global shutter mode to eliminate motion blur. The camera is equipped with an 8mm wide-angle lens with a horizontal viewing angle of approximately 60°, enabling continuous road surface coverage with an image overlap rate greater than 85% at vehicle speeds up to 30 km / h. The camera features automatic aperture and shutter control, with an exposure time range between 5μs and 100ms. Its output image data is transmitted to the central processing unit 22 via a GigE Vision or USB3 Vision standard interface. The camera is fixed approximately 1.0 meter above the vehicle's front bumper and tilted forward by 10° to increase the imaging area.

[0065] The binocular stereo vision module 12 consists of two grayscale industrial cameras mounted at both ends of a precision mechanical guide rail, forming a stereo vision system with a baseline distance of 0.5 meters. The two cameras maintain consistent exposure via a synchronization trigger, and their mounting angle is vertically deflected by approximately 30° to acquire information about road surface elevation variations. The module is equipped with parallax matching software (such as OpenCV StereoBM / SGBM or NVIDIA Isaac SDK) to perform real-time parallax map calculation with a parallax accuracy better than ±0.5 pixels. It can also convert and generate depth maps with a spatial resolution within 5mm (within a 5m line-of-sight). The camera housing has an IP67 protection rating, making it suitable for complex outdoor environments such as rain and snow.

[0066] The infrared thermal imaging module 13 uses an uncooled infrared thermal imager with an operating wavelength of 8–14 μm, a thermal sensitivity (NETD) better than 0.05°C, a resolution of at least 640×480 pixels, and a frame rate of up to 30 frames per second. The device integrates a temperature calibration module and a radiation compensation mechanism, outputting an absolute surface temperature map. The thermal imager is precisely aligned with the RGB image using homography matrix calibration, with a pixel error not exceeding 1 pixel. This module is used to assist in detecting thermal anomaly areas at cracks, improving detection accuracy at night or on contaminated road surfaces.

[0067] The lighting unit 21 includes two sets of LED spotlights positioned on the left and right sides of the camera mount. Each set has a power of no less than 50W and a color temperature of 5000K–6500K, making it suitable for both day and night use. The luminaires are equipped with anti-glare light guides to avoid interfering with surrounding traffic. The system supports PWM dimming and adaptive compensation for ambient brightness, with control based on automatic adjustment of lighting intensity using a photosensor. The luminaires are installed parallel to the camera's optical axis to ensure uniform illumination and good contrast in the imaging area.

[0068] The vibration damping structure 3 includes a mounting bracket 31 and a shock-absorbing structure 32. The RGB camera module, binocular stereo vision module 12, and infrared thermal imaging module 13 are uniformly installed with the lighting unit 21 via the aluminum alloy mounting bracket 31. The bottom of the mounting bracket 31 is equipped with the shock-absorbing structure 32. The entire structure is installed on the roof rack of the vehicle using anti-loosening bolts. The mounting bracket 31 has a vibration resistance of over 20g. The data cables use high-interference military-grade shielded cables and are uniformly laid in a fixed cable tray to avoid tangling or signal interference during driving. Insulating components are installed at the connection between the mounting bracket 31 and the vehicle body to reduce the interference of vehicle electrical noise on image acquisition.

[0069] Mounting bracket 31 is made of aluminum profiles and includes a connecting frame 33 for connecting to the vehicle, a connecting triangle plate 34, and a tilting bracket 35. The tilting bracket 35 is installed between the two connecting triangle plates 34. The RGB camera module, binocular stereo vision module 12, infrared thermal imaging module 13, and lighting unit 21 are installed on the tilting bracket 35. The shock absorption structure 32 includes a connecting block 41 connected to the side wall of the connecting frame 33, a bearing block 42 slidably disposed in the connecting block 41, a connecting shaft 43, a connecting hoop 44, and a clamping plate 45. The connecting shaft 43 is rotatably connected in the bearing block 42, and the connecting hoop 44 is clamped on the connecting shaft 43. The clamping plate 45 includes an upper clamping part 46 and a lower clamping part 47. The upper clamping part 46 is integrally connected to the connecting hoop 44, and the upper clamping part 46 and the lower clamping part 47 are connected together by bolts. The crossbar on the car roof rack passes between the upper clamping part 46 and the lower clamping part 47, and the mounting bracket 31 can be fixed to the crossbar by tightening the bolts.

[0070] It also includes a guide shaft 48 and a buffer 49 that pass through the bearing block 42. The buffer 49 is sleeved on the guide shaft 48 and can be made of buffer rubber or a buffer spring.

[0071] The central processing unit 22 is an embedded GPU edge computing platform, preferably based on the NVIDIA Jetson AGX Orin or Xavier NX system. It integrates an 8-core ARM processor and a 2048-core GPU computing unit, supporting the TensorRT inference acceleration framework. The system is equipped with 32GB of RAM and a 512GB NVMe high-speed solid-state drive, and features multiple high-speed image data interfaces including MIPI, USB 3.1, and GigE. The operating system is based on Ubuntu 20.04 LTS, with built-in runtime environments for detection algorithms such as OpenCV, PyTorch, and ROS. The processing unit also integrates a CAN bus module for acquiring vehicle speed and other driving status parameters via the vehicle's OBD system; the communication module supports 4G / 5G network transmission and uses TCP / IP and MQTT protocols to send detection results to the cloud platform. The entire unit has a 10g shock resistance capability and an operating temperature range of –20°C to 60°C.

[0072] Display module 23 and storage module 24. This module is equipped with a 10.1-inch industrial-grade touchscreen LCD with a resolution of 1920×1200 and a brightness greater than 600 cd / m², ensuring readability even in sunlight. The screen display interface supports functions such as crack image overlay, map navigation, inspection progress display, and information filtering. Inspection data is categorized and stored using both time and location tags, with 1TB industrial-grade SSDs as the storage medium. The system supports multiple data export interfaces, including USB flash drive, Wi-Fi, and Bluetooth, facilitating data transfer and archiving.

[0073] Positioning and Communication Module: This module integrates a GNSS multi-mode satellite positioning system, supporting GPS / BeiDou / GLONASS positioning with an accuracy ≤1.5m. An optional RTK differential enhancement module can be added to achieve centimeter-level positioning accuracy. It also incorporates a high-resolution wheel encoder to acquire real-time driving distance and speed data. The module supports an inertial measurement unit (IMU) to maintain trajectory continuity in tunnels or satellite signal blind spots. Image acquisition timestamps are synchronized with location information to ensure consistency of spatial data.

[0074] The power supply module 26 provides the entire device with power from the testing vehicle's generator system, and is equipped with a DC-DC power conversion module that outputs multiple stable voltages (5V, 12V, 24V). Overcurrent protection and independent fuses are provided for the GPU processor and lighting circuits to enhance system stability. The device is also equipped with a UPS lithium battery module, which can provide at least 15 minutes of runtime in the event of a sudden power outage, ensuring data security and a soft shutdown of the equipment. The overall power consumption is controlled within the range of 300W to 500W, meeting the requirements for long-term low-power operation in the field.

[0075] Example 2:

[0076] refer to Figures 6-10 The software portion of this invention, which is based on machine vision for road surface crack detection, runs within the aforementioned processing unit. The method flow for crack detection is as follows: Figure 2 As shown, the process includes steps such as data acquisition and preprocessing, multimodal fusion, deep learning recognition, result post-processing, and output.

[0077] Step S01: Data Acquisition and Preprocessing. When the inspection vehicle begins its patrol, the processing unit triggers the RGB camera to continuously capture road surface images at a frame rate of approximately 20 Hz, and simultaneously acquires data from the binocular camera and thermal imager. For each acquired frame, distortion correction and perspective transformation are first performed on the RGB image. Using the camera's intrinsic and extrinsic parameters, the original front view is projected into a top-view orthographic view, ensuring the image plane is parallel to the road surface. The perspective transformation uses a pre-calibrated homography matrix to map each pixel to a vehicle-centered planar coordinate system. This process ensures that cracks of the same size in the image are of consistent size at different locations, facilitating subsequent algorithm identification and measurement. Simultaneously, the left and right images acquired by the binocular camera are corrected and matched to calculate a disparity map, which is then converted into a depth map using triangulation. This depth map has the same resolution as the RGB image and undergoes median filtering to smooth out noise points. The infrared thermal imager outputs a temperature field map corresponding to the field of view, with each pixel representing the temperature of that road surface point. The temperature field map is then resampled and matched to the top-view image coordinate system using bilinear interpolation.

[0078] At this point, three aligned images are obtained: a top-view RGB image, a depth image, and an infrared temperature image. Next, optional image enhancements are performed on the RGB images: such as applying adaptive histogram equalization to improve crack texture contrast, or using guided filtering to suppress noise textures. After this step, the multi-channel image matrix is ​​packaged and used as input data for the model.

[0079] Step S02: Multimodal Data Fusion and Input Preparation. Before processing by the deep learning model, the multi-source data needs to be organized in a form usable by the neural network. Specifically, the RGB image, depth image, and temperature image are concatenated and fused along the channel dimension to form a tensor with 6 channels (3 RGB channels + 1 depth grayscale channel + 1 temperature grayscale channel + optional empty channels). Simultaneously, to maintain the consistency of dimensions between different modal data, the depth and temperature channels are normalized separately: the depth value is subtracted from the mean depth of the training set and divided by the standard deviation to ensure it falls within the 0-1 range; the temperature image is also linearly scaled or histogram balanced. Since the value ranges of infrared and visible light differ significantly, this normalization helps the network to train stably and converge. Furthermore, to enable the model to identify the vehicle's driving direction, a positional encoding (e.g., a two-dimensional coordinate matrix of the same size as the image) can be added to the input tensor. In this embodiment, the normalized depth value is superimposed on the RGB image using pseudo-color encoding for human visual viewing, but it still exists as an independent channel when fed into the model. The fused multi-channel data is packaged into a batch and sent to the inference module of the deep learning model.

[0080] Step S03: Crack identification using a deep learning model. The overall architecture of the deep learning model of this invention is as follows: Figure 4 As shown, an improved encoder-decoder architecture is adopted and the detection branch is fused. Specifically:

[0081] The encoder uses ResNet34 as its basic backbone, adding a channel attention module (SE module) and a spatial attention module after each convolutional layer to enhance the detailed features of the cracks. The input 6-channel fused tensor is first downsampled by stem convolution, and then enters the ResNet layer to extract multi-scale features. While outputting the feature map at each backbone stage, a Transformer encoder is introduced to perform global modeling of the feature map: that is, the feature map is unfolded into a sequence, positional encoding is added, and it is input into multiple Transformer encoder layers to capture the global contextual association of the cracks. In particular, the self-attention mechanism of the Transformer enables the model to pay attention to the long-distance continuity of the cracks. For example, even if a crack is briefly occluded, its continuity can be inferred from the features at both ends. The encoder finally produces three feature maps at different scales: a high-resolution detailed feature map F1, a medium-resolution semantic feature map F2, and a low-resolution global feature map F3.

[0082] Decoder: The decoder consists of a semantic segmentation branch and an object detection branch, sharing features extracted by the encoder. The segmentation branch uses U-Net-style stepwise upsampling: F3 is upsampled and concatenated with the F2 feature map, fused via convolution, then upsampled again and fused with F1 to generate a feature map of the same size as the input. An attention gate mechanism is applied during the fusion process to retain only the encoded features that contribute to the crack region. The last convolutional layer uses Sigmoid activation to output a crack probability mask, which, after thresholding, yields a binary crack region. The object detection branch utilizes the encoder's F2 and F3 features to construct a multi-scale feature pyramid using a Feature Pyramid Network (FPN), and then connects to a YOLO-style detection head. Anchor boxes are set on feature maps of different scales to predict the offset and class confidence of the bounding boxes. This invention treats cracks as the object category to be detected (other pavement defects such as potholes and loose surfaces can also be added), therefore the detection head outputs the probability that each anchor box belongs to "crack" and its frame coordinate correction. After non-maximum suppression, several crack bounding box proposals are obtained (rectangular boxes approximate the crack extension range). It is worth noting that the detection branch is mainly used for localization, while the segmentation branch focuses on shape extraction. The combination of the two can achieve instance-level recognition of cracks: In this embodiment, the instance segmentation effect is achieved by mapping the segmentation mask in each detection box to the same instance during training.

[0083] Model Inference and Training: During runtime, the encoder and two decoders simultaneously perform forward computation on the input data, outputting a crack mask and a set of crack detection boxes. Model training employs a multi-task loss function: including pixel-level binary cross-entropy loss for segmentation and bounding box regression and classification losses for detection, which are weighted and summed before backpropagation to update model parameters. In addition to a large number of real-world labeled images, the training data also includes images generated through data augmentation and simulation to improve model robustness. Transfer learning strategies are also employed for the depth and infrared channels during training: the model is first pre-trained on a large amount of synthetic data, enabling it to learn to use depth information to identify three-dimensional deformation of cracks, and then fine-tuned on real data. The model trained in this way demonstrates good adaptability to different environments and a high detection rate for small cracks in this invention.

[0084] Step S04: Post-processing of results and parameter calculation. The deep learning model output includes a crack mask and a list of bounding boxes.

[0085] First, the mask undergoes morphological processing: a single expansion operation connects potential interruptions in narrow cracks, followed by an etching operation to remove burrs and noise. This process can be represented as:

[0086]

[0087] in, For the original mask, For the dilation operator, For the corrosion operator, This is the processed mask image.

[0088] Next, the connected regions of the mask are calculated, with each connected region representing a crack candidate. Using the bounding box information provided by the detection branch, adjacent and overlapping crack instances can be separated: if a connected region falls within two high-confidence detection boxes, the mask is segmented along the corresponding boundary based on the box positions, assigning each segment to a different crack instance; conversely, if a long crack is obscured and divided into two mask segments, but the detection branch outputs a box spanning both segments, these two masks can be merged into a single crack instance. This achieves a complementary advantage between segmentation and detection results.

[0089] Then, parameters are calculated for each final identified crack instance:

[0090] (1) Length calculation: Traverse the skeleton pixels of the crack mask Project it onto the actual road plane coordinates Calculate the cumulative length of the bend. :

[0091] ;

[0092] (2) Width calculation: Perform an Euclidean transformation on the mask relative to the skeleton to obtain each skeleton point. Corresponding crack half width Then the average width and maximum width They are respectively:

[0093] ;

[0094] (3) Depth calculation: Within the crack mask area, depth map is used. Extract minimum depth Then, based on the average road surface depth of the surrounding crack-free area... For reference, the maximum depth is estimated as follows:

[0095] ;

[0096] If the depth map is missing, then set .

[0097] (4) Position calculation: Read the center point of the crack mask. Location data corresponding to image frames This is mapped to geographic coordinates or road station numbers. If road station numbers are used, it can be approximated as:

[0098] ;

[0099] in Given the location of the reference station, Let its coordinates be given.

[0100] After the above parameters are calculated, the system packages the crack number, length, width, depth, location and other information into crack attribute entries for maintenance assessment and subsequent data archiving.

[0101] Step S05: Result Output and Storage. The processing unit passes all crack instances and their attributes to the display module via an interface, and overlays them onto the original image or map for intuitive display (see reference). Figure 3For example, the system displays a real-time image of the road surface on the in-vehicle screen, outlining cracks with red lines and marking their locations with rectangles, while simultaneously updating crack numbers and dimensions in a list. If some crack detection results have low confidence (e.g., potential cracks with only thermal anomalies but not clearly visible in the image), they can be marked with different colors or icons to prompt manual review. All detected crack data is saved chronologically to local storage, with files containing images, crack masks, and JSON / XML descriptions of crack attributes. After a patrol operation is completed, the system summarizes the results and uploads them to a cloud management platform via wireless network. The platform can further correlate the crack data with GIS maps to generate road defect distribution maps and maintenance recommendations. This completes the entire inspection process.

[0102] Example 3:

[0103] Based on smartphone-based crack detection applications, another variation of the device of this invention utilizes a smartphone as a data acquisition and processing terminal to achieve portable road crack detection. Specifically, the multimodal acquisition module in Example 1 can be simplified to the phone's built-in RGB camera (and an optional infrared camera, such as the infrared depth sensor found in some high-end phones for face recognition, which can be used to obtain shallow depth). The user fixes the phone to the vehicle, with the camera facing the road surface. The application runs a lightweight crack detection model (such as a port of Tiny-YOLO and a simplified U-Net) on the phone's GPU, analyzing the video stream recorded during driving frame by frame. Given the limited computing power of the phone, the model can only use the RGB channel, but since a large number of users can provide data from different environments, the model's adaptability is continuously improved through federated learning after cloud aggregation. The detected crack information is uploaded via the phone and integrated into the urban road maintenance database. This implementation method is extremely low-cost and flexible in deployment, and can serve as a supplementary application scenario for the device of this invention in the mass market. Of course, its detection accuracy and reliability are slightly inferior to professional equipment, but it has great potential in smart city road monitoring.

[0104] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A machine vision-based method for detecting road surface cracks, characterized in that, A road surface crack detection device based on machine vision is adopted. The road surface crack detection device based on machine vision includes a multimodal data acquisition module (1), an illumination unit (21), a central processing unit (22), a display module (23) and a storage module (24), a positioning and communication unit (25), a power supply module (26) and a support and vibration reduction structure (3). The multimodal data acquisition module (1) includes a high-definition RGB camera module (11), a binocular stereo vision module (12), and an infrared thermal imaging module (13). The multimodal data acquisition module (1) and the lighting unit (21) are uniformly installed on the support and vibration reduction structure (3), which is fixedly connected to the top or front structure of the vehicle. The central processing unit (22) is electrically connected to the multimodal data acquisition module (1), the positioning and communication unit (25), the display module (23), and the storage module (24); the power supply module (26) provides multiple stable power supplies for the above modules; The method includes the following steps: Data acquisition and preprocessing; control the high-definition RGB camera module to continuously capture road surface images at a frame rate of about 20 Hz to obtain RGB images, and simultaneously acquire data from the binocular stereo vision module and the infrared thermal imaging module. Perform distortion correction and perspective transformation on the RGB images to obtain a top-view RGB image; the left and right images acquired by the binocular stereo vision module (12) are corrected and matched to calculate the disparity map, and then converted into a depth map through triangulation. The resolution is the same as the RGB image, and the noise points are smoothed by median filtering; the infrared thermal imaging module outputs a temperature field map corresponding to the field of view, with each pixel representing the temperature of the road surface point; the temperature field map is resampled and matched to the top-view image coordinate system through bilinear interpolation to obtain an infrared temperature map; The three aligned images are obtained: a top-view RGB image, a depth image, and an infrared temperature image; image enhancement is performed on the top-view RGB image, and the multi-channel image matrix is ​​packaged as the data input to the model; Multimodal fusion and input preparation; The top-view RGB image, depth image, and infrared temperature image are concatenated in the channel dimension to form a six-channel tensor. The six-channel tensor includes: three RGB channels, one depth grayscale channel, one temperature grayscale channel, and a position encoding channel. The depth image and infrared temperature image are normalized and then used as neural network input after the position encoding is added. Crack identification processing: An improved encoder-decoder deep learning model is adopted. The encoder includes a residual network structure and an attention mechanism, and a Transformer structure is introduced for global context modeling. The decoder consists of a semantic segmentation branch and an object detection branch, and outputs a crack probability mask and a set of bounding boxes. Post-processing and parameter extraction: Morphological processing of the mask image is performed, and the bounding box is fused to achieve instance-level segmentation. Further extraction of crack length, width, depth and location information is then performed. Results output and storage: The crack outline and attribute information are overlaid on the original image or map, saved to local storage in time series and uploaded to the cloud platform.

2. The machine vision-based road surface crack detection method according to claim 1, characterized in that, The distortion correction and perspective transformation of the RGB image in the data acquisition and preprocessing are performed by projecting the original image onto a plane coordinate system with the vehicle center as the reference, using a pre-calibrated homography matrix, so that the image size is consistent and facilitates subsequent scale analysis.

3. The machine vision-based road surface crack detection method according to claim 1, characterized in that, The depth map and infrared temperature map are processed by mean-variance normalization or linear mapping.

4. The machine vision-based road surface crack detection method according to claim 1, characterized in that, In the crack identification process: The encoder adopts a ResNet-34 backbone structure with channel attention modules and spatial attention modules, and connects a Transformer module after each layer of the backbone to perform global context modeling of the feature maps; The decoder includes a semantic segmentation branch and a detection branch: The semantic segmentation branch uses a U-Net structure for stepwise upsampling and employs an attention gate mechanism to fuse multi-scale features. The detection branch constructs multi-scale feature maps based on the FPN network and outputs crack bounding boxes using a YOLO-style detection head.

5. The machine vision-based road surface crack detection method according to claim 1, characterized in that, The crack length calculation method in the post-processing and parameter extraction is as follows: extract the skeleton pixels of the crack mask, project them onto the geographic coordinate system, and then accumulate the skeleton curve distance; the crack width is obtained by transforming the distance from the skeleton to the edge to obtain the average width and the maximum width; the crack depth is calculated by the depth difference between the mask area and the surrounding area. Crack location calculation involves matching image frame positioning information with mask center coordinates, mapping it to global positioning coordinates or road station numbers, and then reading the crack mask center point. Location data corresponding to image frames ; The formula for calculating road station numbers is as follows: ; in Given the location of the reference station, Let its coordinates be given.

6. The machine vision-based pavement crack detection method according to claim 1 or 5, characterized in that, The results are displayed in the following way: crack outlines, detection boxes and attribute text are superimposed on the original image or road map, and low-confidence targets are marked with different colors or icons to prompt manual review; the detection data is packaged and stored in the form of images, masks and attribute JSON / XML.

7. The machine vision-based pavement crack detection method according to claim 1 or 5, characterized in that, The deep learning model training employs a multi-task loss function, including: pixel-level binary cross-entropy loss for segmentation masks, L1 regression loss for detecting bounding boxes, and cross-entropy loss for class confidence, which are jointly used to update model parameters via backpropagation. In addition to real labeled images, the training data also includes synthetic images and augmented samples. A transfer learning strategy is used for the infrared and depth channels, which first pre-trains on synthetic data and then fine-tunes on real data to improve the robustness and accuracy of crack detection under different environments.