Crack identification method based on gray-depth different-source image fusion

By simultaneously acquiring and fusing RGB and depth images using onboard equipment, and utilizing an end-to-end network to identify road surface cracks, the problem of low efficiency and limited accuracy in traditional detection methods has been solved, achieving efficient and accurate road surface defect detection.

CN120877227APending Publication Date: 2025-10-31SOUTHEAST UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510841995.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Traditional road surface crack detection relies on manual inspection, which is inefficient and prone to missed or false detections. The accuracy of single grayscale or depth image recognition is limited and cannot meet the needs of road surface inspection.

Method used

The system uses an onboard camera and a depth camera to simultaneously acquire RGB and depth images. Through frame synchronization and grayscale-depth image fusion, a heterogeneous image feature deep fusion recognition network based on an end-to-end network is used for crack identification. A dynamic weighted fusion module is combined to balance the contributions of different modalities.

Benefits of technology

It improves the accuracy and efficiency of crack detection, enables real-time detection of different types of pavement defects, adapts to various pavement scenarios, and reduces the impact of changes in lighting and environmental factors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877227A_ABST
    Figure CN120877227A_ABST
Patent Text Reader

Abstract

The invention provides a crack identification method based on gray-depth different-source image fusion, and belongs to the technical field of road detection. The method comprises the following steps: acquiring data of an RGB image and a depth image of a road surface, converting the RGB image into a grayscale image, creating a grayscale-depth image pair matched with each other, and forming a training set; performing image segmentation labeling on the training set, and setting category labels for the cracks according to the types of the cracks; training the labeled training set by adopting a heterogeneous image feature deep fusion recognition network based on an end-to-end network to obtain a trained segmentation weight; introducing the trained segmentation weight into a heterogeneous image feature deep fusion recognition network based on an end-to-end network to obtain a pixel-level crack segmentation result; according to the method, the grayscale image and the depth image are respectively processed through the double-branch architecture, different-source image data can be effectively fused, the influence of environmental factors such as illumination variation, shadow interference and surface stains is dealt with, and the crack detection precision is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention provides a crack recognition method based on grayscale-depth heterogeneous image fusion, belonging to the field of road detection technology. Background Technology

[0002] Highways are a core component of transportation infrastructure and are crucial to national economic development. By the end of 2023, my country's total highway mileage had reached 5.441 million kilometers. Cracks are the most common type of damage to highway surfaces. If not repaired promptly, they can expand and cause more serious damage, such as potholes. Rainwater seeping into cracks can lead to roadbed subsidence, affecting the service life of highways, increasing repair costs, and threatening driving safety. Therefore, efficient detection of road surface cracks has become an urgent problem for highway maintenance departments.

[0003] Traditional road surface crack detection relies primarily on manual inspection. This method is inefficient, depends on visual inspection of cracks, leading to subjective judgments and a high risk of missed or false detections. Furthermore, manual inspections occupy lanes, disrupting traffic flow and posing certain safety hazards. With the development of automation technology, road inspection vehicles equipped with high-definition cameras or laser sensors can now automatically capture images of road surface cracks.

[0004] However, a single grayscale image or depth image cannot fully reveal the spatial texture or depth information of road surface cracks. Crack identification based on single-source images suffers from common false alarms, limiting the accuracy of crack identification and failing to meet the needs of road surface inspection. Image fusion technology can effectively combine image information from different sources to improve image quality and information content. Depth images and grayscale images each have their advantages and disadvantages, and they are somewhat complementary. Fusing the two types of data can fully leverage their respective strengths and compensate for the limitations of single data.

[0005] Therefore, using depth images, grayscale images, and images from multiple other sources for crack identification can comprehensively reflect the characteristics of pavement cracks. This has significant theoretical and practical value for improving the accuracy and reliability of pavement crack detection and promoting the development of intelligent pavement distress detection technology. Summary of the Invention

[0006] The crack recognition method based on grayscale-depth heterogeneous image fusion provided by this invention can detect road surface defects in real time and determine the number and location of different types of road defects, and is applicable to different road scenarios.

[0007] The present invention adopts the following technical solution:

[0008] The crack recognition method based on grayscale-depth heterogeneous image fusion described in this invention comprises the following steps:

[0009] Step 1: Simultaneously acquire RGB and depth image data of the road surface using the vehicle-mounted camera and depth camera. The RGB image data of the road surface includes time information and size information, and the depth image data includes time information and size information.

[0010] Step 2: Use frame synchronization to synchronize the road surface RGB image data and road surface depth image data obtained in Step 1. Take the road surface RGB image data and road surface depth image data at the same time, convert the RGB image to grayscale image, and create a matching grayscale-depth image pair.

[0011] Step 3: Based on the image conditions, perform selective preprocessing on the images in the matched image pairs from Step 2, select a portion of the images as the training set, and perform data augmentation on the training set to expand the training set.

[0012] Step 4: Use the Labelme tool to perform image segmentation and labeling on the training set.

[0013] Assign category labels to cracks based on their type;

[0014] The category labels include horizontal cracks, vertical cracks, crazing, and pits, and the coordinates of the marked boxes are saved in the corresponding XML files;

[0015] Step 5: Train the training set labeled in Step 4 using a deep fusion recognition network based on heterogeneous image features from an end-to-end network to obtain the trained segmentation weights.

[0016] Step 6: Input the trained segmentation weights into the heterogeneous image feature deep fusion recognition network based on end-to-end network, and use the preprocessed image from step 3 for detection to obtain pixel-level crack segmentation results;

[0017] Step 7: Perform station number matching on the results from Step 6, and count the types and number of cracks.

[0018] The crack identification method based on grayscale-depth heterogeneous image fusion of the present invention includes, in step 1, the road surface RGB image data including time information and size information; and the depth image data including time information and size information.

[0019] The present invention discloses a crack identification method based on grayscale-depth heterogeneous image fusion. In step 1, the vehicle-mounted camera and depth camera are positioned at a certain height relative to the road surface being measured. The formula for calculating the camera height H is as follows:

[0020]

[0021] Where GSD is the ground resolution of the image, f is the lens focal length, and a is the pixel size.

[0022] The crack recognition method based on grayscale-depth heterogeneous image fusion described in this invention includes, in step 3, optional image preprocessing methods, including:

[0023] (1) Adjust image brightness to increase image visibility

[0024] (2) Image cropping reduces model computational load

[0025] The training set data augmentation methods in step 3 include:

[0026] (1) Use geometric transformations to augment images;

[0027] (2) Enhance the image by randomly adjusting the brightness;

[0028] (3) Enhance the image by randomly adjusting the contrast;

[0029] The image in step 3 is preprocessed and augmented using an image preprocessing method that combines one or more methods.

[0030] The crack recognition method based on grayscale-depth heterogeneous image fusion described in this invention includes the following steps in step 5 for calculating the weights:

[0031] 1) In grayscale images Depth image After processing through their respective encoder-decoder paths and upsampling modules, feature maps are removed from the grayscale image modality and the depth image modality.

[0032] 2) Calculate the spatial attention weights for the grayscale image and depth image in each channel, expressed as follows:

[0033]

[0034] 3) Generate the feature map, as expressed below:

[0035] γ I =φ(F I )=σ(W·F I +b),#

[0036] γ R =φ(F R )=σ(W·F R +b),#

[0037]

[0038] Attention log γ is generated using a nonlinear activation function. I and γ R φ represents the shared convolutional layer and activation function, W is the convolutional kernel parameter, b is the bias term, and σ is the Sigmoid activation function;

[0039] Perform this operation on each channel, and the image for each modality can obtain the corresponding channel's attention log value;

[0040] The Softmax function is applied to each channel of each modality to obtain standardized attention weights; the Softmax normalization process ensures that α + β = 1 at each spatial location, as expressed below:

[0041]

[0042] The weighted summation is performed on the corresponding channels of the two modalities to obtain the fused feature map, where ⊙ represents the successive multiplication of elements:

[0043] F fused =α⊙F I +β⊙F R #.

[0044] The crack recognition method based on grayscale-depth heterogeneous image fusion described in this invention includes, in step 6, an end-to-end network-based heterogeneous image feature deep fusion recognition network model comprising an encoder with a dual-branch backbone, a decoder with feature upsampling and dynamic weighted fusion modules, and a segmentation head.

[0045] The encoder with a dual-branch backbone uses a Mix Transformer as its backbone. This architecture employs overlapping block embedding in the convolutional layers to reduce edge information loss when the image is divided into non-overlapping blocks. The following four stages form a hierarchical structure, which increases the number of feature channels while gradually reducing the spatial dimensionality. At the core of each stage of the Transformer are an efficient self-attention module and a Mix-FNN module. The efficient self-attention module can capture long-range dependencies at a low computational cost, thus preserving crack details while ensuring algorithm efficiency. The Mix-FFN integrates deep convolutions into the architecture, enabling the network to implicitly encode positional information and focus on local feature extraction without additional positional embeddings.

[0046] In the decoder, the algorithm integrates a hierarchical frequency-aware upsampling method with multi-resolution features generated by the Mix Transformer in the encoder. The decoder is a progressive hierarchical structure where features from multiple scales are refined one at a time through frequency-aware blocks. This progressive approach allows the decoder to iteratively enhance feature maps. Simultaneously, skip connections are incorporated into the decoder, directly passing information from earlier stages to later stages, combining high-resolution low-level details from earlier stages with semantically rich high-level features from later stages.

[0047] The dynamic weighted fusion module in the decoder prevents a situation where, when an image has a stronger signal or less noise, directly adding or concatenating features from one source image can create an advantage over another, preventing the model from fully utilizing the features from both sources and leading to a decrease in segmentation performance. This module first receives two types of image features from the outputs of the dual-branch decoder and encoder structure, then calculates the attention weights from the two features, and finally fuses them based on these weights.

[0048] Beneficial effects

[0049] Compared with the prior art, the present invention has the following significant advantages:

[0050] 1. This invention uses vehicle-mounted equipment, which does not increase the time cost of data collection compared to traditional road surface data collection, and obtains road surface image data from multiple sources, thereby improving the efficiency of data collection and saving operation time.

[0051] 2. This invention processes grayscale images and depth images separately through a dual-branch architecture, which can effectively fuse heterogeneous image data, cope with the influence of environmental factors such as lighting changes, shadow interference and surface stains, and significantly improve the accuracy of crack detection.

[0052] 3. This invention designs a dynamic weighted fusion module that can dynamically adjust the contribution weights of different modalities based on feature content. Compared with simple addition or stitching fusion methods, this module can effectively balance the feature contributions of intensity images and range images, avoiding the dominance effect of single-source images. Attached Figure Description

[0053] Figure 1 This is a flowchart of the present invention.

[0054] Figure 2 The present invention provides a matching grayscale-depth image pair.

[0055] Figure 3 The images provided by this invention are an enhanced grayscale image and a depth image obtained after preprocessing (flipping), wherein (a) is a grayscale image and (b) is a depth image.

[0056] Figure 4 This invention provides a schematic diagram of road surface defects in grayscale images manually annotated using Labelme.

[0057] Figure 5 The image shows the segmentation result of the heterogeneous image feature deep fusion recognition network based on end-to-end network provided by this invention.

[0058] Figure 6 This is a schematic diagram of the heterogeneous image feature deep fusion recognition network mechanism based on end-to-end network provided by the present invention.

[0059] Figure 7 This is a schematic diagram of the encoder structure of the heterogeneous image feature deep fusion recognition network based on an end-to-end network provided by the present invention.

[0060] Figure 8 This is a schematic diagram of the decoder structure of the heterogeneous image feature deep fusion recognition network based on an end-to-end network provided by the present invention.

[0061] Figure 9 This is a schematic diagram of multimodal feature fusion in a heterogeneous image feature deep fusion recognition network based on an end-to-end network, provided by the present invention. Detailed Implementation

[0062] To make the objectives and technical solutions of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the described embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0063] like Figure 1 As shown, this example provides a crack recognition method based on grayscale-depth heterogeneous image fusion; the specific steps are as follows:

[0064] Step 1: Simultaneously acquire RGB and depth image data of the road surface using the vehicle-mounted camera and depth camera. The RGB image data of the road surface includes time information and size information, and the depth image data includes time information and size information.

[0065] Step 2: Use frame synchronization to synchronize the road surface RGB image data and road surface depth image data obtained in Step 1. Take the road surface RGB image data and road surface depth image data at the same time, convert the RGB image to grayscale image, and create a matching grayscale-depth image pair.

[0066] Step 3: Preprocess the images in the matched image pairs, retaining both the original and preprocessed images to expand the dataset;

[0067] Step 4: Use the Labelme tool to perform image segmentation and annotation on the acquired grayscale and depth images. Set category labels for the cracks according to their types, such as horizontal cracks, vertical cracks, crazing, and pits. Save the coordinates of the bounding boxes and store them in the corresponding XML files.

[0068] Step 5: Train the training set labeled in Step 4 using a deep fusion recognition network based on heterogeneous image features from an end-to-end network to obtain the trained segmentation weights.

[0069] Step 6: Input the trained segmentation weights into the heterogeneous image feature deep fusion recognition network based on end-to-end network, and use the image obtained in Step 3 for detection to obtain pixel-level crack classification and segmentation results;

[0070] Step 7: Perform station number matching on the results from Step 6, and count the types and number of cracks.

[0071] First, experiments were conducted to determine the parameters of the vehicle-mounted camera. These included the camera's focal length, image size, frame rate, and mounting height; the depth camera's image size and frame rate; and the vehicle's speed and route. This was done to prepare for acquiring high-quality RGB and depth images of the road surface.

[0072] The formula for calculating the camera height H is:

[0073]

[0074] Where GSD is the ground resolution of the image, f is the lens focal length, and a is the pixel size.

[0075] In this embodiment, a ground resolution GSD of 4mm / pixel was directly selected. The captured image resolution was 1920*1080, the focal length was 28mm, and the sensor size was one inch (12.8mm*9.6mm). The calculated suspension height was 0.9m. The driving speed was 4m / s. The road width captured was: 0.004*1920 = 7.68m, 0.004*1080 = 4.32m, which meets the lane width requirements. Road surface crack photography and depth image acquisition were conducted under clear daylight conditions to obtain good lighting conditions and a relatively clean road surface.

[0076] Verification of vehicle speed. Before actual testing, the vehicle's route and mode of travel should be planned. For example, when acquiring images of a highway, the vehicle's speed should be reasonably planned based on traffic flow to ensure that it does not affect normal traffic and guarantees the safety of the shooting operation.

[0077] The collected road surface crack images are manually screened to remove poor-quality data (i.e. blurry images), because for network models, the most effective way to improve image recognition accuracy is to improve the quality of the dataset.

[0078] Due to the limited training data, it is necessary to augment the entire image dataset. To expand the training dataset and increase the number of crack images, preprocessing the crack images using various transformations enhances the model's robustness. The main methods include:

[0079] (1) Use geometric transformations (translation, flipping, rotation) to perform data augmentation on images;

[0080] (2) Enhance the image using random brightness adjustment;

[0081] (3) Enhance the image by randomly adjusting the contrast.

[0082] The aim is to prevent overfitting and achieve the training quantity typically required for Transformer networks. This embodiment uses translation, flipping, and scaling to augment the dataset. The final dataset consists of 5000 aerial images of road surface cracks. The effect of the image enhancement transformation is shown below. Figure 3 As shown.

[0083] The LabelMe toolkit was used to label road surface cracks. In this embodiment, crack images were categorized according to crack morphology into transverse cracks, longitudinal cracks, alligator cracks, and potholes. Polygon bounding boxes were used to pinpoint crack locations, category labels were assigned, and the polygon coordinates were saved as an XML file, completing the manual labeling process. If a crack area is too large, multiple polygon bounding boxes can be used for labeling based on the crack type and closure condition to ensure the correct number of positive and negative training samples. This method is based on training experience. Images and labels were summarized to obtain a grayscale-depth image training set and corresponding labels. The manual dataset labeling is as follows: Figure 4 .

[0084] The algorithm backbone is based on the Transformer architecture, integrating an encoder with a dual-branch backbone, a decoder with feature upsampling and dynamic weighted fusion modules, and a segmentation head. The network mechanism is as follows: Figure 6As shown in the diagram. Regarding the encoder, a Mix Transformer is used as the encoder backbone. This architecture employs overlapping block embedding in the convolutional layers to reduce edge information loss when the image is divided into non-overlapping blocks. The following four stages form a hierarchical structure, which increases the number of feature channels while gradually reducing the spatial dimension. The encoder structure is as follows: Figure 7 As shown. At each stage, the core of Transform is the efficient self-attention module and the Mix-FFN module. The efficient self-attention module can capture long-range dependencies at a low computational cost, thus preserving the details of the cracks while ensuring the efficiency of the algorithm; the Mix-FFN integrates deep convolutions into the architecture, enabling the network to implicitly encode positional information and focus on local feature extraction without additional positional embeddings.

[0085] In terms of the decoder, the algorithm integrates a hierarchical frequency-aware upsampling method with the multi-resolution features generated by the Mix Transformer in the encoder. The decoder structure is as follows: Figure 8 As shown, the decoder is a progressive hierarchical structure where features from multiple scales are refined one at a time through frequency-aware blocks. This progressive approach allows the decoder to iteratively enhance feature maps. Simultaneously, skip connections are incorporated into the decoder, directly passing information from earlier stages to later stages, combining high-resolution low-level details from earlier stages with semantically rich high-level features from later stages.

[0086] In the decoder, the algorithm employs a dynamically weighted fusion module to prevent a situation where one image has a stronger signal or less noise, causing the source image features to gain an advantage over another image feature when directly adding or concatenating features. This prevents the model from fully utilizing the image features from both sources, leading to a decrease in segmentation performance. The multimodal feature fusion mechanism is as follows: Figure 9 As shown, this module first receives two types of image features from the outputs of the dual-branch decoder and encoder structure, then calculates the attention weights from the two features, and finally fuses them according to the feature weights.

[0087] The calculation methods for weights and the generation methods for feature maps are as follows. First, in the grayscale image... and depth images After processing through their respective encoder-decoder paths and upsampling modules, feature maps are removed from the grayscale image modality and the depth image modality. Next, to adaptively balance the contribution of each modality, spatial attention weights for the grayscale and depth images are calculated for each channel. and The calculation of attention weights first transforms the features through a convolutional layer, and then generates the attention log γ through a non-linear activation function.I and γ R φ represents the shared convolutional layer and activation function, W is the kernel parameter, b is the bias term, and σ is the sigmoid activation function. This operation is performed for each channel, and the attention log value for each modality's image can be obtained for the corresponding channel.

[0088] γ I =φ(F I )=σ(W·F I +b),#(1)

[0089] γ R =φ(F R )=σ(W·F R +b),#(2)

[0090]

[0091] Then, the Softmax function is applied to each channel of each modality to obtain standardized attention weights. The Softmax normalization process ensures that α + β = 1 at each spatial location.

[0092]

[0093] Finally, weights are calculated on the corresponding channels of the two modalities to obtain the fused feature map, where ⊙ represents the successive multiplication of elements.

[0094] F fused =α⊙F I +β⊙F R #(6)

[0095] The algorithm will finally generate a fused feature map F. fused Before prediction, the segmentation weights need to be trained after inputting the segmentation head. The network training is optimized by adjusting parameters such as the learning rate, stride, and number of training samples per pass, achieving rapid convergence while avoiding overfitting. Only after obtaining the required segmentation weights can crack segmentation be performed on unlabeled grayscale-depth image pairs. The pixel-level crack classification and segmentation results obtained by the algorithm are shown below. Figure 5 .

[0096] Based on the output pixel-level crack classification and segmentation results, the classification and localization results of pavement crack images are statistically analyzed and summarized to provide a basis for maintenance management. The final results show that the category, confidence probability score, and coordinate information of the predicted bounding box for each crack are saved. Since the acquired image information contains temporal information, the actual mileage marker can be calculated from the image coordinates, thus outputting the road crack classification and marker location in text form. This completes the evaluation of the degree of road crack damage, providing a basis for pavement maintenance management.

[0097] The embodiments described above merely illustrate the implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A crack recognition method based on grayscale-depth heterogeneous image fusion, characterized in that: The steps are as follows: Step 1: Simultaneously acquire RGB and depth image data of the road surface using the vehicle-mounted camera and depth camera. The RGB image data of the road surface includes time information and size information, and the depth image data includes time information and size information. Step 2: Use frame synchronization to synchronize the road surface RGB image data and road surface depth image data obtained in Step 1. Take the road surface RGB image data and road surface depth image data at the same time, convert the RGB image to grayscale image, and create a matching grayscale-depth image pair. Step 3: Based on the image conditions, perform selective preprocessing on the images in the matched image pairs from Step 2, select a portion of the images as the training set, and perform data augmentation on the training set to expand the training set. Step 4: Use the Labelme tool to perform image segmentation and labeling on the training set. Assign category labels to cracks based on their type; The category labels include horizontal cracks, vertical cracks, crazing, and pits, and the coordinates of the marked boxes are saved in the corresponding XML files; Step 5: Train the training set labeled in Step 4 using a deep fusion recognition network based on heterogeneous image features from an end-to-end network to obtain the trained segmentation weights. Step 6: Input the trained segmentation weights into the heterogeneous image feature deep fusion recognition network based on end-to-end network, and use the preprocessed image from step 3 for detection to obtain pixel-level crack segmentation results; Step 7: Perform station number matching on the results from Step 6, and count the types and number of cracks.

2. The crack recognition method based on grayscale-depth heterogeneous image fusion according to claim 1, characterized in that: The road surface RGB image data mentioned in step 1 includes time information and size information; the depth image data includes time information and size information.

3. The crack recognition method based on grayscale-depth heterogeneous image fusion according to claim 1, characterized in that: In step 1, the vehicle-mounted camera and depth camera are positioned at a certain height relative to the road surface being measured. The formula for calculating the camera height H is as follows: Where GSD is the ground resolution of the image, f is the lens focal length, and a is the pixel size.

4. The crack recognition method based on grayscale-depth heterogeneous image fusion according to claim 1, characterized in that: The optional image preprocessing methods in step 3 include: (1) Adjust image brightness to increase image visibility (2) Image cropping reduces model computational load The training set data augmentation methods in step 3 include: (1) Use geometric transformations to augment images; (2) Enhance the image by randomly adjusting the brightness; (3) Enhance the image by randomly adjusting the contrast; The image in step 3 is preprocessed and augmented using an image preprocessing method that combines one or more methods.

5. A crack recognition method based on grayscale-depth heterogeneous image fusion according to claim 1, characterized in that: The weight calculation steps in step 5 include the following: 1) In grayscale images Depth image After processing through their respective encoder-decoder paths and upsampling modules, feature maps are removed from the grayscale image modality and the depth image modality. 2) Calculate the spatial attention weights for the grayscale image and depth image in each channel, expressed as follows: 3) Generate the feature map, as expressed below: c I =φ(F I )=σ(W·F I +b),# c R =φ(F R )=σ(W·F R +b),# Attention log γ is generated using a nonlinear activation function. I and γ R φ represents the shared convolutional layer and activation function, W is the convolutional kernel parameter, b is the bias term, and σ is the Sigmoid activation function; Perform this operation on each channel, and the image for each modality can obtain the corresponding channel's attention log value; The Softmax function is applied to each channel of each modality to obtain standardized attention weights; the Softmax normalization process ensures that α + β = 1 at each spatial location, as expressed below: The weighted summation is performed on the corresponding channels of the two modalities to obtain the fused feature map, where ⊙ represents the successive multiplication of elements: F fused =α⊙F I +β⊙F R #。