A method for automatic identification of road defects based on sparse autoencoder networks

Through a multi-channel imaging device that fuses visible light, infrared and depth images, combined with sparse self-coding networks and multi-scale convolution, the lack of illumination and three-dimensional information for disease detection in the prior art is solved, and efficient and accurate road disease recognition is achieved.

CN119963927BActive Publication Date: 2025-08-08RES INST OF HIGHWAY MINIST OF TRANSPORT +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510429604.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-08-08
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

Existing road disease detection methods rely on single-modal visible light images, are susceptible to light, shadows and surface reflections, making it difficult to accurately identify diseases in complex environments. Traditional deep learning methods lack three-dimensional information understanding, resulting in high false detection rates and degraded detection performance.

Method used

A multi-channel imaging device is used to obtain visible light, infrared and depth images, through multi-scale convolution, cross-layer feature fusion and adaptive disease detection optimization, a composite sparse self-coding network is used for feature learning, combined with morphological optimization and spatial clustering, unstable inputs are dynamically corrected, and system robustness is improved.

Benefits of technology

It significantly improves the accuracy, stability and calculation efficiency of disease detection, and can accurately identify road cracks, pits and other diseases in complex environments, reduce false detection rates, and improve the reliability and adaptability of identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963927B_ABST
    Figure CN119963927B_ABST
Patent Text Reader

Abstract

This invention discloses a method for automatically identifying road defects based on a sparse autoencoder network, relating to the field of image processing technology. The method comprises: Step 1: Using a multi-channel imaging device to simultaneously acquire visible light images, infrared images, and depth images to construct a multimodal training set; Step 2: Performing multi-scale region segmentation on the visible light and infrared images in the multimodal training set to obtain segmented regions and outputting a fused image; Step 3: Performing high-frequency noise suppression and mid- and low-frequency feature preservation on the fused image to enhance edge texture and obtain a fused processed image; Step 4: Using the fused processed image, constructing a composite sparse autoencoder network to identify road defects. This invention significantly improves the accuracy, stability, and computational efficiency of defect detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a method for automatically identifying road damage based on a sparse autoencoding network. Background Art

[0002] Currently, road defect detection technologies based on visual analysis fall into two main categories: image processing methods and deep learning methods. Image processing methods primarily extract features and segment road defect areas through traditional computer vision techniques such as edge detection, histogram analysis, and morphological operations. For example, methods such as Canny edge detection and Hough transform are widely used for crack detection, leveraging image gradient information and morphological features to identify crack areas. However, traditional image processing methods are sensitive to lighting conditions, road surface texture, and noise, and detection accuracy is easily affected by drastic lighting changes and complex road conditions (such as mixed asphalt and cement concrete sections). Furthermore, because morphological methods rely on fixed parameter settings, they have poor adaptability to different road scenarios, making it difficult to achieve highly robust defect detection.

[0003] With the development of deep learning technology, data-driven methods such as convolutional neural networks (CNNs) have been gradually applied to road defect detection. For example, existing studies have used deep learning models such as Faster R-CNN, YOLO, and U-Net for target detection or semantic segmentation of road defects. Compared to traditional image processing methods, deep learning can automatically learn features from large datasets, improving detection accuracy and adaptability. However, existing deep learning methods still have some challenges. For example, most deep learning methods only use visible light images for defect detection. Visible light imaging is significantly affected by factors such as illumination, shadows, and surface reflections, making it difficult to accurately identify certain defects in certain environments. For example, in low light conditions, the visibility of cracks and potholes is reduced, and the model may miss or misidentify them. Existing deep learning methods typically use standard CNN architectures for feature extraction. However, road defect morphology is complex and varied, and defects of different scales have significant differences in feature representation. For example, cracks typically exhibit long, thin linear features, while potholes have a large regional distribution. Fixed-size convolution kernels are difficult to adapt to defects of different scales, resulting in reduced detection performance. Because deep learning models are primarily trained on two-dimensional visible light images and lack a 3D understanding of the damaged area, in some cases the model may misidentify high-contrast areas such as shadows and road markings as damaged areas, affecting the final detection accuracy. Furthermore, due to perspective distortion in images taken from different angles, traditional deep learning methods struggle to ensure the alignment of damaged areas across images, impacting subsequent damage measurement and statistical analysis. Summary of the Invention

[0004] In view of this, the present invention provides a method for automatic identification of road defects based on a sparse autoencoder network. By fusing visible light, infrared and depth images, and utilizing multi-scale convolution, cross-layer feature fusion and adaptive defect detection optimization, it can achieve accurate detection and classification of road defects such as cracks, potholes, spalling, and rutting. The method uses a composite sparse autoencoder network for feature learning, and combines morphological optimization, spatial clustering and boundary tracking to extract geometric information of defects, such as crack length, pothole area and rutting depth. At the same time, the present invention introduces an online detection optimization mechanism to dynamically correct unstable inputs and improve the robustness of the system in complex environments. Compared with traditional methods, the present invention significantly improves the accuracy, stability and computational efficiency of defect detection, and is suitable for intelligent road inspection, infrastructure maintenance and urban road management, providing an efficient and accurate automated defect identification technology.

[0005] The technical solution adopted in the present invention is as follows:

[0006] A method for automatically identifying road defects based on a sparse autoencoder network, the method comprising:

[0007] Step 1: Use a multi-channel imaging device to simultaneously acquire visible light images, infrared images, and depth images to construct a multimodal training set;

[0008] Step 2: Perform multi-scale region segmentation on the visible light image and infrared image in the multimodal training set to obtain segmented regions; use the projection correction method to align the depth image and each segmented region in the same coordinate system, and output the fused image;

[0009] Step 3: Suppress high-frequency noise and retain medium- and low-frequency features on the fused image, enhance edge texture, and obtain a fused image;

[0010] Step 4: Use the fused processed images to construct a composite sparse autoencoder network to identify road damage.

[0011] Furthermore, step 4 specifically includes:

[0012] Step 4.1: Extract features from the fused image through multi-layer convolution and downsampling, and reconstruct the image in the decoding stage using a symmetrical upsampling and deconvolution structure. Apply sparse activation constraints to the intermediate hidden layers to reduce background redundancy and highlight road damage morphology.

[0013] Step 4.2: Insert the multi-scale convolution branch and the cross-layer feature fusion branch into the top-level feature map of the composite sparse autoencoder network to identify road damage.

[0014] Furthermore, step 1 specifically includes: using a multi-channel imaging device to simultaneously acquire visible light images, infrared images, and depth images for surface height measurement, and marking each channel image through a unified spatiotemporal synchronization mechanism to generate a multimodal original image sequence arranged in time sequence; dividing the multimodal original image sequence into a number of equidistant image blocks, and recording road surface features in each image block; road surface features include: cracks, potholes, spalling, and rutting; setting independent labels for cracks, potholes, spalling, and rutting in the image blocks to form a multimodal training set.

[0015] Furthermore, step 2 specifically includes: performing multi-scale segmentation processing on the visible light image and the infrared image to generate multiple segmentation areas of different sizes for subsequent detail enhancement and appearance alignment; performing three-dimensional centroid projection and local surface smoothing operations on the depth image to match the coordinate systems of the visible light image and the infrared image to obtain depth features; merging each segmentation area and the corresponding depth feature into a unified coordinate system; through repeated offset measurement and projection correction, ensuring that the spatial correspondence of the segmentation areas in the multimodal training set is consistent, and obtaining a fused image.

[0016] Furthermore, step 3 specifically includes: performing discrete wavelet decomposition and high-frequency noise suppression operations on the fused image to retain the mid-frequency and low-frequency parts that are discriminative for road disease identification; and using noise threshold clipping and gradient-guided enhancement on the high-frequency part to remove random interference when the multi-channel imaging device acquires visible light images, infrared images, and depth images, thereby obtaining a fused processed image.

[0017] Furthermore, step 4.1 specifically includes: constructing an encoding part of a composite sparse autoencoder network composed of multiple layers of convolution and downsampling, which is used to perform channel compression and feature extraction on the fused processed image; highlighting neurons that can recognize the shape and texture of road diseases through a sparse activation mechanism in each hidden layer of the encoding part, and suppressing irrelevant background images; using a deconvolution and upsampling structure symmetrical to the encoding part as the decoding part of the composite sparse autoencoder network, and gradually restoring the extracted features to the same size as the input multimodal training set in the spatial dimension to form a multimodal reconstruction output; by implementing a sparsity strategy that combines local constraints and global constraints in the encoding part and the decoding part, the composite sparse autoencoder network can reduce channel redundancy while ensuring feature resolution.

[0018] Furthermore, step 4.2 specifically includes: embedding a multi-branch disease detection module in the top-level feature map of the composite sparse autoencoder network, which module contains a multi-scale convolution branch and a cross-layer feature fusion branch; the multi-scale convolution branch is used to process road disease areas of different sizes and shapes, and extracts cracks and depressions through parallel convolution channels; the cross-layer feature fusion branch combines the edge and texture features retained in the encoding stage with the high semantic features of the top layer to distinguish areas with similar tones but different disease properties; the multi-branch disease detection module outputs road disease segmentation results with the same size as the multimodal training set, and classifies and labels different types of road diseases.

[0019] Furthermore, the method also includes: Step 5: first performing separate pre-training on the composite sparse autoencoder network, using the difference between the multimodal training set and the corresponding reconstruction result to perform targeted optimization on the encoding weights and decoding weights, thereby enhancing the sparse expression capability of multimodal features; then adding a multi-branch disease detection module for joint training, using disease labels to locate and classify the composite sparse autoencoder network, and performing end-to-end weight iteration based on the joint goal of reconstruction loss and disease recognition loss; setting independent recognition loss weights based on the shape differences between cracks and potholes, thereby enhancing the composite sparse autoencoder network's ability to determine road diseases.

[0020] Furthermore, the method also includes: Step 6: when inferring the unlabeled image, the unlabeled image is input into the encoding part of the composite sparse autoencoder network, and then the multi-branch disease detection module is used to obtain the disease segmentation map and type discrimination result; for the detected disease area, the actual surface height change or area is determined by combining the low-level features and the depth image, and the error areas caused by tiny holes and illumination effects are screened out; by online monitoring of the consistency of the images of each channel of the multi-channel imaging device, dynamic correction of the unstable input scene is achieved.

[0021] Furthermore, the method also includes: step 7: performing a morphological filling operation on the identified diseased area to eliminate pixel-level holes that may exist in narrow areas; performing spatial clustering and boundary tracking on the segmentation results output by the multi-branch disease detection module, and outputting the crack length, pothole area, spalling edge coordinates and rutting depth range; comparing the crack length, pothole area, spalling edge coordinates and rutting depth range with the identified areas in the multimodal training set pixel by pixel, and statistically analyzing the recognition accuracy, false detection rate and missed detection rate, and recording the processing speed and recognition confidence of different road disease types.

[0022] By adopting the above technical solution, the present invention produces the following beneficial effects: the existing road disease detection method mainly relies on a single modality of visible light images for feature extraction and analysis, while the present invention fully utilizes the complementary characteristics of multiple modalities by fusing visible light, infrared and depth images, thereby improving the accuracy and reliability of disease detection. Visible light images can provide information such as texture, cracks, potholes, etc. on the road surface, infrared images are used to detect abnormal road surface temperature, and depth images can capture the three-dimensional morphology of potholes, peeling, and rutting. Through projection correction, appearance alignment and depth data optimization, the present invention ensures the precise spatial matching of different modal data, making the expression of the diseased area in the fused image more complete, and providing more comprehensive input data for subsequent feature extraction and disease classification. Traditional deep learning methods are often interfered with by background noise and irrelevant information in disease detection tasks, resulting in a high false detection rate. The present invention adopts a composite sparse autoencoder network, and through a sparse activation mechanism, ensures that the activation pattern of the neural network is more concentrated on the diseased area, while reducing the influence of the background area. During the training process, the network automatically extracts the core features of road diseases through an unsupervised learning mechanism and imposes sparsity constraints on the hidden layer, making the expression of disease features more compact, reducing computational redundancy, and improving recognition stability. For example, in the crack detection task, the sparse autoencoder network can effectively remove noise interference caused by factors such as road surface texture and shadows, making the crack area more clear and identifiable, thereby improving the final detection accuracy. The present invention embeds a multi-branch disease detection module in the top-level feature map of the composite sparse autoencoder network. The module contains multi-scale convolution branches and cross-layer feature fusion branches, which can effectively improve the detection ability of different types of road diseases. Traditional deep learning methods use fixed-size convolution kernels for feature extraction, which makes it difficult to take into account large-scale potholes and small-scale cracks at the same time. The present invention uses a multi-scale convolution strategy to extract features of the disease area at different scales, ensuring that the boundary information of the cracks is enhanced, while ensuring that the overall morphology of diseases such as potholes and spalling is not ignored. Furthermore, the cross-layer feature fusion branch combines low-level edge information with high-level semantic features, giving the network greater discriminative power when distinguishing areas of similar hue but different damage properties. For example, this method can effectively prevent road joints from being misidentified as cracks, reducing false positives and improving the reliability of damage identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 This is a flow chart of a method for automatically identifying road defects based on a sparse autoencoder network in an embodiment of the present invention. DETAILED DESCRIPTION

[0024] All features disclosed in this specification, or all steps in the disclosed methods or processes, except mutually exclusive features and / or steps, can be combined in any manner.

[0025] Any feature disclosed in this specification (including any appended claims and abstract), unless otherwise stated, may be replaced by other equivalent or similar features. In other words, unless otherwise stated, each feature is only an example of a series of equivalent or similar features.

[0026] Example 1, reference Figure 1 : A method for automatically identifying road damage based on a sparse autoencoder network, the method comprising:

[0027] Step 1: Use a multi-channel imaging device to simultaneously acquire visible light images, infrared images, and depth images to construct a multimodal training set;

[0028] During the entire image acquisition process, visible light imaging is used to capture detailed information about the road surface, including crack morphology, color changes in damaged areas, and texture characteristics. Generally speaking, visible light images provide clear structural information in daytime environments. However, under unstable lighting conditions, shadows, reflections, and variations in light intensity can interfere with recognition accuracy. To address this issue, the present invention incorporates infrared imaging technology, which detects temperature differences on the road surface to aid in the identification of damaged areas. Because damaged areas such as cracks and potholes typically exhibit different thermal radiation characteristics, their temperature distribution often exhibits some abnormality compared to normal road surfaces. Therefore, infrared imaging can provide additional identification criteria, improving the recognition of damaged areas. However, accurately assessing the depth, morphology, and volumetric characteristics of damaged areas using only visible light and infrared information remains difficult. Therefore, the present invention also incorporates depth imaging technology to capture three-dimensional morphological information of the road surface. By measuring the surface's undulations, depth images can identify structural changes caused by subsidence, deformation, or damage. Compared to traditional two-dimensional image analysis, depth information provides a more intuitive spatial measurement image, enabling precise characterization of the geometric characteristics of damaged areas. In addition, in complex environmental conditions, such as at night or in areas with heavy shadows, depth imaging is more stable than visible light imaging and is not affected by changes in lighting, thereby improving the system's adaptability in different scenarios.

[0029] To ensure that the acquired multimodal images can be effectively used for subsequent recognition tasks, the multi-channel imaging device of the present invention needs to have a synchronous acquisition capability to ensure the spatial consistency of different modal images at the same time point. Since the physical positions of visible light, infrared, and depth cameras may differ, their respective imaging ranges and angles are also different, so the directly acquired images may have perspective deviations. If this deviation is not corrected, it will cause spatial misalignment of different modal images during fusion, thereby affecting recognition accuracy. Therefore, during the image acquisition process, it is necessary to use camera calibration technology to calculate the relative position relationship of each sensor, and use geometric transformation methods to pre-process the images to ensure that all modal images can be aligned to the same coordinate system. This allows images of different modalities to be compared with each other, thereby fully leveraging the advantages of multimodal fusion. In addition, in road disease detection tasks, image quality is crucial to recognition performance. During the image acquisition stage, in addition to ensuring the synchronous acquisition and alignment of multimodal images, the present invention also needs to ensure that the image resolution, signal-to-noise ratio, and dynamic range meet the requirements of the recognition task. For example, in depth imaging, the sensor's measurement accuracy determines whether subtle road depressions can be accurately identified, while the infrared image's temperature resolution directly impacts the ability to discern temperature differences in affected areas. Therefore, to improve image quality, the present invention employs an adaptive exposure control strategy during the imaging process, enabling the sensor to automatically adjust imaging parameters based on ambient brightness, target reflectivity, and thermal radiation characteristics, avoiding information loss due to overexposure or underexposure. Furthermore, to reduce interference from ambient noise, the present invention also employs noise suppression algorithms during the image acquisition phase, such as utilizing multi-frame image fusion technology to reduce the impact of random noise or employing a temperature compensation algorithm to correct for temperature drift in infrared images.

[0030] Step 2: Perform multi-scale region segmentation on the visible light image and infrared image in the multimodal training set to obtain segmented regions; use the projection correction method to align the depth image and each segmented region in the same coordinate system, and output the fused image;

[0031] In road defect detection, the shape and size of defects often show large variations. For example, cracks may be elongated and irregular in shape, potholes may have a large area, and bulges may appear in the form of round or irregular protrusions. Traditional fixed window segmentation methods are difficult to effectively adapt to the diversity of defects, which may cause important defect features to be ignored or background noise to be mistaken for defect areas. Therefore, the present invention adopts a multi-scale region segmentation strategy to adapt to defect targets of different scales. On visible light and infrared images, the superpixel segmentation method is first used to divide the image into multiple small areas. Each area is clustered according to color, edge information and texture features to ensure that the segmented area can better fit the boundary of the defect. Compared with traditional pixel-level segmentation methods, the superpixel method can reduce unnecessary computational complexity while retaining important structural information of the image, thereby improving the accuracy and efficiency of segmentation. In order to further enhance the robustness of segmentation, the present invention introduces a region growing method based on superpixel segmentation. By setting an initial seed area, it gradually expands to neighboring pixels with similar features to ensure that the segmentation result can cover the complete defect area. During the region growing process, multiple scale parameters are set to prevent over-extension of smaller cracks and to ensure that larger potholes are fully covered without being fragmented. Simultaneously, an edge detection algorithm is used to optimize the segmentation results, ensuring that region boundaries accurately conform to the disease morphology and avoiding morphological distortion caused by over-segmentation. After this series of processing, the visible and infrared images are segmented into a series of relatively independent disease candidate regions. Subsequent feature analysis can be performed on these segmented regions, reducing computational redundancy and improving recognition accuracy.

[0032] However, accurately capturing the three-dimensional morphology of a defect using only visible light and infrared image segmentation results remains difficult. For example, two cracks may appear similar in a visible light image, but their depths may differ significantly. Misalignment or scale mismatch in depth information can lead to inaccurate estimates of the defect's morphology. Therefore, the present invention employs a projection correction method to align the depth image with the segmented region, ensuring that the defect's three-dimensional morphology matches the two-dimensional texture information. Because the imaging coordinate system of a depth sensor typically differs from that of a visible light camera, directly overlaying the two images can lead to problems such as boundary misalignment and morphological distortion. Therefore, the present method first utilizes calibration techniques to determine the spatial transformation relationship between the depth sensor and visible light camera. Then, a perspective transformation is used to reproject the depth image into the visible light coordinate system, ensuring that defects at the same physical location are correctly aligned in images from different modalities. This allows the depth information to be accurately mapped to the segmented regions in the visible light and infrared images, ensuring precise characterization of the defect's geometric features. Furthermore, given that the resolution of the depth image may differ from that of the visible light image, directly performing a projection transformation can result in information loss or resolution degradation. To address this issue, the present invention introduces a depth interpolation algorithm. During the depth image projection process, interpolation compensates for misaligned pixels to ensure that the projected depth image maintains sufficient spatial consistency with the visible light image. Furthermore, to reduce depth errors caused by sensor noise or environmental interference, the present invention further employs an edge alignment algorithm after the projection transformation to match the projected depth image with the boundaries of the segmented region. This ensures that the depth information is precisely aligned with the diseased area, thereby improving the accuracy of subsequent feature extraction.

[0033] Step 3: Suppress high-frequency noise and retain medium- and low-frequency features on the fused image, enhance edge texture, and obtain a fused image;

[0034] In road damage images, high-frequency components mainly manifest as noise, texture details, and edge mutation areas, while medium and low-frequency components correspond to larger structural information, such as the overall morphology of cracks, the regional contours of potholes, and the gradual trend of bulges. However, due to the physical characteristics of the imaging equipment and the influence of environmental factors, the high-frequency components in the image often contain a large amount of irrelevant noise, such as random noise introduced by factors such as road particles, dust, and lighting changes. These noises may mask the true characteristics of the damage to a certain extent, making it difficult for subsequent deep learning models to effectively distinguish between normal areas and damage areas. Therefore, the present invention first adopts an adaptive high-frequency noise suppression method to smooth the fused image to reduce the interference of noise on the extraction of damage edges. In this process, it is necessary to ensure that the removal of noise does not affect the actual damage contour, otherwise it may cause the crack boundary to be blurred, affecting subsequent feature learning. Therefore, the present invention adopts a noise suppression strategy based on local contrast preservation, which reduces high-frequency noise while maintaining the significance of the edge area, so that the damage area can be more clearly displayed in the image.

[0035] After completing the high-frequency noise suppression, the next key task is to retain and enhance the medium and low-frequency features to highlight the main structural information of the disease. In road disease detection, the edges of diseased areas such as cracks, potholes, and bulges usually show certain texture change characteristics. For example, the edges of cracks are often steeper and have a higher grayscale gradient, while the pothole area shows a smooth grayscale transition. If you only rely on the grayscale value or color information of the original image, you may not be able to accurately distinguish these types of diseases. Therefore, it is necessary to process the medium and low-frequency components of the image to make the characteristics of the diseased area more obvious. The present invention adopts an edge enhancement method. By extracting the gradient information in the image and performing multi-scale fusion on it, the edges of the diseased area are highlighted while reducing the interference of the background area. The image processed in this way can more clearly show the morphological characteristics of the disease and provide a more discriminative input image for the subsequent deep learning model.

[0036] In the process of edge enhancement, the texture characteristics of the diseased area also need to be considered. Since different road materials have different surface textures, for example, asphalt pavements usually appear as a relatively rough granular structure, while cement pavements may have a relatively uniform grayscale distribution, these material characteristics may affect the visual manifestation of the disease. Therefore, in the method of the present invention, by analyzing the local texture features of the image, the enhancement strategy is adaptively adjusted so that different types of road surfaces can be properly processed. For example, for asphalt pavements with a relatively rough surface, stronger edge sharpening may be required to distinguish the diseased area from the normal road texture, while for smoother cement pavements, the edge enhancement strength needs to be appropriately reduced to avoid artifacts caused by over-enhancement. This adaptive texture enhancement method enables the present invention to maintain a high recognition accuracy in different types of road environments.

[0037] Step 4: Use the fused processed images to construct a composite sparse autoencoder network to identify road damage.

[0038] When constructing a sparse autoencoding network, it is first necessary to perform feature learning on the fused image so that the network can autonomously extract effective features that distinguish between diseased areas and normal roads. Unlike traditional manual feature extraction methods, the sparse autoencoding network is an unsupervised learning method that can learn the most discriminative representation directly from the image without relying on manually labeled features. The basic idea is to use a neural network to encode the input image and reconstruct it at the output, so that the network can learn the potential structure of the input image by minimizing the reconstruction error between the input and output. In order to ensure that the features learned by the network have stronger discriminative ability, the present invention adds sparsity constraints during the training process, so that the activation state of the neurons is as sparse as possible, that is, only a few neurons respond to the input image in each calculation. This sparsity enables the network to focus more on key features, reduce the impact of redundant information, and improve the generalization ability of the model, so that it can maintain a high recognition accuracy for different types of road diseases.

[0039] In terms of network structure design, the present invention adopts a composite sparse autoencoder network to fully utilize the complementary characteristics of multimodal images. Since the input image contains visible light, infrared and depth information, simply splicing them together may lead to information redundancy or feature mismatch. Therefore, the present invention adopts a branched network architecture, designs autoencoders for images of different modalities, and performs feature fusion at a high level. This design ensures that each modality of image can fully retain its own characteristics during the encoding process, while enabling the network to capture cross-modal correlation information through high-level fusion. For example, the texture features of the visible light image can be jointly analyzed with the geometric structure features of the depth information to better characterize the morphology of the diseased area, while the infrared image can provide clues to temperature anomalies, helping the network to more accurately distinguish between types of diseases caused by different factors such as material aging, cracks, and potholes. Through such a composite architecture, the network can maximize the advantages of multimodal images and improve the accuracy and stability of recognition.

[0040] Example 2: Step 4 specifically includes:

[0041] Step 4.1: Extract features from the fused image through multi-layer convolution and downsampling, and reconstruct the image in the decoding stage using a symmetrical upsampling and deconvolution structure. Apply sparse activation constraints to the intermediate hidden layers to reduce background redundancy and highlight road damage morphology.

[0042] In step 4.1, feature extraction is first performed on the fused image. This process primarily relies on multi-layer convolution and downsampling operations. The convolution layer extracts local features, extracting edge, texture, and structural information from the image through filters of varying scales. In the lowest-level convolution operation, the network primarily learns basic features, such as local brightness variations on the road surface, the boundary contours of cracks, and the depth gradients of potholes. This information forms the basis for disease detection. As the number of network layers increases, the features extracted by the convolution layer gradually become more complex. For example, the network may learn the directionality of cracks, the grayscale distribution pattern of potholes, and the morphological differences between different disease types. At the same time, downsampling is used to reduce the spatial resolution of the feature map to reduce computational effort and enhance the network's ability to understand global features. The main downsampling methods include maximum pooling and average pooling. Maximum pooling retains the most prominent feature information, helping to enhance the salience of diseased areas, while average pooling helps smooth the feature map and reduce noise interference. The present invention combines these two methods in the downsampling process, ensuring that the extracted features have both good discrimination and sufficient stability.

[0043] After feature extraction is complete, the network enters the decoding phase, where image reconstruction is performed through a symmetrical upsampling and deconvolution structure. The purpose of upsampling is to restore the spatial resolution of the feature map, allowing the network to take into account detailed information during the learning process, while the deconvolution operation is used to restore the structural features of the diseased area, making it clearer and more identifiable. In traditional convolutional neural networks, after multiple layers of downsampling, the spatial information of the feature map is often lost, resulting in deviations in the final recognition results when locating the diseased area. The present invention adopts a symmetrical structure in the decoding phase, namely upsampling and deconvolution, allowing the network to effectively restore the spatial information of the diseased area while reducing the loss of detail caused by downsampling. To further highlight the morphology of the disease, the network imposes sparse activation constraints on the middle hidden layers, so that neuronal activation is more concentrated in key areas rather than being evenly distributed across the entire image. This sparsity helps reduce interference from redundant background information and improves the network's focus on the diseased area. For example, in the crack detection task, the network will preferentially activate neurons near the crack boundary while ignoring irrelevant features of the road texture, thereby improving recognition accuracy.

[0044] Step 4.2: Insert the multi-scale convolution branch and the cross-layer feature fusion branch into the top-level feature map of the composite sparse autoencoder network to identify road damage.

[0045] In step 4.2, the network further inserts a multi-scale convolution branch and a cross-layer feature fusion branch into the top-level feature map to enhance the defect detection capability. Road defects vary greatly in morphology and scale. For example, some cracks may be just tiny lines, while potholes may occupy a larger area. Therefore, relying solely on a fixed-scale convolution kernel for feature extraction may not be able to take into account defect targets of different scales. The present invention designs a multi-scale convolution branch at the top level of the network, that is, using convolution kernels of different sizes in different channels to capture features of different scales. For example, a small-scale convolution kernel is used to extract the edge information of small cracks, while a large-scale convolution kernel is used to capture the overall structural features of potholes. In this way, the network can pay attention to local details and global morphology at the same time, thereby improving the detection capability of defects of different scales. In addition, the present invention also adopts a cross-layer feature fusion strategy, that is, fusing the detail features of the lower layer with the semantic features of the higher layer, so that the network can simultaneously utilize local texture information and global structural information when identifying defects. For example, in some cases, low-level features may provide edge information of the diseased area, while high-level features can reflect the overall shape of the disease. Only by combining these two types of information can more accurate recognition be achieved. Therefore, through cross-layer fusion, the network can more comprehensively understand the characteristics of the diseased area and improve recognition accuracy.

[0046] Example 3: Step 1 specifically includes: using a multi-channel imaging device to simultaneously acquire visible light images, infrared images, and depth images for surface height measurement, and marking each channel image through a unified spatiotemporal synchronization mechanism to generate a multimodal original image sequence arranged in time sequence; dividing the multimodal original image sequence into a number of equidistant image blocks, and recording road surface features in each image block; road surface features include: cracks, potholes, spalling, and rutting; setting independent labels for cracks, potholes, spalling, and rutting in the image blocks to form a multimodal training set.

[0047] Specifically, during the image acquisition process, the multi-channel imaging device simultaneously acquires visible light images, infrared images, and depth images for surface height measurement. Among them, visible light imaging is mainly used to capture the color, texture, and crack morphology characteristics of the road surface, infrared imaging is used to detect the temperature distribution of the road surface, and depth imaging is used to measure the height changes of the road surface to identify the three-dimensional structural characteristics of potholes, ruts, and other defects. In order to ensure that these three modal images can be effectively aligned in the subsequent analysis process, the present invention introduces a unified spatiotemporal synchronization mechanism, that is, during image acquisition, the images of all channels are recorded synchronously and accurately marked by timestamps. The introduction of this mechanism enables the multimodal images at each time point to maintain strict temporal consistency, avoiding the problem of information misalignment caused by acquisition delays or perspective deviations of different sensors. In addition, in terms of spatial synchronization, the present invention corrects the imaging coordinates of each sensor through calibration technology to ensure that images of different modalities can be mapped into the same spatial reference frame, thereby improving the accuracy of image fusion.

[0048] After completing the image acquisition, the system needs to organize the multimodal original images and generate a time-series image sequence. Since the distribution of road diseases has a certain spatial continuity, there may be some overlapping areas in the images collected at different time points. Therefore, the method of the present invention adopts a structured storage method based on time series to ensure the integrity and traceability of the image. In the image organization process, the original image is first divided into blocks, and the continuously collected images are divided into several image blocks of equal distance. Each image block represents a road segment of a fixed length and contains images of all modalities in the segment. The advantage of this division method is that it can improve the efficiency of image storage and processing while ensuring the local consistency of the image, so that subsequent feature extraction and annotation can be performed on smaller computing units, thereby reducing computing overhead and improving the real-time performance of the system.

[0049] In each image block, the system needs to record the characteristic information of the road surface so that these features can be used for supervised learning during subsequent training. The method of the present invention sets independent markers for the four main types of road defects, namely cracks, potholes, spalling and rutting, to ensure that the training image can accurately reflect the distribution of the defects. Among them, cracks are a common form of road surface defects, usually appearing as long and irregular lines, with obvious contrast in visible light images, and may show different brightness distributions in infrared images due to temperature differences. Potholes are large-area depressions caused by damage to the pavement structure. In depth images, they usually appear as obvious areas of reduced height, and may appear as shadows or temperature anomalies in visible light and infrared images. Spalling refers to the localized shedding of road surface materials, which may lead to varying degrees of surface roughness changes. In infrared imaging, it may appear as uneven temperature, and in depth images, it may appear as smaller-scale height changes. Rutting is a depression in the road surface caused by long-term vehicle rolling over. It usually appears as a relatively continuous groove in the depth image, but may present different visual characteristics in visible light and infrared images due to different lighting angles or temperature distributions.

[0050] Example 4: Step 2 specifically includes: performing multi-scale segmentation processing on the visible light image and the infrared image to generate multiple segmentation areas of different sizes for subsequent detail enhancement and appearance alignment; performing three-dimensional centroid projection and local surface smoothing operations on the depth image to match the coordinate systems of the visible light image and the infrared image to obtain depth features; merging each segmentation area and the corresponding depth feature into a unified coordinate system; through repeated offset measurement and projection correction, ensuring that the spatial correspondence of the segmentation areas in the multimodal training set is consistent, and obtaining a fused image.

[0051] Specifically, in the first step of image processing, visible light and infrared images are respectively subjected to multi-scale segmentation to extract disease features of different scales. The morphology and distribution of road diseases have strong scale dependence. For example, cracks may appear as slender texture features, while potholes may occupy a larger area. Therefore, a single-scale segmentation method may not be able to take into account the diversity of these disease targets at the same time. Therefore, the present invention adopts an adaptive multi-scale segmentation strategy, that is, performing segmentation operations at different scales to generate multiple segmented regions of different sizes, so that the network can be optimized for disease targets of different scales in subsequent processing. Specifically, at a smaller scale, the segmentation algorithm will give priority to extracting small disease areas such as cracks, while at a larger scale, it will identify broader potholes, spalling and rutting areas. Through this multi-scale strategy, the complete expression of the disease area at different spatial levels can be ensured, providing richer feature information for the subsequent deep learning model.

[0052] After completing the segmentation of visible light and infrared images, the depth image needs to be geometrically aligned using a three-dimensional centroid projection method. Since the acquisition method of the depth image is different from that of visible light and infrared imaging, its coordinate system is usually in a different reference frame, so directly aligning the depth information with the visible light image may cause deviations. The present invention calculates the three-dimensional centroid coordinates of each depth area and uses projection transformation to map it to the coordinate system of the visible light and infrared images, thereby ensuring the spatial consistency of the three modal images. In this process, local smoothing of the depth image plays a key role. Since the depth sensor may be affected by environmental noise during the measurement process, resulting in local fluctuations in the surface image, direct projection may cause instability of the depth feature. Therefore, before projection, the present invention performs local surface smoothing on the depth image to make the surface height information more coherent, reduce noise interference, and improve matching accuracy. Through this series of processing, the features of the depth image can be accurately matched with the segmented areas of the visible light and infrared images, thereby ensuring the fusion quality of the multimodal image.

[0053] After completing the multi-scale segmentation and depth image projection, each segmented area needs to be merged with the corresponding depth features to form a unified spatial expression. Since the diseased area is expressed differently in different modalities, for example, cracks in visible light images may only appear as color changes, while depth images may provide additional morphological information. Therefore, simply directly splicing the images of each modality is not enough to ensure the fusion quality. The present invention adopts a matching strategy based on spatial consistency, that is, in the segmented areas of visible light and infrared images, the depth feature area that best matches its boundary is found and fused. In this way, each segmented area not only contains the original texture and temperature information, but also combines the corresponding depth features, making the morphological information of the disease more complete. In addition, since depth images usually have a lower resolution and visible light images have a higher resolution, in order to avoid the problem of blurred boundaries after fusion, the present invention adopts a local interpolation optimization method to compensate the resolution of the depth image so that it can maintain the same spatial detail expression as the visible light image, thereby improving the overall recognition accuracy.

[0054] In the entire image fusion process, spatial consistency is a key factor affecting the accuracy of disease recognition. Since there may be perspective offsets in the acquisition of images of different modalities, there may still be slight alignment errors even after projection transformation. If these errors are not handled, they may cause misalignment of the diseased area, thereby affecting the accuracy of recognition. Therefore, the present invention introduces a repeated offset measurement and projection correction mechanism in the generation process of the fused image to ensure that the segmented areas remain spatially consistent in the multimodal training set. Specifically, after the initial projection, the present invention will perform offset measurements based on multiple perspectives, that is, the position of the same diseased area is compared multiple times in different coordinate systems. If a diseased area is found to be offset in a certain modality, the projection correction method is used to adjust it until its spatial position in all modalities is completely consistent. In this way, the spatial offset problem caused by sensor error or calculation error can be effectively reduced, ensuring that all diseased areas have an accurate spatial correspondence in the fused image.

[0055] Example 5: Step 3 specifically includes: performing discrete wavelet decomposition and high-frequency noise suppression operations on the fused image to retain the mid-frequency and low-frequency parts that are discriminative for road disease identification; and applying noise threshold clipping and gradient-guided enhancement to the high-frequency part to remove random interference when the multi-channel imaging device acquires visible light images, infrared images, and depth images, thereby obtaining a fused processed image.

[0056] Specifically, during the entire optimization process, discrete wavelet decomposition is first performed on the fused image to obtain image features at different scales. The core idea of wavelet decomposition is to divide the image signal into different frequency components, where the low-frequency part contains overall structural information, the mid-frequency part contains texture features, and the high-frequency part mainly contains details and noise information. For the task of road defect recognition, the low-frequency part is mainly used to describe the overall morphology of the road surface, such as the global distribution of potholes and ruts, while the mid-frequency part contains the boundary features of cracks and texture information of peeling areas. Therefore, these components have a high degree of discrimination for defect recognition. However, the high-frequency part is often greatly affected by the noise of the imaging equipment, especially in low-light environments or complex scenes. Random interference may cause the features of the defect area to be blurred, or even introduce erroneous defect boundaries. Therefore, after wavelet decomposition, different frequency components need to be processed separately to maximize the retention of defect information while suppressing irrelevant noise.

[0057] In the process of high-frequency noise suppression, the present invention adopts a method based on noise threshold clipping to remove random interference introduced when the multi-channel imaging device acquires visible light images, infrared images and depth images. In traditional image denoising methods, mean filtering or Gaussian filtering is usually used to smooth the image, but these methods often lead to the loss of edge details of the diseased area, thereby affecting the recognition effect. The method of the present invention is different from traditional filtering technology. It is based on the statistical characteristics of wavelet coefficients, sets a noise threshold in the high-frequency part, and clips high-frequency components less than the threshold, thereby eliminating high-frequency interference signals introduced by sensor noise, illumination changes or environmental interference. The advantage of this method is that it can retain the true edge information of the diseased area to the greatest extent while removing noise, so that cracks, pits and other disease forms remain clearly visible after denoising.

[0058] However, simply removing high-frequency noise through threshold clipping may not be enough to fully restore the features of the diseased area. Therefore, the present invention further adopts gradient-guided enhancement technology to optimize the edge features of the diseased area. The basic idea of gradient-guided enhancement is to use image gradient information to enhance the boundaries of the diseased area, so that cracks, pits, peeling and other disease morphologies are more prominent. In this process, the present invention first calculates the gradient distribution after wavelet transform, and determines the significance level of the diseased area through statistical analysis of the gradient amplitude. Subsequently, in the denoised image, adaptive enhancement is applied to the areas with higher gradient amplitudes, so that the contrast of the diseased edges is improved, thereby improving the recognizability of the diseased area. The advantage of this method is that it can further enhance the contour information of the diseased area while denoising, so that the subsequent deep learning model can more accurately learn the morphological characteristics of the disease during feature extraction, thereby improving the final recognition effect.

[0059] After the above processing, the final fused image not only removes high-frequency noise, but also retains the mid-frequency and low-frequency parts that are most discriminatory for disease identification. Compared with traditional noise suppression methods, the discrete wavelet decomposition of the present invention combined with the strategy of noise threshold clipping and gradient-guided enhancement can more accurately separate disease features from random noise, making the disease area clearer in the fused image, and providing a higher-quality training image for the subsequent deep learning model based on the sparse autoencoder network. In addition, due to the multi-scale characteristics of the wavelet transform, the method of the present invention can adapt to disease targets of different scales. Whether it is a small crack or a large area of pits, it can be effectively enhanced in this processing process, thereby improving the disease recognition ability and generalization performance of the entire system.

[0060] Example 6: Step 4.1 specifically includes: constructing an encoding part of a composite sparse autoencoder network composed of multi-layer convolution and downsampling, which is used to perform channel compression and feature extraction on the fused processed image; highlighting neurons that can recognize the shape and texture of road diseases through a sparse activation mechanism in each hidden layer of the encoding part, and suppressing irrelevant background images; using a deconvolution and upsampling structure symmetrical to the encoding part as the decoding part of the composite sparse autoencoder network, and gradually restoring the extracted features to the same size as the input multimodal training set in the spatial dimension to form a multimodal reconstruction output; by implementing a sparsity strategy that combines local constraints and global constraints in the encoding part and the decoding part, the composite sparse autoencoder network can reduce channel redundancy while ensuring feature resolution.

[0061] Specifically, during the feature extraction phase, the encoding portion of the composite sparse autoencoder network consists of multiple layers of convolution and downsampling modules. This aims to perform channel compression on the fused image while extracting the most discriminative defect features. In the road defect detection task, while the fusion of visible, infrared, and depth information provides richer features, directly inputting the raw multi-channel data results in excessive computational effort and may also contain redundant information in some features. Therefore, in the first layer of the encoding portion, the network first extracts basic features such as crack edge information, pothole morphology, and the streamline distribution of ruts through standard convolution operations. Subsequently, the receptive field of the convolution kernel is gradually increased in subsequent layers, enabling the network to learn the global distribution of defect areas at a larger scale while further compressing the number of data channels to reduce computational overhead. Downsampling plays a key role in this process, not only reducing the spatial dimensionality of the feature map and improving computational efficiency, but also enhancing the network's ability to perceive global patterns, allowing the overall morphology of the defect to be more clearly expressed in high-level features.

[0062] To further improve the ability to distinguish disease features and reduce background interference, the present invention introduces a sparse activation mechanism in each hidden layer of the encoding part. That is, by setting sparsity constraints, the network's neuron activation pattern in each layer is made more selective. Specifically, after the convolution operation, the network automatically calculates the response strength of each neuron and, based on the set sparsity target, retains only the neurons most sensitive to the diseased area, while suppressing the response to background information. The introduction of this mechanism allows the network to focus more on diseased areas such as cracks, potholes, and spalling, while reducing the misidentification of irrelevant road surface textures. For example, in visible light images, subtle texture changes on the road surface may be mistaken for cracks. However, through the sparse activation mechanism, the network can automatically filter out this irrelevant information, retaining only the disease edges and morphological features that are truly of diagnostic value. In addition, the temperature distribution in infrared images may be affected by the ambient temperature, resulting in temperature anomalies in some normal areas. Sparsity control can prevent the network from overfitting to these irrelevant temperature changes, thereby improving the robustness of recognition.

[0063] After completing the feature compression of the encoding part, the decoding part adopts a deconvolution and upsampling structure symmetrical to the encoding part to gradually restore the spatial resolution of the feature map and map the learned disease features back to the same size as the input data. Since the downsampling operation of the encoding part will lose some spatial information, the key task of the decoding part is to restore the spatial details of the diseased area through deconvolution operations and ensure that the final output feature map can maintain consistency with the original input data. In traditional neural networks, a simple deconvolution operation may cause the edges of the feature map to be blurred, affecting the accurate positioning of the diseased area. The present invention combines upsampling operations in the decoding process, so that the restored feature map is smoother in the spatial dimension and can be more accurately aligned with the disease boundary. For example, in the crack detection task, upsampling can ensure that the edges of the cracks will not become blurred due to the loss of information in the downsampling process, thereby improving the accuracy of the final disease identification. In addition, in the pit detection task, the combination of deconvolution and upsampling enables the accurate reconstruction of the pit morphology. With the support of multi-scale feature fusion, the depth distribution of the pit can be more accurately characterized, allowing the network to not only identify the presence of diseases, but also to more carefully classify the severity of the diseases.

[0064] During the training of the entire encoder-decoder structure, the present invention employs a sparsity strategy that combines local and global constraints to further optimize network learning and reduce feature channel redundancy. Regarding local constraints, the present invention introduces local sparsity regularization during neuron activation at each layer, allowing feature learning to focus more on areas of diagnostic value while avoiding interference from background information. For example, when learning crack features, the network only activates neurons near the crack edges, while maintaining a lower activation level for large, disease-free areas, thereby reducing computational redundancy and improving the contrast of disease detection. Regarding global constraints, the present invention applies global sparsity control across the entire network hierarchy. This involves analyzing the information contribution of different channels and automatically adjusting the weights of feature channels to ensure that the ultimately learned features remain highly interpretable on a global scale. For example, in a multimodal data fusion scenario, if a channel in a visible light image contributes less to disease features than depth information, the network will automatically reduce the weight of that visible light channel and increase the weight of the depth channel, thereby achieving more optimized feature representation. This mechanism not only improves the effectiveness of feature learning, but also reduces redundancy in the network calculation process, enabling the entire model to achieve more efficient computing performance while ensuring feature resolution.

[0065] Example 7: Step 4.2 specifically includes: embedding a multi-branch disease detection module in the top-level feature map of the composite sparse autoencoder network, which module contains a multi-scale convolution branch and a cross-layer feature fusion branch; the multi-scale convolution branch is used to process road disease areas of different sizes and shapes, and extract cracks and depressions through parallel convolution channels; the cross-layer feature fusion branch combines the edge and texture features retained in the encoding stage with the high semantic features of the top layer to distinguish areas with similar tones but different disease properties; the multi-branch disease detection module outputs a road disease segmentation result that is consistent with the size of the multimodal training set, and classifies and labels different types of road diseases.

[0066] Specifically, the multi-branch defect detection module consists of a multi-scale convolution branch and a cross-layer feature fusion branch. The multi-scale convolution branch primarily processes road defect regions of varying sizes and shapes. In road defect detection tasks, cracks typically appear as elongated, irregular textures, while potholes or depressions exhibit large, regional variations. This makes it difficult for a fixed-size convolution kernel to simultaneously capture defect characteristics of both scales. Therefore, the present invention employs a multi-scale convolutional architecture, designing multiple parallel convolution channels within this branch. Each channel corresponds to a different receptive field range, enabling the extraction of defect regions of varying scales. For example, a convolution channel with a small receptive field is specifically designed to extract edge information of small cracks, while a channel with a large receptive field analyzes the overall morphology of potholes. This enables the network to maintain high detection accuracy for defects of varying scales. Furthermore, in the multi-scale convolutional architecture, the feature extraction results of each channel are fused in a subsequent stage, ensuring that the final output contains information on defects of varying scales, resulting in more precise boundaries for defect regions and reducing missed detections due to varying defect scales.

[0067] The multi-branch defect detection module consists of a multi-scale convolution branch and a cross-layer feature fusion branch. The multi-scale convolution branch primarily processes road defect regions of varying sizes and shapes. In road defect detection tasks, cracks typically appear as elongated, irregular textures, while potholes or depressions exhibit large, regional variations. This makes it difficult for a fixed-size convolution kernel to simultaneously capture defect characteristics of both scales. Therefore, the present invention employs a multi-scale convolutional architecture, designing multiple parallel convolution channels within this branch. Each channel corresponds to a different receptive field range, enabling the extraction of defect regions of varying scales. For example, a convolution channel with a small receptive field is specifically designed to extract edge information of small cracks, while a channel with a large receptive field analyzes the overall morphology of potholes. This enables the network to maintain high detection accuracy for defects of varying scales. Furthermore, in the multi-scale convolutional architecture, the feature extraction results of each channel are fused in a subsequent stage, ensuring that the final output contains information on defects of varying scales. This results in more precise defect boundary information and reduces missed detections due to varying defect scales.

[0068] Throughout the detection module's computational process, the output of the multi-branch defect detection module maintains consistency with the size of the multimodal training set, ensuring that the resulting defect segmentation results correspond one-to-one with the input image, enabling the system to accurately locate defects in practical applications. Furthermore, in the final stage of defect detection, the detection module of the present invention classifies and labels different types of defects based on the extracted feature information. This allows the output to not only provide the spatial distribution of defects but also assign corresponding category labels to each defect region. For example, different types of defects, such as cracks, potholes, spalling, and rutting, will each correspond to a different classification label in the final output image. This allows subsequent road maintenance systems to develop repair strategies based on the specific defect type, rather than relying solely on the identification of defect regions. Compared to traditional defect detection methods, the multi-branch defect detection module of the present invention enhances the network's adaptability to defect targets of varying scales through multi-scale convolution, avoiding the problem of a single-scale convolution kernel's inability to fully extract defect information. Furthermore, the introduction of a cross-layer feature fusion strategy enables the network to more accurately distinguish between different defect types when processing highly similar defect regions, improving recognition accuracy. Furthermore, through the final defect classification labeling, the method of this invention not only outputs the spatial distribution of defects but also provides detailed information on the defect type, making the system more intelligent and automated when applied to actual road inspection tasks. Ultimately, this invention achieves higher accuracy and robustness in the automatic detection and classification of road defects through the use of a multi-layer feature extraction and fusion strategy based on deep learning, providing an efficient and intelligent solution for intelligent transportation, road maintenance, and infrastructure monitoring.

[0069] Example 8: The method also includes: Step 5: first performing separate pre-training on the composite sparse autoencoder network, using the difference between the multimodal training set and the corresponding reconstruction result to perform targeted optimization on the encoding weights and decoding weights, thereby enhancing the sparse expression capability of multimodal features; then adding a multi-branch disease detection module for joint training, using disease labels to locate and classify the composite sparse autoencoder network, and performing end-to-end weight iteration based on the joint goal of reconstruction loss and disease recognition loss; setting independent recognition loss weights based on the shape differences between cracks and potholes, thereby enhancing the composite sparse autoencoder network's ability to determine road diseases.

[0070] Specifically, since deep learning models are easily affected by data distribution during training, if end-to-end training is performed directly, the network may experience slow training convergence due to the imbalance of disease data, the complexity of feature expression, and the difficulty of multimodal fusion, and may even experience overfitting or decreased recognition accuracy in some cases. Therefore, the present invention proposes a strategy of first pre-training the autoencoder network separately and then performing end-to-end optimization in conjunction with the disease detection module to gradually optimize the network's feature extraction capabilities and improve the final disease recognition accuracy. In the first stage, pre-training of the composite sparse autoencoder network is crucial because the network not only needs to learn how to efficiently represent multimodal data under unsupervised conditions, but also needs to have good generalization capabilities in subsequent disease detection tasks. To this end, the present invention first constructs an unsupervised training task, that is, using the difference between the multimodal training set and its corresponding reconstruction results to optimize the weights of the encoder and decoder. In this process, the goal of the encoding part is to convert the input fused processed image into a compact feature representation, while the decoding part needs to restore the original input image as much as possible. By minimizing the error between input and reconstructed output, the network automatically learns the most discriminative features in the data while discarding irrelevant information, thereby improving the sparsity of feature representation. For example, in crack detection, the encoding component prioritizes learning crack boundary information while ignoring irrelevant road background texture. In pothole detection, the encoding component focuses more on morphological changes in the road surface while reducing sensitivity to changes in illumination. Through this pre-training approach, the network can deeply model multimodal data before formally undertaking disease detection tasks, and develop reasonable sparse representation capabilities at the feature extraction level, thereby improving subsequent training efficiency and recognition accuracy.

[0071] After the pre-training is completed, the second stage is entered, that is, end-to-end training is performed with the joint multi-branch disease detection module to ensure that the network's performance in the disease recognition task can be further optimized. In this stage, the disease mark information is formally introduced into the training process. The network not only needs to perform feature reconstruction, but also needs to be further optimized in the recognition and classification tasks of the diseased area. In order to achieve this goal, the present invention designs a joint loss function, which consists of reconstruction loss and disease recognition loss, where the reconstruction loss is used to ensure that the feature expression ability of the encoder will not degenerate during the training process, while the disease recognition loss is used to optimize the accuracy of disease classification. During the optimization process, the parameters of the network will be subjected to end-to-end weight iteration according to the joint objectives of the two loss functions, so that the network can not only maintain good feature representation capabilities, but also achieve high-precision positioning and classification in disease detection tasks. In addition, in order to further enhance the network's recognition ability for different types of diseases, the present invention sets independent recognition loss weights based on the shape differences between cracks and pits. Cracks usually appear as linear structures with long lengths but narrow widths, while pits have obvious regional characteristics. Therefore, in the traditional loss function design, if the same weights are used uniformly, the network may have a better learning effect on a certain type of disease and a poorer learning effect on another type of disease. In order to solve this problem, the present invention introduces category-specific weights in the loss function, so that the network can adaptively adjust the loss contribution according to the morphological characteristics of different disease types. For example, during the training process, if it is found that the recognition accuracy of cracks is low, the loss weights related to cracks will be automatically increased, so that the network pays more attention to the feature learning of cracks during training, thereby improving the final classification effect. Similarly, for pits, if it is found that the network has errors in locating the pit boundaries, the pit loss weights will be increased to strengthen the network's attention to the pit area. This adaptive loss optimization strategy makes the network's recognition ability for different disease types more balanced, thereby improving the overall detection accuracy.

[0072] Ultimately, through this two-stage training strategy, the present invention successfully optimized the feature learning ability of the composite sparse autoencoder network, enabling it to not only learn the optimal expression of multimodal data under unsupervised conditions, but also further optimize the accuracy and robustness of disease detection in the supervised learning stage. Compared with the traditional end-to-end training method, the phased training method of the present invention can effectively avoid the model from falling into local optimality in the early stage, thereby improving the final recognition effect. In addition, through independent recognition loss weight control, the present invention can automatically adjust the learning strategy when facing different forms of road diseases, so that the detection model is more adaptable to different types of disease targets, thereby improving the practicality and intelligence level of the road disease detection system. Ultimately, the method of the present invention not only improves the accuracy of disease identification, but also reduces the instability during the training process, so that the system can perform road disease detection tasks more efficiently in practical applications, providing more accurate technical support for road maintenance and management.

[0073] Example 9: The method also includes: Step 6: When inferring an unlabeled image, the unlabeled image is input into the encoding part of the composite sparse autoencoder network, and then a multi-branch disease detection module is used to obtain a disease segmentation map and a type discrimination result; for the detected disease area, the actual surface height change or area is determined by combining low-level features and depth images, and erroneous areas caused by tiny holes and illumination effects are screened out; by online monitoring of the consistency of the images of each channel of the multi-channel imaging device, dynamic correction of unstable input scenes is achieved.

[0074] Specifically, during inference, the unlabeled image is first input into the encoding portion of the composite sparse autoencoder network. This portion performs feature extraction and channel compression on the input image, preserving the most discriminative defect features while removing irrelevant background information. Because the autoencoder network has learned key patterns of road defects from multimodal data during training, during the inference phase, even if the defect morphology in the input image varies, the network is able to extract core features, effectively representing the defect areas in the high-dimensional feature space. After feature extraction, the data is passed to the multi-branch defect detection module, which uses multi-scale convolution and cross-layer feature fusion to accurately segment the defect areas and classify the defect types. For example, if the input image contains cracks and potholes, the network will identify these defect areas separately and assign corresponding class labels, enabling the subsequent road maintenance system to carry out targeted treatment based on the specific defect type.

[0075] During inference, the unlabeled image is first input into the encoding portion of the composite sparse autoencoder network. This portion performs feature extraction and channel compression on the input image, preserving the most discriminative defect features while removing irrelevant background information. Because the autoencoder network has learned key patterns of road defects from multimodal data during training, during the inference phase, even if the defect morphology in the input image varies, the network is able to extract core features, effectively representing the defect areas in the high-dimensional feature space. After feature extraction, the data is passed to the multi-branch defect detection module, which uses multi-scale convolution and cross-layer feature fusion to accurately segment the defect areas and classify the defect types. For example, if the input image contains cracks and potholes, the network will identify these defect areas separately and assign corresponding class labels, enabling the subsequent road maintenance system to carry out targeted treatment based on the specific defect type.

[0076] In addition to screening for defect areas, the present invention also employs an online monitoring mechanism to ensure input consistency for the multi-channel imaging device under different environments. Because road defect detection systems often operate under a variety of complex environmental conditions, such as strong sunlight, insufficient nighttime lighting, and slippery conditions on rainy days, data from different channels of the imaging device may differ, and these differences can affect the stability of defect detection. To address this issue, the present method monitors the consistency of multi-channel images at different time points and adjusts the detection strategy in real time. For example, if the defect features in a certain area differ significantly between the visible light image and the infrared image, the system automatically adjusts the feature weights to ensure that the detection results are more consistent with the reliable input signal. Furthermore, if the image quality of a channel is poor, such as if the visible light image is too dark due to insufficient illumination, the system weights the contributions of the infrared and depth information to ensure that the final detection results remain highly reliable. Through this online monitoring mechanism, the present invention can maintain high defect recognition accuracy in complex environments and automatically correct for unstable input data, thereby improving the robustness of the overall system. Ultimately, the method of the present invention enables the disease detection system to operate stably in different environments and provide accurate disease segmentation and classification results through multi-level reasoning, screening, and dynamic correction. Compared with traditional road disease detection methods, the present invention not only relies on the end-to-end feature learning capabilities of deep learning, but also combines deep information analysis, low-level feature screening, and online monitoring mechanisms, so that the system can automatically adapt to different environments, reduce false detections, and improve overall recognition accuracy. This method not only improves the intelligence level of road disease detection, but also provides a more efficient and accurate technical means for intelligent transportation and infrastructure maintenance, making road inspections more efficient and reliable, and able to adapt to real-time detection needs in different environments.

[0077] Example 10: The method also includes: Step 7: performing a morphological filling operation on the identified diseased area to eliminate pixel-level holes that may exist in narrow areas; performing spatial clustering and boundary tracking on the segmentation results output by the multi-branch disease detection module, and outputting the crack length, pothole area, spalling edge coordinates and rutting depth range; comparing the crack length, pothole area, spalling edge coordinates and rutting depth range with the identified areas in the multimodal training set pixel by pixel, and statistically analyzing the recognition accuracy, false detection rate and missed detection rate, and recording the processing speed and recognition confidence of different road disease types.

[0078] Specifically, during the defect area optimization process, a morphological filling operation is first used to eliminate pixel-level holes that may exist in narrow areas. Since the neural network may cause small missing or discontinuous areas on the boundaries of the defect area due to factors such as noise, shadows, and lighting changes when segmenting the defect area, especially in crack detection tasks, cracks may be segmented into multiple discontinuous line segments, and in pits and spalling areas, some defect areas may have pixel-level gaps inside. To solve this problem, the present invention uses morphological dilation and closing operations to make the boundaries of the defect area more coherent and fill the tiny holes that may exist inside, thereby improving the integrity of the defect area. For example, in the crack detection task, morphological filling can ensure that the crack is marked as a complete coherent curve rather than multiple scattered pixel blocks, which is crucial for the subsequent calculation of the defect length. In addition, in the detection tasks of pits and spalling areas, this operation can effectively eliminate small areas of misidentification caused by light reflection or sensor noise, making the outline of the defect area clearer and facilitating further analysis.

[0079] Based on the optimized defect area, the present invention further performs spatial clustering and boundary tracking on the segmentation results of the multi-branch defect detection module to extract the geometric feature information of the defect. Since road defects usually have different morphological characteristics, such as cracks that are linear, potholes that are block-shaped, spalling that has complex boundary shapes, and rutting that appears as continuous depressions within a certain depth range, it is difficult to directly obtain the key geometric parameters of the defect by relying solely on pixel-level segmentation results. To solve this problem, the present invention first performs spatial clustering on the defect area, that is, based on connected region analysis, pixels belonging to the same defect target are aggregated together to determine the overall outline of the defect area. For example, in the crack detection task, spatial clustering can merge multiple discontinuous crack fragments to ensure that they can reflect the actual crack extension when calculating the length, and in the pothole detection task, spatial clustering can be used to merge multiple adjacent small pothole areas to obtain more complete information about the defect area. After completing spatial clustering, the system will track the boundaries of the defect area to extract key geometric features. For example, in the crack detection task, the system will extract the starting and ending coordinates of the crack and calculate the overall length of the crack; in the pothole detection task, the system will calculate the area of the pothole and output the center coordinates and boundary shape of the pothole; in the spalling area detection task, the system will record the boundary coordinates of the spalling area in order to assess the impact range of the disease; and in the rutting detection task, the system will calculate the depth range of the rutting to determine whether the rutting has reached a level that affects driving safety.

[0080] After obtaining the geometric parameters of the disease, the present invention further performs a performance evaluation on the detection results to ensure the accuracy of recognition and optimize the parameters of the model. During this process, the system will compare the above detection results with the identified disease areas in the multimodal training set pixel by pixel to calculate the recognition accuracy, false detection rate and missed detection rate. Specifically, the system will calculate the overlap between the detected disease area and the true labeled area, that is, the recognition accuracy is evaluated by calculating the pixel-level IoU indicator; at the same time, the number of pixels that are mistakenly detected as diseases is counted to calculate the false detection rate, and the number of unrecognized pixels in the true disease area is counted to calculate the missed detection rate. This evaluation method can help the system quantify the accuracy of recognition and optimize the weights of the network during the training process, so that the model can more accurately identify road diseases. In addition, the method of the present invention will also record the processing speed and recognition confidence of different road disease types, that is, during the inference process, the average detection time of each disease is calculated, and combined with the confidence score of the model, the recognition difficulty of different disease types and the adaptability of the model in different scenarios are evaluated. For example, if the detection time for a certain type of disease is too long, the system can improve the detection efficiency by adjusting the network structure or optimizing the inference process. If the recognition confidence of a certain type of disease is low, the system can adjust the training weight of the disease to enhance the network's learning ability for this type of disease.

[0081] The present invention is not limited to the aforementioned specific embodiments, but extends to any new features or any new combination disclosed in this specification, as well as any new method or process steps or any new combination disclosed.

Claims

1. A method for automatic identification of road defects based on sparse autoencoder networks, characterized in that: The method comprises: Step 1: Use a multi-channel imaging device to simultaneously acquire visible light images, infrared images, and depth images to construct a multimodal training set; Step 2: Perform multi-scale region segmentation on the visible light image and infrared image in the multimodal training set to obtain segmented regions; use the projection correction method to align the depth image and each segmented region in the same coordinate system, and output the fused image; Step 3: Suppress high-frequency noise and retain medium- and low-frequency features on the fused image, enhance edge texture, and obtain a fused image; Step 4: Use the fused processed images to construct a composite sparse autoencoder network to identify road damage; Step 4 specifically includes: Step 4.1: Extract features from the fused image through multi-layer convolution and downsampling, and reconstruct the image in the decoding stage using a symmetrical upsampling and deconvolution structure. Apply sparse activation constraints to the intermediate hidden layers to reduce background redundancy and highlight road damage morphology. Step 4.2: Insert a multi-scale convolution branch and a cross-layer feature fusion branch into the top-level feature map of the composite sparse autoencoder network to identify road damage; Step 4.1 specifically includes: constructing an encoding part of a composite sparse autoencoder network consisting of multiple layers of convolution and downsampling to perform channel compression and feature extraction on the fused image; using a sparse activation mechanism in each hidden layer of the encoding part to highlight neurons that can recognize the shape and texture of road damage and suppress irrelevant background images; using a deconvolution and upsampling structure symmetrical to the encoding part as the decoding part of the composite sparse autoencoder network to gradually restore the extracted features to the same size as the input multimodal training set in spatial dimensions, forming a multimodal reconstructed output; and implementing a sparsity strategy that combines local and global constraints in the encoding and decoding parts to enable the composite sparse autoencoder network to reduce channel redundancy while ensuring feature resolution. Step 4.2 specifically includes: embedding a multi-branch defect detection module in the top-level feature map of the composite sparse autoencoder network, which contains a multi-scale convolution branch and a cross-layer feature fusion branch; the multi-scale convolution branch is used to process road defect areas of different sizes and shapes, and extracts cracks and depressions through parallel convolution channels; the cross-layer feature fusion branch combines the edge and texture features retained in the encoding stage with the high-semantic features of the top layer to distinguish areas of similar color but different defect properties; the multi-branch defect detection module outputs road defect segmentation results that are consistent with the size of the multimodal training set, and classifies and labels different types of road defects.

2. The method for automatic road damage identification based on a sparse autoencoder network according to claim 1, characterized in that: The step 1 specifically includes: using a multi-channel imaging device to simultaneously acquire a visible light image, an infrared image, and a depth image for surface height measurement, and labeling each channel image through a unified spatiotemporal synchronization mechanism to generate a multimodal original image sequence arranged in time sequence; dividing the multimodal original image sequence into a number of equidistant image blocks, and recording road surface features in each image block; road surface features include: cracks, potholes, spalling, and rutting; and setting independent labels for cracks, potholes, spalling, and rutting in the image blocks to form a multimodal training set.

3. The method for automatic road damage identification based on a sparse autoencoder network according to claim 2, characterized in that: Step 2 specifically includes: performing multi-scale segmentation processing on the visible light image and the infrared image to generate multiple segmentation areas of different sizes for subsequent detail enhancement and appearance alignment; performing three-dimensional centroid projection and local surface smoothing operations on the depth image to match the coordinate systems of the visible light image and the infrared image to obtain depth features; merging each segmentation area and the corresponding depth feature into a unified coordinate system; and through repeated offset measurement and projection correction, ensuring that the spatial correspondence of the segmentation areas in the multimodal training set is consistent, thereby obtaining a fused image.

4. The method for automatically identifying road defects based on a sparse autoencoder network according to claim 3, wherein: Step 3 specifically includes: performing discrete wavelet decomposition and high-frequency noise suppression operations on the fused image to retain the mid-frequency and low-frequency parts that are discriminative for road damage identification; using noise threshold clipping and gradient-guided enhancement on the high-frequency part to remove random interference when the multi-channel imaging device acquires visible light images, infrared images, and depth images, thereby obtaining a fused processed image.

5. The method for automatic road damage identification based on sparse autoencoder network according to claim 4, characterized in that: The method further includes: step 5: first performing separate pre-training on the composite sparse autoencoder network, using the difference between the multimodal training set and the corresponding reconstruction results to perform targeted optimization on the encoding weights and decoding weights, thereby enhancing the sparse expression capability of multimodal features; then adding a multi-branch disease detection module for joint training, using disease labels to locate and classify the composite sparse autoencoder network, and performing end-to-end weight iteration based on the joint goal of reconstruction loss and disease recognition loss; setting independent recognition loss weights based on the shape differences between cracks and potholes, thereby enhancing the composite sparse autoencoder network's ability to determine road diseases.

6. The method for automatic road damage identification based on sparse autoencoder network according to claim 5, characterized in that: The method further includes: step 6: when inferring the unlabeled image, inputting the unlabeled image into the encoding part of the composite sparse autoencoder network, and then obtaining a defect segmentation map and type discrimination result through a multi-branch defect detection module; for the detected defect area, combining low-level features with the depth image to determine the actual surface height change or area, and screening out erroneous areas caused by tiny holes and illumination effects; and realizing dynamic correction of unstable input scenes by online monitoring of the consistency of the images of each channel of the multi-channel imaging device.

7. The method for automatically identifying road defects based on a sparse autoencoder network according to claim 6, characterized in that: The method further includes: step 7: performing a morphological filling operation on the identified defect area to eliminate pixel-level holes that may exist in the narrow area; performing spatial clustering and boundary tracking on the segmentation results output by the multi-branch defect detection module, and outputting crack length, pothole area, spalling edge coordinates, and rutting depth range; comparing the crack length, pothole area, spalling edge coordinates, and rutting depth range with the identified areas in the multimodal training set pixel by pixel, and statistically analyzing the recognition accuracy, false detection rate, and missed detection rate, and recording the processing speed and recognition confidence of different road defect types.

Citation Information

Patent Citations

  • Multi-mode multi-layer fusion deep neural network for face anti-spoofing

    CN110674677A

  • Building extraction method in remote sensing image, electronic equipment and storage medium

    CN115345866A