Asphalt pavement crack identification method based on joint optimization network

By combining an adaptive image denoising network and a spatial convolutional attention U-Net network, the problems of low efficiency and insufficient accuracy in asphalt pavement crack identification in traditional methods are solved, achieving efficient and accurate asphalt pavement crack identification and parameter calculation.

CN120976218AActive Publication Date: 2025-11-18LIAONING TRAFFIC KEXUE RES YUAN
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511491803.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2025-11-18
Estimated Expiration
2045-10-20

AI Technical Summary

Technical Problem

Traditional manual methods for detecting cracks in asphalt pavements are inefficient and highly subjective, while deep learning methods struggle to accurately identify minute cracks when faced with ambient light interference and complex topological structures in pavement images.

Method used

An adaptive image denoising network and a spatial convolutional attention U-Net network are designed. By combining a multi-filter denoising module and a hyperparameter prediction module, an adaptive denoising and semantic segmentation of asphalt pavement cracks are achieved through a joint optimization strategy, thereby improving the ability to identify small cracks.

Benefits of technology

It improves the accuracy and efficiency of asphalt pavement crack identification, reduces the risk of misidentification, enhances image noise reduction and crack segmentation accuracy, and provides a scientific basis for crack morphology identification and parameter calculation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976218A_ABST
    Figure CN120976218A_ABST
Patent Text Reader

Abstract

The invention provides an asphalt pavement crack identification method based on a joint optimization network, and relates to the technical field of pavement disease intelligent detection. The method comprises the following steps: firstly, acquiring an asphalt pavement disease image, acquiring label images of a corresponding image number, and constructing an asphalt pavement crack image data set; secondly, designing an adaptive image noise reduction network, wherein the adaptive image noise reduction network comprises a multi-filter noise reduction module and a hyper-parameter prediction module; designing a spatial convolution attention U-Net network as a semantic segmentation network; combining the adaptive image noise reduction network with the semantic segmentation network by using a joint optimization strategy to form an IDSNet network, and optimizing the IDSNet network according to the training loss of the semantic segmentation network to realize the synchronization of asphalt pavement crack image noise reduction and crack segmentation; and finally, performing fracture form identification and parameter calculation based on a fracture segmentation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent pavement distress detection technology, and in particular to a method for identifying asphalt pavement cracks based on a joint optimization network. Background Technology

[0002] In highway maintenance and management, cracks are one of the most common pavement defects, and their accurate detection is a crucial foundation for assessing pavement technical condition. However, traditional manual inspection methods are inefficient, highly subjective, and difficult to adapt to the needs of large-scale pavement inspection. With the development of automated pavement inspection technology, deep learning-based crack detection methods have gradually become mainstream, but they also face multiple challenges. First, during the acquisition of pavement images, the imaging equipment of automated pavement inspection systems is affected by ambient light interference, leading to a decrease in image quality and the introduction of a large amount of noise. Second, due to the strip-shaped grooves and complex topological structure of pavement cracks, traditional semantic segmentation networks struggle to learn and extract features from small cracks. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to address the shortcomings of the prior art by providing a method for identifying asphalt pavement cracks based on a joint optimization network, thereby realizing the identification of asphalt pavement cracks.

[0004] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: On the one hand, the present invention provides a method for identifying asphalt pavement cracks based on a joint optimization network, comprising: Acquire images of asphalt pavement defects and obtain corresponding labeled images to construct an asphalt pavement crack image dataset; Design an adaptive image denoising network; the adaptive image denoising network includes a multi-filter denoising module and a hyperparameter prediction module; the multi-filter denoising module includes multiple filters, and dynamically adjusts the parameters of each filter through a joint optimization strategy to achieve adaptive denoising under different noise conditions; the hyperparameter prediction module predicts and optimizes the hyperparameters of each filter based on the information of the input image; A spatial convolutional attention U-Net network is designed as a semantic segmentation network; the spatial convolutional attention U-Net network is implemented by introducing spatial and channel reconstruction convolutional modules into the Attention U-Net network; An IDSNet network is formed by combining an adaptive image denoising network and a semantic segmentation network using a joint optimization strategy. The training loss of the semantic segmentation network is used to optimize the IDSNet network, so as to achieve simultaneous denoising and crack segmentation of asphalt pavement crack images. Crack morphology identification and parameter calculation are performed based on crack segmentation results.

[0005] Furthermore, the multi-filter noise reduction module includes an exposure filter, a contrast filter, a tone filter, and a stripe removal filter, which are used to perform exposure filtering, contrast filtering, tone filtering, and stripe removal filtering on the input image, respectively.

[0006] Furthermore, the destriating filter uses the average pixel value of each column in the asphalt pavement crack image to represent the pixel value of each column, and then maps the illuminance values ​​of all columns to the desired illuminance uniformity value through a column-based illuminance balancing operation, as shown in the following formula: , (1); Where γ is a trainable hyperparameter; These are the processed image pixel grayscale values; It is the predetermined average grayscale value of each column after preprocessing; I ij A is the grayscale value of the pixel in the i-th column and j-th row of the k-th asphalt pavement crack image; i k It is the weighted average of the pixels in the i-th column of the k-th asphalt pavement crack image, and its calculation formula is shown in formula (2); k is the number of asphalt pavement crack images; M is the number of pixel columns in the asphalt pavement crack image; N is the number of pixel rows in the asphalt pavement crack image; (2); in, β represents the weighted variable, used to control the weight of the current image, and is also the latest measured illumination information; β represents an empirical constant. is the average pixel value of the i-th column in the k-th asphalt pavement crack image.

[0007] Furthermore, the hyperparameter prediction module is a framework consisting of 5 convolutional layers and 2 fully connected layers. Based on the global content of the asphalt pavement crack image, it predicts and optimizes the hyperparameters required by the multi-filter noise reduction module to achieve adaptive noise reduction of the asphalt pavement crack image.

[0008] Furthermore, the method utilizes a joint optimization strategy to optimize the trainable hyperparameters of each filter in the multi-filter denoising module. Specifically, it optimizes both image denoising and semantic segmentation performance based on a stochastic gradient descent optimizer and a binary cross-entropy loss function. The hyperparameter prediction module initially predicts the hyperparameters used by each filter based on the state of the input asphalt pavement crack image. The asphalt pavement crack image after noise reduction by each filter is semantically segmented by a semantic segmentation network. The hyperparameter prediction module then adjusts the hyperparameters of each filter based on the feedback of the semantic segmentation results to further obtain a semantically segmented binary image of the asphalt pavement crack with better noise reduction effect.

[0009] Furthermore, the crack morphology includes transverse cracks, longitudinal cracks, and alligator cracks; the crack parameters include crack length, crack width, and crack area.

[0010] Furthermore, the specific method for crack morphology recognition based on crack segmentation results is as follows: First, connected components are extracted from the semantically segmented binary image of asphalt pavement cracks. The morphology of the crack is determined by judging the similarity between the trend direction of each crack connected component and the four directions of 0°, 90°, 45° and 135°.

[0011] Furthermore, the crack length is obtained by summing the relative distances of multiple pixels on the crack skeleton curve; the crack skeleton curve is a path composed of a series of individual pixels and is distributed along the central axis of the crack target, reflecting the topological structure and morphological characteristics of the crack.

[0012] Secondly, the present invention proposes a computer-readable storage medium storing executable instructions that, when executed, cause a processor to perform the aforementioned asphalt pavement crack identification method based on a joint optimization network.

[0013] Thirdly, the present invention proposes a computer program product, including a computer program or instructions, which, when executed by a processor, implements the aforementioned asphalt pavement crack identification method based on a joint optimization network.

[0014] The beneficial effects of adopting the above technical solution are as follows: The present invention provides an asphalt pavement crack identification method based on a joint optimization network. (1) In order to process the real noise in the asphalt pavement crack image, an adaptive image denoising network is designed, which consists of two modules: a multi-filter denoising module and a hyperparameter prediction module. In the multi-filter denoising module, a stripe removal filter is designed, which can effectively reduce the influence of stripe noise in the pavement image.

[0015] (2) To address the difficulty in identifying small cracks, a spatial and channel reconstruction convolutional module, namely the SCConv module, was introduced into the semantic segmentation network. The introduction of this module improves the network's ability to capture small cracks and reduces the risk of misidentification.

[0016] (3) To reduce the redundancy of the separate image denoising and crack detection processes and to objectively evaluate the denoising effect on real noise in the image, an IDSNet network integrating denoising and segmentation was proposed using a joint optimization strategy. Under the guidance of the joint optimization strategy, the network was trained with the loss of the crack segmentation result to improve the image denoising effect. At the same time, the denoised image will further improve the crack segmentation effect. Attached Figure Description

[0017] Figure 1 A flowchart of an asphalt pavement crack identification method based on a joint optimization network provided in Embodiment 1 of the present invention; Figure 2 This is a partial image data set of the asphalt pavement crack noise reduction image dataset provided in Embodiment 1 of the present invention, wherein (a) is a partial image of asphalt pavement disease, and (b) is the label image corresponding to the image of asphalt pavement disease; Figure 3 The image shows the filtering effect of the stripe removal filter provided in Embodiment 1 of the present invention, wherein (a) is the original image of asphalt pavement distress, and (b) is the illumination balance image after stripe removal filtering. Figure 4 This is a schematic diagram of the asphalt pavement distress image noise reduction process provided in Embodiment 1 of the present invention; Figure 5 This is a schematic diagram of the spatial convolutional attention U-Net network structure provided in Embodiment 1 of the present invention; Figure 6 This is a schematic diagram of the SCConv module structure provided in Embodiment 1 of the present invention; Figure 7 This is a schematic diagram of the crack length calculation method provided in Embodiment 1 of the present invention; Figure 8 This is a crack morphology recognition result image of different asphalt pavement defects provided in Embodiment 1 of the present invention. Detailed Implementation

[0018] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.

[0019] Example 1:

[0020] In this embodiment, an asphalt pavement crack identification method based on a joint optimization network is described, such as... Figure 1 As shown, it includes the following steps: Step 1: Obtain images of asphalt pavement defects and corresponding number of labeled images to form an asphalt pavement crack noise reduction image dataset. This embodiment utilizes a road inspection vehicle equipped with a line-scan industrial camera to collect over 10,000 asphalt pavement images and store them in a database. By filtering the images in the database, a batch of 5,350 asphalt pavement defect images was obtained. These images underwent preprocessing, including grayscale conversion and resizing. Simultaneously, by manually labeling crack-type defects in the asphalt pavement images, a corresponding number of labeled images were obtained. This resulted in the formation of the Asphalt Pavement Crack Denoising Image Dataset, abbreviated as APCD dataset. Some images in the dataset are shown below. Figure 2 As shown.

[0021] Step 2: Design an adaptive image denoising network to perform adaptive denoising processing on asphalt pavement distress images; This embodiment designs an adaptive image denoising network consisting of a multi-filter denoising module and a hyperparameter prediction module. The multi-filter denoising module includes various filters such as exposure, contrast, hue, and destriping. It dynamically adjusts the filter parameters through a joint optimization strategy to achieve adaptive denoising under different noise conditions. The hyperparameter prediction module predicts and optimizes the hyperparameters of each filter based on the brightness, hue, exposure, and other information of the input image to improve the denoising effect. Images of asphalt pavements are primarily acquired by road inspection vehicles equipped with line-scan industrial cameras and high-power laser illumination devices. For example... Figure 3 As shown in (a), due to changes in lighting conditions during the acquisition process, the acquired road surface image exhibits striped exposure differences; this phenomenon in the image is called stripe noise. Simultaneously, the road surface image also contains numerous overexposed and underexposed areas, which severely impacts the accuracy of semantic segmentation. Therefore, reducing the impact of these exposure differences is crucial for improving the accuracy of semantic segmentation. To mitigate the influence of stripe noise, this invention designs a destriping filter. Figure 3 As can be observed in (a), variations in image illuminance result in uneven brightness levels across different scan columns. However, the illuminance within the same column remains relatively consistent. Therefore, the pixel value of each column can be represented by the average pixel value of each column. Then, a column-based illuminance balancing operation is performed to map the illuminance values ​​of all columns to the desired illuminance uniformity. This process can be summarized by the following formula: (1); in, γ These are trainable hyperparameters; These are the processed image pixel grayscale values; It is the predetermined average gray value of each column after preprocessing; It is the first k The first original image i Liede j The grayscale value of a row of pixels; It is the first k The first image i The weighted average of the column pixels is calculated using the formula shown in formula (2); k It is the number of images; M It is the number of image columns; N It is the number of rows in the image; (2); in, represents the weighted variable, used to control the weight of the current image, and is also the latest measured illumination information; β represents the empirical constant, set to 0.75; is the average pixel value of the i-th column in the k-th asphalt pavement crack image.

[0022] like Figure 3 (b) shows the image of the asphalt pavement after illuminance balancing. It can be observed that the overall brightness of the image has been improved to a certain extent, and the brightness of the image edge area has been significantly enhanced, making the information of cracks better displayed. At the same time, the stripe noise in the image has been effectively controlled, achieving the expected experimental objectives.

[0023] In addition, the multi-filter noise reduction module includes several differentiable filters, including some of the filters for exposure, gamma, contrast, sharpening, hue, and white balance. These filters optimize image quality by adjusting parameters such as exposure and contrast, thereby enhancing the image's representation of the true condition of the asphalt pavement. The filters will process the input pixels... Mapped to output pixels ,in r , g and b These represent the values ​​for the red, green, and blue channels, respectively. Table 1 lists the mapping function for each filter and its corresponding trainable hyperparameters. The differentiability of the filters is also a prerequisite for ensuring that the adaptive image denoising network can be trained through backpropagation. As the network trains, the hyperparameters in the filters will be dynamically adjusted until the image processed by the filters achieves a better image denoising effect.

[0024] Table 1. Filter mapping functions; ; In the contrast filter mapping function The definition is shown in formula (3). Wherein... This represents a brightness function based on the human eye's sensitivity to the three primary colors. The brightness enhancement function is defined as shown in equations (4) and (5). The sharpening filter mapping function... It is the input image. Indicates a Gaussian filter. It is a direct proportionality factor. The sharpening operation process affects... x and Both can be subtle, while the degree of sharpening can be optimized. To adjust; (3); (4); (5); In this embodiment, after ablation testing, it was determined that the multi-filter noise reduction module consists of four types of filters (ECTD): exposure, contrast, hue, and destriping.

[0025] For complex and varied road surface image conditions, multiple filters are needed to work together for image denoising. Different asphalt road surface images have unique image characteristics, such as exposure and contrast. Therefore, the hyperparameters of each filter need to be adjusted specifically to achieve the best denoising effect. Traditionally, adjusting filter hyperparameters requires manual work by experienced engineers. This process is not only time-consuming and labor-intensive, but also difficult to handle all possible image scenarios, resulting in low efficiency and high cost. To overcome this limitation, this invention introduces a hyperparameter prediction module before the multi-filter denoising module. The hyperparameter prediction module is a framework composed of 5 convolutional layers and 2 fully connected layers. It predicts and optimizes the hyperparameters required by the multi-filter denoising module based on the global content of the image, namely brightness, hue, and exposure information, to achieve adaptive denoising. Simultaneously, to reduce computational costs and improve processing efficiency, the input size of the asphalt road surface image is downsampled to a resolution of 256 × 512. This operation not only significantly reduces the amount of data processed by the network, but also preserves sufficient image information for the hyperparameter prediction module to accurately predict hyperparameters. The denoising process of the asphalt road surface image is as follows: Figure 4 As shown, the image is input into the hyperparameter prediction module and the multi-filter denoising module respectively. The two modules run step by step to complete the denoising task of the image.

[0026] Step 3: Design a spatial convolutional attention U-Net network as a semantic segmentation network to extract features from asphalt pavement distress images, thereby achieving crack segmentation and localization. The Spatial Convolutional Attention U-Net network enhances the network's ability to extract features of small cracks by introducing a Spatial and Channel Reconstruction Convolution (SCConv) module into the Attention U-Net network, thereby reducing misidentification and unidentified areas. The denoised image is input into a spatial convolutional attention U-Net network, and image features are extracted step by step through the encoder-decoder structure of the Attention U-Net network; The SCConv module is used to enhance the extraction of features from fine cracks and output crack segmentation results; high-precision mapping of crack locations is achieved through matrix encoding to generate crack location maps. In this embodiment, the Spatial Convolutional Attention U-Net network enhances the network's ability to extract features of small cracks by introducing a Spatial and Channel Reconstruction Convolution (SCConv) module into the Attention U-Net network, thereby reducing misidentification and unidentified areas. The SCConv module processes the spatial and channel information of the image through Spatial Reconstruction Unit (SRU) and Channel Reconstruction Unit (CRU) respectively, capturing richer feature information and improving feature representation capabilities.

[0027] Attention U-Net is an encoder-decoder network. Its core concept is the introduction of an attention gate mechanism in the decoder. This mechanism allows the network to suppress background regions irrelevant to the target layer by layer, improving network accuracy. While the decoder benefits from this improvement, the encoder, responsible for spatial information compression and crack feature extraction, still has significant room for improvement. Therefore, this invention proposes a Spatial Convolutional Attention U-Net network. By embedding Spatial and Channel Reconstruction Convolution (SCConv) modules in each encoder layer, it effectively enhances the network's ability to extract road crack features. Furthermore, compared to the traditional Attention U-Net, Spatial Convolutional Attention U-Net increases the number of network layers to 5, reducing the risk of network degradation and further improving representation capabilities. The structure of the Spatial Convolutional Attention U-Net network is as follows: Figure 5 As shown.

[0028] SCConv is an efficient convolutional strategy that saves computational cost and storage space. It consists of two units: a Spatial Reconstruction Unit (SRU) and a Channel Reconstruction Unit (CRU). It aims to reduce spatial and channel redundancy in convolutional neural networks and improve feature learning capabilities. The detailed structure of the SCConv module is as follows: Figure 6 As shown.

[0029] The purpose of SRU is to reduce redundancy in the spatial dimension, and it consists of separation and reconstruction operations. Specifically, first, group normalization is used to evaluate the performance of the feature map. Second, the normalized relevance weights are used. W γ ∈ R C This demonstrates the importance of different feature maps. Then, it utilizes... Sigmoid The function, by setting a threshold, will reweight the results. W γ Mapped to the range (0,1). In the experiment, the threshold was set to 0.5, and weights above the threshold were set to 1 to obtain information weights. W 1, then set it to 0 to obtain the non-information weight.W 2. Obtain W The entire process can be represented by equation (6). The reconstruction operation adopts a cross-reconstruction strategy, which enhances information exchange while reducing spatial redundancy. However, spatial redundancy is effectively suppressed by SRU, and redundancy still exists on the channel; (6); The purpose of CRU is to reduce channel redundancy, and it consists of segmentation, transformation, and fusion operations. First, the input features are divided into two branches based on the number of channels. X up Include αC aisle, X low The branches contain (1- α ) C Channel. During the conversion phase, for X up Grouped convolution (GWC) and pointwise convolution (PWC) are used, which reduces computational cost while effectively extracting and fusing features, and forms a representative feature map through element-wise summation. Y 1. At the same time, X low By repeatedly applying pointwise convolution to extract hidden information, more feature maps are generated. These feature maps are then concatenated to form another representative feature map. Y 2. During the fusion phase, global average pooling is used to collect global spatial information, including calculating the global average value for each channel. S m = Pooling ( Y m This achieves the goal of reflecting global information for each channel. Simultaneously, a soft attention mechanism is used to process channel statistics. S 1 and S 2. Integrate into feature vectors β 1 and β 2. Finally, by weighting the upper-layer features Y 1 and weighted lower-level features Y 2. Combining and refining the output channel features Y , obtain Y The entire process can be represented by formula (7): (7); in, It is the learnable weight matrix of GWC. and It is the learnable weight matrix of PwC. It is a connection operation.

[0030] Spatial and channel reconstruction convolution operations can effectively reduce the computation of irrelevant features by the network, significantly improving the accuracy and efficiency of crack segmentation. The core of this approach is to establish feature attention focusing in the spatial dimension, reducing redundant computation through compression and recombination techniques; and to implement dynamic feature selection in the channel dimension, using adaptive weighting to strengthen target features. Spatial and channel reconstruction convolutions enhance the network model's ability to extract key crack regions, suppressing background noise interference while improving the capture of subtle features, thus improving both computational efficiency and model robustness.

[0031] On the APCD dataset, the introduction of the SCConv module further improved the results of crack semantic segmentation, with mIoU and F1-Score increasing by 0.50% and 0.93%, respectively.

[0032] Step 4: The adaptive image denoising network and the semantic segmentation network are combined using a joint optimization strategy to form the IDSNet network. The training loss of the semantic segmentation network is used to optimize the IDSNet network, so that image denoising and crack segmentation can be performed simultaneously. Through joint optimization, the quality of the denoised image is improved, which further enhances the accuracy of crack segmentation.

[0033] The semantic segmentation network and the adaptive image denoising network are combined in a cascaded manner to form an end-to-end network, and the overall IDSNet network is optimized based on the loss value of the semantic segmentation network. A joint optimization strategy is used to optimize the trainable hyperparameters of each filter in the multi-filter denoising module. This strategy is based on a stochastic gradient descent (SGD) optimizer and a binary cross-entropy (BCE) loss function to simultaneously optimize the image denoising effect and the semantic segmentation effect. Specifically, the hyperparameter prediction module initially predicts the hyperparameters used by each filter based on the state of the input asphalt pavement crack image. The asphalt pavement crack image denoised by each filter is then semantically segmented by a semantic segmentation network. The hyperparameter prediction module then adjusts the hyperparameters of each filter based on the feedback of the semantic segmentation results to further obtain an asphalt pavement crack image with better denoising effect. Experimental results show that the use of an adaptive image denoising network achieves a 54.60% mIoU and a 59.93% F1-Score for crack semantic segmentation, representing improvements of 6.38% and 5.58% in mIoU and F1-Score, respectively, compared to the undenoised image. IDSNet demonstrates the best crack segmentation performance on the APCD dataset, achieving a 55.10% mIoU and a 60.86% F1-Score. Compared to the traditional Attention U-Net, this represents improvements of 6.88% and 6.51% in mIoU and F1-Score, respectively.

[0034] Step 5: Based on the crack segmentation results, identify the crack morphology and calculate the parameters; Based on the crack segmentation results, crack morphology identification (such as transverse cracks, longitudinal cracks, and alligator cracks) and geometric parameter calculation (such as crack length, width, and area) are performed. In conjunction with the relevant provisions of the "Highway Technical Condition Assessment Standard," statistical analysis of crack damage data is conducted to provide a scientific basis for the formulation of road maintenance strategies.

[0035] In calculating the crack damage rate, the relevant geometric parameters of crack length, width, and area need to be extracted. For both transverse and longitudinal cracks, the affected width is calculated as 0.2m. Therefore, in actual engineering, only the length needs to be calculated, and the calculation of crack length depends on the extraction and analysis of the crack skeleton curve. The skeleton curve is a path composed of a series of individual pixels, distributed along the central axis of the crack target, which can accurately reflect the topological structure and morphological characteristics of the crack. By extracting these key pixels, the centerline of the asphalt pavement crack can be constructed, thereby revealing the spatial distribution pattern of the crack. Figure 7 The graph shows the skeleton curve of a local crack, with black squares representing crack pixels on the skeleton curve. When the relative positions of the crack pixels on the skeleton are points A and B in the graph, assume the coordinates of point A are (…). x 0 ,y 0), the coordinates of point B are ( x 1 ,y 1), then their relative distance d AB This can be expressed by formula (8): (8); Therefore, the length of the entire crack can be obtained by summing the relative distances between multiple pixels. Compared to calculating the crack length, calculating the crack area is more straightforward. Cracks in an image are fundamentally composed of pixels; therefore, by locating the crack area and counting the number of pixels within that area, the crack area can be preliminarily determined. However, the crack area obtained in this step is the total number of pixels in the image. To obtain the true crack area, the total number of pixels needs to be multiplied by the actual physical size corresponding to a single pixel.

[0036] Asphalt pavement cracks can be classified into transverse cracks, longitudinal cracks, alligator cracks, and block cracks. Different types of cracks have different proportions in the crack failure rate calculation, so cracks must be classified when calculating the failure rate. In practical applications, it is relatively easy to distinguish between transverse cracks and longitudinal cracks, but it is more difficult to distinguish between alligator cracks and block cracks. In this embodiment, crack type classification only includes transverse cracks, longitudinal cracks, and alligator cracks.

[0037] First, connected component extraction is performed on the semantically segmented binary image of asphalt pavement cracks. Then, the crack morphology is determined by assessing the similarity between the direction of each crack's connected component and the four directions of 0°, 90°, 45°, and 135°. The recognition results are as follows: Figure 8 As shown.

[0038] To validate the morphology recognition method, this embodiment uses 500 consecutive images from the APCD dataset and employs the gray-level co-occurrence matrix method to identify crack morphologies and statistically analyze the results. Simultaneously, crack categories in the same batch of images are manually identified, and the algorithm's recognition error is calculated based on the manual identification results. As shown in Table 2, the proposed crack morphology recognition method can achieve an accuracy rate of over 92%, demonstrating significant engineering value.

[0039] Table 2 Crack Category Identification Statistics .

[0040] Example 2:

[0041] This embodiment proposes an electronic device, including: one or more processors, and a memory, wherein the memory is used to store instructions, and when the instructions are executed by the one or more processors, the one or more processors execute the asphalt pavement crack identification method based on joint optimization network.

[0042] The electronic device may be a mobile phone, computer, or tablet computer, etc., and includes a memory and a processor. The memory stores a computer program, which, when executed by the processor, implements the asphalt pavement crack identification method based on a joint optimization network as described in the embodiments. It is understood that the electronic device may also include an input / output (I / O) interface and communication components.

[0043] The processor is used to execute all or part of the steps in the asphalt pavement crack identification method based on joint optimization networks as described in the above embodiments. The memory is used to store various types of data, which may include, for example, instructions for any application or method in the electronic device, as well as application-related data.

[0044] The processor can be implemented as an Application Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), controller, microcontroller, microprocessor, or other electronic components, and is used to execute the asphalt pavement crack identification method based on joint optimization network described in the above embodiments.

[0045] Example 3:

[0046] This embodiment proposes a computer-readable storage medium that stores executable instructions. When these instructions are executed, if they are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.

[0047] The computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the asphalt pavement crack identification method based on joint optimization network described in the various embodiments of this application.

[0048] The aforementioned storage media include: flash memory, hard disk, multimedia card, card-type memory (e.g., SD (Secure Digital Memory Card) or DX (Memory Data Register, MDR) memory), random access memory (RAM), static random-access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, disk, optical disk, server, APP (Application) application store, and other media capable of storing program verification codes. These media store computer programs, which, when executed by a processor, can implement the various steps of the asphalt pavement crack identification method based on a joint optimization network described above.

[0049] Example 4:

[0050] This embodiment proposes a computer program product, including a computer program or instructions, which, when executed by a processor, implements the aforementioned asphalt pavement crack identification method based on a joint optimization network.

[0051] Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a computer program product.

[0052] The various embodiments in this application are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0053] The scope of protection of this application is not limited to the embodiments described above. Obviously, those skilled in the art can make various modifications and variations to this disclosure without departing from the scope and spirit of this disclosure. If such modifications and variations fall within the scope of this application and its equivalents, then the intent of this disclosure also includes these modifications and variations.

[0054] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the present invention.

Claims

1. A method for identifying asphalt pavement cracks based on a joint optimization network, characterized in that, include: Acquire images of asphalt pavement defects and obtain corresponding labeled images to construct an asphalt pavement crack image dataset; Design an adaptive image denoising network; The adaptive image denoising network includes a multi-filter denoising module and a hyperparameter prediction module. The multi-filter denoising module includes multiple filters and dynamically adjusts the parameters of each filter through a joint optimization strategy to achieve adaptive denoising under different noise conditions. The hyperparameter prediction module predicts and optimizes the hyperparameters of each filter based on the information of the input image. A spatial convolutional attention U-Net network is designed as a semantic segmentation network; the spatial convolutional attention U-Net network is implemented by introducing spatial and channel reconstruction convolutional modules into the Attention U-Net network; An IDSNet network is formed by combining an adaptive image denoising network and a semantic segmentation network using a joint optimization strategy. The training loss of the semantic segmentation network is used to optimize the IDSNet network, so as to achieve simultaneous denoising and crack segmentation of asphalt pavement crack images. Crack morphology identification and parameter calculation are performed based on crack segmentation results.

2. The asphalt pavement crack identification method based on joint optimization network according to claim 1, characterized in that, The multi-filter noise reduction module includes an exposure filter, a contrast filter, a tone filter, and a stripe removal filter, which are used to perform exposure filtering, contrast filtering, tone filtering, and stripe removal filtering on the input image, respectively.

3. The method for identifying asphalt pavement cracks based on a joint optimization network according to claim 2, characterized in that, The destriating filter uses the average pixel value of each column in the asphalt pavement crack image to represent the pixel value of each column. Then, a column-based illuminance balancing operation is used to map the illuminance values ​​of all columns to the desired illuminance uniformity value, as shown in the following formula: , (1); Where γ is a trainable hyperparameter; These are the processed image pixel grayscale values; It is the predetermined average grayscale value of each column after preprocessing; I ij A is the grayscale value of the pixel in the i-th column and j-th row of the k-th asphalt pavement crack image; i k It is the weighted average of the pixels in the i-th column of the k-th asphalt pavement crack image, and its calculation formula is shown in formula (2); k is the number of asphalt pavement crack images; M is the number of pixel columns in the asphalt pavement crack image; N is the number of pixel rows in the asphalt pavement crack image; (2); in, β represents the weighted variable, used to control the weight of the current image, and is also the latest measured illumination information; β represents an empirical constant. is the average pixel value of the i-th column in the k-th asphalt pavement crack image.

4. The asphalt pavement crack identification method based on a joint optimization network according to claim 1, characterized in that, The hyperparameter prediction module is a framework consisting of 5 convolutional layers and 2 fully connected layers. Based on the global content of the asphalt pavement crack image, it predicts and optimizes the hyperparameters required by the multi-filter noise reduction module to achieve adaptive noise reduction of the asphalt pavement crack image.

5. The method for identifying asphalt pavement cracks based on a joint optimization network according to claim 1, characterized in that, The method utilizes a joint optimization strategy to optimize the trainable hyperparameters of each filter in the multi-filter denoising module. Specifically, it optimizes both image denoising and semantic segmentation performance based on a stochastic gradient descent optimizer and a binary cross-entropy loss function. The hyperparameter prediction module initially predicts the hyperparameters used by each filter based on the state of the input asphalt pavement crack image. The asphalt pavement crack image after noise reduction by each filter is semantically segmented by a semantic segmentation network. The hyperparameter prediction module then adjusts the hyperparameters of each filter based on the feedback of the semantic segmentation results to further obtain a semantically segmented binary image of the asphalt pavement crack with better noise reduction effect.

6. The method for identifying asphalt pavement cracks based on a joint optimization network according to claim 1, characterized in that, The crack morphology includes transverse cracks, longitudinal cracks, and alligator cracks; the crack parameters include crack length, crack width, and crack area.

7. The asphalt pavement crack identification method based on a joint optimization network according to claim 5, characterized in that, The specific method for crack morphology recognition based on crack segmentation results is as follows: First, connected components are extracted from the semantically segmented binary image of asphalt pavement cracks. The morphology of the crack is determined by judging the similarity between the trend direction of each crack connected component and the four directions of 0°, 90°, 45° and 135°.

8. The method for identifying asphalt pavement cracks based on a joint optimization network according to claim 6, characterized in that, The crack length is obtained by summing the relative distances of multiple pixels on the crack skeleton curve; the crack skeleton curve is a path composed of a series of individual pixels and is distributed along the central axis of the crack target, reflecting the topological structure and morphological characteristics of the crack.

9. A computer-readable storage medium for executing the asphalt pavement crack identification method based on a joint optimization network as described in any one of claims 1-8, characterized in that, It stores executable instructions that, when executed, cause the processor to perform the asphalt pavement crack identification method based on a joint optimization network.

10. A computer program product for executing the asphalt pavement crack identification method based on a joint optimization network as described in any one of claims 1-8, characterized in that, This includes a computer program or instructions that, when executed by a processor, implement the aforementioned asphalt pavement crack identification method based on a joint optimization network.

Citation Information

Patent Citations

  • Dam crack detection method based on unmanned aerial vehicle image and deep learning

    CN118072193A

  • Medical image semantic segmentation method based on attention mechanism optimization

    CN120496757A

  • Pavement crack semantic segmentation method based on Transform and CNN architecture

    CN120807916A