An Unsupervised Pavement Crack Detection Method and System Based on Denoising Diffusion Model

By applying denoising diffusion model and pseudo-crack synthesis technology in pavement crack detection, the problem of relying on manual labeling data and low detection accuracy in the existing technology is solved, and efficient and accurate unsupervised pavement crack detection is achieved.

CN119784764BActive Publication Date: 2025-05-27HARBIN INST OF TECH AT WEIHAI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510285813.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-05-27
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

The existing pavement crack detection methods rely on manual labeling data, which are costly and inefficient, and the pavement crack detection based on the diffusion model has problems such as low training efficiency, high calculation overhead, and insufficient detection accuracy.

Method used

Unsupervised road surface crack detection method based on denoising diffusion model is adopted, pseudo-label is generated through pseudo-crack synthesis technology, and a segmentation model is built for training using the denoising diffusion implicit model and target-guided reconstruction mechanism, which reduces data labeling costs and improves detection accuracy and robustness.

Benefits of technology

It realizes efficient training of road surface crack detection models without manual labeling, reduces training costs and human resource consumption, improves the accuracy and robustness of crack detection, and performs well in different lighting and complex background conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119784764B_ABST
    Figure CN119784764B_ABST
Patent Text Reader

Abstract

The present invention provides an unsupervised road surface crack detection method and system based on a denoising diffusion model, which relates to the technical field of road surface crack detection. The method includes: adding noise forward to a pseudo-crack image to generate two images with different noise levels; inputting one of the noise-added images into a diffusion model to reconstruct a target image, and using the target image to guide the reconstruction process of the other noise-added image to obtain an image with defects removed; constructing a segmentation model based on a pseudo-crack data set, with the input being the splicing of a pseudo-crack image and its reconstructed image, and outputting a predicted crack result, and training by minimizing a loss function. The present invention uses pseudo-crack synthesis technology to provide pseudo-label information for the segmentation model, greatly reducing the cost of data annotation and improving the accuracy and robustness of crack detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pavement crack detection, and particularly to an unsupervised pavement crack detection method and system based on a denoising diffusion model. Background Art

[0002] Pavement cracks are the most common type of damage in road facilities, usually caused by various factors such as temperature changes, traffic loads, moisture erosion, and material aging. Pavement cracks not only affect the service life of the road but also seriously threaten traffic safety. Therefore, timely and effectively detecting pavement cracks is crucial for road maintenance and safety management. Traditional crack detection methods mainly rely on manual inspection, which is inefficient, costly, and easily limited by the experience and subjective judgment of inspectors. With the development of technology, automated crack detection systems have gradually become a research hotspot.

[0003] Traditional image processing-based methods, such as threshold methods, local binary pattern (LBP), Fourier transform, and wavelet transform, etc., although perform well in specific scenarios, these methods rely on manual feature extraction and have poor generalization ability under complex backgrounds or environmental changes. To overcome these problems, supervised learning methods based on machine learning have gradually been applied, especially convolutional neural networks (CNNs) have been widely used in pavement crack detection. By training a classifier using labeled data, supervised learning methods can automatically extract features from images and perform crack segmentation. However, these methods also have certain limitations, mainly reflected in the dependence on a large amount of labeled data. The collection of labeled data is not only costly but also requires a large amount of human resources, restricting its popularization in practical applications.

[0004] In recent years, unsupervised learning methods, as an emerging solution, have begun to be applied in the field of crack detection. The unsupervised methods based on reconstruction generally first reconstruct the crack image to generate an ideal image after removing the cracks, and then calculate the difference between the crack image and the ideal image, and mark the regions with large differences as crack defects. However, these methods generally have problems such as poor reconstruction quality, inability to effectively process large-scale crack regions, and insufficient positioning accuracy.

[0005] In recent years, the diffusion model has achieved great success in the field of natural image generation. By gradually adding noise to the data and restoring the data through the reverse denoising process, this model has been proven to have strong capabilities in industrial product defect image repair and anomaly detection tasks. However, due to the large domain differences between the data characteristics of pavements and industrial products or natural images, directly using the diffusion model for pavement crack detection methods still faces problems such as low training efficiency, large computational overhead, and insufficient detection accuracy. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide an unsupervised road crack detection method and system based on a denoising diffusion model, which uses pseudo-crack synthesis technology to provide pseudo-label information for the segmentation model, greatly reducing the cost of data annotation and improving the accuracy and robustness of crack detection.

[0007] To solve the above technical problems, the technical solution of the present invention is as follows:

[0008] In the first aspect, an unsupervised road crack detection method based on a denoising diffusion model, the method includes:

[0009] Step 1, use the unlabeled normal road image dataset D1 to train the denoising diffusion implicit model, and normalize the input image to a unified size W×H;

[0010] Step 2, construct a road surface defect texture dataset D2, add noise to the normal road image and input it into the diffusion model for reconstruction, and add random perturbations to generate abnormal road surface images;

[0011] Step 3, generate a random Perlin noise image and binarize it to obtain a mask image, fuse the abnormal texture image and the normal road image based on the mask image to generate a pseudo-crack road surface image and its pseudo-label;

[0012] Step 4, perform forward noise addition on the pseudo-crack image to generate two images with different noise levels;

[0013] Step 5, input one of the noise-added images into the diffusion model to reconstruct it into a target image, and use the target image to guide the reconstruction process of the other noise-added image to obtain an image with defects removed;

[0014] Step 6, construct a segmentation model based on the pseudo-crack dataset, with the input being the concatenation of the pseudo-crack image and its reconstructed image, output the predicted crack result, and train it by minimizing the loss function;

[0015] Step 7, replace the pseudo-crack image with the real road surface image to be detected to obtain a reconstructed image, concatenate them and input them into the trained segmentation model to obtain the final defect detection result.

[0016] Further, constructing the road surface defect texture dataset D2, adding noise to the normal road image and inputting it into the diffusion model for reconstruction, and adding random perturbations to generate abnormal road surface images, includes:

[0017] Construct a road surface defect texture dataset D2, for each normal road image Add noise, input it into the diffusion model trained in Step 1 for reconstruction, and add random perturbations during the reverse denoising reconstruction process of the diffusion model to make it deviate from the normal distribution and generate abnormal road surface images .

[0018] Furthermore, a random Berlin noise image is generated and binarized to obtain a mask image. The abnormal texture image and the normal road image are fused based on the mask image to generate a pseudo-crack pavement image and its pseudo-label, including:

[0019] Construct a pavement pseudo-crack dataset D3. For each pavement image in dataset D1 , generate a random two-dimensional Berlin noise image P with a size of W×H, and randomly select a threshold for P for image binarization to obtain a mask image with pixel-by-pixel labels;

[0020] Use the constructed pavement defect texture dataset D2 as the abnormal texture source A. Randomly select an abnormal texture image from A and, based on the mask image and the normal road image for fusion to obtain a pseudo-crack pavement image . The defect pseudo-label of the pseudo-crack pavement image is the obtained label image .

[0021] Furthermore, forward noise is added to the pseudo-crack image to generate two images with different noise levels, including:

[0022] For dataset D3, for each image respectively use the diffusion time steps and , and based on the noise model in the diffusion model in step 1, perform forward noise addition on the pseudo-crack image to obtain an image and an image .

[0023] Furthermore, input a noisy image into the diffusion model to reconstruct it into a target image, and use the target image to guide the reconstruction process of another noisy image to obtain an image with defects removed, including:

[0024] For each image in dataset D3 , input the image into the diffusion model, and define the reconstructed image as the target image ;

[0025] Use the target image to guide the reverse reconstruction process of the diffusion model, and reconstruct the image into an image with defects removed .

[0026] Further, a segmentation model is constructed based on the pseudo-crack dataset. The input is the concatenation of the pseudo-crack image and its reconstructed image, and the predicted crack result is output. It is trained by minimizing the loss function, including:

[0027] Construct and train a segmentation model based on the dataset D3. The input of the segmentation model is the pseudo-crack image concatenated in the channel dimension and its reconstructed image , and the output is a single-channel predicted crack result image;

[0028] It is trained by minimizing the loss function between the predicted output of the segmentation model and the pseudo-label to obtain a segmentation model for pavement crack detection.

[0029] Further, replace the pseudo-crack image with the real pavement image to be detected to obtain a reconstructed image. After concatenation, input it into the trained segmentation model to obtain the final defect detection result, including:

[0030] Replace the pseudo-crack image in steps 4 and 5 with the real pavement image to be detected, and obtain the reconstructed image in the manner of steps 4 to 5 ;

[0031] Concatenate the pavement image to be detected and the corresponding reconstructed image in the channel dimension and input them into the trained segmentation model to obtain the final defect detection result.

[0032] In the second aspect, an unsupervised pavement crack detection system based on a denoising diffusion model includes:

[0033] The first training module is used to train a denoising diffusion implicit model using the unlabeled normal road image dataset D1, and the input image is normalized to a unified size W×H;

[0034] The generation module is used to construct a pavement defect texture dataset D2. By adding noise to the normal road image and inputting it into the diffusion model for reconstruction, random perturbations are added to generate abnormal pavement images;

[0035] The processing module is used to generate a random Perlin noise image and binarize it to obtain a mask image, and fuse the abnormal texture image and the normal road image based on the mask image to generate a pseudo-crack pavement image and its pseudo-label;

[0036] The noise addition module is used to perform forward noise addition on the pseudo-crack image to generate two images with different noise levels;

[0037] The reconstruction module is used to input one of the noise-added images into the diffusion model to reconstruct it into a target image, and use the target image to guide the reconstruction process of the other noise-added image to obtain an image with defects removed;

[0038] The second training module is used to construct a segmentation model based on the pseudo-crack dataset. The input is the concatenation of the pseudo-crack image and its reconstructed image, and the output is the predicted crack result. It is trained by minimizing the loss function.

[0039] The splicing module is used to replace the pseudo-crack image with the real pavement image to be detected to obtain the reconstructed image. After splicing, it is input into the trained segmentation model to obtain the final defect detection result.

[0040] In a third aspect, a computing device includes:

[0041] One or more processors;

[0042] A storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method.

[0043] In a fourth aspect, a computer-readable storage medium stores a program that implements the method when executed by a processor.

[0044] The above solution of the present invention has at least the following beneficial effects:

[0045] The present invention only relies on normal pavement images for training, avoiding the need for manually labeled data, and greatly reducing the training cost and human resource consumption. Secondly, by combining the denoising diffusion implicit model (DDIM) with the target-guided reconstruction mechanism, the present invention utilizes the advantages of two noise scales, enabling the generated reconstructed images to have both pixel quality and semantic quality, being able to more precisely restore the details of the crack images, and at the same time improving the accuracy of crack localization.

[0046] The pseudo-crack synthesis technology proposed by the present invention can provide rich training samples, enhancing the robustness of the model under different lighting, shadow, and complex background conditions, making the crack detection perform excellently in various environments. The method provided by the present invention has been tested on the publicly available Crack500 dataset, and has reached 44.71%, 66.08%, and 93.66% respectively in key indicators such as IoU, F1-Score, and AUROC, all of which are better than existing unsupervised crack detection methods, showing its efficient and accurate performance. Description of the Drawings

[0047] Figure 1 is the overall flowchart of the embodiment of the present invention;

[0048] Figure 2 is the architecture diagram of the training process of the present invention;

[0049] Figure 3 is the architecture diagram of the testing process of the present invention;

[0050] Figure 4 is a flowchart of the pseudo-crack synthesis method in the present invention;

[0051] Figure 5 is an example diagram of the detection result of the present invention. Specific embodiments

[0052] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0053] As Figure 1 shown, an embodiment of the present invention proposes an unsupervised road surface crack detection method based on a denoising diffusion model, and the method includes the following steps:

[0054] Step 1, use the unlabeled normal road image dataset D1 to train the denoising diffusion implicit model (as Figure 2 shown on the left). In this embodiment, the publicly available Crack500 dataset is used. 1896 normal road images are cropped from its training set as the dataset D1. Before training, the input road images are first normalized to the same size of 256×256. The total number of time steps in the diffusion process is set to , the batch size is set to 32, the initial learning rate is 3×10 -4 , the weight decay is set to 0.05, and 4000 epochs are trained. For the denoising diffusion implicit model, an improved Unet framework is adopted. The specific structure is: the first layer is the input layer, and the size of each input sample is , the first layer uses convolution kernels, and the number of output channels is 64; the downsampling part includes six convolutional layers, which are downsampled layer by layer, and the resolutions are: 256→128→64→32→16→8 in sequence. Each convolutional layer contains 2 residual blocks (each residual block includes two Convolution), attention modules (applied at resolutions 32, 16, and 8), after each downsampling, the resolution is halved and the number of channels increases (doubling factors are 1, 1, 2, 2, 4, 4); followed by a Middle Block at the lowest resolution layer, containing 2 residual blocks and 1 multi-head attention block for processing the bottleneck features of the network; the upsampling part contains six convolutional layers, upsample layer by layer, with resolutions 8→16→32→64→128→256 in sequence, each layer contains 2 residual blocks and attention modules (applied at resolutions 8, 16, and 32). The upsampling part uses nearest neighbor interpolation for resolution magnification and adjusts the number of channels through convolution, and uses skip connections to concatenate the features of each layer with the corresponding downsampled features. The last layer is the output layer, using convolution kernels, with the number of channels being 3, to generate an image with a final size of ;

[0055] Step 2, construct the road surface defect texture dataset D2. Add noise to each normal road image , and then input it into the diffusion model trained in Step 1 for reconstruction. Add random perturbations during the reverse denoising reconstruction process of the diffusion model to make it deviate from the normal distribution and generate abnormal road surface images , the specific process is as follows:

[0056] During the reverse denoising process, the diffusion model gradually restores the image through the following conditional probability distribution:

[0057] ;

[0058] where, is the time step, , represents the reconstructed image at time , refers to the original input image predicted by the noise model of the diffusion model based on the noise image , is the noise scheduling parameter in the diffusion process, , is an equally spaced sequence from 0.0001 to 0.02, is the noise prediction value of the diffusion model, is a random noise term sampled from the standard Gaussian distribution.

[0059] To generate abnormal road surface images , introduce additional perturbation terms during the reverse denoising process, including two core steps:

[0060] (a) Add a random perturbation s to the noise model, specifically expressed as follows:

[0061]

[0062] Among them, , is the abnormal image at the moment, is the control parameter of the abnormal intensity, and its value ranges from [0, 1]. In this implementation, is set.

[0063] (b) On the basis of the previous step, the prior information of the crack is introduced to further enhance the realism of the defect texture. Generally, the inside of the crack is wider and sunken downward, and dust and dirt are accumulated inside, resulting in the color of the crack being generally darker, and there is a large color difference visually from the surrounding normal road surface. The specific implementation method is as follows:

[0064] ;

[0065] Among them, is randomly sampled in the interval so that the image texture reconstructed based on the diffusion model is generated in a darker direction to approximate the visual characteristics of the road surface crack;

[0066] Step 3, as Figure 4 shown, construct the road surface pseudo-crack dataset D3. For each road surface image (normal texture) in the dataset D1, generate a random two-dimensional Berlin noise image P with a size of W×H, and randomly select a threshold for P to perform image binarization to obtain a mask image with pixel-by-pixel labels. Use the constructed road surface defect texture dataset D2 as the abnormal texture source A, randomly select abnormal texture images from A, and based on the mask image and the normal road image to perform fusion to obtain the pseudo-crack road surface image . The defect pseudo-label of this image is the label image obtained previously. The above process can be expressed by the following formula:

[0067]

[0068] Among them, is the inversion of the pseudo-label , represents the element-wise multiplication operation, and is the parameter that controls the opacity of the mixing process and is randomly sampled from a given interval (i.e., ). The finally synthesized training samples include two parts: the road surface image containing pseudo-cracks, and the pixel-level pseudo-label ;

[0069] Step 4, for the dataset D3, for each image use a smaller and a larger diffusion time step respectively and , and perform forward noise addition on the pseudo-crack image based on the noise model in the diffusion model in step (1) to obtain an image with small noise and an image with large noise ;

[0070] Step 5, as shown in Figure 2 (middle), for each image in the dataset D3 , input the small-noise image into the trained diffusion model, and through a shorter diffusion trajectory, reconstruct it into an image that retains more fine-grained textures but may not fully restore the crack defects . Define the image as the target image, and use this target to guide the reverse process of the diffusion model to reconstruct the large-noise image into an image with cracks removed . To guide 's reconstruction result to approximate the target image , by adding noise consistent with the image at the moment to the target image , an intermediate target image with a similar distribution to can be generated. This enables to be gradually guided towards approximation in each step of denoising. The target image is calculated through a linear combination of the prediction of the noise by the diffusion model in each step of reverse diffusion :

[0071] ;

[0072] where is the noise scheduling parameter in the diffusion process. Then, according to the conditional score term , further correct the noise estimate value of the model. The corrected noise term is expressed as:

[0073] ;

[0074] where is a hyperparameter that adjusts the strength of the conditional constraints. Finally, in each step of reverse diffusion denoising, the model updates the denoising result based on the corrected noise term as follows:

[0075] ;

[0076] Through the target-guided mechanism, the reconstructed image can not only completely remove the crack defect area but also retain more detailed textures of the original image. For the Crack500 dataset, the noise scale is set to 200, is set to 550. The target-guided mechanism enables the reconstructed image to not only filter out crack defects but also be very similar to the original normal image ;

[0077] Step 6, as shown in Figure 2 (right), constructs and trains a segmentation model based on dataset D3. The input of this model is the pseudo-crack image stitched in the channel dimension and its reconstructed image, and the output is a single-channel predicted crack result image. The model is trained by minimizing the loss function between the predicted output of the segmentation model and the pseudo-label.

[0078] The segmentation model adopts a U-Net architecture, which includes an encoder, a decoder, and skip connections. Among them, the encoder gradually reduces the spatial resolution and extracts features. It consists of six convolutional blocks with a basic channel number of 64. Each convolutional block contains two convolutional layers and a max-pooling layer with a kernel size of and a stride of ; the decoder upsamples through nearest neighbor interpolation and consists of five convolutional blocks with a kernel size of . After upsampling each decoder convolutional block, it is concatenated with the output of the encoder convolutional block for feature fusion, and then processed through two convolutional layers; finally, a segmentation result probability map at the pixel level is generated through an output layer with a kernel size of and a Sigmoid activation function. Manually or automatically set the binarization threshold. Here, the manual threshold of 0.5 is adopted, which can convert the result probability map into a binary mask segmentation result map. The training parameter batch size is set to 64, the initial learning rate is 1×10 −4 , and it is trained for 100 epochs;

[0079] Step 7, in the actual road crack detection application, the pseudo-crack images in steps 4 and 5 Replace it with the actual pavement image to be detected. Obtain the reconstructed image in the manner of Steps 4 to 5. After concatenating the pavement image to be detected and the corresponding reconstructed image in the channel dimension, input them into the trained segmentation model, and finally generate the predicted defect mask as the final detection result.

[0080] As Figure 2 As shown in the left figure, use the unlabeled normal road image dataset D1 to train the Denoising Diffusion Implicit Model (DDIM). In this embodiment, the publicly available Crack500 dataset is used. Cut out 1896 normal road images from its training set as the dataset D1. Before training, first normalize the input road images to the same size of 256×256. The total number of time steps in the diffusion process is set to , the batch size is set to 32, the initial learning rate is 3×10 -4 , the weight decay is set to 0.05, and train for 4000 epochs. For the Denoising Diffusion Implicit Model, an improved Unet framework is adopted. The specific structure is: The first layer is the input layer, and the size of each input sample is , the first layer uses convolution kernels, and the number of output channels is 64; the downsampling part includes six convolutional layers, which are downsampled layer by layer, and the resolutions are: 256→128→64→32→16→8 in sequence. Each convolutional layer contains 2 residual blocks (each residual block includes two convolutions) and attention modules (applied at resolutions of 32, 16, and 8). After each downsampling, the resolution is halved and the number of channels increases (the multiplication factors are 1, 1, 2, 2, 4, 4); then there is a Middle Block located at the lowest resolution layer, which contains 2 residual blocks and 1 multi-head attention block for processing the bottleneck features of the network; the upsampling part contains six convolutional layers, which are upsampled layer by layer, and the resolutions are: 8→16→32→64→128→256 in sequence. Each layer contains 2 residual blocks and attention modules (applied at resolutions of 8, 16, and 32). The upsampling part uses the nearest neighbor interpolation method to enlarge the resolution and adjusts the number of channels through convolution, and uses skip connections to splice the features of each layer with the corresponding downsampled features. The last layer is the output layer, which uses convolution kernels, the number of channels is 3, and generates an image with a final size of ;

[0081] (2) Construct the pavement defect texture dataset D2. In this embodiment, the publicly available DTD texture dataset is used, and randomly select data augmentation methods (random rotation, color jitter, histogram equalization, pixel value flipping, sharpness enhancement, hue change) for abnormal textures to obtain the pavement defect texture dataset D2.

[0082] (3) AsFigure 4 As shown, construct the pavement pseudo-crack dataset D3. For each pavement image in dataset D1 (normal texture), generate a random two-dimensional Perlin noise image P of size W×H, and randomly select a threshold for P for image binarization to obtain a mask image with pixel-by-pixel labels . Use the constructed pavement defect texture dataset D2 as the abnormal texture source A, randomly select abnormal texture images from A, and based on the mask image and the normal road image for fusion to obtain a pseudo-crack pavement image . The defective pseudo-label of this image is the label map obtained previously . The above process can be expressed by the following formula:

[0083] ;

[0084] where, is the inversion of the pseudo-label , represents the element-wise multiplication operation, and is a parameter that controls the opacity of the mixing process, randomly sampled from a given interval (i.e., ). The finally synthesized training samples include two parts: the pavement image containing pseudo-cracks , and the pixel-level pseudo-label .

[0085] (4) For dataset D3, for each image , use a smaller and a larger diffusion time step and respectively, and perform forward noise addition on the pseudo-crack image based on the noise model in the diffusion model in step (1) to obtain an image with small noise and an image with large noise .

[0086] (5) As Figure 2 (in the middle) shows, for each image in dataset D3, input the small-noise image into the trained diffusion model, and through a shorter diffusion trajectory, reconstruct it into an image that retains more fine-grained textures but may not fully restore the crack defects. Define the image as the target image, and use this target to guide the reverse process of the diffusion model to reconstruct the large-noise image into an image without cracks. To guide Reconstruction result Approximates to the target image By adding to the target image Noise that is consistent with the Image at the moment An intermediate target image with a similar distribution to the Can be generated . This enables, in each step of denoising, To be gradually guided towards Approximation. The target image In each step of reverse diffusion is calculated through a linear combination of the noise prediction of the diffusion model As follows: :

[0087] ;

[0088] Wherein, Is the noise scheduling parameter in the diffusion process. Then, according to the conditional score term , The noise estimate of the model is further corrected , The corrected noise term Is expressed as:

[0089] ;

[0090] Wherein, Is the hyperparameter that adjusts the strength of the conditional constraint. Finally, in each step of reverse diffusion denoising, the model updates the denoising result Based on the corrected noise term As:

[0091] ;

[0092] Through the target guidance mechanism, the reconstructed image Can not only completely remove the crack defect area, but also retain more detailed textures of the original image. For the Crack500 dataset, the noise scale Is set to 200, Is set to 550. The target guidance mechanism enables the reconstructed image To not only filter out crack defects, but also be very similar to the original normal image .

[0093] (6) As shown in Figure 2 (right), a segmentation model is constructed and trained based on the dataset D3. The input of this model is the pseudo-crack image And its reconstructed image , the output is a single-channel predicted crack result image. The training is carried out by minimizing the loss function between the predicted output of the segmentation model and the pseudo-label.

[0094] The segmentation model adopts a U-Net architecture, which includes an encoder, a decoder, and skip connections. Among them, the encoder gradually reduces the spatial resolution and extracts features. It consists of six convolutional blocks with a basic number of channels of 64. Each convolutional block contains two convolutional layers and a max-pooling layer with a convolutional kernel size of and a stride of ; the decoder performs upsampling through nearest-neighbor interpolation and consists of five convolutional blocks with a convolutional kernel size of . After upsampling each decoder convolutional block, it is concatenated with the output of the encoder convolutional block for feature processing, and then processed through two convolutional layers; finally, a pixel-level segmentation result probability map is generated through an output layer with a convolutional kernel size of and a Sigmoid activation function. Manually or automatically set the binarization threshold. Here, a manual threshold of 0.5 is adopted, which can convert the result probability map into a binary mask segmentation result map. The training parameter batch size is set to 64, and the initial learning rate is 1×10 −4 , and train for 100 epochs.

[0095] (7) In the actual road crack detection application, replace the pseudo-crack images in steps (4) and (5) with real pavement images to be detected. Obtain the reconstructed image in the manner of steps (4) to (5). Concatenate the pavement image to be detected and the corresponding reconstructed image in the channel dimension and input them into the trained segmentation model, and finally generate a predicted defect mask as the final detection result.

[0096] An unsupervised road crack detection system based on a denoising diffusion model, including:

[0097] The first training module is used to train the denoising diffusion implicit model using the unlabeled normal road image dataset D1, and the input images are normalized to a unified size of W×H;

[0098] The generation module is used to construct the pavement defect texture dataset D2. By adding noise to the normal road images and inputting them into the diffusion model for reconstruction, random perturbations are added to generate abnormal pavement images;

[0099] The processing module is used to generate a random Perlin noise image and binarize it to obtain a mask image, and fuse the abnormal texture image and the normal road image based on the mask image to generate a pseudo-crack pavement image and its pseudo-label;

[0100] A noise addition module, configured to perform forward noise addition on the pseudo-crack image to generate two images with different noise levels;

[0101] A reconstruction module, configured to input one of the noise-added images into a diffusion model to reconstruct it into a target image, and use the target image to guide the reconstruction process of the other noise-added image to obtain an image with defects removed;

[0102] A second training module, configured to build a segmentation model based on a pseudo-crack data set, with the input being the splicing of the pseudo-crack image and its reconstructed image, and the output being the predicted crack result, and perform training by minimizing a loss function;

[0103] A splicing module, configured to replace the pseudo-crack image with a real pavement image to be detected to obtain a reconstructed image, and input the spliced image into the trained segmentation model to obtain the final defect detection result.

[0104] It should be noted that this system corresponds to the above method. All implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0105] An embodiment of the present invention further provides a computing device, including: a processor and a memory storing a computer program. When the computer program is run by the processor, it executes the method as described above. All implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0106] An embodiment of the present invention further provides a computer-readable storage medium storing instructions. When the instructions are run on a computer, the computer is caused to execute the method as described above. All implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

Claims

1. An unsupervised pavement crack detection method based on a denoising diffusion model, characterized in that: The method comprises: Step 1: Use the unlabeled normal road image dataset D1 to train the denoising diffusion implicit model, and the input image is normalized to a uniform size of W×H; Step 2, construct a road surface defect texture dataset D2, add noise to normal road images and input the diffusion model for reconstruction, add random disturbance to generate abnormal road surface images; Step 3, generate a random Perlin noise image and binarize it to obtain a mask image, fuse the abnormal texture image with the normal road image based on the mask image, and generate a pseudo-crack road surface image and its pseudo-label, including: constructing a road surface pseudo-crack dataset D3, for each road surface image in the dataset D1 , generate a random two-dimensional Perlin noise image P of size W×H, and randomly select a threshold for P to binarize the image to obtain a mask map with pixel labels; take the constructed road defect texture dataset D2 as the abnormal texture source A, randomly select abnormal texture images from A, and based on the mask map Normal road image Fusion is performed to obtain a pseudo-crack road surface image , Pseudo-crack pavement image The defect pseudo label is the obtained label map ; Step 4: Forward-add noise to the pseudo-crack image to generate two images with different noise levels, including: for each image in the dataset D3, Both use the diffusion time step and , based on the noise model in the diffusion model in step 1, the pseudo crack image Perform forward noise addition to obtain an image and an image ,in, The diffusion time step is less than ; Step 5: Input a noisy image into the diffusion model to reconstruct it into a target image, and use the target image to guide the reconstruction process of another noisy image to obtain an image with defects removed, including: for each image in the data set D3 , the image Input the diffusion model and define the reconstructed image as the target image ; Using the target image Guide the inverse reconstruction process of the diffusion model to transform the image Reconstruct the image to remove defects ; Step 6: construct a segmentation model based on the pseudo-crack dataset, with the input being the concatenation of the pseudo-crack image and its reconstructed image, and outputting the predicted crack result, which is trained by minimizing the loss function; Step 7: Replace the pseudo crack image with the real road surface image to be detected to obtain a reconstructed image, which is then stitched and input into the trained segmentation model to obtain the final defect detection result.

2. The unsupervised pavement crack detection method based on denoising diffusion model according to claim 1 is characterized in that: Construct a road surface defect texture dataset D2 by adding noise to normal road images and inputting the diffusion model for reconstruction, and add random disturbances to generate abnormal road surface images, including: Construct a road defect texture dataset D2, for each normal road image Add noise and input it into the diffusion model trained in step 1 for reconstruction. Add random disturbances to the diffusion model inverse denoising and reconstruction process to make it deviate from the normal distribution and generate abnormal road surface images. .

3. The unsupervised pavement crack detection method based on denoising diffusion model according to claim 2 is characterized in that: A segmentation model is built based on the pseudo-crack dataset. The input is the concatenation of the pseudo-crack image and its reconstructed image. The output is the predicted crack result. The training is performed by minimizing the loss function, including: The segmentation model is constructed and trained based on the dataset D3. The input of the segmentation model is the pseudo crack image spliced ​​in the channel dimension. and its reconstructed image , the output is a single-channel predicted crack result image; The training is performed by minimizing the loss function between the predicted output of the segmentation model and the pseudo label to obtain a segmentation model for pavement crack detection.

4. The unsupervised pavement crack detection method based on denoising diffusion model according to claim 3 is characterized in that: The real road surface image to be inspected replaces the pseudo crack image to obtain the reconstructed image, which is then stitched and input into the trained segmentation model to obtain the final defect detection results, including: The pseudo crack images in step 4 and step 5 are Replace it with the real road image to be detected, and obtain the reconstructed image by following steps 4 to 5. ; The road surface image to be inspected and the corresponding reconstructed image are spliced ​​in the channel dimension and input into the trained segmentation model to obtain the final defect detection result.

5. An unsupervised pavement crack detection system based on a denoising diffusion model, characterized in that: Applied to the method according to any one of claims 1 to 4, comprising: The first training module is used to train the denoising diffusion implicit model using the unlabeled normal road image dataset D1, and the input image is normalized to a uniform size W×H; The generation module is used to construct the road surface defect texture dataset D2 by adding noise to the normal road image and inputting the diffusion model for reconstruction, and adding random disturbance to generate abnormal road surface images; A processing module is used to generate a random Perlin noise image and binarize it to obtain a mask image, fuse the abnormal texture image with the normal road image based on the mask image, and generate a pseudo-crack road surface image and its pseudo-label; The noise adding module is used to perform forward noise adding on the pseudo crack image to generate two images with different noise levels, including: for each image in the dataset D3, Both use the diffusion time step and , based on the noise model in the diffusion model in step 1, the pseudo crack image Perform forward noise addition to obtain an image and an image ; The reconstruction module is used to input a noisy image into the diffusion model to reconstruct the target image, and use the target image to guide the reconstruction process of another noisy image to obtain an image with defects removed, including: for each image in the data set D3 , the image Input the diffusion model and define the reconstructed image as the target image ; Using the target image Guide the inverse reconstruction process of the diffusion model to transform the image Reconstruct the image to remove defects ; The second training module is used to build a segmentation model based on the pseudo-crack dataset. The input is the concatenation of the pseudo-crack image and its reconstructed image. The output is the predicted crack result. The training is performed by minimizing the loss function. The splicing module is used to replace the pseudo crack image with the real road surface image to be detected to obtain a reconstructed image, which is then input into the trained segmentation model to obtain the final defect detection result.

6. A computing device, characterized in that include: one or more processors; A storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program, which implements the method according to any one of claims 1 to 4 when executed by a processor.

Citation Information

Patent Citations

  • Remote sensing image semantic segmentation method and system based on diffusion model and knowledge distillation

    CN117152427A

  • Concrete crack image generation method based on de-noising diffusion probability model

    CN118657719A