Road surface crack segmentation method based on weak supervision

By combining the integrated data set, self-attention mechanism and dense condition random field, the problem of high labeling cost and insufficient accuracy of weak supervision methods in road surface crack detection is solved, and low-cost and efficient crack segmentation is achieved, especially in complex backgrounds to accurately locate the crack boundaries, enhancing the adaptability and robustness of the model.

CN120259650APending Publication Date: 2025-07-04KUNSHAN TRANSPORTATION ENG TEST CENT CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510288919.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing weak supervision methods have problems such as high labeling cost, insufficient accuracy, noise interference and poor generalization capabilities in road surface crack detection. Especially when the cracks are small or the background is complex, the heat map generated by the class activation map (CAM) is blurred and vulnerable to noise interference.

Method used

By integrating the data set for image-level annotation, the CycleGAN generator that adds a self-attention mechanism and CAM loss function generates an initial activation map, and uses dense condition random fields to refine the crack area to generate the final crack segmentation image.

Benefits of technology

It realizes low-cost, accurate and efficient pavement crack segmentation, which can accurately locate crack boundaries in complex backgrounds, reduce background noise interference, and enhance the model's adaptability and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259650A_ABST
    Figure CN120259650A_ABST
Patent Text Reader

Abstract

The invention relates to a pavement crack segmentation method based on weak supervision, and the method comprises the following steps: S1, carrying out the integration of a data set, training the integrated data set, and carrying out the image-level marking; s2, adding a self-attention mechanism module and a CAM loss function, and generating an initial activation graph of a crack region by using a CyclGAN generator; s3, marking a crack region for the initial activation graph, refining the crack region through a dense conditional random field, and carrying out refined marking on the refined crack region; and S4, outputting a segmentation result, and generating a final crack segmentation image. According to the method, the advantage of low marking cost required by weak supervision is combined with the requirement of pavement crack segmentation, and low-cost, accurate and efficient pavement crack segmentation is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of pavement crack segmentation and the field of image processing technology, and specifically to a pavement crack segmentation method based on weak supervision. Background Technique

[0002] Road crack detection is an important task in infrastructure management. The traditional manual inspection method has low efficiency, high cost, and is easily affected by human factors. With the development of computer vision and deep learning technologies, automated crack detection has gradually become the mainstream. At present, common automated detection methods include image classification, object detection, and semantic segmentation, etc. Among them, semantic segmentation can provide pixel-level accurate crack localization, but it also faces some challenges.

[0003] Traditional semantic segmentation methods usually require a large amount of pixel-level labeled data, which not only consumes time and energy, but also has a very high labeling cost in large-scale crack detection tasks. To solve this problem, weak supervision learning methods have emerged. It reduces the dependence on accurate labeling by using image-level labels or rough region annotations, thereby reducing labor costs. However, existing weak supervision methods still have some problems. Especially when using the class activation map (CAM), the generated heat map usually cannot accurately calibrate the boundary of the crack. Especially when the crack is small or the background is complex, it is easily affected by noise interference.

[0004] The class activation map (CAM) generates a heat map reflecting the network's attention area by performing weighted summation on the feature map of the last layer of the convolutional neural network. In the crack segmentation task, CAM can help the network focus on the crack area, but the boundary of the heat map it generates is usually relatively blurred and cannot accurately segment the crack. In addition, the CAM method is sensitive to background noise and may cause the model to mis-focus on irrelevant areas, affecting the segmentation effect.

[0005] In summary, how to overcome the problems of high labeling cost, insufficient accuracy, noise interference, and poor generalization ability existing in the weak supervision method and CAM technology in the crack detection task in the prior art is an important task in the current field of pavement crack segmentation. Summary of the Invention

[0006] To solve the above problems in the prior art, the present invention provides a pavement crack segmentation method based on weak supervision, including the following steps:

[0007] S1. Integrate the data set, train the integrated data set, and perform image-level annotation.

[0008] S2. Add a self-attention mechanism module and a CAM loss function, and use the CycleGAN generator to generate an initial activation map of the crack area.

[0009] S3. Mark the crack area on the initial activation map, refine the crack area through a dense conditional random field, and perform refined annotation on the refined crack area.

[0010] S4. Output the segmentation result to generate the final crack segmentation image.

[0011] Furthermore, the dataset integration in S1 includes: using the Crack500 dataset, CFD dataset, Deepcrack dataset, and custom dataset as the integrated datasets.

[0012] Furthermore, the training of the integrated dataset in S1 includes: randomly dividing the integrated dataset into a training set and a validation set according to a preset ratio, and training and validating the dataset.

[0013] Furthermore, the preset ratio is 99:1 - 9:1.

[0014] Furthermore, the image-level annotation in S1 includes: dividing the pictures in the trained dataset into pavement images with cracks and normal pavement images without cracks, and generating image-level weakly supervised labels.

[0015] Furthermore, adding the self-attention mechanism module in S2 includes: adding a self-attention module to the CycleGAN generator network to form a network structure including a convolutional module, a residual module, a deconvolutional module, and a self-attention module.

[0016] Furthermore, adding the CAM loss function in S2 includes: modifying the CycleGAN loss function, adding the CAM loss function on the basis of the original loss function, and the added CAM loss function is:

[0017]

[0018] where H and W are the height and width of the image respectively, is the value of the generated class activation map at position (i, j), and CAM target (i, j) is the value of the class activation map of the target image at position (i, j).

[0019] Furthermore, the loss function Loss includes adversarial loss, cycle consistency loss, identity loss, and CAM loss, and the loss function Loss is:

[0020] Loss = λ1L GAN (G, D Y , X, Y) + λ2L GAN (F, D X , Y, X) + λ3L cyc (G, F) + λ4L Idt (G, F) + λ5L CAM

[0021] Among them, λ1, λ2, λ3, λ4, λ5 are weight parameters used to balance the contributions of different losses, and L GAN is the adversarial loss, and L cyc is the cycle consistency loss, and L Idt is the identity loss, G is the generator, D Y is the discriminator of the target domain, X is the source domain input image, Y is the target domain input image, F is the inverse generator, D X is the discriminator of the source domain.

[0022] Furthermore, the dense conditional random field in S3 is refined by the following energy function:

[0023]

[0024] Among them, ψ u (x i ) is the unary term, ψ p (x i , x j ) is the pairwise term, i and j are the numbers of image pixels, x i , x j are the pixel values or feature values at i and j; the unary term ψ u (x i ) is derived from the classification confidence of the CAM.

[0025] Furthermore, the pairwise term ψ p (x i , x j ) uses a Gaussian kernel function to refine the annotation of the crack area:

[0026]

[0027] Among them, μ(x i , x j ) is a similarity function used to measure whether pixels x i and x j belong to the same category, w1 and w2 are weighting parameters, p i , p j are the position coordinates of x i and x j respectively, I i , I j are the intensity information of x i and x j respectively, θ α controls the influence range of the spatial distance (position) in the similarity calculation, θ β controls the influence range of the intensity information in the similarity calculation, θ γControl the influence range of the Gaussian kernel based only on spatial position.

[0028] The present invention adapts to the development trend of automatic detection of road surface cracks, combines the advantages of low cost of weak supervision annotation with the requirements of road surface crack segmentation, and proposes a weak supervision-based road surface crack segmentation method. By introducing a self-attention mechanism into CycleGAN, it can effectively enhance the saliency of the crack area, capture the long-range dependence relationship of the cracks, and significantly improve the accuracy of crack segmentation. Especially in the case of fine cracks and complex backgrounds, it can accurately locate the boundaries of the cracks. By using image-level labels instead of precise pixel-level annotations, the workload and cost of manual annotation are reduced. By optimizing the generated class activation map (CAM), the network can still effectively perform crack segmentation in the case of scarce labeled data. By introducing a dense conditional random field (Dense CRF), the present invention can effectively reduce the interference of background noise, refine the crack boundaries, thereby improving the accuracy of the segmentation results, and enhancing the adaptability and robustness of the model to complex road surface scenarios. The present invention combines the advantages of low annotation cost required by weak supervision with the requirements of road surface crack segmentation to achieve low-cost, accurate, and efficient road surface crack segmentation. Brief Description of the Drawings

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0030] Figure 1 It is the overall work flow chart of the weak supervision road surface crack segmentation method based on the improved CycleGAN implemented by the present invention;

[0031] Figure 2 It is a representative sample of the custom dataset in the embodiment of the present invention;

[0032] Figure 3 It is the structure diagram of the generator of the improved CycleGAN network in the embodiment of the present invention;

[0033] Figure 4 It is the output comparison diagram after introducing the self-attention mechanism and the CAM loss in the embodiment of the present invention;

[0034] Figure 5 It is the comparison diagram of the dense conditional random field road surface crack segmentation results in the embodiment of the present invention. Detailed Embodiments

[0035] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0036] As Figure 1 shown, the road surface crack segmentation method based on laser point cloud data in this embodiment includes the following steps:

[0037] S1. Integrate the data set, train the integrated data set, and perform image-level annotation.

[0038] S2. Add a self-attention mechanism module and a CAM loss function, and use the CycleGAN generator to generate an initial activation map of the crack area.

[0039] S3. Mark the crack area on the initial activation map, refine the crack area through a dense conditional random field, and perform refined annotation on the refined crack area.

[0040] S4. Output the segmentation result to generate the final crack segmentation image.

[0041] Preferably, step S1 includes:

[0042] S1-1. Use the Crack500 data set, CFD data set, Deepcrack data set, and custom data set as the overall data set.

[0043] S1-2. For the data set generated in step S1-1, immediately divide the data set into a training set and a validation set in a ratio of 9:1. If the amount of data is sufficient, more data sets can also be divided to obtain more accurate results.

[0044] S1-3. Divide the pictures in the data set into two categories: road surface images with cracks and normal road surface images, and generate image-level weak supervision labels.

[0045] Preferably, step S2 includes:

[0046] S2-1. Add a self-attention module to the CycleGAN generator network to form a network structure including a convolutional module, a residual module, a transposed convolutional module, and a self-attention module.

[0047] S2-2. Modify the CycleGAN loss function, and add a CAM loss function on the basis of the original loss function. The CAM loss function is:

[0048]

[0049] where H and W are the height and width of the image respectively, is the value of the generated class activation map at position (i, j), CAM target is the value of the class activation map of the target image at position (i, j).

[0050] S2-3. The loss function is obtained by weighting the adversarial loss, cycle consistency loss, identity loss, and CAM loss. The loss function Loss is:

[0051] Loss = λ1L GAN (G, D Y , X, Y) + λ2L GAN (F, D X , Y, X) + λ3L cyc (G, F) + λ4L Idt (G, F) + λ5L CAM

[0052] where λ1, λ2, λ3, λ4, λ5 are weight parameters used to balance the contributions of different losses, L GAN is the adversarial loss, L cyc is the cycle consistency loss, L Idt is the identity loss.

[0053] S2-4. Based on steps S2-1 and S2-3, train the CycleGAN model and use the trained generator to generate the initial class activation map.

[0054] Preferably, step S3 includes:

[0055] S3-1. For the initial activation map obtained in S2, mark the approximate position of the crack and use it as the initial estimate.

[0056] S3-2. The dense conditional random field optimizes the following energy function:

[0057]

[0058] where the unary term ψ u (x i ) comes from the classification confidence of CAM, and the pairwise term ψ p (x i , x j ) then uses the Gaussian kernel function to refine the annotation of the crack area:

[0059]

[0060] Figure 2 are the representative pavement images with cracks and normal pavement images after arrangement, Figure 3To improve the generator structure of the CycleGAN network, Figure 4 It is a comparison diagram of CAM obtained based on the CycleGAN network and the improved CycleGAN network, Figure 5 It is a segmentation effect diagram based on the dense conditional random field. It can be seen from the figure that compared with the traditional fully supervised method, the present invention only requires image-level annotation, greatly reducing the consumption of manpower and time; at the same time, by adding a self-attention mechanism in CycleGAN, the ability to model long-range dependencies is enhanced, the generation and reconstruction effects of crack details are improved, and the overall stability is also enhanced.

[0061] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.

Claims

1. A pavement crack segmentation method based on weak supervision, characterized in that, It includes the following steps: S1. Integrate the dataset, train the integrated dataset, and perform image-level annotation; S2. Add a self-attention mechanism module and a CAM loss function, and use the CycleGAN generator to generate an initial activation map of the crack area; S3. Mark the crack area in the initial activation map, refine the crack area through a dense conditional random field, and perform refined annotation on the refined crack area; S4. Output the segmentation result to generate the final crack segmentation image.

2. The pavement crack segmentation method based on weak supervision according to claim 1, wherein The dataset integration in S1 includes: using the Crack500 dataset, CFD dataset, Deepcrack dataset, and custom dataset as the integrated dataset.

3. The pavement crack segmentation method based on weak supervision according to claim 2, wherein The training of the integrated dataset in S1 includes: randomly dividing the integrated dataset into a training set and a validation set according to a preset ratio, and training and validating the dataset.

4. The pavement crack segmentation method based on weak supervision according to claim 3, characterized in that The preset ratio is 99:1 - 9:

1.

5. The pavement crack segmentation method based on weak supervision according to claim 3, wherein, The image-level annotation in S1 includes: dividing the pictures in the trained dataset into pavement images with cracks and normal pavement images, and generating image-level weak supervision labels.

6. The pavement crack segmentation method based on weak supervision according to claim 1, characterized in that Adding a self-attention mechanism module in S2 includes: adding a self-attention module to the CycleGAN generator network to form a network structure including a convolutional module, a residual module, a deconvolutional module, and a self-attention module.

7. The pavement crack segmentation method based on weak supervision according to claim 6, characterized in that Adding a CAM loss function in S2 includes: modifying the CycleGAN loss function, and adding a CAM loss function on the basis of the original loss function. The added CAM loss function is: where H and W are the height and width of the image, respectively, is the value of the generated class activation map at position (i, j), CAM target (i, j) is the value of the class activation map of the target image at position (i, j).

8. The pavement crack segmentation method based on weak supervision according to claim 7, characterized in that, The loss function Loss includes adversarial loss, cycle consistency loss, identity loss, and CAM loss. The loss function Loss is: Loss=λ1L GAN (G,D Y ,X,Y)+λ2L GAN (F,D X ,Y,X)+λ3L cyc (G,F)+λ4L Idt (G,F)+λ5L CAM Among them, λ1, λ2, λ3, λ4, λ5 are weight parameters used to balance the contributions of different losses, and L GAN is the adversarial loss, and L cyc is the cycle-consistency loss, and L Idt is the identity loss, G is the generator, D Y is the discriminator of the target domain, X is the source-domain input image, Y is the target-domain input image, F is the inverse generator, D X is the discriminator of the source domain.

9. The pavement crack segmentation method based on weak supervision according to claim 1, wherein The dense conditional random field in S3 is refined through the following energy function: where, ψ u (x i ) is a single-variable term, ψ p (x i , x j ) is a pair-variable term, i and j are the numbers of image pixels, x i , x j are the pixel values or feature values at i and j; the single-variable term ψ u (x i ) is derived from the classification confidence of the CAM.

10. The pavement crack segmentation method based on weak supervision according to claim 9, wherein For the variable term ψ p (x i ,x j ), use the Gaussian kernel function to refine the annotation of the crack area: Among them, μ(x i ,x j ) is a similarity function used to measure whether pixels x i and x j belong to the same category. w1 and w2 are weighting parameters. p i , p j are the position coordinates of x i and x j respectively. I i , I j are the intensity information of x i and x j respectively. θ α controls the influence range of the spatial distance (position) in the similarity calculation. θ β controls the influence range of the intensity information in the similarity calculation. θ γ controls the influence range of the Gaussian kernel based only on the spatial position.