E-FCM-CNN Pavement Crack Image Recognition Method and System Integrated with Unsupervised Learning

Through the E-FCM-CNN method that integrates unsupervised learning and convolutional neural networks, the existing pavement crack recognition method has solved the problem of low recognition efficiency and low accuracy in complex environments, and efficient and automated pavement crack recognition is achieved.

CN118015341BActive Publication Date: 2025-05-30HUAZHONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410063013.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-16
Publication Date
2025-05-30
Estimated Expiration
2044-01-16

AI Technical Summary

Technical Problem

The existing pavement crack identification methods have problems with low identification efficiency and low accuracy, especially in complex urban and township road environments, manual labeling is time-consuming and has limited applicability.

Method used

The E-FCM-CNN method with fusion unsupervised learning is adopted, and the artificial marking process is eliminated through the unsupervised learning algorithm, combined with the convolutional neural network model (feature extraction network, region suggestion network and object detection network) to identify road surface cracks, and the fuzzy clustering algorithm is used to optimize the unsupervised learning process.

Benefits of technology

It significantly improves the accuracy and efficiency of road crack image recognition, reduces the time cost of manual marking, is suitable for complex road environments, and improves the automation level of the recognition system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118015341B_ABST
    Figure CN118015341B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field related to pavement quality detection, and discloses an E-FCM-CNN pavement crack image recognition method and system integrating unsupervised learning. The method includes collecting pavement images, where the pavement images at least include images of single background without cracks, cluttered background with cracks, single background with cracks, and cluttered background with cracks; preprocessing and augmenting the pavement images to obtain a data set; using an unsupervised learning algorithm to classify the data set, and generating target bounding boxes for the pavement crack regions in the images; constructing a convolutional neural network model, inputting the data set into the feature extraction network to obtain feature maps, and inputting the feature maps into the region proposal network to obtain candidate bounding boxes containing cracks; comparing the candidate bounding boxes with the target bounding boxes to obtain valid candidate boxes; and inputting the valid candidate boxes and the feature maps into the ROI pooling layer and the fully connected layer in the target detection network to obtain the image recognition result. This application can achieve fast and accurate recognition of pavement cracks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field related to pavement quality detection, and more specifically, relates to a method and system for identifying pavement crack images by fusing unsupervised learning with E-FCM-CNN. Background Art

[0002] Under the action of loads such as vehicles, cracks, potholes and other cracks also occur in traffic pavements, which have a great impact on road safety and comfort. Among them, urban and rural roads are severely damaged due to heavy traffic pressure, low maintenance level and factors such as emergency braking of vehicles.

[0003] The rectification of pavement cracks is premised on identification. The existing identification schemes are mainly manual identification or image recognition based on deep learning. Among them, manual identification is the most widely used, but it has the deficiencies of slow identification efficiency and low identification accuracy; image recognition based on supervised learning requires a lot of time and effort for manual annotation, and most of them are applicable to ideal pavements with a relatively clean background such as highways. The recognition effect is greatly affected by the pavement background and noise. Compared with highways, the pavement environment in urban and rural roads is more complex, and the presence of fallen leaves and the like will interfere with the images.

[0004] The present invention proposes a method for identifying pavement crack images by fusing unsupervised learning with Faster R-CNN. Unsupervised learning can eliminate the manual marking link in the image training process and reduce the loss of labor costs. Using Faster R-CNN can effectively identify cracks and background parts in the image, separate the object from the background, and enable the image recognition to be directly applied to scenarios with a relatively messy pavement background (such as fallen leaves by the urban road side), improving the accuracy of image recognition. Summary of the Invention

[0005] Aiming at the above defects or improvement requirements of the prior art, the present invention provides a method and system for identifying pavement crack images by fusing unsupervised learning with E-FCM-CNN, which can achieve fast and accurate identification of pavement cracks.

[0006] To achieve the above object, according to one aspect of the present invention, an E-FCM-CNN pavement crack image recognition method integrating unsupervised learning is provided, including: S1: Collect pavement images, where the pavement images include at least images of a single background without cracks, a cluttered background with cracks, a single background with cracks, and a cluttered background with cracks; S2: Preprocess and augment the pavement images to obtain a dataset; S3: Use an unsupervised learning algorithm to classify the dataset and generate target bounding boxes for the pavement crack regions in the images; S4: Construct a convolutional neural network model, where the convolutional neural network model includes a feature extraction network, a region proposal network, and an object detection network. Input the dataset into the feature extraction network to obtain a feature map, and input the feature map into the region proposal network to obtain candidate region bounding boxes containing cracks; S5: Compare the candidate region bounding boxes with the target bounding boxes to obtain valid candidate boxes; S6: Input the valid candidate boxes and the feature map into the ROI pooling layer and the fully connected layer in the object detection network in sequence to obtain an image recognition result.

[0007] Preferably, the unsupervised learning algorithm is a fuzzy clustering algorithm integrating the elbow method.

[0008] Preferably, in step S3, the specific steps of using an unsupervised learning algorithm to classify the dataset are as follows: Represent the sample data in the dataset as D-dimensional vectors, and obtain the optimal number of clusters, that is, the number of clusters C, through the elbow method; Obtain the membership matrix of each sample data in the dataset belonging to the i-th cluster, 1 ≤ i ≤ C; Construct an objective function based on the membership matrix, and perform iterative optimization with the minimization of the objective function as the goal to obtain the target number of clusters C*, where the objective function J FCM is:

[0009]

[0010] where, U = {u Qk} ∈ R N×C , d ik = ||x k - v Q || A is the distance between the k-th sample and the Q-th cluster center v Q , v Q ∈ V, V = {v 1 , v 2 , … v c} ∈ R C×D , V is a set containing C cluster centers, ‖·‖ A represents performing the A-norm processing, which is usually used as the Euclidean distance.

[0011] Preferably, the feature extraction network and the region proposal network share convolutional layers.

[0012] Preferably, in step S5, the effective candidate boxes are manually compared with the target marker boxes.

[0013] Preferably, the loss function of the convolutional neural network model includes classification loss and regression loss, where:

[0014] The classification loss is:

[0015]

[0016] The regression loss is:

[0017]

[0018] Where N cls is the size of the mini-batch (a smaller subset of data selected from the training dataset, usually a power of 2), i is the i-th target marker box, p i is the probability that the recommended region is the object of study (such as a crack), is the probability calculated corresponding to the GT. When the IoU between the i-th target marker box and the GT is greater than the preset value, the target marker box is a crack; otherwise, the target marker box is the background. N reg is the size of the feature map, L reg is the Smooth L1 Loss (used to correct the anchor box), t i is the predicted border of the recommended region, is the offset between the GT box and the anchor box.

[0019] Preferably, the preprocessing in step S2 includes image enhancement and normalization processing of the road surface image.

[0020] On the other hand, this application provides an E-FCM-CNN pavement crack image recognition system integrating unsupervised learning, including: Acquisition module: used to acquire pavement images, where the pavement images at least include images of single background without cracks, cluttered background with cracks, single background with cracks, and cluttered background with cracks; Processing module: used to preprocess and expand the pavement images to obtain a dataset; Classification module: used to classify the dataset using an unsupervised learning algorithm and generate target bounding boxes for the pavement crack areas in the images; Construction module: used to construct a convolutional neural network model, where the convolutional neural network model includes a feature extraction network, a region proposal network, and a target detection network. Input the dataset into the feature extraction network to obtain a feature map, and input the feature map into the region proposal network to obtain candidate region bounding boxes containing cracks; Comparison module: used to compare the candidate region bounding boxes with the target bounding boxes to obtain valid candidate boxes; Recognition module: used to input the valid candidate boxes and the feature map into the ROI pooling layer and the fully connected layer in the target detection network in sequence to obtain an image recognition result.

[0021] Generally speaking, compared with the prior art by the above technical solution conceived by the present invention, the E-FCM-CNN pavement crack image recognition method and system integrating unsupervised learning provided by the present invention mainly have the following beneficial effects:

[0022] 1. This application uses a convolutional neural network including a feature extraction network, a region proposal network, and a target detection network. After feature extraction, input it into the region proposal network to generate crack candidate bounding boxes, and screen and compare the bounding boxes obtained by unsupervised learning with the candidate region bounding boxes to obtain valid candidate boxes, which can effectively distinguish the target object from the environmental background, reduce the influence of the cluttered background on the recognition result, and significantly improve the image recognition accuracy.

[0023] 2. This application uses an unsupervised learning method to identify and mark images containing pavement cracks, effectively reducing the time cost of manual marking.

[0024] 3. In the unsupervised learning stage, a fuzzy clustering algorithm (E-FCM) integrating the elbow method is proposed to optimize the unsupervised learning algorithm.

[0025] 4. The convolutional neural network uses a joint training form in the feature extraction network and the region construction network part, and shortens the calculation amount and training time by sharing convolutional layers. Description of the Drawings

[0026] Figure 1 It is a step schematic diagram of the E-FCM-CNN pavement crack image recognition method integrating unsupervised learning of this application;

[0027] Figure 2It is the flowchart of the steps of the E-FCM-CNN pavement crack image recognition method integrating unsupervised learning in this application;

[0028] Figure 3 It is the working principle diagram of the convolutional neural network model in this application. Specific embodiments

[0029] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0030] The first aspect of the present invention provides an E-FCM-CNN pavement crack image recognition method integrating unsupervised learning, as Figure 1 and Figure 2 shown, the method includes the following steps S1 to S6.

[0031] S1: Collect pavement images, and the pavement images at least include images of a single background without cracks, a cluttered background with cracks, a single background with cracks, and a cluttered background with cracks.

[0032] Collect highway images through in-vehicle cameras or professional shooting tools, and the collected pavement images at least include images of a single background without cracks, a cluttered background with cracks, a single background with cracks, and a cluttered background with cracks.

[0033] S2: Preprocess and expand the pavement images to obtain a dataset.

[0034] Perform contrast enhancement and normalization processing on the pavement images to reduce the influence of irrelevant factors such as light on the subsequent process.

[0035] In a further preferred solution, histogram equalization-based image enhancement and normalization processing are adopted.

[0036] In terms of histogram equalization-based image enhancement, it mainly includes the following five steps:

[0037] (1) Calculate the gray probability density function PDF according to the image gray level;

[0038] (2) Calculate the cumulative probability distribution function CDF;

[0039] (3) Normalize the CDF to the original image gray level value range, such as [0, 255];

[0040] (4) Round the CDF to the nearest integer to obtain the gray conversion function s k= T rk ;

[0041] (5) Use the CDF as the conversion function to convert the point with grayscale r k to s k grayscale.

[0042] Perform the following operations in the feature layer data normalization process:

[0043]

[0044]

[0045]

[0046] y i ← γx i + β

[0047] where x i is the i-th pavement image as input, and y i is the processed pavement image; m is the number of pavement images input in the current training, is the mean value, is the variance, is the result after the normalization process of the pavement image features. ∈ is a small positive number used to avoid division by zero. γ and β are variables that change with the network gradient update.

[0048] Process the picture by randomly rotating the pavement image, etc., and expand the picture to 3 times the original size to increase the training samples for subsequent model training and obtain the data set.

[0049] S3: Use an unsupervised learning algorithm to classify the data set and generate target bounding boxes for the pavement crack areas in the image.

[0050] Unsupervised learning takes an unlabeled sample data set as the research object, learns the potential laws and structural information contained in the sample data to obtain the corresponding label category information, and then divides the unlabeled sample data information into different clusters according to the label category information.

[0051] The clustering algorithm can save the labeling link that consumes a lot of time and energy, realize the statistical analysis of multi-dimensional features such as grayscale, texture and gradient in the image, and complete the accurate feature extraction of the image through iterative optimization.

[0052] This application designs an unsupervised learning algorithm, using the fuzzy clustering algorithm (Elbow Fuzzy Cluster Means, abbreviated as E-FCM) that combines the elbow method: the data set X = {x 1 , x 2,…x n} ∈ R N×D , each sample data is represented as a D-dimensional vector, and the optimal number of clusters is calculated by the elbow method, denoted as C clusters. Specifically, the optimal number of clusters C for the task is selected using a graphical tool based on the calculated SSE (Sum of Squared Errors / Cluster Inertia).

[0053]

[0054] where, μ (i) is the representative point (centroid) of cluster j, x (i) is the i-th sample, w (i,j) is used to distinguish whether the sample is in the cluster. Assume that the sample x (i) is in cluster j, then w (i,j) = 1, otherwise 0.

[0055] Construct the membership matrix u k (k = 1, 2, …, N) of each sample data x Qk in the dataset X belonging to the Q-th (1 ≤ Q ≤ C) cluster, where u Qk ∈ [0, 1]). Then, the clustering result is represented as U = {u Qk} ∈ R N×C , where U is a fuzzy membership matrix. According to the criteria of fuzzy theory, for each sample in the dataset, it follows the rule shown in the formula.

[0056]

[0057] The clustering process of the FCM algorithm is the process of solving the minimum value of the objective function J FCM , and the solution method of the objective function J FCM is shown in the following formula:

[0058]

[0059] In the above formula, U = {u Qj} ∈ R N×C , d ik = ||x k - v Q || A is the distance between the k-th sample and the Q-th cluster center v Q , v Q ∈ V, V = {v 1 , v 2 , … v c} ∈ R C×D , V is the set containing C cluster centers, and ‖·‖ A represents the A-norm processing, usually used as the Euclidean distance.

[0060] The above membership matrix U and cluster center set V are iteratively optimized in multiple rounds according to the following formula until the minimum value of the objective function J FCM is obtained, that is, the final optimal clustering result is obtained.

[0061]

[0062]

[0063] Among them, is the membership degree of each sample data x k (k = 1, 2, …, N) belonging to the Qth (1 ≤ Q ≤ C) cluster after iteration, is the Qth cluster center after iteration.

[0064] S4: Construct a convolutional neural network model. The convolutional neural network model includes a feature extraction network, a region proposal network, and an object detection network. Input the dataset into the feature extraction network to obtain a feature map, and input the feature map into the region proposal network to obtain candidate region boxes containing cracks.

[0065] As Figure 3 shown, the convolutional neural network model includes a feature extraction network, a region proposal network (RPN, Region Proposal Network), and an object detection network.

[0066] In a further preferred solution, the feature extraction network and the region proposal network are partially jointly trained, and the computational amount and training time are shortened by sharing convolutional layers.

[0067] The loss function of the convolutional neural network model includes a classification loss and a regression loss, where:

[0068] The classification loss is:

[0069]

[0070] The regression loss is:

[0071]

[0072] Among them, N cls is the size of the mini-batch (a smaller subset of data selected from the training dataset, generally a power of 2), i is the ith target bounding box, p i is the probability that the recommended region is the object of study (such as a crack), To calculate the probability corresponding to GT, when the IoU between the i-th target bounding box and GT is greater than the preset value, the target bounding box is a crack; otherwise, the target bounding box is the background, N reg is the size of the feature map, L reg is the Smooth L1 Loss (used to correct the anchor box), t i is the predicted bounding box of the recommended region, is the offset between the GT box and the anchor box.

[0073] S5: Manually compare the candidate region box with the target bounding box to obtain valid candidate boxes.

[0074] S6: Input the valid candidate boxes and the feature map into the ROI pooling layer and the fully connected layer in the object detection network in sequence to obtain the image recognition result.

[0075] The E-FCM-CNN pavement crack image recognition method integrating unsupervised learning proposed in this application. Unsupervised learning can eliminate the manual labeling link in the image training process, reduce the loss of labor costs. Compared with the current computer vision object recognition network based on Faster-R-CNN, the training time is shorter, the number of images processed per second (FPS) is higher, and at the same time, a relatively high accuracy rate of highway pavement crack recognition is maintained.

[0076] The second aspect of this application provides an E-FCM-CNN pavement crack image recognition system integrating unsupervised learning, including:

[0077] Acquisition module: used to acquire pavement images, and the pavement images at least include images of a single background without cracks, a cluttered background with cracks, a single background with cracks, and a cluttered background with cracks;

[0078] Processing module: used to preprocess and expand the pavement images to obtain a data set;

[0079] Classification module: used to classify the data set using an unsupervised learning algorithm and generate target bounding boxes for the pavement crack regions in the images;

[0080] Construction module: used to construct a convolutional neural network model. The convolutional neural network model includes a feature extraction network, a region proposal network, and an object detection network. Input the data set into the feature extraction network to obtain a feature map, and input the feature map into the region proposal network to obtain candidate region boxes containing cracks;

[0081] Comparison module: used to compare the candidate region boxes with the target bounding boxes to obtain valid candidate boxes;

[0082] Recognition module: configured to input the valid candidate boxes and the feature map into the ROI pooling layer and the fully connected layer in the target detection network in sequence to obtain an image recognition result.

[0083] Those skilled in the art can easily understand that the above is only a preferred embodiment of the present invention, and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A pavement crack image recognition method integrating unsupervised learning with E-FCM-CNN, characterized in that: include: S1: collecting road surface images, wherein the road surface images include at least images with a single background without cracks, images with a cluttered background with cracks, images with a single background with cracks, and images with a cluttered background with cracks; S2: preprocessing the road surface image and expanding it to obtain a data set; S3: using an unsupervised learning algorithm to classify the data set and generate a target marking box for the road crack area in the image; S4: constructing a convolutional neural network model, wherein the convolutional neural network model includes a feature extraction network, a region proposal network and a target detection network, inputting the data set into the feature extraction network to obtain a feature map, and inputting the feature map into the region proposal network to obtain a candidate region box containing cracks; S5: Compare the candidate region frame with the target mark frame to obtain a valid candidate frame; S6: Input the valid candidate box and feature map into the ROI pooling layer and the full connection layer in the target detection network in sequence to obtain the image recognition result.

2. The method according to claim 1, characterized in that The unsupervised learning algorithm is a fuzzy clustering algorithm integrated with the elbow method.

3. The method according to claim 1, characterized in that In step S3, the specific steps of using an unsupervised learning algorithm to classify the data set are as follows: The sample data in the data set is represented as a D-dimensional vector, and the optimal number of clusters, i.e., the number of clusters C, is obtained by the elbow method; Get the membership matrix of each sample data in the data set belonging to the i-th cluster, 1≤i≤C; The objective function is constructed according to the membership matrix, and iterative optimization is performed with the objective of minimizing the objective function to obtain the target cluster number C*, where the objective function J FCM for: Among them, U={u Qk }∈R N×C , d ik =||x k -v Q || A is the kth sample and the Qth cluster center v Q The distance between Q ∈V, V={v1,v2,…v c }∈R C×D , V is a set of C cluster centers, ‖·‖ A Indicates A-norm processing, which is usually used as Euclidean distance.

4. The method according to claim 1, characterized in that: The feature extraction network and the region proposal network share convolutional layers.

5. The method according to claim 1, characterized in that In step S5, the valid candidate frame is manually compared with the target marked frame.

6. The method according to claim 1, characterized in that The loss function of the convolutional neural network model includes classification loss and regression loss, where: The classification loss is: The regression loss is: Among them, N cls is the size of the mini-batch, i is the i-th target marker box, p i is the probability that the recommended area is the research object, To calculate the probability of the corresponding GT, when the IoU between the i-th target marking box and the GT is greater than the preset value, the target marking box is a crack, otherwise the target marking box is the background, N reg is the feature map size, L reg is Smooth L1Loss, t i Predict bounding boxes for the recommended regions, is the offset between the GT box and the anchor box.

7. The method according to claim 1, characterized in that The preprocessing in step S2 includes image enhancement and normalization processing on the road surface image.

8. An E-FCM-CNN pavement crack image recognition system integrating unsupervised learning, characterized in that: include: Acquisition module: used for acquiring road surface images, wherein the road surface images include at least images with a single background without cracks, images with a cluttered background with cracks, images with a single background with cracks, and images with a cluttered background with cracks; Processing module: used for preprocessing the road surface image and expanding it to obtain a data set; Classification module: used for classifying the data set using an unsupervised learning algorithm and generating a target marking frame for the road crack area in the image; Construction module: used to construct a convolutional neural network model, which includes a feature extraction network, a region proposal network and a target detection network. The data set is input into the feature extraction network to obtain a feature map, and the feature map is input into the region proposal network to obtain a candidate region box containing cracks. Comparison module: used for comparing the candidate area frame with the target mark frame to obtain a valid candidate frame; Recognition module: used to input the valid candidate boxes and feature maps into the ROI pooling layer and the full connection layer in the target detection network in sequence to obtain image recognition results.

Citation Information

Patent Citations

  • Road surface image crack detection method based on improved Faster RCNN

    CN116109560A

  • Method for processing images using artificial intelligence and apparatus thereof

    KR1020230053272A