A road defect detection method based on deep learning
Road defect detection is carried out through deep learning methods, and all-round image data preprocessing and multi-scale feature extraction are used, combined with attention mechanism and reinforcement learning, and the problems of low efficiency and poor accuracy in traditional methods are solved, and automated identification and precise classification of road defects are achieved.
Patent Information
- Application Number
- CN202410949985.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-16
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-07-16
AI Technical Summary
Traditional road defect detection methods are inefficient, costly and have safety hazards, poor real-time performance, difficult to adapt to road defects of different scales and types, and lack in-depth distinction between defect types and damage level judgments.
Using a deep learning-based method, a comprehensive image data preprocessing, multi-scale feature extraction and object recognition algorithm is used to combine attention mechanisms and reinforcement learning to identify and classify road defects and make a judgment on the degree of damage.
It improves the accuracy and robustness of road defect detection, can automatically identify and classify different types of defects, provide accurate damage level evaluation, reduce noise impact, and enhances the detection accuracy and complex scenario understanding of the model.
Smart Images

Figure CN118918551B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and intelligent transportation systems, and particularly to a road defect detection method based on deep learning for automatically identifying and classifying different types of obstacles on roads, so as to reduce the errors and costs of manual detection and provide support for road maintenance and traffic safety management. Background Art
[0002] Computer vision is solving more and more practical applications, and the accuracy in the field of road defect detection in the field of intelligent transportation depends on efficient image preprocessing techniques. There are various types of road defects with different shapes and sizes, which makes the detection difficult. Effective preprocessing of road defect images plays a crucial role in improving the accuracy of identifying and classifying different defect types and judging the damage degree of road defects. With the continuous advancement of urbanization, the increasing road usage frequency and traffic density make road defect detection and maintenance more important. The timely discovery and repair of road defects are of great significance for ensuring traffic safety, extending the service life of roads, and optimizing traffic management. However, traditional road defect detection methods mainly rely on manual inspections or formulaic image detections, which are inefficient, costly, and have certain safety hazards. With the rapid development of computer vision and artificial intelligence technologies, deep learning-based road defect detection technologies can automatically detect and classify different types of road defects from road video streams, providing key information support for road maintenance and traffic management.
[0003] Traditional road defect detection methods also have problems of poor real-time performance and recognition accuracy. Usually, traditional road defect detection methods use a single preprocessing operation for preprocessing different defect images, and there is still noise in the processed images, which affects subsequent recognition. At the same time, the image defect features are affected by acquisition, and traditional feature extraction methods are single, making it difficult to adapt to different scales and different types of road defects. Moreover, traditional defect classification methods are simple and lack in-depth differentiation of defect types. In addition, traditional methods usually do not classify the detected defects and judge the damage level, and cannot provide more effective intelligent information support for users. Summary of the Invention
[0004] The object of the present invention is to provide a road defect detection method based on deep learning, which uses omnidirectional image data preprocessing and an efficient target recognition algorithm to identify, classify different types of road defects, judge the damage degree, and synchronously upload it to the server.
[0005] The road defect detection method based on deep learning includes the following steps:
[0006] S1: According to the original image data, construct an image preprocessor, and perform preprocessing operations of multi-frame smoothing and noise reduction based on time series;
[0007] S2: Based on the improved multi-scale feature extraction algorithm, extract the eigenvalues of the processed image to obtain different features of low, medium, and high levels of current road surface cracks, potholes and other defects;
[0008] S3: Construct an object detection algorithm to identify road defects in the feature map;
[0009] S4: Construct a road defect classifier to classify the detected objects according to pothole defects and crack defects;
[0010] S5: According to different road defect categories, calculate their damage degrees respectively and make damage level divisions;
[0011] S6: Transmit the identified road defect information back to the vehicle display for warning, and upload the relevant information to the cloud for analysis and statistics;
[0012] Furthermore, the step S1 specifically includes:
[0013] S101: For each video stream to be processed, decompose the video into individual frames;
[0014] S102: Intercept k consecutive images, select the middle frame as the key frame, filter out the noise frames according to the inter-frame similarity, and perform pixel weighted replacement on the noise frames;
[0015] S103: Perform a smoothing operation on the received image using an improved adaptive weighted average filtering method;
[0016] S104: Arrange the smoothed images in chronological order and perform temporal noise reduction processing based on the jump structure;
[0017] Furthermore, the step S102 specifically includes:
[0018] First, calculate the similarity between each frame image and the key frame image. Among them, the similarity calculation formula is:
[0019]
[0020] Among them, i, respectively represent the current frame and the middle frame image of the video, μ i , respectively represent the means of the pixels of the two frame images, respectively represent the variances of the pixels of the two frame images, cov(·) is the covariance function, and σ i is the standard deviation of the Gaussian function of the i-th frame image, and C1, C2 are stability constants;
[0021] Then, set the similarity threshold τ, and filter out the noise frames whose similarity to the key frame is less than the threshold τ;
[0022] Furthermore, calculate the pixel weights of the noise frame images. Among them, the pixel weight calculation formula is:
[0023]
[0024] Among them, Rmse(·) is the root mean square error function, and R(·) is the range function;
[0025] Finally, perform weighted replacement on the noise frame pixels according to the weights. Among them, the pixel weighted replacement formula is:
[0026]
[0027] Among them, I w (x, y) is the weighted pixel value, I i (x, y), I j (x, y) represent the pixel values of the i-th noise image and the j-th normal frame image closest to it at the coordinate (x, y), w i is the weight of the i-th image, and i is the current time point;
[0028] Furthermore, the step S103 specifically includes:
[0029] (1) For the given filtering window S xy , calculate the minimum gray value Z xy , the maximum gray value Z min , and the average gray value Z max of all pixels of each pixel point (x, y) in S mean , and set the noise judgment threshold T;
[0030] (2) Judge whether the pixel mean value of the filtering window is a noise point. If T min <(Z min +Z max ) / Z mean <T max holds, then the mean value Z mean is a normal point, otherwise expand the window to continue searching for the normal mean value;
[0031] (3) Judge whether Z min <I w (x, y)<Z max is satisfied. If the condition is satisfied, output the pixel value of this point, otherwise calculate the difference value between each pixel of the input image and the pixels in its neighborhood using the Euclidean distance;
[0032] (4) Calculate the pixel weights according to the difference value. The pixel weight calculation formula is:
[0033]
[0034] Among them, W(x, y) represents the weight of the pixel point (x, y), indicating the difference value between the pixel at this point and the central pixel (c x , c y );
[0035] (5) Multiply the weight by the corresponding pixel value to calculate the filtered pixel value. The calculation method of the filtered pixel value is:
[0036]
[0037] Among them, FP(x, y) represents the filtered pixel value, and I w (x, y) represents the pixel value in the original image;
[0038] (6) Repeat the steps until all pixel points are traversed;
[0039] (7) When there are still noise pixels after the window reaches the maximum size, use Gaussian filtering to replace the noise pixels;
[0040] Furthermore, the specific steps of step S104 include:
[0041] First, input k consecutive images into the encoder, and perform downsampling operations through three convolution kernels with different sizes in sequence to extract feature information, obtaining the encoded feature Z E Expressed as:
[0042]
[0043] Then, divide the encoded feature Z E into groups of n each in chronological order, with a total of k - n + 1 groups, which are used as the input of the bidirectional recurrent neural network that combines residuals in the skip structure to learn the forward and backward time series correlations in the time series;
[0044] Finally, input the k - n + 1 groups of time series into the decoder for upsampling operations, and perform matching splicing with the corresponding positions in the encoder to obtain the final denoised image;
[0045] Among them, the steps of learning by the bidirectional recurrent neural network that combines residuals are:
[0046] First, calculate the reset gate and the propagation gate
[0047]
[0048]
[0049] Among them, are the reset gates in the forward and reverse directions respectively, which determine to what extent the hidden state information at the previous moment is reset, thus allowing the network to forget irrelevant information; are the propagation gates in the forward and reverse directions respectively, which control to what extent the new candidate hidden state information is updated into the current hidden state; chx(·) and shx(·) are the hyperbolic cosine activation function and hyperbolic sine activation function used for gating degree scaling in the reset gate and propagation gate respectively; W o , U o , b o , W s , U s , b s are the gating parameters for the forward and reverse propagation of the model. W o and W s represent the weight matrices between the reset gate and the propagation gate layers respectively, U o and U s represent the time step matrices of the reset gate and the propagation gate respectively, and b o and b s represent the bias matrices of the reset gate and the propagation gate respectively;
[0050] Then, calculate the candidate hidden states in the forward and reverse directions respectively
[0051]
[0052] Among them, are the candidate hidden states in the forward and reverse directions respectively, which contain the information of the current input and the hidden state at the previous moment; W f , W b , U f , U b , b f , b b are the forward and reverse gating parameters of the model. W f , W b are the weight matrices for forward and reverse learning, U f , U b are the time step matrices for forward and reverse learning, and b f , b b are the bias matrices for forward and reverse learning; H is the H-Swith(·) activation function, and the formula of the H-Swith activation function is as follows:
[0053]
[0054] ReLU6 = min(ReLU, 6)
[0055] Finally, based on the propagation gates and candidate hidden states, the forward and backward hidden states at the current moment are concatenated into the final hidden state h t , and a residual connection is made with the input encoding feature Z E to obtain the final temporal feature Z R :
[0056]
[0057] Z R = Z E + h t
[0058] where h t is the final hidden state, which is the concatenation of the forward and backward hidden states; Z R is the final temporal feature, which is the residual connection between the input encoding feature Z E and the hidden state h t ; [:] is the concatenation operation, and ⊙ is the element-wise multiplication;
[0059] Furthermore, the specific steps of step S2 include:
[0060] S201: Apply the dilated convolution operation with improved multi-dilation rate fusion on the input image to obtain the preliminary feature maps C1 - C n of the input image and the corresponding sub-feature maps D1 - D n with different depths;
[0061] S202: Introduce a cross-attention layer between each layer of the feature maps C1 - C n to learn the correlation between its sub-feature maps through the designed AN network;
[0062] S203: Perform multi-scale fusion on the respective different sub-feature maps D n in the feature maps C1 - C * to obtain the final feature pyramid P;
[0063] S204: Apply the multi-scale G-LBP operator on each layer of the obtained feature pyramid P to perform feature extraction to obtain the feature maps;
[0064] Furthermore, the specific steps of step S201 include:
[0065] First, use n convolutional layers of different sizes to extract the multi-scale features of the image to obtain n different levels of preliminary feature maps C1 - C n ;
[0066] Then, between the feature maps C1 - C nPerform convolution using the dilated convolution operation with multi-dilation rate fusion respectively to obtain sub-feature maps D1 - D with different depths corresponding to them n ;
[0067] Among them, the dilated convolution operation with multi-dilation rate fusion is defined as follows:
[0068] The dilation rate Rate of the i-th layer convolution i represents the distance dilation rate between adjacent elements in the convolution kernel of this layer, and Rate i is defined as:
[0069] Rate i = f(i)
[0070] Among them, f(i) is a hierarchical function, and f(i) = 2 i-1 ;
[0071] The actual effective convolution size V of the i-th layer convolution kernel with size k i can be expressed as:
[0072] V i = (k - 1) * Rate i + 1
[0073] The operation during the i-th layer image convolution can be expressed as:
[0074]
[0075] Among them, I(i, j) and K(i, j) are the pixel values at the (i, j) coordinates on the image and the convolution kernel respectively;
[0076] Furthermore, the step S202 specifically includes:
[0077] First, calculate the attention weights of different feature maps:
[0078] W = softmax(AN(D i , D j ))
[0079] Among them, D i , D j represent different sub-feature maps;
[0080] Then, apply the weights to the class feature map to obtain the fused feature map D:
[0081] D = W * D i
[0082] Finally, merge the fused feature map D with the original class sub-feature map D i to obtain a new feature map D * :
[0083] D * = Merge(D, D i )
[0084] where D * represents the final sub - feature map, and Merge(·) represents the splicing operation;
[0085] Furthermore, step S204 specifically includes:
[0086] (1) Divide the image into n grids of m×m, and each grid includes different radii r and different sampling points p;
[0087] (2) Calculate the eigenvalue M of the MC - LBP operator when r is r1 and p is p1 in each local grid. The formula for calculating the eigenvalue M is as follows:
[0088]
[0089] where M p,r (x, y) represents the eigenvalue of the MC - LBP operator, p is the number of sampling points, r is the radius, (x, y) are the coordinates of the center point, a i is the angle information of the i - th sampling point, and s(·) is the threshold function. The formula for the threshold function is as follows:
[0090]
[0091] where θ i is the angle of the i - th sampling point, and g(x, y), g(x + rcos(θ i ), y + rsin(θ i )) are the gray values of the central pixel and the sampling point respectively;
[0092] (3) Update the overall eigenvalue M of the i - th local grid according to each eigenvalue M p,r (x, y) in the local grid. The formula for calculating the eigenvalue M i is as follows: i The formula for calculating the eigenvalue M
[0093]
[0094] where p represents the number of sampling points;
[0095] (4) Calculate the eigenvalue T of the MC - LBP operator when r is r2 and p is p2 in the whole - image range according to the updated local - grid eigenvalue M i . The formula for calculating the eigenvalue T m is as follows: m The formula for calculating the eigenvalue T
[0096]
[0097] T m = softmax(M * p,r (x, y))
[0098] where M * p,r (x, y) is the eigenvalue of the global MC-LBP operator, g c is the gray value of the central pixel, and M i is the gray value of the local grid where the sampling point is located; T m is the normalized eigenvalue of the global MC-LBP operator;
[0099] (5) Convolve the Gabor filter with each pixel position of the image to obtain the response value at that position. The specific method for the Gabor filter to extract features is as follows:
[0100] T g = softmax(conv(G, I))
[0101] where T g represents the normalized Gabor feature, G represents the Gabor filter, and I represents the input image;
[0102] (6) Initialize the Q-table and select the initial attention weight ξ;
[0103] (7) According to the current attention weight ξ, explore how to increase or decrease the weight ξ with a certain probability;
[0104] (8) Define the reward signal Q to guide the model to learn how to select feature weights. The calculation formula for defining the reward signal Q is:
[0105]
[0106] where Q(s, a) is the Q value of performing action a in state s, α is the learning rate, r is the immediate reward obtained after performing action a, γ is the discount factor, representing the importance of future rewards, s' is the next state transferred to after performing action a, and Q(s, a) is updated according to the current reward signal r and the expected future reward maxQ(s, a) to guide the model to select the best action in a given state;
[0107] (9) Weightedly fuse the features T m and T g extracted by the MC-LBP operator and the Gabor filter to obtain the G-LBP feature T * :
[0108] T * = ξ * T m+(1 - ξ)*T g
[0109] Wherein, T * represents the G - LBP feature, and T m , T g are the normalized MC - LBP and Gabor features respectively;
[0110] Furthermore, the step S3 specifically includes:
[0111] S301: Randomly perform rotation and cropping operations on the input feature map and adjust it to a feature map of the same size to complete data augmentation;
[0112] S302: Construct a deep - learning - driven object - detection module to accurately identify and locate road surface defects on the feature map;
[0113] S303: Further filter overlapping detection boxes through the non - maximum suppression algorithm to ensure that each defect is detected only once;
[0114] Furthermore, the step S302 specifically includes:
[0115] First, detect objects of different scales on the obtained feature map through a convolutional attention module;
[0116] (1) Calculate the channel attention weight CA i and the spatial attention weight SA i between different time frames for each scale of the feature map;
[0117] (2) For different feature maps, use the attention weights to adjust their original feature maps, and the adjustment formula is as follows;
[0118]
[0119] In the formula, F i represent the adjusted feature map and the original feature map respectively; CA i , SA i represent the channel and spatial attention weights respectively;
[0120] (3) Add the adjusted feature maps, and the addition formula is as follows;
[0121]
[0122] In the formula, F * is the channel of the final feature map; ε1 and ε2 are weights for adjusting the fusion of different feature maps and satisfy:
[0123] ε1 + ε2 ≤ 1
[0124] (4) Set multiple prior boxes with different aspect ratios as the basis for subsequent predicted bounding boxes;
[0125] (5) Flatten the regional features corresponding to each prior box into a vector, and predict the offset of the bounding box for the generated prior boxes in the fully connected layer of regression prediction;
[0126] Among them, the default prior box generated at position (i, j) on the feature map can be expressed as:
[0127] B i,j =(x, y, w, h)
[0128] Among them, B i,j has (x, y) as the center coordinates, and w, h as the length and width of the prior box;
[0129] The offset prediction l k can be expressed as:
[0130] l k =(Δx, Δy, Δw, Δh)
[0131] Among them, Δx, Δy represent the offsets of the horizontal and vertical coordinates of the center coordinates of the prior box, and Δw, Δh are the offsets of the scaling ratios of the width and height of the prior box respectively;
[0132] The actual offsets of the center coordinates, width, and height can be expressed as:
[0133] t x =(x - x a ) / w a
[0134] t y =(y - y a ) / h a
[0135] t w =ln(w / w a )
[0136] t h =ln(h / h a )
[0137] Among them, t x , t y represent the actual offsets of the center coordinates, and t w , t h represent the actual offsets of the length and width of the prior box. x a , y a represent the actual center coordinates respectively, and w a , h a represent the length and width of the actual prior box respectively;
[0138] The predicted offsets of the center coordinates and the width and height can be expressed as:
[0139] t' x =(x - Δx) / Δw
[0140] t' y =(y - Δy) / Δh
[0141] t' w =ln(w / Δw)
[0142] t' h =ln(h / Δh)
[0143] where t' x , t' y represents the predicted offset of the center coordinate, and t' w , t' h is the predicted offset of the length and width of the prior box;
[0144] The position error is defined as follows:
[0145]
[0146] where N is the number of samples, and HL(·) is the Loss function, which is defined as follows:
[0147]
[0148] where δ is the threshold parameter of the loss function;
[0149] Furthermore, the step S4 specifically includes:
[0150] S401: Construct a road defect classifier, classify the road defects in the detection box output by the object detection module into pothole defects and crack defects, and output the classification result R1;
[0151] S402: Input the classified result R1 into the neural network and perform Dropout operation to prevent overfitting, and output the classification result R2;
[0152] S403: Input the two classification results R1 and R2 into the voter, set the weights and use the soft voting mechanism to obtain the final road defect classification result R * ;
[0153] Furthermore, the step S401 specifically includes:
[0154] (1) Calculate the distance between each sample point and other sample points and their corresponding weights. The weight calculation formula is as follows:
[0155]
[0156] Among them, k is the number of K nearest neighbors of sample i, S(i) is the set of k nearest neighbor points of sample i, and d i,j is the distance between sample points i and j;
[0157] (2) Calculate the local density of each sample point according to its weight and the number of neighbors. The local density of this point represents the degree of density of this point in the feature space, and calculate the local density of weighted neighbors The formula is as follows:
[0158]
[0159] (3) Calculate the relative distance η between each pair of sample points and other samples with a density higher than it i , and the relative distance calculation formula is as follows:
[0160]
[0161] (4) Calculate the sample decision value θ i , and select the set of cluster centers, assign cluster labels. The sample decision value calculation formula is as follows:
[0162]
[0163] (5) Select the k samples with the highest decision value as the initial cluster centers;
[0164] (6) Calculate the similarity between samples and construct a similarity matrix. The similarity between samples is defined as follows:
[0165]
[0166] Among them, sim i,j and w i,j are the similarity and similarity weight between sample points i and j respectively;
[0167] (7) Traverse the samples with assigned cluster labels, find the unassigned sample with the greatest similarity to it, and assign the unassigned sample to the cluster where the assigned sample is located. Repeat this operation until all samples have been assigned, or the similarity between an assigned sample and an unassigned sample is 0;
[0168] (8) If there are still unassigned samples, assign the remaining samples to the cluster where the sample with a higher density and the closest distance to it is located;
[0169] (9) Take the sample with the highest sample decision value θ i in the cluster as the new cluster center, repeat the above steps until the cluster center no longer changes, and output the classification result R1;
[0170] Furthermore, step S5 specifically includes:
[0171] S501: Calculate the damage degrees of pothole defects and crack defects respectively according to the classification result R * and the number of pixels quantified by the pothole defect target, p
[0172] S502: Classify the damage levels of pothole defects and crack defects;
[0173] Among them, the calculation formulas for the damage degrees of potholes and cracks are as follows:
[0174]
[0175] where p a is the number of pixels quantified by the pothole defect target, p s is the number of pixels quantified by the entire road pothole defect image, and a is the number of pixels within the minimum circumscribed rectangle of the crack; the damage degree D p value of the pothole is used to measure the severity of the pothole image damage, and D n is used to measure the severity of the crack image damage;
[0176] The damage level classification of potholes and cracks is as follows:
[0177] Damage level of road pothole defects: D p ≤C1 is a light pothole; C1≤D p ≤C2 is a medium pothole; C2≤D p is a heavy pothole;
[0178] Damage level of road crack defects: D a ≤C3 is a light crack defect; C3≤D a is a heavy crack defect;
[0179] Furthermore, step S6 specifically includes:
[0180] S601: Use the classifier to identify the obstacle types and locations of pothole or crack defects on the road, and transmit them back to the vehicle's display in a flashing manner through different graphics and colors;
[0181] S602: Upload the road defect information and the actions taken by the driver to the cloud server.
[0182] The beneficial effects of the present invention include:
[0183] (1) The image preprocessing of the present invention is based on the idea of video stream data processing. First, each time the collected video data is intercepted with consecutive image frames as the input, which not only reduces the computational amount of image preprocessing, preserves the sequential relationship and temporal information between image frames, but also uses the image data information of the frames before and after a certain moment as a reference, which can better eliminate various noises on the collected images and provide a solid foundation for the subsequent road defect detection work;
[0184] (2) The present invention proposes a new noise judgment method: First, judge whether an entire frame of image is a noise image, and then judge whether a certain point in the image is a noise point. The frame - to - frame weighted replacement and improved weighted average filtering are respectively used for noise reduction processing, which can effectively reduce the influence of noise on road defect detection at different levels and improve the accuracy and robustness of the detection algorithm;
[0185] (3) The present invention proposes a learning method for continuous grouping of video streams. While using a carefully designed activation function for non - linear mapping, it uses the temporal correlation before and after video stream grouping and combines the residual idea for temporal noise reduction, which can better remove the noise of images and improve the accuracy of image recognition;
[0186] (4) The present invention makes full use of the advantages of the attention mechanism, designs a variety of different attentions and applies them to different steps in the process of road defect detection respectively. It can better complete the following tasks from different angles in this work: learning of feature maps at different levels, fusion of different reverse feature values, and recognition of targets at different scales. Replacing manual work with the attention mechanism of machines imitating humans can perform hierarchical learning in an automated manner, ensuring the comprehensiveness and accuracy of road defect detection.
[0187] (5) In the process of detecting road defects based on deep learning, the present invention applies the multi - scale circular operator MC - LBP and Gabor filter to extract features on each layer of the obtained feature pyramid. It designs an attention mechanism based on reinforcement learning to fuse the features extracted by MC - LBP and Gabor, which can effectively capture the local texture information and direction information in the road image to provide a more comprehensive feature representation, dynamically adjust the weights of features according to the importance of defect features, enabling the model to allocate different attentions to features at different scales and different levels, and can better obtain the comprehensive features of road defects, further improving the accuracy of image recognition and classification;
[0188] (6) In the process of constructing the road defect classifier, the present invention can effectively capture the subtle differences between pothole defects and crack defects by calculating the distances and weights between sample points, as well as local density and relative distance, providing more accurate classification results. In addition, by inputting the classified results into a neural network and using a soft voting mechanism to fuse the two classification results, the decisions of multiple classifiers can be comprehensively considered, further improving the robustness and accuracy of the model;
[0189] (7) In the method for road defect detection based on deep learning, the present invention designs a damage degree calculation method, namely pothole damage degree and crack damage degree, which can accurately quantify the damage of the road surface. In addition, this method innovatively classifies the damage levels of pothole defects and crack defects, making the detection results more refined, providing auxiliary decision-making support for the subsequent maintenance of the road, enabling the staff to repair the road at the lowest cost, and having significant technical advantages and application value in the field of road defect detection.
[0190] Other advantages, objectives and features of the present invention will be elaborated in detail in the subsequent specification. Through in-depth research on the following text, those skilled in the art will be able to clearly recognize these advantages and features and obtain valuable teachings from the practice of the present invention. The objectives and other advantages of the present invention can be realized and embodied in the following specification and the previously mentioned claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0191] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings, where:
[0192] Figure 1 is a schematic flow chart of a method for road defect detection based on deep learning according to the present invention;
[0193] Figure 2 is an example diagram of filtering a normal image by the improved adaptive weighted average filter according to the present invention;
[0194] Figure 3 is an example diagram of filtering a noisy image by the improved adaptive weighted average filter according to the present invention;
[0195] Figure 4 is a schematic diagram of time-series noise reduction of a road defect image according to the present invention;
[0196] Figure 5 is a schematic diagram of obtaining road defect features according to the present invention;
[0197] Figure 6 is a schematic diagram of MC-LBP feature extraction according to the present invention;
[0198] Figure 7Schematic diagram of the working process of the road defect classifier of the present invention;
[0199] Figure 8 Example diagram for grading road crack defects of the present invention;
[0200] Figure 9 Example diagram for grading road pothole defects of the present invention. Specific implementation manners
[0201] The preferred implementation manners of the present invention will be described in detail below. It should be clear that the preferred embodiments are only for illustrating the present invention, rather than for limiting the protection scope of the present invention.
[0202] Considering the defects of traditional road defect detection methods, the present invention makes full use of image preprocessing technology and the attention mechanism of deep learning, and with the help of omnidirectional image data preprocessing and efficient target recognition algorithms, identifies and classifies different types of road defects, and judges the damage degree and synchronously uploads it to the server. Among them, image preprocessing is a prerequisite for accurate target recognition. It can calculate the similarity between frames, screen out noise frames, perform pixel weighted replacement on them, and perform smoothing processing on the overall video stream using adaptive weighted average filtering after excluding noise frames; the encoder learns the forward and backward correlations of the video stream time and performs temporal noise reduction with structural jumps; the omnidirectional noise reduction mechanism can greatly reduce the negative impact of different noises such as salt and pepper noise and Gaussian noise generated during the acquisition of different defects on subsequent defect recognition.
[0203] The attention mechanism is an important component in deep learning. It can improve the expression ability of the model by combining the multi-scale feature extraction attention mechanism with reinforcement learning, learn how to select feature weights, so as to achieve more effective feature extraction while ignoring irrelevant background information; through the cross-attention layer in the feature pyramid network, the high-level features can pay attention to the detailed information in the low-level features, while the low-level features can pay attention to the context information in the high-level features, further enhancing the expression ability of the model. Through the object detection attention mechanism, calculate the channel attention weight and spatial attention weight between different time frames for the feature map of each scale, so as to better detect objects of different scales; different attention mechanisms used in different periods have unique advantages in road defect detection; using different attention mechanisms can adaptively adjust the weights of the feature map, enhance the sensitivity of the model to defect features, reduce the interference of background noise, not only improve the detection accuracy and robustness of the model, but also enhance the model's understanding ability of complex scenes.
[0204] As Figure 1 shown, a road defect detection method based on deep learning of the present invention includes the following steps:
[0205] Step S1: Based on the original image data, construct an image pre-processor to perform pre-processing operations of multi-frame smoothing and noise reduction based on time series;
[0206] Step S2: Based on the improved multi-scale feature extraction algorithm, extract eigenvalues from the processed image to obtain different features of low, medium, and high levels of current road surface cracks, potholes and other defects;
[0207] Step S3: Construct an object detection algorithm to perform object recognition of road defects in the feature map;
[0208] Step S4: Construct a road defect classifier to classify the detected objects into pothole defects and crack defects;
[0209] Step S5: According to different road defect categories, calculate their damage degrees respectively and make damage level divisions;
[0210] Step S6: Transmit the recognized road defect information back to the vehicle display for early warning, and upload the relevant information to the cloud for analysis and statistics.
[0211] It should be noted that any process or method description in the flowchart of the present invention or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process, and the scope of the preferred embodiments of the present invention includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present invention belong.
[0212] As Figures 2 to 9 shown below, the specific steps of the above method will be further elaborated through a specific embodiment.
[0213] The said Step S1 includes the following sub-steps:
[0214] Step S101: Collect video image data, perform real-time collection using an in-vehicle camera with 25fps, and intercept each second of the video and decompose it into individual frames;
[0215] For each video frame image data collected in Step S101, the data information included in the video frame image is: {Video frame: D frame , Pixel information: D color , Image resolution: D resolution , Frame rate: D rate , Compression coding format: D encoding}, and the width D weight of the collected image is 256 pixels, and the height D height=256 pixels;
[0216] Step S102, intercepting 25 consecutive frames of images within one second, and selecting the middle frame of the video stream: the 13th frame is used as the key frame, filtering the noise frames according to the similarity between the frames, and performing pixel weighted replacement on the noise frames;
[0217] In step S102, the specific process of performing pixel weighted replacement on the noise frame is as follows:
[0218] (1) Calculate the similarity between each of the 24 non-key frame images and the key frame. The similarity calculation formula is:
[0219]
[0220] Among them, i, Represent the current frame and the middle frame of the video, and are 1, 2, ..., 12, 14, ..., 25 and 13 respectively; μ i , Represent the mean of two frames of images, Represent the variance of two frames of images, cov(·) is the covariance function, σ i is the standard deviation of the Gaussian function of the i-th frame image, C1 and C2 are stability constants, which are 0.05 respectively; the similarities between the 24 non-key frame images and the key frame are as follows: the similarities are as follows: [0.79, 0.93, 0.91, 0.90, 0.88, 0.92, 0.88, 0.97, 0.98, 0.87, 0.95, 0.90, 0.91, 0.98, 0.81, 0.78, 0.78, 0.96, 0.95, 0.97, 0.99, 0.95, 0.88, 0.95];
[0221] (2) The image frames whose similarity is less than the threshold value τ = 0.8 are screened as noise frames, among which the image frames 1, 17, and 18 whose similarity with the key frame is less than the threshold value 0.8 are judged as noise frames;
[0222] (3) Calculate the pixel weight of the filtered noise frame. The pixel weight calculation formula is:
[0223]
[0224] Among them, Rmse(·) is the root mean square error function, R(·) is the range function, and the noise frame weights are: [0.40, 0.38, 0.38];
[0225] (4) Perform pixel weighted replacement on the noise frame according to the weights. The calculation formulas for pixel weighted replacement are:
[0226]
[0227] Among them, I w (x, y) is the weighted pixel value, I i (x, y), I j (x, y) represents the pixel value at the coordinate (x, y) of the i-th frame of noise image and the j-th normal frame image closest to it, w i is the weight of the i-th frame image; The three frames of noise images are respectively weighted and replaced with the 2nd, 16th, and 19th normal frame images closest to them;
[0228] Step S103: Smooth the image after receiving noise replacement using improved adaptive weighted average filtering;
[0229] In step S103, the specific steps of smoothing the image using the improved filtering are as follows:
[0230] In this embodiment, the smoothing process of the improved adaptive weighted average filtering is as Figure 2 、 3 shown;
[0231] (1) For each pixel point (x, y) in the given 3×3 filtering window S xy find the minimum gray value Z min , the maximum gray value Z max and the average gray value Z mean of all pixels, and set the noise judgment threshold T∈[1.7, 2.1];
[0232] (2) Judge whether the mean value of the filtering window is a noise point, that is, if T min <(Z min +Z max ) / Z mean <T max holds, then the mean value Z mean is a normal point and use this mean value, otherwise expand the window side length by 2 and continue to find the normal mean value;
[0233] (3) Judge whether Z min <I(x, y)<Z max is satisfied. If the condition is satisfied, output the pixel value of this point, otherwise calculate the difference value between each pixel of the input image and its neighboring pixels using the Euclidean distance;
[0234] (4) Calculate the pixel weight according to the difference value. The calculation formula of the pixel weight is:
[0235]
[0236] Among them, W(x, y) represents the weight of the pixel point (x, y), represents the difference value between this pixel and the pixel at the center point (128, 128);
[0237] (5) Multiply the weights by the corresponding pixel values to calculate the weighted average pixel value. The calculation method of the weighted average pixel value is as follows:
[0238]
[0239] where FP(x, y) represents the pixel value after filtering, and I w (x, y) represents the pixel value in the original image;
[0240] (6) Repeat the steps until all pixel points are traversed;
[0241] (7) When there are still noisy pixels after the window reaches the maximum size, use Gaussian filtering for replacement;
[0242] In this embodiment, the pixel information in the 3×3 filtering window: I(1, 1) = 100, I(1, 2) = 120, …, I(3, 3) = 150, and the difference values between each point and the central pixel: diff(1, 1) = 10, diff(1, 2) = 5, diff(1, 3) = 15, …, diff(3, 3) = 0. Calculate the weights of each pixel point according to the calculation formula of pixel weights. The standard deviation σ of the Gaussian function is 5, and the pixel weights are calculated as follows:
[0243]
[0244]
[0245] Calculate the weighted average pixel value of each pixel point according to the calculation formula of the weighted average pixel value:
[0246] W(1, 1) ≈ 107.2
[0247] W(1, 2) ≈ 110.7
[0248] W(1, 3) ≈ 102.2
[0249] …
[0250] W(3, 3) ≈ 150
[0251] Step S104: Arrange the smoothed images in chronological order and perform temporal noise reduction processing based on the jump structure;
[0252] In step S104, the specific steps of the temporal noise reduction processing based on the jump structure are as follows:
[0253] In this embodiment, the model of the temporal noise reduction based on the jump structure is as Figure 4 shown;
[0254] (1) Feed seven consecutive frames of images: I1, I2, I3, …, I7 into the encoder in sequence, and perform convolution operations on each frame of image using three convolutional kernels in sequence to obtain encoded features:
[0255]
[0256] In this embodiment, the number and size of the convolutional kernels in the encoder are:
[0257] The first convolutional kernel: N1 = 32, size 7×7;
[0258] The second convolutional kernel: N2 = 64, size 5×5;
[0259] The third convolutional kernel: N3 = 128, size 3×3;
[0260] (2) Divide the encoded features Z E into three groups with every five images in chronological order as a group, and input them into the bidirectional recurrent neural network to learn the forward and backward time series correlations in the time series to obtain the time series features Z R ;
[0261] (3) Input the image time series features Z R into the decoder for upsampling, and perform matching splicing with the corresponding positions in the encoder to obtain the final denoised images I'1, I'2, I'3, …, I'7;
[0262] Among them, the steps of learning with the bidirectional recurrent neural network combined with residuals are:
[0263] (1) Calculate the reset gates and propagation gates
[0264]
[0265]
[0266] at time t in the forward and backward directions respectively. Among them, are the reset gates in the forward and backward directions respectively, which determine to what extent the hidden state information of the previous moment is reset, so as to allow the network to forget irrelevant information; o 、U o 、b o 、W s 、U s 、b sis the gating parameter for the forward and backward propagation of the model, W o and W s represent the weight matrices between the reset gate and the propagation gate respectively, U o and U s represent the time-step matrices of the reset gate and the propagation gate respectively, b o and b s represent the bias matrices of the reset gate and the propagation gate respectively; the elements of the matrices W o , U o , W s , U s all follow a uniform distribution U(-0.1, 0.1); the elements of the matrices b o , b s take 0 and 0.01 respectively;
[0267] (2) Calculate the candidate hidden states in the forward and backward directions respectively
[0268]
[0269] Among them, are the candidate hidden states in the forward and backward directions respectively, containing information about the current input and the hidden state at the previous moment; W f , W b , U f , U b , b f , b b are the forward and backward gating parameters of the model, W f , W b are the weight matrices for forward and backward learning, U f , U b are the time-step matrices for forward and backward learning, b f , b b are the bias matrices for forward and backward learning; the elements of the matrices b f , b b take 0.01; H is the H-Swith(·) activation function, and the formula of the H-Swith activation function is as follows:
[0270]
[0271] ReLU6 = min(ReLU, 6)
[0272] (3) According to the propagation gate and the candidate hidden states, concatenate the forward and backward hidden states at the current moment into the final hidden state h t , and perform a residual connection with the input encoding feature Z E to obtain the final temporal feature Z R :
[0273]
[0274] Z R = Z E + h t
[0275] where h t is the final hidden state, which is the concatenation of the forward and backward hidden states; Z R is the final temporal feature, which is the residual connection of the input encoding feature Z E and the hidden state h t ; [:] is the concatenation operation, and ⊙ is the element-wise multiplication; the model learning rate lr = 0.0001, the number of iterations epoch = 100, and the batch size batch_size = 64;
[0276] The specific steps of step S2 include the following steps:
[0277] Step S201: By applying dilated convolutions with improved multi-dilation rate fusion to the input image, three preliminary feature maps C1 - C3 of the input image and their three corresponding sub-feature maps D1 - D3 with different depths are obtained. The specific steps of using the dilated convolutions with improved multi-dilation rate fusion operation are as follows:
[0278] In this embodiment, the operation process of the dilated convolutions with improved multi-dilation rate fusion is as Figure 5 shown;
[0279] First, use 3 convolutional layers with different sizes to extract multi-scale features of the image, obtaining 3 preliminary feature maps C1 - C3 at different levels;
[0280] In this embodiment, the number and size of the convolutional kernels are:
[0281] The first convolutional kernel: N1 = 64, size 3×3, stride 1;
[0282] The second convolutional kernel: N2 = 128, size 5×5, stride 1;
[0283] The third convolutional kernel: N3 = 256, size 7×7, stride 1;
[0284] After connecting the ReLU(·) activation function and performing batch normalization after the convolutional layer, the preliminary feature maps C1 - C3 are obtained;
[0285] where the normalization formula is:
[0286]
[0287] where x is the input data, μ, σ 2They are the mean and variance of the data respectively; γ, β, and ∈ are scaling parameters, taking 0.1, 0, and 10 respectively -5
[0288] Then, dilated convolution operations with multi-dilation rate fusion are respectively performed on the feature maps C1 - C3 to obtain sub-feature maps D1 - D3 with different depths corresponding to them;
[0289] Among them, the dilated convolution operation with multi-dilation rate fusion is defined as follows:
[0290] The dilation rate Rate of the i-th layer of convolution i represents the distance dilation rate between adjacent elements in the convolution kernel of this layer, and Rate i can be expressed as:
[0291] Rate i = f(i)
[0292] where i is the layer number, and f(i) is the layer function, f(i) = 2 i-1 ;
[0293] The actual effective convolution size V of the i-th layer of convolution kernel with size k i can be expressed as:
[0294] V i = (k - 1) * Rate i + 1
[0295] The operation during the i-th layer of image convolution can be expressed as:
[0296]
[0297] where I(i, j) and K(i, j) are the values at the i-th and j-th positions on the image and the convolution kernel respectively;
[0298] In this embodiment, on the feature maps C1 - C3, according to the number of layers 1, 2, and 3, dilated convolutions with dilation rates of 1, 2, and 4 are respectively used to obtain their respective sub-feature maps: C1 - D1, C1 - D2, C1 - D3,..., C3 - D1, C3 - D2, C3 - D3. Among them, the receptive field sizes of D1 - D3 are respectively:
[0299] D1: (3 - 1)×1 + 1 = 3;
[0300] D2: (3 - 1)×2 + 1 = 5;
[0301] D3: (3 - 1)×4 + 1 = 9;
[0302] Step S202: Introduce a cross-attention layer in each layer between the feature maps C1 - C3, and learn the correlation between their sub-feature maps through the designed AN network. The specific learning process of the AN network is as follows:
[0303] First, calculate the attention weights of different feature maps:
[0304] W = softmax(AN(D i , D j )))
[0305] Then, apply the weights to the class feature map to obtain the fused feature map:
[0306] D = W * D i
[0307] Finally, merge the fused feature map D with the original class sub-feature map D i to obtain a new feature map D * :
[0308] D * = Merge(D, D i )
[0309] In this embodiment, the 9 sub-feature maps are divided into 27 groups: {(C1 - D1, C2 - D1), (C2 - D1, C3 - D1), (C3 - D1, C1 - D1), …, (C3 - D3, C1 - D3)}, which are respectively input into the layer attention network AN. After obtaining the fused feature sub-maps, they are merged with the original feature sub-maps to obtain the new feature map D * 1, …, D * 9 and are divided into three groups according to the source feature maps C1 - C3:
[0310] Feature maps generated from C1: D1, D2, D3;
[0311] Feature maps generated from C2: D4, D5, D6;
[0312] Feature maps generated from C3: D7, D8, D9;
[0313] Step S203: Perform multi-scale fusion on the different sub-feature maps D in the feature maps C1 - C3 to obtain the final feature pyramid P;
[0314] In this embodiment, according to the source, the feature sub-maps are fused at a single scale through average pooling to obtain feature maps: F_C1, F_C2, F_C3; F_C1, F_C2, F_C3 are concatenated and connected to the ReLU(·) activation function to obtain the feature pyramid P;
[0315] Step S204: Apply the multi-scale G-LBP operator on each layer of the obtained feature pyramid P to extract features and obtain a feature map. The multi-scale feature extraction operator is specifically as follows:
[0316] In this embodiment, the calculation process of the MC-LBP operator is as Figure 6 shown;
[0317] (1) Divide each layer of the feature pyramid P into 9 2×2 grids;
[0318] (2) Calculate the MC-LBP value with r = 1 and p = 8 for each grid as the feature value of the grid:
[0319]
[0320] (3) According to each feature value M p,r (x, y) in the local grid, update the overall feature value M i of the i-th local grid. The calculation formula for the feature value M i is as follows:
[0321]
[0322] (4) According to the updated local grid feature value M i , calculate the MC-LBP operator feature value T m in the full image range with r = r2 and p = p2. The calculation formula for the feature value T m is as follows:
[0323]
[0324] T m = softmax(M * p,r (x, y))
[0325] (5) Apply a Gabor filter with a frequency of 0.2, an azimuth angle of 0 degrees, a standard deviation of 2, a phase offset of 0, and an aspect ratio of 0.5 to the image. By performing a convolution operation with the image at each pixel position, obtain the response value at that position and normalize it. The specific process of feature extraction by the Gabor filter is as follows:
[0326] T g = softmax(conv(G, I))
[0327] (6) Define two actions as: increase a1, decrease a2, and an ∈-greedy policy, and set the initial attention weight a = 0.5, discount factor γ = 0.9, learning rate lr = 0.1, and exploration factor ∈ = 0.1;
[0328] (7) Initialize the action value function table Q - table: Initialize each state - action pair to 0;
[0329] (8) According to the current attention weight ξ, explore how to increase or decrease the weight ξ with a certain probability
[0330] (9) Execute the action and update the Q - value:
[0331]
[0332] Among them, Q(s,a) is the Q - value of executing action a in state s, α is the learning rate, r is the immediate reward obtained after executing action a, γ is the discount factor, indicating the importance of future rewards, s' is the next state transferred to after executing action a, and Q(s,a) is updated according to the current reward signal r and the expected future reward maxQ(s,a), guiding the model to select the best action in a given state;
[0333] (8) Calculate the initial G - LBP feature according to the formula:
[0334] T * = ξ * T m +(1 - ξ) * T g
[0335] In this embodiment, the weight ξ takes 0.65;
[0336] The specific steps of step S3 are as follows:
[0337] Step S301: Randomly perform rotation and cropping operations on the input feature map and adjust it to a feature map of the same size to complete data augmentation;
[0338] In this embodiment, each frame of the image is randomly rotated by an angle: [-10°, 10°] and cropped to a size of 80% to 100%;
[0339] Step S302: Construct a deep - learning - driven target detection module to accurately identify and locate the road surface defects on the feature map. The specific steps for constructing the target detection algorithm are as follows:
[0340] (1) For each scale of the feature map, calculate the channel attention weight CA i and the spatial attention weight SA i ;
[0341] (2) Adjust the original feature map for different feature maps using the attention weights respectively. The adjustment formula is as follows;
[0342]
[0343] (3) Add the adjusted feature maps, and the addition formula is as follows;
[0344]
[0345] In this embodiment, ε1 and ε2 are taken as 0.3 and 0.7 respectively;
[0346] (4) Set multiple prior boxes with different aspect ratios in each unit as the benchmark for the subsequent predicted bounding boxes;
[0347] In this embodiment, prior boxes with an aspect ratio of 1:1 suitable for square targets, prior boxes with an aspect ratio of 1:2 suitable for taller targets, and prior boxes with an aspect ratio of 2:1 suitable for horizontally longer targets are established respectively;
[0348] (5) Flatten the regional features corresponding to each prior box into a vector, and predict the offset of the generated prior box in the fully connected layer of regression prediction;
[0349] Among them, the default prior box generated at the position (i, j) on the feature map can be expressed as:
[0350] B i,j =(x, y, w, h)
[0351] Among them, B i,j takes the road defect target (54, 105) as the center coordinate, and w and h are the length and width of the prior box: 16, 32;
[0352] The offset prediction l k can be expressed as:
[0353] l k =(Δx, Δy, Δw, Δh)
[0354] Among them, Δx and Δy represent the horizontal and vertical coordinate offsets of the center coordinate of the prior box, and Δw and Δh are the offset amounts of the width and height scaling ratios of the prior box respectively;
[0355] The actual offsets of the center coordinate and the width and height can be expressed as:
[0356] t x =(x - x a ) / w a
[0357] t y =(y - y a ) / h a
[0358] t w =ln(w / w a )
[0359] t h = ln(h / h a )
[0360] where t x , t y represents the actual offset of the center coordinates of dimension: 2, 8, t w , t h represents the actual offset of the length and width of the prior box: 0.08, 0.09, x a , y a respectively represent the actual center coordinates: (54, 105), w a , h a respectively represent the length and width of the actual prior box: 16, 32;
[0361] The predicted offsets of the center coordinates and the width and height can be expressed as:
[0362] t' x = (x - Δx) / Δw
[0363] t' y = (y - Δy) / Δh
[0364] t' w = ln(w / Δw)
[0365] t' h = ln(h / Δh)
[0366] where t' x , t' y represents the predicted offset of the center coordinates: 3, 8, t' w , t' h is the predicted offset of the length and width of the prior box: 0.1, 0.09;
[0367] The position error is defined as follows:
[0368]
[0369] where N is the number of samples, and HL(·) is the Loss function, which is defined as follows:
[0370]
[0371] where δ is the threshold parameter of the loss function, taking 0.8. In this embodiment, the HL(·) error is 1.028;
[0372] Step S303: Further filter the overlapping detection boxes through the non-maximum suppression algorithm with the IoU threshold set to 0.5 to ensure that each defect is detected only once;
[0373] Step S4 specifically includes the following steps:
[0374] Step S401: Construct a classifier to classify the road defects in the detection boxes output by the target detection module into pothole defects and crack defects. The steps to construct the classifier are as follows:
[0375] (1) Calculate the distance between each sample point and other sample points and their corresponding weights. The weight calculation formula is as follows:
[0376]
[0377] where k is the number of K-nearest neighbors of sample i, S(i) is the set of k nearest neighbor points of sample i, and d i,j is the distance between sample points i and j;
[0378] (2) Calculate the local density of each sample point based on its weight and the number of neighbors. The local density of this point represents the degree of density of this point in the feature space. Calculate the local density of weighted neighbors The formula is as follows:
[0379]
[0380] (3) Calculate the relative distance η i for each pair of sample points and other samples with a density higher than it. The relative distance calculation formula is as follows:
[0381]
[0382] (4) Calculate the sample decision value θ i and select the set of cluster centers and assign cluster labels. The sample decision value calculation formula is as follows:
[0383]
[0384] (5) Select the k samples with the highest decision values as the initial cluster centers;
[0385] (6) Calculate the similarity between samples and construct a similarity matrix. The similarity between samples is defined as follows:
[0386]
[0387] where sim i,j and w i,j are the similarity and similarity weight between sample points i and j respectively;
[0388] (7) Traverse the samples with assigned cluster labels, find the unassigned sample with the greatest similarity to them, and assign the unassigned sample to the cluster where the assigned sample is located. Repeat this operation until all samples have been assigned, or the similarity between the assigned samples and the unassigned samples is 0;
[0389] (8) If there are still unassigned samples, assign the remaining samples to the cluster where the sample with a higher density and the closest distance to it is located;
[0390] (9) Take the sample with the highest decision value θ i in the cluster as the new cluster center, repeat the above steps until the cluster center no longer changes, and output the classification result R1;
[0391] Step S402: Input the classified results into the neural network and perform Dropout operation to prevent overfitting;
[0392] Step S403: Input the two classification results into the voter, set the weights and use the soft voting mechanism to obtain the final road defect classification result;
[0393] In this embodiment, k is taken as 5. After calculation, the weights, relative distances, and decision values of the first 5 samples are:
[0394] Sample 1: η1 = 0.2934, θ1 = 0.1264;
[0395] Sample 2: η2 = 0.3121, θ2 = 0.1208;
[0396] Sample 3: η3 = 0.2756, θ3 = 0.0983;
[0397] Sample 4: η4 = 0.2934, θ4 = 0.1232;
[0398] Sample 5: η5 = 0.3001, θ5 = 0.1045;
[0399] The neural network consists of three convolutional layers of 1×1, 3×3, and 5×5, and the dropout rate of Dropout is taken as 0.1; the weights are taken as 0.65 and 0.35 respectively, and the classification results R1, R2, R * are all {1, 0, 1, 1, 0}, where 0 represents pothole defects and 1 represents crack defects;
[0400] The specific steps of the said step S5 are as follows:
[0401] Step S501: Calculate the damage degrees of pothole defects and crack defects respectively. The calculation formulas for the damage degrees of potholes and cracks are as follows:
[0402]
[0403] Where p a is the number of pixel points quantified for the pothole defect target, p s is the number of pixel points quantified for the entire road pothole defect image, and a is the number of pixels within the minimum circumscribed rectangle of the crack; the damage degree D p value is used to measure the severity of the pothole image damage, and D n is used to measure the severity of the crack image damage;
[0404] Step S502: Classify the damage levels of pothole defects and crack defects. The damage level classification is as follows:
[0405] Damage level of road pothole defects: D p ≤17% is a light pothole, 17%≤D p ≤21% is a medium pothole, 21%≤D p is a heavy pothole;
[0406] Damage level of road crack defects: D a ≤18% is a light crack defect, 18%≤D a is a heavy crack defect;
[0407] Step S6 specifically includes the following steps:
[0408] Step S601: Transmit the obstacle types and positions of potholes or crack defects identified on the road by the classifier back to the vehicle's display in a flashing manner through different graphics and colors;
[0409] Step S602: Upload the road defect information and the actions taken by the driver to the cloud server;
[0410] In this embodiment, the road defect image data is fully preprocessed, Figure 2 and Figure 3 respectively show the working principles of the improved adaptive weighted average filtering on normal images and noisy images; Figure 4 The temporal noise reduction structure in Figure 5 can effectively utilize the temporal correlation to perform multi-frame noise reduction on road defect videos; Figure 6 uses dilated convolutions with multiple dilation rates combined with an attention mechanism to obtain road defect features; Figure 7 shows the feature operators extracted by the MC-LBP operator of the present invention at multiple scales; Figure 8 andFigure 9 This is an example of the road defect level classification of the present invention; the present invention improves the target recognition ability through rich image preprocessing work and the self-attention mechanism, which helps to add more accurate and rapid recognition of road defects and includes richer data information.
[0411] In the present invention, specific embodiments are applied to elaborate the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
[0412] Those of ordinary skill in the art will realize that the embodiments described herein are for helping the reader understand the principle of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.
Claims
1. A road defect detection method based on deep learning, characterized in that: It includes the following steps: Step S1: Based on the original image data, construct an image pre-processor and perform pre-processing operations of multi-frame smoothing and noise reduction based on time series. Step S2: Based on the improved multi-scale feature extraction algorithm, extract the eigenvalues of the processed image to obtain different features of low, medium, and high levels of the current road surface cracks and pothole defects; the specific steps of step S2 include: Step S201: Apply the improved dilated convolution operation with multi-dilation rate fusion on the input image to obtain the preliminary feature maps C1-C n and their corresponding sub-feature maps D1-D with different depths n ; Step S202: Introduce a cross-attention layer between each layer of the feature maps C1-C n to learn the correlation between its sub-feature maps through the designed AN network; Step S203: Perform multi-scale fusion on the different sub-feature maps D n in the feature maps C1-C * to obtain the final feature pyramid P; Step S204: Apply the multi-scale G-LBP operator on each layer of the obtained feature pyramid P to perform feature extraction to obtain the feature maps. The multi-scale feature extraction method is specifically as follows; (1) Divide the image into n grids of m×m, and each grid includes different radii r and different sampling points p. (2) Calculate the eigenvalue M of the MC-LBP operator with r = r1 and p = p1 in each local grid. The calculation formula of the eigenvalue M is as follows: Among them, M p, r (x, y) represents the eigenvalue of the MC-LBP operator, p is the number of sampling points, r is the radius, (x, y) are the coordinates of the center point, s(•) is the threshold function, and the formula of the threshold function is as follows: where θ i is the angle of the i-th sampling point, and g(x, y) and g(x + rcos(θ i ), y + rsin(θ i )) are the gray values of the central pixel and the sampling point, respectively; (3) According to each eigenvalue M in the local grid p, r(x, y), Update the overall eigenvalue M of the i-th local grid i, Eigenvalue M i The calculation formula is as follows: where p represents the number of sampling points. (4) According to the updated local grid eigenvalue M i , calculate the eigenvalue T of the MC-LBP operator with r = r2 and p = p2 in the whole image range m , the eigenvalue T m The calculation formula is as follows: Among them, M * p, r(x, y) is the eigenvalue of the global MC-LBP operator, and g c is the gray value of the central pixel, and M i is the gray value of the local grid where the sampling point is located; T m is the normalized eigenvalue of the global MC-LBP operator; (5) Perform a convolution operation between the Gabor filter and each pixel position of the image to obtain the response value at that position. The specific method for the Gabor filter to extract features is: Among them, represents the normalized Gabor feature, where G represents the Gabor filter and I represents the input image; (6) Apply the attention mechanism to fuse the features T m and T g extracted by the MC-LBP operator and the Gabor filter to obtain the G-LBP feature; Step S3: Construct an object detection algorithm to perform object recognition of road defects in the feature map. Step S4: Construct a road defect classifier to classify the detected objects into pothole defects and crack defects. Step S5: According to different road defect categories, calculate their damage degrees respectively and make damage level divisions. Step S6: Transmit the recognized road defect information back to the vehicle display for warning, and upload the relevant information to the cloud for analysis and statistics.
2. The method for detecting road defects based on deep learning according to claim 1, wherein: The specific content of step S1 includes: Step S101: For each video stream to be processed, decompose the video into individual frames. Step S102: Intercept k consecutive images and select the middle frame as the key frame. Filter out noise frames according to the inter-frame similarity and perform pixel weighted replacement on the noise frames. The specific operation steps are as follows: First, calculate the similarity between each frame image and the key frame image. The similarity calculation formula is: Among them, i and represent the current frame and the intermediate frame image of the video respectively, and μ i , represent the means of the pixels of the two frame images respectively, , represent the variances of the pixels of the two frame images respectively, cov(•) is the covariance function, and σ i is the standard deviation of the Gaussian function of the i-th frame image, and C1 and C2 are stability constants; Then, set the similarity threshold , and filter out the noise frames whose similarity to the key frame is less than the threshold ; Furthermore, calculate the pixel weights of the noise frame images. The pixel weight calculation formula is: where Rmse(•) is the root mean square error function, R(•) is the range function, and this formula and subsequent exp(•) are all exponential functions with the natural constant e as the base. Finally, perform weighted replacement on the noise frame pixels according to the weights. The pixel weighted replacement formula is: Among them, I w (x, y) is the weighted pixel value, I i (x, y), I j (x, y) represents the pixel value at the coordinate (x, y) of the i-th frame of noise image and the j-th normal frame image closest to it, w i is the weight of the i-th frame image, and i is the current time point; Step S103: Perform a smoothing operation on the received image using an improved adaptive weighted average filtering method. The specific steps for using the improved filter to smooth the image are: (1) For a given filtering window S xy , calculate the minimum gray value Z xy , the maximum gray value Z min , and the average gray value Z max of all pixels for each pixel point (x, y) within S mean , and set the noise judgment threshold T; (2) Determine whether the pixel mean of the filtering window is a noise point. If it holds, then the mean value Z mean is a normal point; otherwise, expand the window and continue to find the normal mean value. (3) Judgment , if the condition is met, output the pixel value of this point, otherwise calculate the difference value between each pixel of the input image and the pixels in its neighborhood using the Euclidean distance; (4) Calculate the pixel weights according to the difference value. The pixel weight calculation formula is: Among them, W(x, y) represents the weight of the pixel point (x, y), and diff(x, y) = represents the difference value between the pixel at this point and the central pixel (c x , c y ); (5) Multiply the weights by the corresponding pixel values to calculate the filtered pixel values. The calculation method of the filtered pixel values is: FP(x, y)= Among them, FP(x, y) represents the filtered pixel value, and I w (x, y) represents the pixel value in the original image; (6) Repeat the steps until all pixel points are traversed. (7) When there are still noise pixels after the window reaches the maximum size, use Gaussian filtering to replace the noise pixels. Step S104: Arrange the smoothed images in chronological order and perform time series noise reduction processing based on the jump structure. The specific steps for time series noise reduction processing based on the jump structure are: First, input k consecutive images into the encoder, and successively perform downsampling operations through three convolutional kernels of different sizes to extract feature information, obtaining the encoded feature Z E It is expressed as: Then, the encoded feature Z E is divided into groups of every n in chronological order, with a total of k - n + 1 groups, which serve as the input to the bidirectional recurrent neural network that combines residuals in the skip structure, learning the forward and backward time series correlations in the time series; Finally, input k - n + 1 groups of time series into the decoder for upsampling operation, and perform matching splicing with the corresponding positions in the encoder to obtain the final denoised image.
3. The method for detecting road defects based on deep learning according to claim 2, wherein: In step S104, the steps for learning with a bidirectional recurrent neural network combined with residuals are: First, calculate the forward and reverse reset gates at time t , and the propagation gates , : Among them, , are the reset gates in the forward and reverse directions respectively, determining to what extent the hidden state information at the previous moment is reset, thus allowing the network to forget irrelevant information; , are the propagation gates in the forward and reverse directions respectively, controlling to what extent the new candidate hidden state information is updated into the current hidden state; chx(•) and shx(•) are the hyperbolic cosine activation function and hyperbolic sine activation function used for gating degree scaling in the reset gate and propagation gate respectively; W o , U o , b o , W s , U s , b s are the gating parameters for the forward and reverse propagation of the model. W o and W s represent the weight matrices between layers of the reset gate and propagation gate respectively. U o and U s represent the time step matrices of the reset gate and propagation gate respectively. b o and b s represent the bias matrices of the reset gate and propagation gate respectively; Then, calculate the candidate hidden states in the forward and reverse directions respectively , : Among them, and are the candidate hidden states in the forward and reverse directions, respectively, containing the information of the current input and the hidden state at the previous moment; W f , W b , U f , U b , b f , and b b are the forward and reverse gating parameters of the model. W f , W b are the weight matrices for forward and reverse learning. U f , U b are the time-step matrices for forward and reverse learning. b f , and b b are the bias matrices for forward and reverse learning; H is the H-Swith(•) activation function, and the formula of the H-Swith activation function is as follows: This formula and subsequent min(•) are all minimum value functions. Finally, according to the propagation gates and the candidate hidden states, the forward and backward hidden states at the current time , are concatenated into the final hidden state h t , and a residual connection is made with the input encoding feature Z E to obtain the final temporal feature Z R : Among them, h t is the final hidden state, which is the concatenation of the forward and backward hidden states; Z R is the final temporal feature, which is the residual connection of the input encoding feature Z E and the hidden state h t ; [:] is the concatenation operation, and ⊙ is the element-wise multiplication.
4. The method for detecting road defects based on deep learning according to claim 1, wherein: In step S201, the specific steps for using the dilated convolution operation with improved multi-dilation rate fusion are: First, use n convolutional layers of different sizes to extract multi-scale features of the image, obtaining n preliminary feature maps C1 - C at different levels n ; Then, perform convolutions on the feature maps C1-C n using the dilated convolution operation with multi-dilation rate fusion respectively to obtain the sub-feature maps D1-D n ; Among them, the dilated convolution operation with multi-dilation rate fusion is defined as follows: Dilation rate Rate of the i-th layer convolution i Represents the dilation rate of the distance between adjacent elements in the convolution kernel of this layer, Rate i Is defined as: where f(i) is a hierarchical function, and f(i) = 2 i-1 ; The actual effective convolution size V of the convolutional kernel of size k in the i-th layer i Is expressed as: V i =(k - 1)*Rate i + 1 The operation during the convolution of the i-th layer image is expressed as: where I(i, j) and K(i, j) are the pixel values at the (i, j) coordinates on the image and the convolution kernel respectively. Both this formula and the subsequent conv(•) are product summation functions; In the step S202, the specific learning process of the AN network is as follows: First, calculate the attention weights of different feature maps: W = softmax(AN(D i , D j )) Among them, D i , D j represent different sub-feature maps; Then, apply the weights to the class feature map to obtain the fused feature map D: Finally, merge the fused feature map D with the original class sub-feature map D i to obtain a new feature map D * : D * =Merge(D, D i ) Among them, D * represents the final sub-feature map, Merge(•) represents the splicing operation, and both this formula and the subsequent softmax(•) are normalized exponential functions.
5. A method for detecting road defects based on deep learning according to claim 1, characterized in that: In the step S204, the application steps of the attention mechanism are: First, initialize the action value function table Q-table and select the initial attention weight ξ; Then, according to the current attention weight ξ, explore how to increase or decrease the weight ξ with a certain probability; Furthermore, define the reward signal Q to guide the model to learn how to select feature weights. The calculation formula for the reward signal Q is defined as: where Q(s, a) is the Q value of performing action a in state s, α is the learning rate, r is the immediate reward obtained after performing action a, γ represents the discount factor, indicating the importance of future rewards, s' is the next state transferred to after performing action a, and Q(s, a) is updated according to the current reward signal r and the expected future reward maxQ(s, a) to guide the model to select the best action in a given state. Both this formula and the subsequent max(•) are maximum value functions; Finally, perform weighted fusion to obtain the final G-LBP feature. The G-LBP feature calculation formula is: Among them, T * represents the G-LBP feature, and T m , T g are the normalized MC-LBP and Gabor features respectively.
6. The method for detecting road defects based on deep learning according to claim 1, characterized in that: The step S3 specifically includes: Step S301: Randomly perform rotation and cropping operations on the input feature map and adjust it to a feature map of the same size to complete data augmentation; Step S302: Construct a deep learning-driven target detection module to accurately identify and locate the road surface defects on the feature map. The specific construction steps of the target detection module are: First, detect targets at different scales on the feature map through a convolutional attention module; Then, set multiple prior boxes with different aspect ratios as the benchmark for subsequent predicted bounding boxes; Finally, flatten the regional features corresponding to each prior box into a vector, and predict the offset of the generated prior box in the fully connected layer of the regression prediction; where the default prior box generated at the position (i, j) on the feature map is expressed as: Among them, B i, j has the center coordinates of (x, y), and w and h are the length and width of the prior box; Offset prediction l k Expressed as: where Δx and Δy represent the horizontal and vertical offsets of the center coordinates of the prior box, and Δw and Δh are the offsets of the width and height scaling ratios of the prior box respectively; The actual offsets of the center coordinates and the width and height are expressed as: where t x , t y represents the actual offset of the center coordinate of the dimension, and t w , t h represents the actual offset of the length and width of the prior box. x a , y a represent the actual center coordinates respectively, and w a , h a represent the length and width of the actual prior box respectively; The predicted offsets of the center coordinates and the width and height are expressed as: where \(t'\) x , \(t'\) y represents the predicted offset of the center coordinates, and \(t'\) w , \(t'\) h is the predicted offset of the length and width of the prior box; The position error is defined as follows: HL(x, y) where N is the number of samples, and HL(•) is the Loss function, which is defined as follows: = where δ is the threshold parameter of the loss function; Step S303: Further filter the overlapping detection boxes through the non-maximum suppression algorithm to ensure that each defect is detected only once.
7. A method for detecting road defects based on deep learning according to claim 6, characterized in that: In the step S302, the steps for designing the attention mechanism are: First, calculate the channel attention weights CA between feature maps of each scale at different time frames i and the spatial attention weights SA i ; Then, for different feature maps, adjust their original feature maps using the attention weights. The adjustment formula is as follows; Among them, , F i represent the adjusted feature map and the original feature map respectively; CA i , SA i represent the channel and spatial attention weights respectively; Finally, add the adjusted feature maps, and the addition formula is as follows; Where F * is the channel of the final feature map; ε1 and ε2 are the weights for adjusting the fusion of different feature maps and satisfy: ε1 + ε2 ≤ 1.
8. A method for detecting road defects based on deep learning according to claim 1, characterized in that: The specific steps of step S4 include: S401: Construct a road defect classifier to classify the road defects in the detection boxes output by the target detection module into pothole defects and crack defects. The specific construction steps of the road defect classifier are as follows: (1) Calculate the distance between each sample point and other sample points and their corresponding weights. The weight calculation formula is as follows: where k is the number of K-nearest neighbors of sample i, S(i) is the set of k nearest neighbor points of sample i, d i, j为样本点i和j之间的距离; (2) Calculate the local density of each sample point based on its weight and the number of neighbors. The local density of this point represents the degree of density of this point in the feature space, and calculate the local density φ of the weighted nearest neighbor. i , and the formula is as follows: (3) Calculate the relative distance η for each pair of sample points and other samples with higher density than it i , and the relative distance calculation formula is as follows: (4) Calculate the sample decision value , and select the set of cluster centers, assign cluster labels. The calculation formula for the sample decision value is as follows: (5) Select the k samples with the highest decision values as the initial cluster centers; (6) Calculate the similarity between samples and construct a similarity matrix. The definition of the similarity between samples is as follows: where, sim i, j and w i, j are the similarity and similarity weight between sample points i and j, respectively; (7) Traverse the samples with assigned cluster labels, find the unassigned sample with the greatest similarity to it, and assign the unassigned sample to the cluster where the assigned sample is located. Repeat this operation until all samples are assigned, or the similarity between an assigned sample and an unassigned sample is 0; (8) If there are still unassigned samples, assign the remaining samples to the cluster where the sample with a higher density and the closest distance to it is located; (9) Take the sample with the highest decision value in the cluster as the new cluster center, and repeat the above steps until the cluster center no longer changes, and output the classification result R1; Take the sample with the highest decision value in the cluster as the new cluster center, and repeat the above steps until the cluster center no longer changes, and output the classification result R1; Step S402: Input the classified result R1 into the neural network and perform Dropout operation to prevent overfitting, and output the classification result R2; Step S403: Input the two classification results R1 and R2 into a voter, set the weights φ1 and φ2, and use the soft voting mechanism to obtain the final road defect classification result , where R 坑洼 = 0, R 裂缝 = 1.
9. A road defect detection method based on deep learning according to claim 1, characterized in that: The specific steps of step S5 include: Step S501: According to the classification result R * Calculate the damage degrees of pothole defects and crack defects respectively. The calculation formulas for the damage degrees of potholes and cracks are as follows: Among them, p a is the number of pixel points quantified for the pothole defect target, and p s is the number of pixel points quantified for the entire road pothole defect image. a is the number of pixels within the minimum bounding rectangle of the crack. By using the pothole damage degree D p value to measure the severity of damage to the pothole image, and by using D a to measure the severity of damage to the crack image; Step S502: Classify the damage levels of pothole defects and crack defects. The damage level classification is as follows: Road pothole defect damage level: D p ≤ C1 is a light pothole; C1 ≤ D p ≤ C2 is a medium pothole; C2 ≤ D p is a heavy pothole; Damage level of road crack defects: D a ≤C3 is a minor crack defect; C3≤D a is a major crack defect.
10. A method for detecting road defects based on deep learning according to claim 1, characterized in that: The specific steps of step S6 include: Step S601: Use the classifier to identify the obstacle types and locations of pothole or crack defects on the road, and transmit them back to the vehicle's display in a flashing manner through different graphics and colors; Step S602: Upload the road defect information and the actions taken by the driver to the cloud server.
Citation Information
Patent Citations
Treadmill with a track-type walking belt
US10220249B1
Window shade attachment
US2350236A
Adhesives
US2380239A
Preservation of food products
US2550256A
Road surface defect identification method and system based on video deep learning
CN112184625A