Pavement crack identification and classification method based on multi-feature scale fusion Faster-R-CNN

Through the multi-feature fusion of the Faster-R-CNN deep learning framework and ZFNet and RPN full convolutional networks, the problem of insufficient pavement crack detection accuracy is solved, and automated and real-time pavement disease identification and classification is realized, which reduces detection costs and improves detection accuracy.

CN120339219APending Publication Date: 2025-07-18NINGXIA UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510415813.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing pavement crack detection methods rely on single mode data, lack the ability to integrate multi-dimensional features, and are difficult to take into account the geometric characteristics, material properties and dynamic changes of the disease, resulting in insufficient detection accuracy and cannot meet the comprehensive demand of smart transportation systems for real-time, robustness and economy.

Method used

The Faster-R-CNN deep learning framework is adopted, combined with ZFNet and RPN full convolution network for multi-feature fusion, optimize detection accuracy through the Soft-NMS algorithm, build an automated pavement crack detection framework, use high-definition camera to obtain image data, perform preprocessing and feature extraction, and realize disease classification and border regression.

Benefits of technology

It reduces the detection cost, meets the real-time detection needs, improves the detection accuracy and automation level, realizes the automatic identification and positioning of road surface diseases, and provides a scientific basis for road maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339219A_ABST
    Figure CN120339219A_ABST
Patent Text Reader

Abstract

The invention discloses a pavement crack recognition and classification method based on multi-feature scale fusion Faster-R-CNN, and particularly relates to the technical field of road detection, and the method comprises the steps: obtaining pavement image data of a research region through an automobile carrying a high-definition camera, and carrying out the preprocessing of the pavement image data to form a required image data set; a ZFNet is used as a crack feature extraction module to perform multi-feature fusion and enhance feature information, an RPN full convolutional network is combined to perform detection by optimizing an anchor box generation mode, and back propagation and stochastic gradient descent are used to optimize and extract a disease candidate region to realize disease classification and frame regression; according to the Faster-R-CNN target detection algorithm, the detection precision is improved through a Soft-NMS algorithm, the Soft-NMS algorithm is the same as an NMS algorithm in the execution process, but function operation is used for original confidence score, the objective is to reduce the confidence score, reduce the omission ratio and improve the detection precision, and construction of the Faster-R-CNN target detection algorithm is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of road detection, and particularly relates to a method for pavement crack recognition and classification based on multi-feature scale fusion Faster-R-CNN. Background Art

[0002] At present, the detection of pavement cracks in the industry is mainly carried out by professional technicians and high-cost facilities. However, in some areas, the roads are located in special and harsh conditions, resulting in high costs for manual detection, difficulty in meeting real-time requirements, and difficulty in adapting to large-scale road networks. The automatic detection methods for road diseases mainly include disease detection methods based on infrared thermal imaging, geological radar, ultrasonic, laser scanning, and image acquisition. Although the existing detection methods have their own advantages, they still have significant limitations.

[0003] Among them, the method based on infrared thermal imaging relies on the temperature difference of the road surface to detect cracks or cavities, but it is easily interfered by environmental temperature fluctuations, lighting conditions, and the heat conduction characteristics of road materials. Especially in rainy, snowy weather or when the temperature difference between day and night is large, the false detection rate increases significantly. The ground-penetrating radar technology identifies structural layer diseases through electromagnetic wave reflection. Although it can detect deep defects, the equipment cost is high, the operation requires high professionalism, the data processing is complex, it is difficult to achieve real-time detection, and the resolution of shallow cracks is low. The ultrasonic detection technology uses the characteristics of sound wave propagation to evaluate diseases. Although it can quantify the crack depth, it requires contact measurement and has low detection efficiency, making it difficult to meet the large-scale and rapid detection requirements of vehicle-mounted mobile scenarios. The laser scanning technology realizes high-precision surface deformation analysis through point cloud modeling, but its equipment is large in size and high in cost, and it is sensitive to road surface interferences such as dust and water stains. The processing of point cloud data is time-consuming and it is difficult to meet the real-time requirements in a dynamic environment. The vision detection method based on image acquisition (such as traditional image segmentation or early deep learning models) has a low cost, but it highly depends on the quality of manually labeled data, is easily interfered by factors such as uneven lighting, shadow occlusion, and road surface stains, and has limited ability to distinguish multi-scale cracks and complex backgrounds, presenting a balance problem between missed detection and false detection.

[0004] In addition, the above methods mostly rely on single-modal data and lack the ability of multi-dimensional feature fusion, making it difficult to take into account the geometric characteristics, material properties, and environmental dynamic changes of diseases, resulting in insufficient comprehensive detection accuracy and inability to meet the comprehensive requirements of real-time, robustness, and economy of the intelligent transportation system. Summary of the Invention

[0005] To solve the problems of low efficiency, high labor time cost, and insufficient automation in current road surface crack detection methods, the Faster-R-CNN deep learning framework is introduced to construct a general framework for automatic and integrated road surface crack detection, which can automatically identify and locate road surface diseases, providing a scientific basis for highway maintenance and repair. This method not only reduces the detection cost, meets the real-time detection requirements, but also significantly improves the current situation of the existing detection methods in terms of automation, poor expressiveness, and insufficient application. In addition, it can intuitively and accurately represent the road scene with cracks, thus providing a bottom-layer platform and data support for the digital representation of road surface information.

[0006] Therefore, the present invention provides a method for identifying and classifying road surface cracks based on multi-feature scale fusion Faster-R-CNN to solve the problems raised in the background technology.

[0007] To achieve the above object, the present invention provides the following technical solution: A method for identifying and classifying road surface cracks based on multi-feature scale fusion Faster-R-CNN, including: S1. Obtain the road surface image data of the research area through a vehicle equipped with a high-definition camera and perform preprocessing to form the required image data set;

[0008] S2. Use ZFNet as the crack feature extraction module for multi-feature fusion to enhance the feature information, and then combine the RPN fully convolutional network to perform detection by optimizing the anchor box generation method, and use backpropagation and stochastic gradient descent to optimize the extraction of disease candidate regions to achieve disease classification and bounding box regression;

[0009] Among them, the attention pooling module in ZFNet fuses the original coordinate attention mechanism and the aggregation attention pooling mechanism, and uses parallel pooling operations along the same direction for feature aggregation. Considering the geometric distance and feature distance of pixel points, the dimensionality increase and enhancement of the overall features are carried out, and then the feature maximum value and the neighborhood mean value are aggregated to reflect the image feature information;

[0010] S3. Use the Soft-NMS algorithm to improve the detection accuracy. During the execution of the Soft-NMS algorithm, it is the same as NMS, but the original confidence score is used for function operation, and the goal is to reduce the confidence score to reduce the missed detection rate and improve the detection accuracy, realizing the construction of the Faster-R-CNN object detection algorithm.

[0011] Preferably, the specific process of the attention pooling module in ZFNet is as follows:

[0012] First, according to the feature vector g(i) of any point i and the feature vectors g(k) of the k nearest neighbor points around it, calculate the mean value of the two through the L1 norm to obtain the feature distance The formula is as follows:

[0013]

[0014] In the formula: |·| represents the L1 norm; Ave(·) represents the mean function; g(i) and g(k) respectively represent the feature vectors of the i-th and the k-th nearest neighbor points;

[0015] Then, according to the geometric distance and the feature distance to determine the attention weight based on the relationship between the neighboring pixels reflected, take the negative of the geometric distance and the feature distance and perform weighted summation through the normalized exponential function softmax to merge the results and obtain the distance feature Then, fuse the distance feature with the local feature to obtain the fused distance feature and perform weight learning. The formula is as follows:

[0016]

[0017] In the formula: and are the geometric distance and the feature distance; and are the distance feature and the local feature; is the merging operation;

[0018] Finally, calculate the aggregated local features, use the convolution function and the activation function to automatically learn the attention weights to select important features and remove unimportant features, and perform weighted summation on the learned weights and the local feature to calculate the local average feature f iLAve , as shown in Equation (4), and then calculate the local maximum feature f iLmax among the k nearest neighbor points. Finally, merge the local average feature f iLAve and the local maximum feature f iLmax as the local aggregated feature f iL , as shown in Equation (5):

[0019]

[0020] Preferably, the training process of the RPN fully convolutional network for extracting the disease candidate region network is end-to-end. The optimization methods used are back-propagation (BP) and Stochastic Gradient Descent (SGD), and the loss function is the combined loss of classification error and regression error. The loss curve is as Figure 7 shown, as Equation (6) is as follows:

[0021]

[0022] where: i represents the i-th needle point, p i represents the possibility that the needle point i contains the target; is the label of the needle point, indicating that the i-th needle point contains the disease, indicating that the i-th needle point does not contain the disease; t i is a vector representing the predicted border abscissa x of the predicted needle point; the predicted border ordinate y; the predicted border width w; the predicted border height h, t i represents the true border of the needle point containing the disease to be measured; λ is used to adjust the relative importance of the two sub-loss functions; L cls represents the logarithmic loss of disease and non-disease; L reg is the regression loss of the anchor border containing the disease to be measured;

[0023] where smooth L1 is a robust regression loss function, as shown in Equation (7):

[0024]

[0025] The purpose of boundary regression is to predict the precise position of the bounding box, and the parameterization of the coordinates of the bounding box is shown in Equation (8):

[0026]

[0027] where: x, y, w, and h represent the center abscissa, ordinate, width of the border, and height respectively; x, x a and x * represent the abscissas of the predicted border, the anchor border, and the true border of the target; y, y a , y* represent the ordinates of the predicted border, the anchor border, and the true border of the target, h, h a , h* represent the widths of the predicted border, the anchor border, and the true border of the target, w, w a , w* represent the heights of the predicted border, the anchor border, and the true border of the target..

[0028] Preferably, the specific function of the Soft-NMS algorithm is as follows:

[0029]

[0030] However, Equation (9) is discontinuous, which will cause a break in the scores in the bounding box set. Therefore, Equation (9) is written in the Gaussian weighted form shown in Equation (10):

[0031]

[0032] Where: M represents the detection box with the maximum score, b i is other detection boxes, and iou(M, b i ) is the intersection over union of M and b i , S i is the score of b i , σ controls the attenuation rate, and D is the set of other detection boxes.

[0033] The present invention has the following advantages:

[0034] Preprocess the collected road surface image data, screen and label it to finally form an image data set, embed ZFNet as a crack feature extraction module, and then combine the RPN fully convolutional network to extract disease candidate regions, realize disease classification and bounding box regression. Through the performance improvement of the Soft-NMS algorithm and the training of the data set, the Soft-NMS algorithm is used to replace the NMS algorithm and the data set is trained, reducing the computational amount and improving the detection accuracy. The training results of the multi-feature scale fusion Faster-R-CNN object detection algorithm for the road surface image data set are obtained. Compared with the prior art, it realizes the automatic extraction of road surface diseases, the accurate identification of disease regions, and the accurate classification and positioning of disease categories, and finally achieves the purpose of detection and prevention. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 is a schematic diagram of the ZFNet feature extraction structure provided in this embodiment;

[0036] Figure 2 is a flowchart of the Faster-R-CNN algorithm provided in this embodiment;

[0037] Figure 3 is the training data set provided in this embodiment;

[0038] Figure 4 is a schematic diagram of the attention pooling model provided in this embodiment;

[0039] Figure 5 is a schematic diagram of the initial anchor box provided in this embodiment;

[0040] Figure 6 is a schematic diagram of the optimized anchor box provided in this embodiment;

[0041] Figure 7 is a training loss curve diagram of the RPN provided in this embodiment;

[0042] Figure 8 and Figure 9 is a schematic diagram of the training results of the Faster-R-CNN road surface diseases provided in this embodiment;

[0043] Figure 10 This is the flow chart provided by the present invention. Specific implementation manner

[0044] The following is a description of the implementation manner of the present invention by specific specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of protection of the present invention.

[0045] As Figure 10 shown, this embodiment provides a method for pavement crack recognition and classification based on multi-feature scale fusion Faster-R-CNN. S1. Obtain pavement image data in the research area through a vehicle equipped with a high-definition camera and perform preprocessing to form the required image data set.

[0046] Specifically, S1 is as follows:

[0047] S1.1. Use a vehicle equipped with a high-definition camera lens to obtain highway pavement images within the coverage area.

[0048] S1.2. According to requirements, preprocess the collected image data, including screening and annotation, and finally form an image data set.

[0049] S2. Use ZFNet as a crack feature extraction module for multi-feature fusion to enhance feature information, and then combine with the RPN fully convolutional network to detect by optimizing the anchor box generation method, and use backpropagation and stochastic gradient descent to optimize the extraction of disease candidate regions to achieve disease classification and border regression.

[0050] Among them, ZFNet maps the feature map from the high level back to the input image space through transposed convolution, records the feature information during pooling, and its structural diagram is as Figure 1 shown. First, input image features, perform max pooling and average pooling respectively in the attention pooling module, then record the feature information obtained by pooling, merge the max pooling feature and the average pooling feature through a merging operation, and finally output the fused feature information through a transposed convolution operation. And RPN, as a fully convolutional neural network, takes the feature map output by the previous feature extraction network as input through the anchor point mechanism and shared convolutional features, outputs rectangular candidate regions, and each rectangular candidate region border has a score. Therefore, the entire algorithm flow of Faster-R-CNN is as Figure 2As shown, the entire process first inputs the image dataset, completes feature extraction through the disentangled attention network, then enters the feature fusion module, completes attention pooling and fuses multi-scale features through the balanced pyramid model, then constructs defect candidate bounding boxes and inputs the fused features, completes classification and bounding box regression through the dual integral channel, and finally inputs the classification and bounding box regression results into the fully connected layer for mapping and outputs the final road surface crack recognition and classification results.

[0051] The attention pooling module in ZFNet fuses the original coordinate attention mechanism and the aggregation attention pooling mechanism, and uses parallel pooling operations along the same direction for feature aggregation. Considering the geometric distance and feature distance of pixel points to perform dimension elevation and enhancement of the overall features, and then aggregating the feature maximum value and the neighborhood mean value to reflect the image feature information, as Figure 4 shown.

[0052] S3. Improve the detection accuracy through the Soft-NMS algorithm. Since the biggest problem in the traditional Non-Maximum Suppression (NMS) algorithm is that it forcibly zeros the scores of adjacent detection boxes (i.e., removes detection boxes with an overlap greater than the overlap threshold). In this case, if a real object appears in the overlapping area, it will lead to the failure of detecting this object and reduce the average detection rate of the algorithm. While Soft-NMS does not simply delete detection boxes with an IoU (Intersection over Union) greater than the threshold during the algorithm execution, but reduces the score. The execution process of the Soft-NMS algorithm is the same as that of NMS, but it uses a function operation on the original confidence score, aiming to reduce the confidence score to reduce the missed detection rate and improve the detection accuracy, and realize the construction of the Faster-R-CNN object detection algorithm.

[0053] Based on the training results of the multi-feature scale fusion Faster-R-CNN object detection algorithm for the road surface image dataset, and on this basis, realize the automatic extraction of road surface diseases, the accurate recognition of disease areas, and the accurate classification and positioning of disease categories. Among them, the training dataset is as Figure 3 shown.

[0054] Specifically, the specific process of the attention pooling module in ZFNet is as follows:

[0055] First, according to the feature vector g(i) of any point i and the feature vectors g(k) of the k-th nearest neighbor points around it, calculate the feature distance by calculating the mean value of the two through the L1 norm The formula is as follows:

[0056]

[0057] In the formula: |·| is the L1 norm; Ave(·) is the mean value function; g(i) and g(k) respectively represent the feature vectors of point i and the k-th nearest neighbor point;

[0058] Then, according to the geometric distance and the feature distance to determine the attention weight based on the relationship between the pixel neighboring points reflected, take the negative values of the geometric distance and the feature distance , and through weighted summation by the normalized exponential function softmax and merging the results, obtain the distance feature Then, fuse the distance feature with the local feature to obtain the fused distance feature and perform weight learning. The formula is as follows:

[0059]

[0060] In the formula: and are the geometric distance and the feature distance; and are the distance feature and the local feature; is the merging operation;

[0061] Finally, calculate the aggregated local feature, use the convolution function and the activation function to automatically learn the attention weight to select important features and remove unimportant features, and perform weighted summation on the learned weight and the local feature to calculate the local average feature f iLAve , as shown in Equation (4), and then calculate the local maximum feature f iLmax among k neighboring points. Finally, merge the local average feature f iLAve and the local maximum feature f iLmax as the local aggregated feature f iL , as shown in Equation (5):

[0062]

[0063] The anchor box mechanism is the core of the Region Proposal Network (RPN) in the fully convolutional network. The size and ratio of the anchor box have a very important impact on the candidate region generation part. Appropriate anchor boxes can detect more target objects to be measured. If the size of the anchor box is quite different from the size of the target, the generated candidate regions will be inaccurate, thus seriously affecting the model performance and detection effect. Therefore, appropriate anchor box parameters are crucial for the detection effect of the model.

[0064] In the original Faster RCNN model, three aspect ratios of [0.5, 1, 2] and three scaling ratios of [8, 16, 32] are combined to generate as Figure 5The shown set of anchor boxes has poor detection effect. Therefore, in this embodiment, the generation method of anchor boxes is optimized. Based on the original three aspect ratios, two aspect ratios of 0.2 and 5 are added, two scaling ratios of 16 and 32 are removed, and the scaling ratio of 2 is added. Finally, each pixel point generates the anchor boxes as shown in Figure 6 with five aspect ratios of [0.2, 0.5, 1.0, 2.0, 5] and two scaling ratios of [2, 8].

[0065] Specifically, the training process of the RPN fully convolutional network for extracting the disease candidate region network is end-to-end. The optimization methods used are back-propagation (BP) and Stochastic Gradient Descent (SGD). The loss function is the joint loss of classification error and regression error. The loss curve is as shown in Figure 7 . In the training rounds from 0 to 50, the loss value of the model training decreases rapidly, indicating that the model can converge well. In the rounds from 50 to 150, the decreasing rate of the loss curve decreases significantly, indicating that the model gradually tends to be stable. After 150 rounds, the loss value of the model fluctuates between 0.005 and 0.01, indicating that the model training is stable at this time and the training result is optimal at this time, as shown in Equation (6) as follows:

[0066]

[0067] In the formula: i represents the i-th anchor point, p i represents the possibility that the anchor point i contains the target; is the label of the anchor point, represents that the i-th anchor point contains the disease, represents that the i-th anchor point does not contain the disease; t i is a vector representing the 4 parameterized coordinates of the predicted anchor point, t i represents the true bounding box of the anchor point containing the disease to be detected; λ is used to adjust the relative importance of the two sub-loss functions; L cls represents the logarithmic loss of disease and non-disease; L reg is the regression loss of the bounding box containing the disease anchor point;

[0068] Among them, smooth L1 is the robust regression loss function, as shown in Equation (7):

[0069]

[0070] The purpose of boundary regression is to predict the accurate position of the bounding box. The parameterization of the coordinates of the bounding box is shown in Equation (8):

[0071]

[0072] In the formula: x, y, w, and h respectively represent the central abscissa, ordinate, width, and height of the border; x, x a and x * represent the abscissas of the predicted border, the anchor border, and the target true border; the annotation methods of y, h, and w are the same as those of x.

[0073] The specific function of the Soft-NMS algorithm is as follows:

[0074]

[0075] However, formula (9) is discontinuous, which will lead to a break in the scores in the border set. Therefore, formula (9) is written in the Gaussian weighted form shown in formula (10):

[0076]

[0077] In the formula: M represents the detection box with the maximum score, b i is other detection boxes, iou(M, b i ) is the intersection over union of M and b i , S i is the score of b i , σ controls the attenuation rate, and D is the set of other detection boxes.

[0078] In step S4, combining the contents of steps S1 to S3, a multi-feature scale fusion Faster-R-CNN object detection algorithm is constructed, and through this algorithm, the automatic extraction of road surface diseases, the accurate recognition of disease areas, and the accurate classification and positioning of disease categories are realized. Among them, the training result diagrams are as shown in Figure 8 and Figure 9 . Figure 8 is the extraction result of the road surface crack image trained by the Faster-R-CNN network. Figure 9 Among them, the detected cracks include transverse cracks (marked with red borders), longitudinal cracks (marked with blue borders), and crazing (marked with purple borders). The confidence scores of each crack detection are also indicated in the upper right of the corresponding border. As can be seen from Figure 9 , the confidence scores all exceed 0.7, and the detection effect is good.

[0079] Although the present invention has been described in detail above with general descriptions and specific embodiments, based on the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection required by the present invention.

Claims

1. A pavement crack recognition and classification method based on multi-feature scale fusion Faster-R-CNN, comprising: S1. Obtain pavement image data of the research area through a vehicle equipped with a high-definition camera and perform preprocessing to form the required image dataset, characterized in that: S2. Use ZFNet as a crack feature extraction module for multi-feature fusion, and then combine it with the RPN fully convolutional network to use the backpropagation and stochastic gradient descent optimization methods to extract disease candidate regions, realizing disease classification and bounding box regression; Among them, the attention pooling module in ZFNet fuses the original coordinate attention mechanism and the aggregation attention pooling mechanism, and uses parallel pooling operations along the same direction for feature aggregation; considering the geometric distance and feature distance of pixel points for the upsampling and strengthening of the overall features, and then aggregating the feature maximum value and the neighborhood mean value to reflect the image feature information; S3. Improve the detection accuracy through the Soft-NMS algorithm. The Soft-NMS algorithm performs function operations on the original confidence score to reduce the missed detection rate and improve the detection accuracy, realizing the construction of the Faster-R-CNN object detection algorithm.

2. The pavement crack recognition and classification method based on multi-feature scale fusion Faster-R-CNN according to claim 1, wherein: The specific process of the attention pooling module in ZFNet is as follows: First, according to the feature vector g(i) of any point i and the feature vectors g(k) of its k-th nearest neighbors, the feature distance is calculated by taking the L1 norm of the average of the two vectors. The formula is as follows: In the formula: |·| is the L1 norm; Ave(·) is the mean function; g(i) and g(k) respectively represent the feature vectors of the i-th and the k-th nearest neighbor points; Then, according to the geometric distance and the feature distance to determine the weight of attention based on the relationship between pixel neighboring points reflected, take the negative of the geometric distance and the feature distance and perform weighted summation through the normalized exponential function softmax to merge the results and obtain the distance feature Then, fuse the distance feature with the local feature to obtain the fused distance feature and perform weight learning. The formula is as follows: Wherein: and are the geometric distance and the feature distance; and are the distance feature and the local feature; is the merging operation; Finally, calculate the aggregated local features. Using the convolution function and activation function, automatically learn the attention weights to select important features, remove unimportant features, and use the learned weights and the local features to perform weighted summation and calculate the local average feature f iLAve , as shown in Equation (4). Then, calculate the local maximum feature f iLmax among k neighboring points. Finally, merge the local average feature f iLAve and the local maximum feature f iLmax as the local aggregated feature f iL , as shown in Equation (5):

3. A pavement crack recognition and classification method based on multi-feature scale fusion Faster-R-CNN according to claim 1, characterized in that: The optimization method used in the training process of the RPN fully convolutional network for extracting disease candidate regions is backpropagation and stochastic gradient descent, and the loss function is the joint loss of classification error and regression error, as shown in formula (6) below: where: i represents the i-th needle point, p i represents the possibility that the needle point i contains the target; is the label of the needle point; t i is a vector representing the predicted abscissa x of the predicted bounding box of the predicted needle point; the predicted ordinate y of the predicted bounding box; the predicted width w of the predicted bounding box; the predicted height h of the predicted bounding box, t i represents the true bounding box of the needle point containing the disease to be measured; λ is used to adjust the relative importance of the two sub-loss functions; L cls represents the logarithmic loss between the disease and non-disease; L reg is the regression loss of the anchor bounding box containing the disease to be measured; where smooth L1 is a robust regression loss function, as shown in Equation (7): The parameterization of the coordinates of the bounding box is shown in formula (8): Where: x, y, w, and h represent the central abscissa, ordinate, width, and height of the bounding box, respectively; x, x a and x * represent the abscissas of the predicted bounding box, the anchor bounding box, and the ground truth bounding box; y, y a , y * represent the ordinates of the predicted bounding box, the anchor bounding box, and the ground truth bounding box; h, h a , h* represent the widths of the predicted bounding box, the anchor bounding box, and the ground truth bounding box; w, w a , w* represent the heights of the predicted bounding box, the anchor bounding box, and the ground truth bounding box.

4. A pavement crack recognition and classification method based on multi-feature scale fusion Faster-R-CNN according to claim 1, characterized in that: The specific function of the Soft-NMS algorithm is as follows: Write formula (9) in the Gaussian weighted form shown in formula (10): Where: M represents the detection box with the maximum score, b i is other detection boxes, iou(M, b i ) is the intersection over union of M and b i , S i is the score of b i , σ controls the rate of attenuation, and D is the set of other detection boxes.

Citation Information

Cited By

  • Multi-scale road crack detection method and device in extreme weather and medium

    CN120747078A

  • Crack depth measuring system and method for constructional engineering

    CN121612202A