An Online Visual Defect Detection Method for Traffic Facility Inspection
By combining deep learning algorithms and traditional algorithms, defect detection models and measurement learning models are trained, and feature extraction and matching is used for local binary mode operators and attention mechanisms, the problems of multiple visual defect detection in real-time inspection of traffic facilities are solved, and the real-time and highly adaptable defect detection effect is achieved.
Patent Information
- Application Number
- CN202210001738.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-04
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-01-04
AI Technical Summary
The prior art is difficult to effectively detect multiple types of visual defects in real-time inspection of traffic facilities, and the general object detector of deep learning algorithms cannot directly meet the requirements of defect detection.
A visual defect online detection method for traffic facilities inspection is adopted, and the defect detection model and measurement learning model are obtained through offline training. Combined with deep learning algorithms and traditional algorithms, local binary mode operators are used to extract features, and the detection model is embedded through attention mechanism to detect defect types and locations. During online inspection, motion characteristics and appearance characteristics are used to match and update defect trajectories to realize real-time inspection.
Real-time online detection of visual defects in traffic facilities is realized, and can detect multiple types of defects, have real-time and continuous detection capabilities, strong adaptability and high performance.
Smart Images

Figure CN114549401B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a visual defect detection method, in particular to an online visual defect detection method for traffic facility inspection tours. Background Art
[0002] Early traditional computer vision techniques were used for defect detection. For example, in Document 1: Q. Zou, Y. Cao, Q. Li, Q. Mao, and S. Wang, “Cracktree: Automatic crack detection from pavement images,” Pattern Recognition Letters, vol. 33, no. 3, pp. 227–238, 2012. 275–291, 1993. After preprocessing and noise reduction of the image, a threshold is used to obtain candidate crack regions, and then methods in morphology are used to optimize the crack regions. Most subsequent methods are based on handcrafted features and block-based classification methods. For example, in Document 2: M. Quintana, J. Torres, and J. M. Menéndez, “A simplified computer vision system for road surface inspection and maintenance,” IEEE Transactions on Intelligent Transportation Systems, vol. 17, no. 3, pp. 608–619, 2016. This method uses the local binary pattern operator to extract the features of image patches, and then uses a support vector machine for classification. It regards defect detection as a block-based classification task, which requires cutting the image into multiple small patches and then training for classification. These traditional methods all have obvious disadvantages. It is difficult to detect defects with complex morphologies and large illumination changes, and they can only perform classification and are sensitive to the scale size of the blocks. Therefore, they have gradually been replaced by deep learning methods.
[0003] With the increasing progress of deep learning in the field of vision, many works have also applied deep learning techniques to defect detection tasks. For example, in Reference 3: L. Zhang, F. Yang, Y. D. Zhang, and Y. J. Zhu, “Road crack detection using deep convolutional neural network,” in ICIP, 2016, pp. 3708–3712. It constructs a neural network to classify defect images divided into blocks, and the classification accuracy has exceeded that of defect detection methods based on traditional algorithms. Subsequent methods mostly regard defect detection as a segmentation task. For example, in Reference 4: Yang, Fan, et al. "Feature pyramid and hierarchical boosting network for pavement crack detection." IEEE Transactions on Intelligent Transportation Systems 21.4 (2019): 1525-1535. A feature pyramid and a hierarchical boosting structure are added to the segmentation network to obtain better detailed and semantic features for segmenting the defect area. However, this method can only detect single defects and cannot detect multiple types of defects. At the same time, there are also those using general object detectors based on deep learning for defect detection tasks. For example, in Reference 5: Naddaf-Sh, Sadra, et al. "An efficient and scalable deep learning approach for road damage detection." 2020 IEEE International Conference on Big Data (Big Data). IEEE, 2020. The general detector EfficientDet is used, combined with data augmentation for defect detection. However, there are significant differences between the detection objects of general object detectors and visual defects of traffic facilities. General object detectors usually detect objects with clear semantics and determined boundaries, while visual defects of traffic facilities are usually semantically ambiguous and have uncertain boundaries. Therefore, general object detectors cannot directly meet the requirements of defect detection.
[0004] These methods are all for visual defect detection of static images only and cannot be applied to real-time inspection of traffic facilities. In fact, online inspection of traffic facilities pays more attention to the location and size of visual defects, but the detection process must be real-time and continuous, and have the adaptability to detect various defects. Moreover, the current popular defect detection algorithms based on deep learning are all simple improvements of general-object deep learning algorithms, without considering the particularity of defects for improvement, and the performance is low. Summary of the Invention
[0005] Objective of the Invention: The technical problem to be solved by the present invention is to provide an online visual defect detection method for traffic facility inspection in view of the deficiencies of the prior art.
[0006] To solve the above technical problem, the present invention discloses an online visual defect detection method for traffic facility inspection, including the following steps:
[0007] Step 1, obtaining a defect detection model through offline training:
[0008] Construct a defect detection model and a metric learning model;
[0009] Use the defect data set and annotations to train the types and positions of defects in traffic facility inspection images;
[0010] Use the local binary pattern operator to extract features, embed the features into the detection model in the form of an attention mechanism, detect the types and positions of defects in the image, and fuse the deep learning algorithm and the traditional algorithm to obtain a defect detection model;
[0011] Cut the defect image blocks according to the annotations, use the cut defect image blocks of different types and annotations to train the metric learning model, and extract the appearance features of the defects;
[0012] Step 2, online inspection of defects:
[0013] For the input video frame, use the defect detection model trained in Step 1 to detect defects;
[0014] Match and update the defect trajectories based on the motion features of the defects and the appearance features obtained by the metric learning model;
[0015] Statistically obtain the real-time inspection results.
[0016] Step 1 in the present invention includes the following steps:
[0017] Step 1-1, training the defect detection model: Improve the defect detection model, define the objects of general object detection as high-level abstract concepts. High-level abstract concepts emphasize semantic information, de-emphasize texture information, have clear relationships between local and global parts, and have clear semantic boundaries; define visual defects as low-level abstract concepts. Low-level abstract concepts de-emphasize semantic information, emphasize texture information, and have no clear semantic boundaries; enhance the proportion of low-level features in the final detection feature map; reuse low-level features using dense connections, extract features using local binary operators, embed them into the deep network with an attention mechanism, and fuse deep learning algorithms and traditional algorithms; use the K-means clustering algorithm to cluster the anchor box sizes, add mosaic data augmentation, improve the loss function, and use the defect dataset and its annotations to train the improved defect detection model to detect the types and positions of defects in the input image;
[0018] Step 1-2, cutting the defects in the dataset to obtain defect image patches: Cut the defects in the image according to the defect annotations in the dataset, and attach the class label as the training dataset for the metric learning model;
[0019] Step 1-3, training the metric learning model: Use the training dataset obtained in Step 1-2 to train the metric learning model. After removing the last classification layer, use this model as the appearance feature extractor.
[0020] In Step 1-1 of the present invention, it includes: Reusing low-level features using dense connections, extracting features using local binary operators, and embedding them into the deep network with an attention mechanism; dividing the dataset into a training set and a test set, and using the K-means clustering algorithm to cluster the anchor boxes; adopting random horizontal flipping, random cropping, and mosaic as data augmentation methods, improving the loss function, and performing training.
[0021] In Step 1-2 of the present invention, it includes: Cutting each defect in the defect dataset as the dataset for the metric learning model.
[0022] In Step 1-3 of the present invention, defects of a single type in the defect dataset are used as the same target in the same sequence for training the metric learning model.
[0023] Step 2 in the present invention includes the following steps:
[0024] Step 2-1, defect detection: Use the detection model trained in Step 1 to detect the video frames during the inspection process, and obtain the positions, labels, and confidences of all defects therein; where the defect position is a rectangular bounding box indicating the position of the defect in the image; the defect label is a string indicating the defect category; the defect confidence is a decimal number in the range of 0 to 1, which is the degree of confidence in this detection result obtained by the detection model;
[0025] Step 2-2, calculate the appearance features of the detected defects: Use the defect positions detected in Step 2-1 to crop the input video frames to obtain the cropped defect image blocks of the detected defects. Use the metric learning model trained in Step 1 to calculate the cropped defect image blocks to obtain the appearance features of the defects;
[0026] Step 2-3, defect motion feature calculation: Use Kalman filtering to perform motion prediction on the already tracked defect trajectories; Predict the motion information at the current moment based on the motion information at the previous moment to obtain the predicted motion features of the trajectory;
[0027] Step 2-4, trajectory matching: Match the tracking trajectories and detection results using the appearance features obtained in Step 2-2 and the motion features predicted by Kalman filtering; Use the Mahalanobis distance to evaluate the motion matching degree between the tracking trajectories and the detection results; At the same time, calculate the minimum cosine distance of the appearance features of the tracking trajectories and the detection results; Fuse the above two features as the matching cost matrix, and use the Hungarian algorithm for cascade matching; Calculate the IoU distance between the trajectories with only one-frame matching and the detection results for the matching of the Hungarian algorithm to obtain the final matching results;
[0028] Step 2-5, update the state of the trajectory: For the successfully matched trajectories, use their corresponding detection results to update Kalman filtering. For the trajectories that are not successfully matched, mark them as lost. For the detection results that are not successfully matched, initialize them as new trajectory segments;
[0029] Step 2-6, statistical output: For the real-time input video frames, determine the defect type of the entire trajectory according to the majority defect category in each trajectory. Count the types and quantities of the current trajectory segments that appear, and output the types and quantities of all defects that have appeared in the current video so far.
[0030] The calculation of the defect motion features in Step 2-3 of the present invention is as follows: Each defect trajectory has two states, the mean and the covariance; The mean represents the position information of the target, which consists of the center coordinates (cx, cy) of the bounding box, the aspect ratio r, the height h, and their respective velocity change values to form an 8-dimensional vector, denoted as x = [cx, cy, r, h, vx, vy, vr, vh], where vx is the velocity change value on the x-axis, vy is the velocity change value on the y-axis, vr is the velocity change value of the aspect ratio, and vh is the velocity change value of the height. Each velocity value is initialized to 0; The covariance matrix represents the uncertainty of the target position; When performing motion prediction, each defect trajectory predicts the state at the next moment according to its state at the previous moment:
[0031] x′ = Fx
[0032] P′ = FPF T +Q
[0033] Where x is the mean value of the trajectory at the previous moment, F is the state transition matrix, x' is the mean value at the prediction moment, and the mean value prediction is a uniform motion model, where F is:
[0034]
[0035] P is the covariance matrix of the trajectory at the previous moment, Q is the noise matrix of the system, representing the reliability of the entire system, P' is the covariance matrix at the prediction moment, and dt is the time change.
[0036] In steps 2-4 of the present invention, the appearance features obtained through steps 2-2 and 2-3 and the motion features predicted by the Kalman filter are used to match the tracking trajectory and the detection result; the Mahalanobis distance is used to evaluate the motion matching degree between the tracking trajectory and the detection result:
[0037]
[0038] Where d (1) (i, j) represents the motion matching degree between the j-th detection result and the i-th trajectory, z j is the motion state of the j-th detection result, and x i is the predicted motion mean vector of the i-th trajectory at the current moment, and P i is the covariance matrix of the motion information of the predicted space of the i-th trajectory at the current moment;
[0039] Calculate the minimum cosine distance of the appearance features of the i-th trajectory and the j-th detection result:
[0040]
[0041] Where d (2) (i, j) represents the cosine distance between the j-th detection result and the i-th trajectory, r j is the appearance feature of the j-th detection result, and R i is the set of appearance features stored in the i-th trajectory. Then, these two features are fused as:
[0042] c(i, j) = λd (1) (i, j) + (1 - λ)d (2) (i, j)
[0043] Among them, c(i,j) represents the matching cost between the j-th detection result and the i-th trajectory, and λ is the balance coefficient. The matching process is carried out according to the matching cost matrix of the trajectory and the detection result: First is the cascade matching. The cascade matching preferentially uses the Hungarian algorithm to match the most recently appeared trajectories to obtain a preliminary matching result; then is the IoU matching. The trajectory segments with only one-frame matching are used as candidates, and the IoU distance between them and the remaining unmatched detection results is calculated; the Hungarian algorithm is used again for matching to obtain the final matching result.
[0044] In step 2-5 of the present invention, the state of the trajectory segment is updated. For the successfully matched trajectories, the corresponding detection results are used to update the Kalman filter; based on the defects detected at the prediction moment, the state of the trajectories associated with them is corrected to obtain a more accurate result:
[0045] y = z - Hx′
[0046] S = HP′H T +R
[0047] K = P′H T S -1
[0048] x = x′ + Ky
[0049] P = (I - KH)P′
[0050] Among them, z is the mean vector of the detected defects and does not include the velocity change amount; H is the measurement matrix, which maps the mean vector of the trajectory to the detection space and calculates the mean error between the detection result and the trajectory;
[0051] R is the noise matrix of the detector; the covariance matrix of the trajectory is first mapped to the detection space and then the noise matrix is added; K is the Kalman gain, which is used to estimate the importance of the error; the updated mean vector x and covariance matrix P are calculated;
[0052] In step 2-5 of the present invention, the state of the trajectory segment is updated. For the trajectories that are not successfully matched, they are marked as lost, and for the detection results that are not successfully matched, they are initialized as new trajectory segments.
[0053] Beneficial effects:
[0054] 1. The present invention can improve the general object detector according to the characteristics of the defects, integrating the advantages of deep learning algorithms and traditional algorithms, and being more adaptable to the characteristics of defect detection.
[0055] 2. The present invention can perform on-line inspection of defects, mark the positions and their category labels of all defects, and count the defects appearing in the entire video.
[0056] 3. The present invention can detect multiple types of defects through data configuration, rather than just detecting a single type of defect. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.
[0058] Figure 1 It is a schematic diagram of the processing flow of the present invention.
[0059] Figure 2 It is a schematic diagram of the cropped defect image block of the embodiment.
[0060] Figure 3 It is a schematic diagram of the input video of the embodiment.
[0061] Figure 4 It is a schematic diagram of the detection result of the embodiment.
[0062] Figure 5 It is a schematic diagram of the tracking result of the embodiment.
[0063] Figure 6 It is a schematic diagram of the defect statistics result of the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] The present invention will be further described below in conjunction with the drawings and embodiments.
[0065] As Figure 1 shown, the present invention discloses a method for online defect detection for traffic facility inspection, specifically including the following steps:
[0066] Step 1, offline training of a defect detection model and a metric learning model (see the literature: Wojke, N., & Bewley, A. (2018, March). Deep cosine metric learning for person re-identification. In 2018 IEEE winter conference on applications of computer vision (WACV) (pp. 748 - 756). IEEE.): Use the defect data set and annotations to train CrackDet, and at the same time use traditional algorithms to extract features, and embed them into CrackDet with an attention mechanism to integrate the advantages of deep learning algorithms and traditional algorithms. Use different types of cropped defect image blocks and annotations to train the metric learning model for subsequent extraction of the appearance features of defects;
[0067] Step 2, Online inspection of defects: For the input video frames, use the detection model trained in Step 1 to detect defects. Based on the motion features of the defects and the metric learning model to obtain appearance features for matching and updating the defect trajectories, and finally statistically obtain the real-time inspection results.
[0068] The input of the training stage (Step 1) of the present invention is a defect data set with annotations, and the input of the inspection stage (Step 2) is an arbitrary traffic inspection video input.
[0069] The main processes of each step are specifically introduced below:
[0070] 1. Offline training of the defect detection model and the metric learning model
[0071] Train the defect detection model: Improve the detection model. Generally, the objects of general object detection are high-level abstract concepts, emphasizing semantic information, neglecting texture information, with a clear relationship between local and global parts and a clear semantic boundary; while visual defects are generally low-level abstract concepts, neglecting semantic information, emphasizing texture information, without a clear semantic boundary, and local and global parts are similar. Therefore, it is necessary to enhance the proportion of low-level features in the final decision as much as possible. Use dense connections to reuse low-level features multiple times, and at the same time use traditional algorithms to extract features, and embed them into the deep network with an attention mechanism to integrate the advantages of deep learning algorithms and traditional algorithms. Use the K-means clustering algorithm to cluster the anchor box sizes, add mosaic data augmentation (see the literature: Bochkovskiy, A., Wang, C.Y., & Liao, H.Y.M. (2020). Yolov4: Optimal speed and accuracy of object detection. arXiv preprint arXiv:2004.10934.), improve the loss function, and use the defect data set and its annotations to train the improved detection model. This model can detect the types and positions of defects in the input image, and we name it CrackDet. It includes the following steps:
[0072] Step 11, Use dense connections to reuse low-level features multiple times, and at the same time use traditional algorithms to extract features, and embed them into the deep network with an attention mechanism. Divide the data set into a training set and a test set, and use the K-means clustering algorithm to cluster the anchor boxes. Improve the loss function and use various data augmentation methods for training.
[0073] In the present invention, the training set has a total of 18,930 images, the validation set has a total of 2,111 images, the test set has a total of 5,295 images, the hyperparameter of the number of anchor boxes in the K-means clustering algorithm is set to 7, and the data augmentation methods include image flipping, histogram equalization, central cropping, etc.
[0074] Step 12, cutting defects in the dataset to obtain defect image patches: Cut the defects in the image according to the defect annotations in the dataset, and attach the category labels as the training data for the metric learning model.
[0075] In the present invention, the numbers of the four types of defect image patches obtained by cutting are as follows: transverse cracks (5918), longitudinal cracks (4014), reticular cracks (7535), and potholes (5103).
[0076] Step 13, since there is no labeled target tracking sequence annotation for the defects, use the defects of a single type in the defect dataset as the same target in the same sequence for the training of the metric learning model. At the same time, rotation operations that cause ambiguity need to be removed during the data augmentation operation.
[0077] 2. Online inspection of defects
[0078] For the input video frames, use the detection model trained in Step 1 to detect defects, match and update the defect trajectories based on the motion features of the defects and the appearance features obtained from the metric learning model, and finally obtain the real-time inspection results through statistics. It includes the following steps:
[0079] In the present invention, the detection result of each defect is a 6-tuple, including three pieces of information, namely the defect location (4-tuple), the defect category label, and the defect category confidence.
[0080] Step 22, calculating the appearance features of the detected defects: Use the defect locations detected in Step 2-1 to crop the original video frames to obtain the detected defect image patches, and use the metric learning model trained in Step 1 to calculate the cropped defect image patches to obtain the appearance features of the defects;
[0081] Step 23, defect motion feature calculation: Each defect trajectory has two states: mean and covariance. The mean represents the position information of the target, which is composed of the center coordinates (cx, cy) of the bounding box, the aspect ratio r, the height h, and their respective velocity change values, and is represented by an 8-dimensional vector as x = [cx, cy, r, h, vx, vy, vr, vh], and each velocity value is initialized to 0; the covariance matrix represents the uncertainty of the target position. When performing motion prediction, each defect trajectory predicts the state at the next moment based on its state at the previous moment.
[0082] x′ = Fx#(1)
[0083] P′ = FPF T +Q#(2)
[0084] Among them, x is the mean of the trajectory at the previous moment, F is the state transition matrix, and x′ is the mean at the prediction moment. Here, the mean prediction is a uniform motion model. Among them, F is:
[0085]
[0086] In formula (2), P is the covariance matrix of the trajectory at the previous moment, Q is the noise matrix of the system, representing the reliability of the entire system, and P' is the covariance matrix at the prediction moment;
[0087] Before detecting each new video input frame, it is necessary to perform motion prediction on each trajectory segment and update the mean vector x' and covariance matrix P' of each trajectory segment.
[0088] Step 24, trajectory matching: Here, the matching of the tracking trajectory and the detection result requires using the appearance features obtained in Steps 2-2 and 2-3 and the motion features predicted by the Kalman filter. The Mahalanobis distance is used to evaluate the motion matching degree between the tracking trajectory and the detection result.
[0089]
[0090] Formula (3) represents d (1) (i,j) is the motion matching degree between the j-th detection result and the i-th trajectory, z j is the motion state of the j-th detection result, and x i is the predicted motion mean vector of the i-th trajectory at the current moment, and P i is the covariance matrix of the motion information in the predicted space of the i-th trajectory at the current moment.
[0091] At the same time, we will calculate the minimum cosine distance of the appearance features between the i-th trajectory and the j-th detection result.
[0092]
[0093] Among them, d (2) (i,j) represents the cosine distance between the j-th detection result and the i-th trajectory, and r j is the appearance feature of the j-th detection result, and R i is the set of appearance features stored in the i-th trajectory. Then, these two features are fused into:
[0094] c(i,j) = λd (1) (i,j) + (1 - λ)d (2) (i,j) #(5)
[0095] Among them, c(i,j) represents the matching cost between the j-th detection result and the i-th trajectory, and λ is the balance coefficient. With the matching cost matrix of the trajectory and the detection result, the following is the matching process. First is the cascade matching. The cascade matching preferentially uses the Hungarian algorithm to match the most recently appeared trajectories to obtain a preliminary matching result. Then is the IoU matching. The trackers that are only matched in one frame are used as candidates, and the IoU distance between them and the remaining unmatched detection results is calculated, and the Hungarian algorithm (see the literature: H.W. Kuhn, "The Hungarian method for the assignment problem," Naval Research Logistics Quarterly, vol. 2, pp. 83–97, 1955.) is used for matching to obtain the final matching result;
[0096] Step 25, updating the state of the trajectory segment includes: for the successfully matched trajectories, use their corresponding detection results to update the Kalman filter (see the literature: R. Kalman, "A New Approach to Linear Filtering and Prediction Problems," Journal of Basic Engineering, vol. 82, no. Series D, pp. 35–45, 1960.). Based on the defects detected at the prediction moment, correct the state of the trajectory associated with it to obtain a more accurate result.
[0097] y = z - Hx' #(6)
[0098] S = HP'H T + R #(7)
[0099] K = P'H T S -1 #(8)
[0100] x = x' + Ky #(9)
[0101] P = (I - KH)P' #(10)
[0102] In formula (6), z is the mean vector of the detected defects, but does not include the velocity change amount. H is the measurement matrix, which maps the mean vector of the trajectory to the detection space, and this formula calculates the mean error between the detection result and the trajectory. R in formula 7 is the noise matrix of the detector. Formula 7 first maps the covariance matrix of the trajectory to the detection space and then adds the noise matrix. Formula 8 calculates the Kalman gain K, and the Kalman gain is used to estimate the importance of the error. Formulas 9 and 10 obtain the updated mean vector x and covariance matrix P;
[0103] For the trajectories that do not match successfully, mark them as lost. For the detection results that do not match successfully, initialize them as new trajectory segments.
[0104] Embodiment
[0105] In this embodiment, as Figure 3 shown is the input inspection video (several frames are intercepted for display). Through the online defect detection method described in the present invention, it is possible to Figure 3 detect 10 defects of 3 types as shown in the inspection video, and mark their respective positions and class labels. The specific implementation process is as follows: Figure 6 shown, and mark their respective positions and class labels. The specific implementation process is as follows:
[0106] Step one is the offline training process. Figure 2 is a display of the cropped defect data set, which is used to train the metric learning model for subsequent extraction of the appearance features of detected defects.
[0107] In step two, perform online detection on the Figure 3 shown inspection video. Figure 4 is the detection result of the intercepted frames of the inspection video. In the first selected video frame, the detection model detected two defects, which are two mesh cracks, with class ID D20. The defect positions are marked with blue bounding boxes, and the confidence scores are 0.5309 and 0.3581 respectively. In the second video frame, the detection model detected one mesh crack defect, with a confidence score of 0.3055. In the third video frame, two mesh crack defects were detected, with confidence scores of 0.5071 and 0.5167 respectively. Figure 5 is the tracking result of the intercepted frames of the inspection video. In the first selected video frame, the model tracked two defects, with defect IDs 1 and 2 respectively. The types of both defects are mesh cracks, and their average detection confidence scores are 0.5189 and 0.3675 respectively. In the second selected video frame, the model tracked one defect, with defect ID 2, which means this visual defect is the same as the defect with ID 2 in the first frame, and its average confidence score is 0.3765. In the third selected video frame, the model tracked two defects. The one with ID 36 is a mesh crack, with an average confidence score of 0.4699, and the one with ID 20 is a longitudinal crack, with an average confidence score of 0.4789. In the same picture, this visual defect is a mesh crack in the detection result, but in the tracking result, the frames being tracked vote to correct the class of this defect. Figure 6 The upper left corner of the picture shows the statistical results of the inspection video so far. In this 20 - second inspection video, a total of 4 mesh cracks, 4 longitudinal cracks, and 2 pits were detected.
[0108] The present invention provides an idea and method for an online visual defect detection method for traffic facility inspections. There are many methods and ways to specifically implement this technical solution. The above description is only a preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented using existing technologies.
Claims
1. An online visual defect detection method for traffic facility inspection, characterized in that, it includes the following steps: Step 1, offline training to obtain a defect detection model: Construct a defect detection model and a metric learning model; Use the defect dataset and annotations to train the types and locations of defects in traffic facility inspection images; Use the local binary pattern operator to extract features, embed the features into the detection model in the form of an attention mechanism, detect the types and locations of defects in the image, and fuse the deep learning algorithm and the local binary pattern operator to obtain a defect detection model; Cut the defect image blocks according to the annotations, and use the cut defect image blocks of different types and annotations to train the metric learning model to extract the appearance features of the defects; Step 2, online inspection of defects: For the input video frames, use the defect detection model trained in Step 1 to detect defects; Match and update the defect trajectories based on the motion features of the defects and the appearance features obtained by the metric learning model; Statistically obtain the real-time inspection results; Among them, Step 1 includes the following steps: Step 1-1, training the defect detection model: Improve the defect detection model, define the objects of general object detection as high-level abstract concepts, where high-level abstract concepts emphasize semantic information, de-emphasize texture information, have a clear relationship between local and global parts, and have clear semantic boundaries; define visual defects as low-level abstract concepts, where low-level abstract concepts de-emphasize semantic information, emphasize texture information, and have no clear semantic boundaries; enhance the proportion of low-level features in the final detection feature map; reuse low-level features using dense connections, use the local binary pattern operator to extract features, embed them into the deep network with an attention mechanism, and fuse the deep learning algorithm and the local binary pattern operator; use the K-means clustering algorithm to cluster the anchor box sizes, add mosaic data augmentation, improve the loss function, and use the defect dataset and its annotations to train the improved defect detection model to detect the types and locations of defects in the input image; Step 1-2, cutting the defects in the dataset to obtain defect image blocks: Cut the defects in the image according to the defect annotations in the dataset, and attach the category labels as the training dataset for the metric learning model; Step 1-3, training the metric learning model: Use the training dataset obtained in Step 1-2 to train the metric learning model, remove the last classification layer, and use this model as an appearance feature extractor; Step 2 includes the following steps: Step 2-1, defect detection: Use the detection model trained in Step 1 to detect the video frames during the inspection process, and obtain the positions, labels, and confidence levels of all defects; among them, the defect position is a rectangular bounding box indicating the position of the defect in the image; the defect label is a string indicating the defect category; the defect confidence level is a decimal number in the range of 0 to 1, which is the degree of confidence in the detection result obtained by the detection model; Step 2-2, calculate the appearance features of the detected defects: Use the defect positions detected in Step 2-1 to crop the input video frames to obtain the cropped defect image patches of the detected defects, and use the metric learning model trained in Step 1 to calculate the cropped defect image patches to obtain the appearance features of the defects; Step 2-3, defect motion feature calculation: Use Kalman filtering to perform motion prediction on the already tracked defect trajectories; Predict the motion information at the current moment based on the motion information at the previous moment to obtain the predicted motion features of the trajectory; Step 2-4, trajectory matching: Match the tracking trajectories and detection results using the appearance features obtained in Step 2-2 and the motion features predicted by Kalman filtering; Use the Mahalanobis distance to evaluate the motion matching degree between the tracking trajectory and the detection result; At the same time, calculate the minimum cosine distance of the appearance features of the tracking trajectory and the detection result; Fuse the above two features as the matching cost matrix and use the Hungarian algorithm for cascaded matching; Calculate the intersection-over-union distance between the trajectory with only one-frame match and the detection result for the matching of the Hungarian algorithm to obtain the final matching result; Step 2-5, update the state of the trajectory: For the successfully matched trajectories, use their corresponding detection results to update the Kalman filter. For the trajectories that are not successfully matched, mark them as lost. For the detection results that are not successfully matched, initialize them as new trajectory segments; Step 2-6, statistical output: For the real-time input video frames, determine the defect type of the entire trajectory according to the majority defect category in each trajectory, count the types and quantities of the currently appearing trajectory segments, and output the types and quantities of all defects that have appeared in the current video so far; In Step 2-3, the defect motion feature calculation is as follows: Each defect trajectory has two states, mean and covariance; The mean represents the position information of the target, which consists of the center coordinates (cx, cy) of the bounding box, the aspect ratio r, the height h, and their respective velocity change values to form an 8-dimensional vector, denoted as x = [cx, cy, r, h, vx, vy, vr, vh], where vx is the velocity change value on the x-axis, vy is the velocity change value on the y-axis, vr is the velocity change value of the aspect ratio, vh is the velocity change value of the height, and each velocity value is initialized to 0; The covariance matrix represents the uncertainty of the target position; When performing motion prediction, each defect trajectory predicts the state at the next moment based on its state at the previous moment: x′ = Fx P′ = FPF T +Q where x is the mean of the trajectory at the previous moment, F is the state transition matrix, x′ is the mean at the predicted moment, and the mean prediction is a constant velocity model, where F is: P is the covariance matrix of the trajectory at the previous moment, Q is the system noise matrix, representing the reliability of the entire system, P′ is the covariance matrix at the predicted moment, and dt is the time change amount; In Step 2-4, use the appearance features obtained in Step 2-2 and the motion features predicted by Kalman filtering to match the tracking trajectories and detection results; Use the Mahalanobis distance to evaluate the motion matching degree between the tracking trajectory and the detection result: where d (1) (i, j) represents the motion matching degree between the j-th detection result and the i-th trajectory, z j is the motion state of the j-th detection result, x i is the predicted motion mean vector of the i-th trajectory at the current moment, P i is the covariance matrix of the motion information of the predicted space of the i-th trajectory at the current moment; Calculate the minimum cosine distance of the appearance features between the i-th trajectory and the j-th detection result: where d (2) (i, j) represents the cosine distance between the j-th detection result and the i-th trajectory, r j is the appearance feature of the j-th detection result, and R i is the set of appearance features stored in the i-th trajectory. Then, these two features are fused into: c(i,j) = λd (1) (i,j) + (1 - λ)d (2) (i,j) Among them, c(i,j) represents the matching cost between the j-th detection result and the i-th trajectory, and λ is a balance coefficient; the matching process is carried out according to the matching cost matrix of the trajectory and the detection result: first is the cascade matching, which preferentially uses the Hungarian algorithm to match the most recently appeared trajectories to obtain a preliminary matching result; then is the intersection over union (IoU) matching, taking the trajectory segments with only one-frame matching as candidates and calculating the IoU distance between them and the remaining unmatched detection results; and then using the Hungarian algorithm for matching again to obtain the final matching result.
2. A vision defect online detection method for traffic facility inspection according to claim 1, characterized in that, Step 1-1 includes: using dense connections to reuse the underlying features, using the local binary pattern operator to extract features, and embedding them into the deep network with an attention mechanism; dividing the dataset into a training set and a test set, using the K-means clustering algorithm to cluster the anchor boxes; adopting random horizontal flipping, random cropping, and mosaic as data augmentation methods, improving the loss function, and performing training.
3. A vision defect online detection method for traffic facility inspection according to claim 2, characterized in that, Step 1-2 includes: cutting each defect in the defect dataset as the dataset of the metric learning model.
4. A vision defect online detection method for traffic facility inspection according to claim 3, characterized in that, In step 1-3, the defects of a single type in the defect dataset are used as the same target of the same sequence for training the metric learning model.
5. A vision defect online detection method for traffic facility inspection according to claim 4, characterized in that, In step 2-5, update the state of the trajectory segment. For the successfully matched trajectory, use its corresponding detection result to update the Kalman filter; based on the defect detected at the prediction moment, correct the state of the trajectory associated with it to obtain a more accurate result: y = z - Hx′ S = HP′H T + R K = P'H T S -1 x = x′ + Ky P = (I - KH)P′ Among them, z is the mean vector of the detected defect and does not include the velocity change; H is the measurement matrix, which maps the mean vector of the trajectory to the detection space to calculate the mean error between the detection result and the trajectory; R is the noise matrix of the detector; first map the covariance matrix of the trajectory to the detection space and then add the noise matrix; K is the Kalman gain, which is used to estimate the importance of the error; calculate the updated mean vector x and covariance matrix P.
6. A vision defect online detection method for traffic facility inspection according to claim 5, characterized in that, In step 2-5, update the state of the trajectory segment. For the unsuccessfully matched trajectory, mark it as lost, and for the detection result that fails to match, initialize it as a new trajectory segment.
Citation Information
Patent Citations
A ship target tracking method based on depth learning
CN109509214A
Real-time pedestrian tracking method applied to unmanned vehicle
CN111488795A