Rail Surface Defect Detection Method Based on Improved Feature Pyramid Network and Metric Learning
By adopting improved feature pyramid network and metric learning methods in rail surface defect detection, the problems of small sample size and different defect sizes are solved, and more accurate defect detection results are achieved.
Patent Information
- Application Number
- CN202211209175.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-30
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-09-30
AI Technical Summary
In the detection of rail surface defects, the problem of small sample sizes leading to overfitting, and the problem of different defect sizes leading to low recognition rate.
Using improved feature pyramid network (FPN) and metric learning methods, the model is pre-trained on a large data set by adding deformable convolution and convolution attention modules (CBAM) to the ResNet50 network, and using transfer learning to pre-train the model on a large data set, and then migrating to rail surface defect detection. Multimodal feature extraction and distance measurement learning are used to realize defect classification recognition.
It effectively solves the overfitting problem caused by small sample size, and improves the ability to identify defects at different scales, achieving more accurate detection of rail surface defects.
Smart Images

Figure CN115526864B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to surface defect detection technology, and in particular to a method for detecting rail surface defects based on an improved feature pyramid network and metric learning. Background Art
[0002] In recent years, the railway department has implemented a major speed-up strategy, aiming to significantly increase the running speed and frequency of trains, especially high-speed trains. As a result, the track is more likely to be worn to varying degrees, leading to surface defects. As an important part of the railway track, rail surface defects can cause great damage to the wheels and bearings of rail vehicles. When the wheels move on a track with a defective surface, periodic collisions will cause coupled vibrations of the entire vehicle and track system. This will shorten the service life of train components and is also an important cause of vehicle derailment, rollover, and axle burning. Rail surface flaw detection is an important means for detecting rail defects. Usually, such inspection work mainly relies on mechanical detection, supplemented by manual inspection and visual inspection. However, this not only causes serious waste of human resources but also brings problems such as time-consuming, low accuracy, and subjective evaluation. At present, this method is gradually being replaced by automatic detection methods, such as ultrasonic detection, eddy current detection, magnetic flux leakage detection, etc. This requires high reliability and accuracy of the sensor itself and is inconvenient to operate. Machine learning can focus on image features to quickly and accurately detect rail defects. Therefore, rail surface defect detection based on machine vision has important application prospects and academic research significance.
[0003] With the rapid development of deep learning technology, deep convolutional neural networks (DCNNs) are used to extract and identify features, and have good applications in fields such as object recognition and object detection. Researchers have proposed many methods. For example, Kang et al. proposed a new surface defect detection system based on deep convolutional neural networks to solve the problems of visual complexity and small quantity of defects. L Shang et al. proposed a new two-stage railway detection method, including using traditional image processing methods and extracting features by convolutional neural networks (CNNs) for image classification. Jin et al. established a deep multi-modal track detection system for surface defects using an improved Gaussian mixture model, Markov random field, and CNN. Yu et al. proposed a defect recognition model for identifying defects of different scales from coarse to fine. Liu et al. proposed a detection model based on DCNN and a new sample generation method to solve the problem of few samples. Gibert et al. applied DCNN to the automatic detection of fastener states, but the training complexity is high and a large number of training samples are required. Feng et al. improved the YOLO algorithm and feature pyramid network (FPN), and used the backbone network and detection layer of MobileNet to detect rail defects. Yang et al. proposed a fast rail surface defect detection method, including track extraction, defect segmentation based on edge pixels, and YOLOv2 for precise positioning and defect detection. Ni et al. proposed an attention network for detecting rail surface defects by the intersection over union on the joint consistency of center point estimation. Liu et al. used a pyramid feature extraction module to extract multi-scale feature maps, and then trained by a lightweight network for rail surface defect detection.
[0004] For rail defect detection, many deep learning-based detection methods have good effects, but the main difficulties are as follows: 1) There are few defect samples, and it is easy to overfit when training traditional convolutional neural networks. 2) The defect scales are different, and the network has a low recognition rate for tiny defects.
[0005] Regarding Difficulty 1, researchers have conducted in-depth research on generative adversarial networks, transfer learning, and meta-learning. For example, Goodfellow et al. proposed generative adversarial networks (GANs), and there have been some studies and applications in image generation. Zhang et al. established an adversarial data augmentation model based on feature reconstruction and deformation information, using data augmentation for few-shot learning. Weiss et al. reviewed transfer learning and discussed the application issues of transfer learning in the few-shot scenario. Vinyals et al. proposed a matching network using one-shot learning to learn classification. Snell et al. proposed a prototype network for few-shot learning, learning a metric space through the prototype network, in which classification can be performed by calculating the distance to the prototype representation of each class. Gao et al. proposed a prototype network for few-shot relationship classification with hybrid attention to address the problems of being vulnerable to noisy instances and feature sparsity in few-shot learning. Lv et al. proposed a few-shot learning method combining an attention mechanism, using CNN to extract image features, and a relationship network to calculate the similarity between images, predicting the image category through the similarity.
[0006] Regarding Difficulty 2, researchers have carried out defect detection research from the perspective of multi-scale feature fusion. For example, to solve the problem of multi-scale defect detection where features disappear as the network deepens, Xu et al. proposed a bidirectional attention feature pyramid network structure, effectively realizing multi-scale defect detection. FPN can effectively perform multi-scale fusion. Li et al. combined Faster RCNN and FPN, increasing the use of refined shallow features to achieve better detection results for small targets. Li et al. proposed a PCB defect detection algorithm based on an extended feature pyramid network model, with the backbone constructed from a part of ResNet-101 and using the FPN network structure to obtain the final feature layer. Dong et al. proposed a pyramid feature fusion and global context attention network for the complexity problem of inner surface defects. Yang et al. proposed a pipeline magnetic flux leakage image detection algorithm based on a multi-scale SSD network to solve the problem of low detection accuracy for small targets in the SSD detection algorithm. Wu et al. designed an extended convolutional module, using multi-scale convolutional kernels to adapt to different sizes of defects to enhance the feature extraction ability of the network.
[0007] Regarding the above two difficulties, the present invention proposes a method for detecting rail surface defects based on an improved feature pyramid network and metric learning; using the transfer learning method to transfer the trained parameter model to solve the problem of training the network in the few-shot scenario, and using metric learning to solve the problem of detection categories. Summary of the Invention
[0008] To address the deficiencies in the existing technology, the objective of the present invention is to provide a method for detecting rail surface defects based on an improved Feature Pyramid Network (FPN) and metric learning.
[0009] To achieve the objective of the present invention, the technical solution adopted by the present invention is as follows:
[0010] A method for detecting rail surface defects based on an improved Feature Pyramid Network (FPN) and metric learning, comprising the following steps:
[0011] (1) Add deformable convolution (DC) and convolutional block attention module (CBAM) to the ResNet50 network to obtain an improved Feature Pyramid Network (FPN) as the rail surface defect feature extraction network.
[0012] (2) Pre-train the improved Feature Pyramid Network (FPN) on the MS COCO dataset, and then transfer the trained network parameters and network model to the rail surface defect detection model.
[0013] (3) Use the improved FPN network to extract and locate the rail surface defect features from the rail surface defect dataset to obtain the feature map ROI.
[0014] (4) Input the feature map ROI into the RepMet network, which includes a distance metric learning (DML) embedding module; use the multi-modal network structure to obtain the ROI features to get the corresponding feature vectors, and then classify and identify the defects by calculating the distances between the representatives of each modality and the feature vectors obtained by the embedding module.
[0015] Further, in step (1), for the improved Feature Pyramid Network (FPN), first perform deformable convolution (DC) on the image, focus on the rail defect features through the convolutional block attention module (CBAM), after the 1×1 convolution operation, upsample the feature map of the previous layer and stack and fuse it with the feature map of the current layer, and obtain feature maps of different scales (p2, p3, p4, p5) after the 3×3 convolution. Then obtain the feature map ROI from the corresponding regional proposal network (RPN), and finally perform ROI pooling to complete the extraction of the rail surface defect features.
[0016] Further, the deformable convolution (DC) adds a learnable offset ΔP to each point of the convolution operation. n , after learning the target through the offset, the size and position of the deformable convolution kernel are adjusted according to the current image, and the sampling points of the convolution kernel at different positions change adaptively according to the image content; the deformable convolution (DC) is defined as:
[0017]
[0018] In the formula, y(P0) represents the value of the point P0 in the image after the convolution operation; Pn It represents all positions in R.
[0019] Furthermore, the convolutional attention module CBAM includes a channel attention module CAM and a spatial attention module SAM. These two modules focus on the important features of the image in the channel dimension and the spatial dimension respectively, aggregate the channel information of the feature map through two pooling operations, and then connect and convolve them through a standard convolutional layer;
[0020] Given the input feature map F ∈ R C×H×W , through the one-dimensional channel attention map M c ∈ R C×1×1 and the two-dimensional spatial attention map M s ∈ R 1×H×W perform serial calculations to obtain the output F”, that is:
[0021]
[0022]
[0023] where F' is the feature of the input feature after passing through the channel attention module; F” is the feature finally output after passing through the attention module; represents element-wise multiplication; C is the number of channels; H is the height of the feature map; W is the width of the feature map.
[0024] Furthermore, when performing RPN processing on the feature map, the bounding box regression loss function is used to optimize the bounding box.
[0025] Furthermore, step (4) includes the steps:
[0026] (4.1) Use the multi-modal network structure to perform another convolution operation to extract the defective ROI features, and then extract the feature vectors E corresponding to each modality through the fully connected layer ij ; Each modality network consists of 2 fully connected layers, and a ReLU activation function is added after each fully connected layer;
[0027] (4.2) After non-linear processing by the embedding module DML, the feature vector E i is obtained; The DML embedding module consists of 3 fully connected layers, and a ReLU activation function is added after each fully connected layer;
[0028] (4.3) Calculate the distance between the feature vector E i and the feature vector E ij , which represents the distance between the i-th class extracted by FPN in the embedded feature vector E i and each modality feature vector E ijThe distances are used to calculate the probabilities of a given ROI in each class i in each modality j, so as to conduct the classification determination of the ROI.
[0029] Further, in step (4.3), E i and E ij The distance calculation formula between them:
[0030]
[0031] In the formula, n represents the number of categories, and D j (E i ) is used to calculate and extract the probabilities of each class i in each modality j, expressed as the formula:
[0032]
[0033] Further, it also includes the calculation of the posterior probability of the defective ROI category and the calculation of the posterior probability of the background category.
[0034] Further, when performing classification and recognition, a loss function is used for optimization, and an embedding loss function and a cross-entropy loss function are adopted.
[0035] The beneficial effect of the present invention is that, compared with the prior art, in view of the problem of different sizes of defect scales, the present invention proposes to add CBAM and change the traditional convolution in FPN to dynamic convolution, and use RPN to extract features and locate the bounding boxes of defects, so as to solve the problem of extracting surface defects of different scales.
[0036] In view of the problem of few samples, the present invention pre-trains the improved FPN on a large dataset, and then uses transfer learning to transfer the trained model structure to few-shot defect detection. It can avoid the overfitting phenomenon, obtain better network parameters, and thus obtain more accurate defect features.
[0037] The present invention establishes a multi-modal network structure to ensure that the features extracted from the same category feature maps are as consistent as possible, while the features of different categories are as different as possible; then metric learning is used to complete few-shot rail surface defect detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is the schematic diagram of the rail surface defect detection method based on the improved feature pyramid network and metric learning;
[0039] Figure 2 is the structure diagram of the improved feature pyramid network;
[0040] Figure 3 is the structure diagram of CBAM;
[0041] Figure 4It is the structural diagram of the RepMet model;
[0042] Figure 5 It is a schematic diagram of the types of rail defects that can be detected;
[0043] Figure 6 It is a schematic diagram of the training loss values of different models;
[0044] Figure 7 It is the CBAM attention heat map;
[0045] Figure 8 It is the effect diagram of the detection samples. Specific implementation manners
[0046] The technical solution of the present invention will be further described below in conjunction with the accompanying drawings and embodiments. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of this application.
[0047] As Figure 1 shown, the rail surface defect detection method based on the improved feature pyramid network and metric learning of the present invention includes the following steps:
[0048] (1) Construct an improved feature pyramid network (feature pyramid network, FPN). The present invention is designed based on the ResNet50 network, and the deformable convolution DC and the convolutional attention module CBAM are added to obtain the improved feature pyramid network FPN as the rail surface defect feature extraction network;
[0049] As Figure 2 shown, for the improved feature pyramid network FPN, in order to increase the weight of defect features in network training, a convolutional block attention module CBAM is added to the convolutional layer of ResNet50.
[0050] As Figure 3 shown, the convolutional attention module CBAM contains 2 independent sub-modules, namely the channel attention module (CAM) and the spatial attention module (SAM). CAM uses the maximum pooling output and the average pooling output of the shared network, and SAM uses similar two outputs and passes them along the channel axis to the convolutional layer. These two modules focus on the important features of the image in the channel dimension and the spatial dimension respectively. The channel information of the feature map is aggregated through two pooling operations, and then they are connected and convolved through a standard convolutional layer to generate an attention map. The two modules are in series. Given the input feature map F ∈ R C×H×W , through the one-dimensional channel attention mapping M c ∈ R C×1×1 and the two-dimensional spatial attention mapping M s ∈ R 1×H×WSerial calculation is performed to obtain the output F”, that is:
[0051]
[0052]
[0053] Among them, F' is the feature of the input feature passing through the channel attention module; F” is the feature finally output after passing through the attention module; represents element-wise multiplication; C is the number of channels; H is the height of the feature map; W is the width of the feature map.
[0054] The traditional convolution operation adopted in the basic ResNet50 network can be defined as:
[0055]
[0056] In the formula: y(P0) represents the value after the convolution operation is completed at point P0 in the image; P n represents all positions in R, usually taking integers.
[0057] It can be seen from the formula that since the traditional convolution operation can only extract fixed geometric structures with the same receptive field, its generalization ability for the problem of different scales of defects is limited.
[0058] To address the problem of rail defects of different scales, deformable convolution DC is introduced in ResNet50. Since the deformable convolution DC adds an offset, after learning the target through this offset, the size and position of the deformable convolution kernel will be adjusted according to the current image, and the convolution kernel sampling points at different positions will adaptively change according to the image content. The definition of DC can be expressed by the following formula. On the basis of the definition of the traditional convolution operation, a learnable offset ΔP is added to each point n , usually a decimal.
[0059]
[0060] As Figure 2 shown, in order to obtain feature maps of different scales of rail surface defects, the present invention is based on the ResNet50 network for FPN design. The region proposal networks (RPN) in Faster RCNN are used to extract ROIs from feature maps of different scales. The anchor sizes corresponding to feature maps of different scales are also different, with sizes of 16×16, 32×32, 64×64, and 128×128 respectively. Since the size ratios of defect targets are different, 3 types of anchor structures with ratios of 1∶2, 1∶1, and 2∶1 are used, and 12 anchors can be obtained by combining with the anchor sizes.
[0061] Improved Feature Pyramid Network Structure. First, perform deformable convolution operation on the image, focus on the rail defect features through CBAM, upsample after 1×1 convolution operation, and obtain feature maps of different scales (p2, p3, p4, p5) after 3×3 convolution. Then, obtain the ROI from the corresponding RPN respectively, and finally perform ROI pooling to complete the feature extraction of the rail surface defects.
[0062] Due to the complexity and different scales of the rail surface defects, it is necessary to solve the problem of extracting surface defects of different scales. The defect feature extraction on the rail surface adopts the FPN structure, which can detect larger, second-largest, and smaller target defects in the feature maps of high, middle, and low layers respectively. First, extract feature maps from different convolutional layers, then perform 2-fold upsampling on the feature map of the previous layer and superimpose and fuse the feature map of the current layer to achieve the information fusion of the shallow feature map and the deep feature map.
[0063] When performing RPN processing on the feature map, the more consistent the predicted box of the defect is with the ground truth box, the better the defect extraction effect. Therefore, the bounding box regression function can be used to optimize the bounding box. The position of the box is mainly determined by the height, width, and center coordinates of the box together. The following formula is the formula for calculating the deviation of the predicted bounding box, the ground truth bounding box, and the anchor.
[0064]
[0065] Among them, (x, y, w, h) represent the center point coordinates, width, and height of the predicted bounding box; (x*, y*, w*, h*) represent the center point coordinates, width, and height of the ground truth bounding box; (x a , y a , w a , h a ) are the center point coordinates, width, and height of the anchor.
[0066] The following formula is the regression loss function adopted by RPN.
[0067]
[0068]
[0069]
[0070] In the formula, N is the number of anchors, and R represents the smooth L1 function.
[0071] During the training process, use the overlap rate IOU between the predicted bounding box and the actual bounding box as the threshold to divide the positive samples and negative samples. Among them, the region of interest with IOU > 0.7 is used as the positive sample, and the region of interest with IOU < 0.3 is used as the negative sample, and represents the offset predicted by the anchor. Indicates the offset of the anchor from the actual border.
[0072] (2) Pre-train the improved Feature Pyramid Network (FPN) on the MS COCO dataset, and then transfer the trained network parameters and network model to the rail surface defect detection model;
[0073] The rail surface defect detection model includes an improved FPN network and a RepMet network.
[0074] The MS COCO dataset is a dataset constructed by Microsoft. The pictures in COCO contain natural pictures and common object pictures in life. The background is relatively complex, the number of objects is relatively large, and the object size is smaller.
[0075] Rail surface defect detection belongs to few-shot detection. If there are not enough samples for network training, the network will have the phenomenon of overfitting. The present invention uses the transfer learning method to pre-train the improved FPN on the MS COCO dataset, and transfer the trained network parameters and model to the rail surface defect detection model as the defect feature extraction part. This can avoid the overfitting phenomenon and obtain better network parameters, and then obtain more accurate defect features.
[0076] (3) Fine-tune the rail surface defect dataset, and use the improved FPN network to extract and locate the rail surface defect features to obtain the feature map ROI;
[0077] Fine-tuning the rail surface defect dataset includes cropping and classifying the images.
[0078] (4) Input the feature map ROI into the RepMet network for defect classification and recognition; calculate the probability of rail defect classification and recognition through the multi-modal network structure and feature vectors to achieve the goal recognition and classification of defect detection.
[0079] As Figure 4 shown, the RepMet model is a new distance metric learning (DML) method for classification and few-shot detection models. The feature vectors of the regions of interest (ROIs) of the input image obtained through the FPN are divided into two branches.
[0080] On the one hand, the feature vector outputs the vector E through the DML embedding module, which consists of several fully connected layers (FC) with batch normalization (BN) and rectified linear unit (ReLU).
[0081] On the other hand, the feature vector is transformed into the representations of each category through a fully connected layer for the feature vector of the region of interest in the image. A set of "representative" feature vectors R can be obtained from the multimodal mixture distribution ij , and each vector R ij represents the center of the j-th mode of the discriminative mixture distribution learned in the embedding space for the i-th class among N classes. Assume that there are a fixed number K of modes (peaks) in the distribution of each class, then 1 ≤ j ≤ K. For a given image (or ROI) and its corresponding embedded feature vector E, the distances between E and the representative R ij are calculated, and these distances are used to calculate the probabilities of the given image (or ROI) under each mode j of each class i.
[0082] In order to achieve defect recognition by training with a small number of sample defects, the present invention is improved based on the RepMet network. First, a multimodal network structure is used to obtain ROI features to get the corresponding feature vectors, and then the defect category is recognized by calculating the distances between the representatives of each modality and the feature vectors obtained by the embedding module.
[0083] The improved FPN network extracts the feature map ROI as the input of the RepMet network, and uses the multimodal network structure to perform another convolution operation to further extract the defect ROI features, ensuring that the features extracted from the feature maps of the same category are as consistent as possible, while the features of different categories are as different as possible.
[0084] F i = f(x i ), i ∈ R
[0085] where: f is the convolution operation; x i is the feature map for extracting the ROI region of the i-th class; R represents the set of all ROI regions; F i is the feature map after the convolution operation.
[0086] The DML embedding module E consists of 3 fully connected layers, and a ReLU activation function is added after each fully connected layer for non-linear processing. The sizes of the fully connected layers are 512, 256, and 128 respectively. After non-linearly processing the ROI by this module, the feature vector E i = E(F i ) ∈ R e is obtained, and then the common characteristics of all ROI feature information can be extracted using the same non-linear structure.
[0087] The establishment of the multimodal network enables the network to extract as much the same feature information for the same feature categories as possible, and extract feature information with larger differences for different feature categories. Similar to the embedding module, the multimodal structure processes the feature map Fi The convolution operation is upsampled to the fully connected layer FC m , and the size of this layer is set to 1024 to extract richer feature information, which is beneficial for each modality to learn feature information. Each modality network consists of 2 fully connected layers, where FC n1 = 256d, FC n2 = 128d, (n = 1, …, N), n is the number of categories to be classified. After the fully connected layer, the ReLU activation function is added for non-linear processing, and then the corresponding feature vectors E ij (i = 1,2, …,n; j = 1,2, …,k) are obtained for each modality.
[0088] Finally, calculate the distance between E i and E ij , which represents the distance of the i-th class extracted by FPN in the embedding vector E i and the feature vectors E ij of each modality. As shown in the formula:
[0089]
[0090] In the formula: n represents the number of categories.
[0091] D j (E i ) is used to calculate the probability of each i-th class extracted in each modality j, which is expressed as the formula:
[0092]
[0093] In the formula: It is assumed here that the distributions of all categories follow an isotropic multivariate Gaussian distribution with variance σ 2 . The classification determination of the ROI adopts the formula:
[0094]
[0095] In the formula: C = i represents the i-th class, and its minimum value takes the minimum distance calculated for all modalities.
[0096] After calculating the posterior probability of the defective ROI category, it is necessary to further calculate the posterior probability of the background category. Here, the same as RepMet, the foreground probability is used to calculate the background probability, as shown in the formula:
[0097] P(B|X) = P(B|E) = 1 - minD j (E i )
[0098] The loss function of the classification and recognition part adopts two loss functions, namely the embedding loss function and the cross-entropy loss function. The embedding loss function is to ensure that the smaller distance d is closest to the correct modality class. The larger the distance d, it means that it does not belong to the class learned by this modality. The embedding loss function is as shown in the formula:
[0099]
[0100] where i* is the label of the correct class.
[0101] The cross-entropy loss function is as shown in the formula:
[0102]
[0103] The embedding loss function L em and the cross-entropy loss function L CE The sum L t = L em + L CE Performs backpropagation adjustment on the network parameters for few-shot defect classification and recognition.
[0104] The experimental environment of the present invention is completed under PyTorch 1.8 built based on Python 3.7. The method proposed by the present invention is evaluated using the miniImageNet dataset, and defect detection and recognition are performed on the constructed rail surface defect detection dataset.
[0105] The miniImageNet dataset is a benchmark dataset for meta-learning and few-shot learning, which contains 100 categories, and each category contains 600 samples. All samples in the miniImageNet dataset come from the ImageNet dataset.
[0106] The rail surface defect detection dataset consists of rail surface images with at least one defect. There are two types of images, one is the image of the express track, and the other is the image of the ordinary or heavy-haul track. To facilitate the study of rail surface image data, the images are cropped. After analysis and research, the rail surface defect images are classified, as Figure 5 shown, it can be divided into 5 categories: crack, regular circle, irregular, small, and unclear. The crack shape refers to a long and narrow crack across the rail surface; the regular circle refers to a circular defect on the rail surface; the irregular shape means that the surface defect of the rail may be caused by many fine-grained shapes; the small dot-like defect refers to a very fine defect that appears on the rail surface, and the type of defect can be seen when magnifying the image; the fuzzy shape is that there is a defect on the rail surface, but the outline of the defect cannot be clearly seen by the eye.
[0107] I. Experimental Results of the miniImageNet Dataset
[0108] Training is carried out using two modes: 5-way 1-shot and 5-way 5-shot. That is, in each experiment, 1 sample (1-shot) or 5 samples (5-shot) are drawn from each category in the defect dataset for training. The evaluation criterion is accuracy (Acc): Acc = (TP + TN) / Total, where TP + TN represents the number of correctly predicted samples and Total represents the total number of samples.
[0109] The loss curves under different training sampling times are as Figure 6 shown. During the training process, while the loss value of the training set of the network decreases, the performance on the validation set continuously improves, indicating that the network can learn classification performance suitable for defect detection. In addition, to better reflect the performance difference between the network and traditional deep learning, the network performance is generally tested in two ways: training on the miniImageNet dataset and learning from scratch. The network of the present invention adopts the first learning method, and the experimental results are shown in Table 1.
[0110] Table 1 mAP results of the miniImageNet dataset
[0111] Model 5-way 1-shot 5-way 5-shot Matching Network 43.56 55.31 Prototypical Network 49.42 68.20 MAML 48.70 63.11 Relation Network 50.44 65.32 RepMet 56.90 68.80 Ours 58.70 73.42
[0112] From the experimental results in Table 1, it can be seen that the improved FPN of the present invention can better focus on target features, and the experimental performance obtained by using the metric learning method in the experiment is higher than that of other comparison methods, fully proving the effectiveness of this method in few-shot object detection.
[0113] An ablation experiment is conducted on the model of the present invention to verify the influence of the improved part of the present invention on the overall model performance. The "non-CBAM" experiment means that CBAM is not adopted in the model, and the rest of the settings remain unchanged. The "non-FPN" experiment means that FPN is not used, and only the features of the last layer of the network are used to complete ROI extraction, and the rest of the settings remain unchanged. The "non-DC" experiment means that the deformable convolution operation is not performed on the network feature map, and the rest of the settings remain unchanged. The "non-FT" experiment means that the pre-trained model is not fine-tuned during the ROI extraction process. The "IOU-DML" experiment means that the DML subnet structure proposed by RepMet is combined to extract ROI. The ablation experiment results are shown in Table 2.
[0114] Table 2 mAP evaluation results of ablation experiments on the miniImageNet dataset
[0115] 5-way 1-shot 5-way 5-shot non-FPN 31.45 40.27 non-CBAM 58.02 71.72 non-DC 48.53 61.01 non-FT 57.13 69.84 IOU-DML 56.98 69.21 Ours 58.70 73.42
[0116] From the results in Table 2, when neither FPN nor DC operation is used, the performance of the network on the miniImageNet dataset with 5-way few-shot learning drops severely. When CBAM is not used, the model performance slightly decreases compared to the model of the present invention. When the ROI extraction module is not fine-tuned using the loss function, the model performance slightly decreases compared to the model of the present invention. After replacing the ROI recognition module with the DML sub-network in RepMet, the model performance is improved. It can be seen from the results that FPN and DC operations have a greater impact on the model, and the ROI module, fine-tuning operation, and CBAM can improve the model performance to a certain extent.
[0117] II. Detection Results of Rail Surface Defect Dataset
[0118] The method of the present invention is used to experiment on the rail surface defect dataset to verify the effectiveness of the method of the present invention in rail surface defect detection.
[0119] CBAM is an important component in the model of the present invention. In order to study the influence of the attention module on the performance, relevant experiments are carried out on CBAM. The experimental results are as Figure 7 shown. It can be seen that the network with CBAM added has a better learning effect on the defective part.
[0120] The experiment adopts the task division plan of 5-way 1-shot and 5-way 5-shot, trains the network respectively, and evaluates the model performance on the test set. Table 3 shows the mAP evaluation results of the rail surface defect dataset. Figure 8 For the effect diagram of the detection samples.
[0121] Table 3 mAP Evaluation Results of Rail Dataset
[0122]
[0123]
[0124] As can be seen from Table 3, as the sample size increases, in the 5-way 5-shot case, the performance of the method of the present invention is higher than that of other few-shot detection methods including the RepMet network. The performance of the same network in the 5-way 5-shot case is higher than that in the 5-way 1-shot case, indicating that more samples can enable the network to learn more information and the network performance can also be greatly improved.
[0125] The experimental results show that the performance of this method on the miniImageNet public dataset is better than that of other models, with an average accuracy of 73.42% under 5-way 5-shot; on the rail surface defect dataset, the average accuracy of 5-way 5-shot can reach up to 63.29%. When experimenting on the miniImageNet public dataset and the rail surface defect detection dataset, the present invention has better performance improvement compared with other few-shot methods. With the increase of samples, there is further room for improvement in the method of the present invention.
[0126] The applicant of the present invention has made a detailed description and illustration of the embodiments of the present invention in combination with the accompanying drawings of the specification. However, those skilled in the art should understand that the above embodiments are only the preferred implementation schemes of the present invention, and the detailed description is only to help readers better understand the spirit of the present invention, rather than a limitation on the protection scope of the present invention. On the contrary, any improvement or modification based on the spirit of the present invention should fall within the protection scope of the present invention.
Claims
1. A method for detecting rail surface defects based on an improved feature pyramid network and metric learning, characterized in that, It includes the following steps: (1) Add the deformable convolution DC and the convolutional attention module CBAM to the ResNet50 network to obtain the improved Feature Pyramid Network FPN as the rail surface defect feature extraction network; For the improved Feature Pyramid Network FPN, first perform deformable convolution DC on the image, focus on the rail defect features through the convolutional attention module CBAM, after the 1×1 convolution operation, upsample the feature map of the previous layer and superimpose and fuse the feature map of the current layer, and obtain feature maps of different scales (p2, p3, p4, p5) after the 3×3 convolution. Get the feature map ROI from the corresponding Region Proposal Network RPN, and finally perform ROI pooling to complete the feature extraction of the rail surface defects; The deformable convolution DC adds a learnable offset ΔP to each point of the convolution operation n , after learning the target through the offset, the size and position of the deformable convolution kernel are adjusted according to the current image, and the sampling points of the convolution kernel at different positions change adaptively according to the image content; the deformable convolution DC is defined as: where y(P0) represents the value of the P0 point in the image after the convolution operation; P n represents all positions in R; (2) Pre-train the improved Feature Pyramid Network FPN on the MS COCO dataset, and then transfer the trained network parameters and network model to the rail surface defect detection model; (3) Use the improved FPN network to extract and locate the rail surface defect features from the rail surface defect dataset to obtain the feature map ROI; (4) Input the feature map ROI into the RepMet network. The RepMet network includes a distance metric learning DML embedding module; use a multi-modal network structure to obtain the ROI features to get the corresponding feature vectors, and then calculate the distances between the representatives of each modality and the feature vectors obtained by the DML embedding module for defect classification and recognition; Specifically, it includes: (4.1) Use the multi-modal network structure to perform another convolution operation to extract the defect ROI features, and then obtain the feature vectors E corresponding to each modality through the fully connected layer. ij Each modality network consists of 2 fully connected layers, and the ReLU activation function is added after each fully connected layer. (4.2) The feature vector E is obtained after non-linear processing by the DML embedding module i ; The DML embedding module consists of 3 fully connected layers, and the ReLU activation function is added after each fully connected layer; (4.3) Calculate the eigenvector E i The distance between the eigenvector E ij represents the distance between the i-th class extracted by FPN in the embedded eigenvector E i and each modal eigenvector E ij These distances are used to calculate the probability of a given ROI in each class i in each modality j, so as to perform the classification determination of the ROI.
2. The rail surface defect detection method based on the improved feature pyramid network and metric learning according to claim 1, wherein The convolutional attention module CBAM includes a channel attention module CAM and a spatial attention module SAM. These two modules focus on the important features of the image in the channel dimension and the spatial dimension respectively, aggregate the channel information of the feature map through two pooling operations, and then connect and convolve them through a standard convolutional layer; Given the input feature map F ∈ R C×H×W , through the one-dimensional channel attention map M c ∈ R C×1×1 and the two-dimensional spatial attention map M s ∈ R 1×H×W Serial calculation yields the output F", that is: Among them, F' is the feature of the input feature passing through the channel attention module; F" is the feature finally output after passing through the attention module; represents element-wise multiplication; C is the number of channels; H is the height of the feature map; W is the width of the feature map.
3. The rail surface defect detection method based on the improved feature pyramid network and metric learning according to claim 1, characterized in that, When performing RPN processing on the feature map, optimize the bounding box with the bounding box regression loss function.
4. The rail surface defect detection method based on the improved feature pyramid network and metric learning according to claim 1, characterized in that In step (4.3), E i and E ij The distance calculation formula between them is: where n represents the number of categories, D j (E i ) is used to calculate and extract the probability of each category i in each modality j, expressed as the formula: Where σ is the variance.
5. The rail surface defect detection method based on the improved feature pyramid network and metric learning according to claim 1, characterized in that It also includes the calculation of the posterior probability of the defect ROI category and the calculation of the posterior probability of the background category.
6. The rail surface defect detection method based on the improved feature pyramid network and metric learning according to claim 1, wherein When performing classification and recognition, optimize with loss functions, using the embedding loss function and the cross-entropy loss function.
Citation Information
Patent Citations
Multi-scale target detection model method based on metric learning
CN111652216A
Unsupervised remote sensing image change detection method, storage medium and computing device
CN112541904A