Railway track defect detection method based on MFGA-YOLO model
By improving the backbone structure and neck structure of the YOLO11N model, combining adaptive fine-grained channel attention and Wise-ShapeIoU loss function, the accuracy and efficiency of railway track defect detection are improved, and the problem of insufficient accuracy in track defect detection of YOLO algorithm is solved.
Patent Information
- Application Number
- CN202510265939.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-07
AI Technical Summary
The existing YOLO algorithm fails to fully utilize the characteristic information related to track defects in railway track defect detection, resulting in low detection accuracy and problems of missed detection and missed detection.
The railway track defect detection method based on the MFGA-YOLO model is adopted, and the backbone structure of the YOLO11N model is replaced by a hybrid aggregation network MANet, and the adaptive fine-grained channel attention mechanism AFGCA is added, and the fast convolutional gating unit Faster-CGLU is used in the neck structure, and the model weight is optimized in combination with the Wise-ShapeIoU loss function to improve detection accuracy and efficiency.
It significantly improves the accuracy and efficiency of railway track defect detection, reduces the cost of manual inspection, and enhances the model's ability to pay attention to key features and adaptability to complex environments.
Smart Images

Figure CN120298307A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular to a method for detecting railway track defects based on the MFGA-YOLO model. Background Art
[0002] With the continuous progress of social economy, the development of rail transit systems not only marks the acceleration of the urbanization process, but also reflects the increasing demand of the public for efficient and environmentally friendly travel modes. However, rail transit faces many challenges during operation; for example, subway trains operate in underground tunnels with poor lighting conditions for a long time, and EMU trains operate under various adverse weather conditions, and natural factors such as rain, snow, and fog also have a serious impact on train maintenance; at the same time, high-intensity and high-frequency operations increase the risk of equipment wear. Therefore, ensuring the safety and reliability of subway trains has become particularly crucial, and an important link among them is to effectively detect track defect faults.
[0003] Traditional railway track defect detection methods often rely on manual inspections and sensors. Manual inspections are not only inefficient and costly but also difficult to ensure accuracy in complex and variable environments. Although sensor-based detection methods can replace manual inspections, they also have significant drawbacks. For example, natural factors such as temperature and humidity have a great impact on the detection results of sensors. Therefore, in order to reduce railway maintenance costs and improve fault repair efficiency, it has become an inevitable trend in the industry to use advanced technical means to achieve automated and intelligent railway track defect detection. With the rapid development of deep learning, intelligent detection technology has become increasingly mature, and deep learning technology has performed very well in image classification and object detection. As a result, many very classic object detection algorithms have emerged. For example, the RCNN algorithm and Fast R-CNN algorithm proposed by Girshick et al. in the early days opened the precedent of two-stage object detection. The idea of RCNN is to first generate candidate regions and then classify samples through a convolutional neural network, using a deep convolutional neural network for image detection. Then, He et al. proposed the Faster R-CNN algorithm based on RCNN, which greatly improved the speed and accuracy of object detection and also simplified the model training and deployment process. It has been widely used due to its high accuracy and better performance. However, although these two-stage object detections have high detection accuracy, their real-time performance is very low, so their development in the industrial field is very limited. Due to the obvious drawback of low real-time performance of two-stage object detection, Redmon proposed the You Only Look Once (YOLO) object detection algorithm, opening the precedent of single-stage object detection. The YOLO algorithm regards the object detection recognition problem as a regression problem rather than a classification problem, does not need to generate candidate regions, and directly predicts the bounding box and class probability from the image. This feature greatly improves the real-time performance of object detection. However, the early YOLO algorithm sacrificed detection accuracy for real-time performance. Therefore, the YOLOv3 algorithm proposed by Redmon et al. improved the problem of low accuracy of the original YOLO algorithm by adding a residual network structure, adaptive anchor boxes, etc. At the same time, due to the idea of the residual network, the inference speed of the model was greatly accelerated while improving the accuracy.
[0004] Meanwhile, due to the emergence of the Transformer, object detection algorithms based on the Transformer have also emerged one after another. For example, the DETR series of object detection algorithms of Free-NMS have opened up a new way of object detection algorithms. Subsequently, the RT-DETR end-to-end real-time object detection algorithm has emerged. However, due to the relatively immature development of the DETR series of algorithms and their relatively high number of parameters and computational complexity, their deployment in industry is relatively limited. Therefore, the YOLO series of object detection algorithms are very mature and popular in industrial applications. Many researchers have introduced YOLO object detection into industrial fault detection and defect detection. For example, Wei Ruoyu et al. used YOLOv3 for defect detection of track fasteners. Ma Zhipeng et al. used the YOLOv4 algorithm for intelligent detection of the suspension state of rigid catenaries in urban rail transit. Zou Yiming et al. used the YOLOv5 algorithm for detecting the missing bolts of subway vehicle bogies. Chen Ting et al. used the YOLOv7 algorithm for detecting the apparent diseases of subway tunnel linings. Li Xianwang et al. used the YOLOv8 algorithm for detecting the weld defects of subway trains.
[0005] Although the above YOLO algorithms have shown certain competitiveness in track-related defect detection, these object detection algorithms do not fully utilize the feature information related to track defects. For example, there are problems such as lack of consideration of the edge information of track defects and some uncommon defects, resulting in low detection accuracy of track defects and incomplete detection of track defects, and situations such as missed detection and false detection occur.
[0006] Therefore, the fault detection algorithm for railway track defects still needs to be improved, and there is still great potential for the improvement of YOLO-related algorithms. Summary of the Invention
[0007] The purpose of the present invention is to provide a railway track defect detection method based on the MFGA-YOLO model to improve the detection efficiency and accuracy in complex environments such as railway train operation.
[0008] To achieve the above purpose, the present invention provides the following technical solutions:
[0009] On the one hand, the present invention provides a railway track defect detection method based on the MFGA-YOLO model, including the following steps:
[0010] Input the dataset containing railway track defect images into the MFGA-YOLO model for training. After the training is complete, input the collected railway track defect images into the MFGA-YOLO model to detect railway track defects and output the defect detection results.
[0011] The MFGA-YOLO model is based on the YOLO11N model architecture. First, it uses the Mixed Aggregation Network (MANet) to replace some of the C3k2 structures in the backbone of the YOLO11N model, and adds an Adaptive Fine-Grained Channel Attention mechanism (AFGCA) at the end of the backbone.
[0012] Secondly, in the neck structure Neck of the YOLO11N model, it uses the Faster Convolutional Gated Unit (Faster-CGLU) to replace the C3k2 structure, and adds the Adaptive Fine-Grained Channel Attention mechanism (AFGCA) after some of the Faster Convolutional Gated Units (Faster-CGLU).
[0013] Finally, it uses the Wise-ShapeIoU loss function to adjust the adaptive weights of the MFGA-YOLO model.
[0014] In some embodiments, the dataset includes the RSDDs railway track surface defect dataset, the railway track defect detection dataset from GitHub, and the dataset on railway track fastener and bolt anomalies. The RSDDs railway track surface defect dataset contains 195 images, divided into Class I and Class II, and due to the large background noise, defect detection is quite challenging. The railway track defect detection dataset on GitHub contains approximately more than 4,200 images with common railway track defects, and the dataset for railway track fastener and bolt anomaly detection has approximately 7,000 images. These images are manually labeled to ensure that each type of defect is accurately classified.
[0015] In some embodiments, the Mixed Aggregation Network (MANet) includes a 1×1 bypass convolution, a depthwise separable convolution, and a C2f module;
[0016] The 1×1 bypass convolution is used to proofread the channels to adapt to each feature map after convolution;
[0017] The depthwise separable convolution includes a depthwise convolution and a pointwise convolution, which are used to reduce the number of parameters and the computational complexity;
[0018] The C2f module is used to extract features at different levels and degrees of abstraction from the input data, and branch out the input data again to increase the non-linearity of the network to improve the network's ability to model complex data. The C2f module enhances the integration of feature levels of different branches by performing feature concatenation in the channel dimension.
[0019] In some embodiments, the Faster-CGLU (Fast Convolutional Gated Unit) includes the Faster-Block module in FasterNet and the Convolutional GLU module in TransNeXt;
[0020] The Faster-Block module includes a partial convolution module and a hierarchical scaling module. The partial convolution module divides the input channels into two parts. One part of the channels performs convolution operations, and the other part of the channels remains unchanged, thereby reducing the computational amount. The hierarchical scaling module scales the features through trainable scaling parameters to enhance the learning ability of the model;
[0021] The Convolutional GLU module is an improved channel mixer that combines the channel attention mechanism and convolution operations. At the same time, a 3×3 depthwise separable convolution is used before the activation function GELU in the GLU gating branch to enhance the local modeling ability of the model.
[0022] In some embodiments, the Wise-ShapeIoU loss function is as follows:
[0023] L = IoU + d + 0.5 * shape_cost;
[0024] Where
[0025]
[0026] In the formula, d is the distance between the center points of the predicted box and the target box; A inter is the intersection of the target box and the predicted box; A union is the union of the target box and the predicted box; IoU is the intersection over union of the predicted box and the target box; shape_cost is the non-linear transformation of the shape similarity cost, using the exponential function and the fourth power to enhance the penalty for large differences and ensure that the cost is positive; ω w 、ω h are used to measure the relative differences in width and height between the predicted box and the target box.
[0027] On the other hand, the present invention provides a railway track defect detection system based on the MFGA-YOLO model, applying the above method, including:
[0028] Dataset module: A storage module for the railway track defect image dataset, and inputting the railway track defect image dataset into the MFGA-YOLO module;
[0029] Training module: Adopting the Pytorch deep learning framework for training the MFGA-YOLO model;
[0030] MFGA-YOLO Module: It is used to detect whether there are defects in the input railway track defect images. Its architecture is based on the YOLO11N model architecture. First, it uses the Mixed Aggregation Network (MANet) to replace part of the C3k2 structure in the backbone of the YOLO11N model, and adds an Adaptive Fine-Grained Channel Attention mechanism (AFGCA) at the end of the backbone;
[0031] Secondly, in the neck structure Neck of the YOLO11N model, it uses the Faster Convolutional Gated Unit (Faster-CGLU) to replace the C3k2 structure, and adds an Adaptive Fine-Grained Channel Attention mechanism (AFGCA) after some Faster Convolutional Gated Units (Faster-CGLU);
[0032] Finally, it uses the Wise-ShapeIoU loss function to adjust the adaptive weights of the MFGA-YOLO model;
[0033] Image Acquisition Module: It is used to acquire railway track defect images and input the railway track defect images into the MFGA-YOLO module;
[0034] Output Module: It is used to output the detection results of the MFGA-YOLO module.
[0035] Compared with the prior art, the present invention has the following beneficial effects:
[0036] Based on the YOLO11N architecture, the present invention integrates multi-scale features through the Mixed Aggregation Network, accelerates calculations and maintains accuracy through the Faster Convolutional Gated Unit, focuses on key regions through the Adaptive Fine-Grained Channel Attention mechanism, and accurately evaluates target matching through the Wise-ShapeIoU loss function, thus significantly improving the detection performance. By inputting the image to be detected into the model to detect the railway track, it improves the detection efficiency and greatly reduces the cost of manual detection. Description of the Drawings
[0037] Figure 1 It is a schematic diagram of the overall structure of the YOLO11N model;
[0038] Figure 2 It is a schematic diagram of the overall structure of the MFGA-YOLO model in the first embodiment of the present invention;
[0039] Figure 3 It is a schematic diagram of the structure of the Mixed Aggregation Network (MANet) in the first embodiment of the present invention;
[0040] Figure 4 It is a schematic diagram of the structure of the Adaptive Fine-Grained Channel Attention (AFGCA) in the first embodiment of the present invention;
[0041] Figure 5Schematic diagram of the structure of the Faster-CGLU (Fast Convolutional Gated Unit) in Embodiment 1 of the present invention;
[0042] Figure 6 Visualization schematic diagram of the railway track defect detection data set in Embodiment 1 of the present invention;
[0043] Figure 7 Comparison experiment diagram of the railway track fastener and bolt anomaly data set in Embodiment 1 of the present invention. Detailed implementation manners
[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0045] Embodiment 1:
[0046] Please refer to Figures 1 - 5 , a railway track defect detection method based on the MFGA-YOLO model. The data set containing railway track defect images is input into the MFGA-YOLO model for training. After the training is completed, the collected railway track defect images are input into the trained MFGA-YOLO model to detect railway track defects and output the defect detection results.
[0047] The MFGA-YOLO model is based on the YOLO11N model architecture. As Figure 1 shown, the YOLO11N model includes the following modules:
[0048] Input end:
[0049] The Input end in the overall structure of the YOLO11N model mainly preprocesses the input data, normalizes and scales the size to make it adaptable to the input of the network. At the same time, a variety of data augmentation methods are adopted to increase the robustness of the model.
[0050] BackBone structure:
[0051] The Backbone module is mainly used for feature extraction and mainly includes the Conv structure, C3k2 structure, SPPF structure, and C2PSA structure. The Conv structure first uses a 3×3 convolutional kernel and a standard 2D convolution with a stride of 2, then performs batch normalization, and finally uses SiLU as the activation function. The C3k2 structure extends the C2f structure of YOLOv8. The C3k2 structure can be set as the initial C2f block. The C3k2 structure is a faster and more efficient variant of the CSP bottleneck. It uses two convolutions instead of one large convolution, thus accelerating the feature extraction speed. The SPPF (Spatial Pyramid Pooling Fast Module) structure first performs multiple convolutions to reduce the number of channels, then performs max pooling at different scales, and connects the results to capture features at multiple spatial scales. The newly added C2PSA (Cross-Stage Partial Spatial Attention) structure in YOLO11N enhances the spatial attention in the feature map, improving the model's attention to important parts of the image. By spatially pooling features, the model can more effectively focus on specific regions of interest. This structure can add attention to the feature map, helping the model focus on important regions of the image.
[0052] Neck structure:
[0053] The Neck structure is mainly responsible for receiving feature information from the backbone structure, aggregating features of different resolutions, and passing them to the head for prediction. It also uses the faster and more efficient C3k2 structure, which can further enhance the model's attention mechanism and robustness. Through multi-level feature fusion and processing, effective detection of multi-scale targets is achieved. Through convolutional layers, upsampling and downsampling operations, and the concatenation of feature maps, the neck network can capture target information at different scales and pass it to the head network for final prediction.
[0054] Head structure
[0055] The Head structure is the detection head and is set to output predictions from three feature maps (P3, P4, and P5), corresponding to different granularity levels in the image. Specifically, P3 is a feature pyramid with 8x downsampling for processing smaller targets, P4 is a feature pyramid with 16x downsampling for processing medium targets, and P5 is a feature pyramid with 32x downsampling for processing large targets. This method ensures that small objects are detected more finely (P3), while larger objects are captured by higher-level features (P5). Among them, depthwise separable convolution (DWConv) is used on the classification branch. Compared with standard convolution, this depthwise separable convolution performs depth convolution in each channel dimension and then pointwise convolution to restore the feature map size, greatly reducing the computational complexity and model parameters. At the same time, classification loss, regression loss, and DFL (Distribution Focal Loss Module) for bounding box refinement are defined respectively for classification and prediction.
[0056] In view of the problems such as the difficulty in extracting the characteristic information of the subtle defects of railway tracks and the insufficient attention to the key characteristic information, in order to achieve the efficient and high-precision detection of railway track defects, the MFGA-YOLO model of this embodiment improves the relevant structure on the basis of the YOLO11N model.
[0057] As Figure 2 shown, the MFGA-YOLO model uses the mixed aggregation network MANet (Mixed Aggregation Network) in Hyper-YOLO to improve part of the C3k2 structure in the backbone structure of the original YOLO11N model, and adds an adaptive fine-grained channel attention mechanism AFGCA (Adaptive Fine-Grained Channel Attention) to the backbone network (BackBone structure) to improve the Backbone structure in YOLO11N. Its purpose is to efficiently select the key features highly relevant to the target task from multi-scale information and enhance the feature extraction ability of the backbone network. At the same time, the Faster-Block module in FasterNet and the Convolutional GLU module in TransNeXt convolution are introduced to form Faster-CGLU to improve the neck structure of YOLO11N. Some convolutional structures in the Faster-Block module can more effectively extract spatial features by reducing redundant calculations and memory access. It selects the characteristics of a part of the channels for conventional convolution, and the characteristics of the remaining part of the channels remain unchanged, reducing the computational complexity, thus realizing a fast and efficient neural network.
[0058] The Convolutional GLU module adds a 3×3 depth convolution before the gating branch of GLU, so that each token obtains a unique gating signal based on its adjacent fine-grained features, to alleviate the potential model depth degradation and achieve information perception close to human foveal vision. The Wise-ShapeIoU loss function is adopted to enable the model to adaptively adjust the weights, automatically adjust the loss weights of positive and negative samples, and accelerate the training and convergence of the model. The MFGA-YOLO model of the present invention has a better detection effect on its railway track defect-related data set.
[0059] As Figure 3As shown, the Mixed Aggregation Network (MANet) is a network architecture for enhancing feature extraction capabilities. By fusing multiple convolution operations, it improves the diversity and information flow of visual features, thereby enhancing the model's feature representation and semantic depth at different levels. The Mixed Aggregation Network MANet combines three typical convolution variants: 1×1 bypass convolution, depthwise separable convolution, and the C2f module, and improves the diversity and information flow of visual features by using them in combination.
[0060] The 1×1 bypass convolution is used to proofread the channels to adapt to each feature map after convolution. The depthwise separable convolution is divided into depthwise convolution and pointwise convolution. Compared with the traditional standard convolution, the depthwise separable convolution can significantly reduce the number of parameters and computational complexity. Since the traditional convolution needs to apply a large convolution kernel to each position on all input channels, increasing the computational complexity of the model, while the depthwise separable convolution decomposes this process into two steps, and each step uses a smaller convolution kernel.
[0061] The convolution operation of the C2f module extracts features at different levels and abstraction degrees from the input data, and branches out the input data to increase the non-linear ability of the network to improve the network's modeling ability for complex data. The C2f module enhances the integration of feature levels of different branches by performing feature concatenation in the channel dimension.
[0062] As Figure 4 shown, the Adaptive Fine-Grained Channel Attention (AFGCA) mechanism aims to improve the performance of visual tasks by enhancing the model's understanding and processing ability of image features. In traditional convolutional neural networks, all channels are usually treated equally regardless of their importance for a specific task. However, the Adaptive Fine-Grained Channel Attention (AFGCA) mechanism breaks this limitation by introducing a dynamic weight allocation strategy that adjusts the importance of each channel according to its contribution to the current task. This method can not only emphasize the features that are crucial for completing the task, but also suppress or ignore those that are irrelevant or may introduce noise. It can focus more precisely on the key information in the image, thereby improving the model's performance in various visual tasks, making the model more flexible and robust when dealing with different scenarios and tasks. At the same time, through fine-grained attention allocation, it can better capture the detailed features in the image, further improving the model's recognition accuracy and generalization ability, and also enhancing the model's ability to understand and interpret data.
[0063] As Figure 5The figure shows the structural schematic diagram of the Faster Convolutional Gated Unit (Faster-CGLU). The Faster-Block module is a key module in FasterNet, including two key modules: partial convolution and hierarchical scaling. They are used together to reduce redundant calculations and memory access, thereby improving the calculation speed without sacrificing accuracy. Among them, the partial convolution module reduces the amount of calculation by dividing the input channels into two parts, where one part of the channels performs convolution operations and the other part remains unchanged. The hierarchical scaling module scales the features through trainable scaling parameters to enhance the learning ability of the model.
[0064] The Convolutional Gated Unit (CGLU) is an improved channel mixer that combines the channel attention mechanism and convolution operations. A 3×3 depthwise separable convolution is used before the activation function in the GLU gating branch to enhance the local modeling ability of the model.
[0065] The channel attention mechanism can enhance the attention to important features of the model while suppressing unimportant channels, thereby improving the model's perception ability of key information. By learning the importance weights of each channel in the input feature map, the model can better adapt to different input data, thus improving the generalization ability of the model. The computational complexity of this attention mechanism is relatively low and will not significantly increase the computational burden of the model, which can improve the efficiency of the model.
[0066] At the same time, by combining the advantages of the Faster-Block module and the Convolutional Gated Unit CGLU, efficient calculation of the model can be achieved. The partial convolution technology can significantly reduce redundant calculations and memory access. This combination can not only reduce the time of forward propagation and backward propagation, improve the calculation efficiency, but also reduce the model complexity without sacrificing performance, thereby reducing memory occupancy and the demand for computing resources, and enhancing the model's perception ability of the hardware. In addition, the built-in gating mechanism of the Convolutional Gated Unit CGLU helps to alleviate the vanishing gradient problem in deep networks, enhances the model's ability to learn complex non-linear relationships, and ensures good training dynamics during fast calculation.
[0067] The Wise-Power-IOU loss function adjusts the weights of different loss components dynamically, giving different parts of the focus, and automatically adjusts the loss weights of positive and negative samples according to the difficulty of the samples. Combining multi-scale information fusion can better handle multi-scale features, improve the detection ability of small objects, emphasize the information on the edges of the bounding boxes, and contribute to more accurate target localization. Among them, a weighting mechanism is introduced to better estimate the overlap degree between the predicted box and the ground truth box, which can reduce the limitations of the traditional single IoU (Intersection over Union) when the overlap area is small or the position is inaccurate. The Wise-Power-IOU loss function is robust to predicted boxes with different distributions, reduces the influence of common small errors on the boundary on the model training process, and enables the model to learn the features of the target more precisely. It can help improve the detection accuracy of the model in various scenarios, especially in the case of complex backgrounds, the detection effect is very obvious.
[0068] The Wise-ShapeIoU loss function is defined as follows:
[0069] The horizontal and vertical coordinates of the predicted box are: (x 11 , y 11 ), (x 12 , y 12 ), and the horizontal and vertical coordinates of the ground truth box are: (x 21 , y 21 ), (x 22 , y 22 ).
[0070] Calculate the width and height of the predicted box and the ground truth box respectively:
[0071] The width and height of the predicted box:
[0072]
[0073] The width and height of the ground truth box:
[0074]
[0075] In the formula, ∈ is an adjustment coefficient to prevent division by zero.
[0076] Calculate the shape distance weight:
[0077] Width weight:
[0078]
[0079] Height weight:
[0080]
[0081] In the formula, ww and h h are weights, aiming to adapt to target boxes of different shapes. Since different objects may have different aspect ratios, the loss function can more intelligently measure the positional difference between the predicted box and the target box. Especially when dealing with non-square or objects with a large aspect ratio, it can improve the accuracy of the model in locating these objects.
[0082]
[0083] In the formula, c w and c h are the width and height of the rectangle that minimally covers two bounding boxes respectively.
[0084]
[0085] In the formula, c is the diagonal length of the minimum covering rectangle box and is an adjustment coefficient to prevent division by zero.
[0086]
[0087] In the formula, d is the distance between the centers of the predicted box and the target box; d x and d y are used to measure the difference between the centers of two bounding boxes in the x-axis and y-axis directions. Here, the position of the center point is calculated by taking the average, and the square of the difference between them is calculated.
[0088]
[0089] In the formula, h h and w w are used to measure the relative difference in width and height between the predicted box and the target box. After being weighted by h h and w w and then divided by the maximum value of the two, it is used to emphasize the difference in shape.
[0090]
[0091] Shape_cost is a non-linear transformation of the shape similarity cost, using the exponential function and the fourth power to enhance the penalty for large differences and ensure that the cost is positive.
[0092] The regression loss function of the final bounding box is:
[0093] L = IoU + d + 0.5shape_cost(13);
[0094] IoU is the intersection over union of the predicted box and the target box, defined as:
[0095]
[0096] Wherein, is the intersection of the target box and the predicted box, and is the union of the target box and the predicted box.
[0097] The dataset of this embodiment is detected using the RSDDs railway track surface defect dataset, the railway track defect detection dataset from GitHub, and the dataset on railway track fastener and bolt anomalies. For all three datasets, each dataset is divided into a training set, a validation set, and a test set according to the ratio of 7:2:1, so as to systematically evaluate the performance of the MFGA-YOLO model.
[0098] The Python version of the experiment is all based on 3.10.16, the Pytorch version is 2.3.0, and it is trained with CUDA 12.6. During the training process, Batch_size = 32, and 400 epochs of training are performed. The SGD optimizer is used, and the learning rate adopts the cosine annealing algorithm. The input image is 640×640. The specific experimental hardware configuration is shown in Table 1:
[0099] Table 1 Experimental hardware environment configuration table
[0100]
[0101] To verify the effectiveness of the MFGA-YOLO model, a series of mainstream object detection algorithms are selected for comparative analysis on the RSDDS dataset. By comparing precision (P), recall (R), mean average precision (mAP@0.5), F1 score, and the number of parameters as evaluation indicators, the specific results are shown in Table 2:
[0102] Table 2 Comparison table of various indicators of each algorithm model on the RSDDs dataset
[0103]
[0104]
[0105] As can be seen from Table 2, although the MFGA-YOLO model is improved based on the YOLO11N model, its various indicators have basically corresponded to those of the YOLO11s model. Therefore, it can be verified that the various indicators of the MFGA-YOLO model are basically completely superior to the basic model YOLO11N model.
[0106] To illustrate the performance of the MFGA-YOLO model on other datasets, this embodiment conducts training and testing on the railway track defect detection dataset to compare the various indicators of the YOLO11N model and the MFGA-YOLO model. The results on the test set are shown in Table 3.
[0107] Table 3 Comparative Experiment Table of Railway Track Defect Detection Datasets
[0108]
[0109] As can be seen from Table 3, the MFGA-YOLO model shows significant advantages in the performance on the railway track defect detection dataset. Not only does it achieve a precision of 0.975, which is better than 0.959 of YOLO11N, but it also has a higher recall rate of 0.936, greatly reducing the possibility of missed detections. In addition, the MFGA-YOLO model also performs extremely well in the mean average precision (mAP@0.5), reaching 0.979, and in a wider IoU range (mAP@[0.5:0.95]), its mAP value is 0.707, both exceeding the corresponding values of YOLO11N, which proves that it can maintain excellent performance in object detection tasks of different sizes and complexities. Although the MFGA-YOLO model has slightly higher parameter quantity and computational complexity, for scenarios with extremely high requirements for defect detection accuracy, these additional resource consumptions are worthwhile. At the same time, its detection results are visualized, as Figure 6 shown.
[0110] At the same time, experiments are carried out in another railway track fastener and bolt anomaly detection dataset to illustrate that the MFGA-YOLO model can show excellent performance in various defect detections of railway tracks. The results of the comparative experiments on the test set are shown in Table 4 as follows.
[0111] Table 4 Comparative Experiment Table of Railway Track Fastener and Bolt Anomaly Datasets
[0112]
[0113] For the railway track fastener and bolt anomaly dataset, its categories are relatively single and easy to detect. Therefore, both the YOLO11N model and the MFGA-YOLO model show good experimental results, which also shows that the MFGA-YOLO model can show very good performance in defect detection whether under complex conditions or in a relatively simple environment. The visualization of the detection results is as Figure 7 shown.
[0114] Based on the YOLO11N architecture, the present invention significantly improves the detection performance by introducing a hybrid aggregation network to integrate multi-scale features, a fast convolutional gated unit to accelerate calculations and maintain accuracy, an adaptive fine-grained channel attention to focus on key regions, and a Wise-ShapeIoU loss function to accurately evaluate object matching. Inputting the picture to be detected into the model to detect railway tracks improves the detection efficiency and greatly reduces the cost of manual detection.
[0115] Example 2
[0116] A railway track defect detection system based on the MFGA-YOLO model, comprising:
[0117] Dataset module: A storage module for the railway track defect image dataset, and input the railway track defect image dataset into the MFGA-YOLO module.
[0118] Training module: Adopt the Pytorch deep learning framework to train the MFGA-YOLO model.
[0119] MFGA-YOLO module: Used to detect whether there are defects in the input railway track defect images. Its architecture is based on the YOLO11N model architecture. First, use the hybrid aggregation network MANet to replace part of the C3k2 structure in the backbone structure Backbone of the YOLO11N model, and add the adaptive fine-grained channel attention mechanism AFGCA at the end of the backbone structure Backbone.
[0120] Secondly, in the neck structure Neck of the YOLO11N model, use the faster convolutional gated unit Faster-CGLU to replace the C3k2 structure, and add the adaptive fine-grained channel attention mechanism AFGCA after some faster convolutional gated units Faster-CGLU.
[0121] Finally, adopt the Wise-ShapeIoU loss function to adjust the adaptive weights of the MFGA-YOLO model;
[0122] Image acquisition module: Use cameras, high-definition cameras, etc. to collect railway track defect images, and input these railway track defect images to be detected into the MFGA-YOLO module.
[0123] Output module: Used to output the detection results of the MFGA-YOLO module.
[0124] A railway track defect detection system based on the MFGA-YOLO model of the present invention can be installed in a computer device. The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a railway track defect detection program based on the MFGA-YOLO model. Among them, the memory includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. The processor is the control core of the electronic device, connecting various components of the entire computer device through various interfaces and lines, and by running or executing the programs or modules stored in the memory, and calling the data stored in the memory, to perform various functions of the computer device and process data.
[0125] The module described in the present invention refers to a series of computer program segments that can be executed by the processor of a computer device and can complete fixed functions, and are stored in the memory of the computer device.
[0126] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A railway track defect detection method based on the MFGA-YOLO model, characterized in that, It includes the following steps: Input the dataset containing railway track defect images into the MFGA-YOLO model for training. After the training is complete, input the collected railway track defect images into the MFGA-YOLO model to detect railway track defects and output the defect detection results; The MFGA-YOLO model is based on the YOLO11N model architecture. First, use the Mixed Aggregation Network (MANet) to replace part of the C3k2 structure in the Backbone of the YOLO11N model, and add the Adaptive Fine-Grained Channel Attention mechanism (AFGCA) at the end of the Backbone; Secondly, in the Neck structure of the YOLO11N model, use the Faster Convolutional Gated Unit (Faster-CGLU) to replace the C3k2 structure, and add the Adaptive Fine-Grained Channel Attention mechanism (AFGCA) after some of the Faster Convolutional Gated Units (Faster-CGLU); Finally, use the Wise-ShapeIoU loss function to adjust the adaptive weights of the MFGA-YOLO model.
2. The railway track defect detection method based on the MFGA-YOLO model according to claim 1, characterized in that, The dataset includes the RSDDs railway track surface defect dataset, the railway track defect detection dataset from GitHub, and the dataset on railway track fastener and bolt anomalies.
3. A railway track defect detection method based on the MFGA-YOLO model according to claim 1, characterized in that, The Mixed Aggregation Network (MANet) includes a 1×1 bypass convolution, a depthwise separable convolution, and a C2f module; The 1×1 bypass convolution is used to proofread the channels; The depthwise separable convolution includes a depthwise convolution and a pointwise convolution, which are used to reduce the number of parameters and computational complexity; The C2f module is used to extract features of different levels and abstraction degrees from the input data and branch out the input data again.
4. A method for detecting railway track defects based on the MFGA-YOLO model according to claim 1, characterized in that, The Faster Convolutional Gated Unit (Faster-CGLU) includes the Faster-Block module in FasterNet and the Convolutional GLU module in TransNeXt; The Faster-Block module includes a partial convolution module and a hierarchical scaling module. The partial convolution module divides the input channels into two parts. One part of the channels performs convolution operations, and the other part of the channels remains unchanged, thus reducing the computational amount; The hierarchical scaling module scales the features through trainable scaling parameters to enhance the learning ability of the model; The Convolutional GLU module is an improved channel mixer, which combines the channel attention mechanism and convolution operations. At the same time, a 3×3 depthwise separable convolution is used before the activation function GELU of the GLU gating branch to enhance the local modeling ability of the model.
5. A railway track defect detection method based on the MFGA-YOLO model according to claim 1, characterized in that The Wise-ShapeIoU loss function is: L = IoU + d + 0.5shape_cost; Where, shape_cost=(1 - e -ωw ) 4 +(1 - e -ωh ) 4 ; Where d is the distance between the centers of the predicted box and the target box; A inter is the intersection of the target box and the predicted box; A union is the union of the target box and the predicted box; IoU is the intersection over union of the predicted box and the target box; shape_cost is a non-linear transformation of the shape similarity cost, using the exponential function and the fourth power to enhance the penalty for large differences and ensure that the cost is positive; ω w and ω h are used to measure the relative differences in width and height between the predicted box and the target box.
6. A railway track defect detection system based on the MFGA-YOLO model, applying the method according to any one of claims 1-5, characterized in that, It includes: Dataset module: A storage module for the railway track defect image dataset, and input the railway track defect image dataset into the MFGA-YOLO module; Training module: Adopt the Pytorch deep learning framework to train the MFGA-YOLO model; MFGA-YOLO Module: It is used to detect whether there are defects in the input railway track defect images. Its architecture is based on the YOLO11N model architecture. First, it uses the Mixed Aggregation Network (MANet) to replace part of the C3k2 structure in the backbone of the YOLO11N model, and adds an Adaptive Fine-Grained Channel Attention mechanism (AFGCA) at the end of the backbone; Secondly, in the neck structure of the YOLO11N model, it uses the Faster Convolutional Gated Unit (Faster-CGLU) to replace the C3k2 structure, and adds an Adaptive Fine-Grained Channel Attention mechanism (AFGCA) after some of the Faster Convolutional Gated Units (Faster-CGLU); Finally, it uses the Wise-ShapeIoU loss function to adjust the adaptive weights of the MFGA-YOLO model; Image Acquisition Module: It is used to acquire railway track defect images and input the railway track defect images into the MFGA-YOLO module; Output Module: It is used to output the detection results of the MFGA-YOLO module.
Citation Information
Patent Citations
Improved YOLO v8-based mushroom stick cultivation mushroom grading method and system
CN119206711A
Lightweight track surface defect real-time detection method and device
CN119313663A
High-precision fabric defect detection method based on lightweight hybrid aggregation network and enhanced generalized feature fusion
CN119444714A
Method and apparatus for computer-vision-based object detection
US20240169740A1
Cited By
Abrasion measuring method for conductor rail of metro overhead line system
CN120953243A
Railway wagon floor damage fault detection method and device and electronic equipment
CN121661059A
Railway facility inspection method and system integrating GNSS positioning and image recognition
CN121919772A
Railway wagon part anomaly detection method and system
CN122290091A
An anomaly detection system and method based on a large pre-trained model for contrastive language images
CN122574480A