Track surface defect detection and identification method based on improved YOLOv9
By improving the YOLOv9 detection model, combining online reparameterized convolution, SCDown downsampling, grouped convolution and ECA attention mechanism, the problem of accuracy and efficiency of track surface defect detection is solved, and more efficient track surface defect recognition is achieved.
Patent Information
- Application Number
- CN202510642204.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-07-04
AI Technical Summary
The existing track surface defect detection methods have problems such as low detection accuracy and low identification efficiency, especially in automated inspections, which are affected by weather and vehicle speed, and manual reviews can easily lead to missed inspections or missed inspections.
Using the improved YOLOv9 detection model, a small object detection layer was added through online reparameterized convolution OREPA, SCDown downsampling module, grouping convolution and ECA attention mechanism, and combined with data enhancement technology, a multi-scale feature map was constructed to improve detection accuracy and efficiency.
It significantly improves the detection accuracy and efficiency of track surface defects, can more accurately identify small targets at long distances, deal with overlapping problems between targets, and enhances the adaptability and scalability of the model.
Smart Images

Figure CN120259846A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of railway surface defect detection, and particularly to a method for detecting and identifying railway surface defects based on improved YOLOv9. Background Art
[0002] In modern society, with the rapid economic growth and rising transportation demands, the rail transit system has been continuously expanding and upgrading. The safe operation of the railway depends on the integrity and reliability of its key components. As an important part of the rail transit system, the railway surface undertakes the important task of supporting the train operation. The good condition of the railway surface is the key to ensuring the stable and safe operation of the rail transit system.
[0003] Traditional methods for detecting railway surface defects usually rely on professional personnel to conduct visual inspections on site. However, railway lines often traverse diverse terrains, including mountains, tunnels, and urban areas, which makes the detection work time-consuming and laborious, and there are certain safety risks. Moreover, when using automated inspection vehicles to perform inspection tasks, some technical challenges are also encountered. For example, weather conditions and vehicle driving speeds may affect the quality and resolution of the collected images, resulting in the inability to comprehensively and clearly identify tiny defects on the railway surface in real time, such as cracks, wear, or potholes. Usually, it is necessary to manually review the pictures collected by the inspection vehicle to find defects, but the long-term review work is prone to fatigue, leading to missed or misdetected defects, and the detection results are greatly affected by the experience and subjective judgment of quality inspection personnel. Summary of the Invention
[0004] To solve the problems of low accuracy and low efficiency in detecting and identifying railway surface defects, this application provides a method for detecting and identifying railway surface defects based on improved YOLOv9.
[0005] The above-mentioned invention objective of this application is achieved through the following technical solutions:
[0006] A method for detecting and identifying railway surface defects based on improved YOLOv9, including the steps of:
[0007] S1: Collect high-definition images of the railway surface through the China Railway High-Speed Comprehensive Inspection Train, and obtain the initial data set after quality screening;
[0008] S2: Use the LabelImg tool to annotate the initial data set, construct an annotated data set including "normal" and "defect" categories, and divide it into a training set, a validation set, and a test set according to a preset ratio (8:1:1);
[0009] S3: Construct an improved YOLOv9 detection model, train and optimize the improved YOLOv9 detection model, and expand the training set through data augmentation techniques;
[0010] S4: Evaluate the model performance using the test set, and output the precision, recall, average precision, and intersection over union data.
[0011] In a preferred example of the present application: The process of constructing the improved YOLOv9 detection model in step S3 specifically includes the steps:
[0012] S31: In the RepNCSPELAN4 module, use the online reparameterized convolution OREPA structure to replace the original RepNCSP module;
[0013] S32: Use the SCDown downsampling module to reconstruct the network downsampling process;
[0014] S33: Integrate grouped convolution and ECA attention mechanism in the SPPELAN module;
[0015] S34: Add a small object detection layer with a resolution of 160×160 in the prediction stage of the model to capture small objects at a long distance.
[0016] In a preferred example of the present application: The online reparameterized convolution OREPA structure includes:
[0017] Adopt a composite structure containing multiple parallel convolution branches in the training stage;
[0018] Generate an equivalent single convolution layer through parameter fusion in the inference stage.
[0019] In a preferred example of the present application: The SCDown downsampling module includes:
[0020] Adjust the number of channels through 1×1 pointwise convolution;
[0021] Perform downsampling through a 3×3 depthwise separable convolution layer with a stride of 2.
[0022] In a preferred example of the present application: Step S33 specifically includes:
[0023] S331: Divide the input feature map and convolution kernel of the SPPELAN module into multiple groups to implement, and each group performs convolution operations independently;
[0024] S332: Perform global average pooling on the input feature map to obtain a channel descriptor x with a dimension of (C, 1, 1), where C is the number of channels;
[0025] S333: Use a one-dimensional convolution of size k to capture local cross-channel information;
[0026] S334: The output of the one-dimensional convolution passes through a Sigmoid function to generate a weight between 0 and 1, which is used to adjust the importance of each channel;
[0027] S335: Place the channel descriptor x and the weight w of the one-dimensional convolution into the preset formula y = σ(w * x) to calculate and generate the attention weight y, where * represents the convolution operation, σ is the Sigmoid activation function, and the dimension of y is also (C, 1, 1);
[0028] S336: Multiply the attention weight by the original input feature map to obtain the final output feature map.
[0029] In a preferred example of the present application: The intersection over union is used to measure the overlap degree between the predicted bounding box and the ground truth bounding box. The calculation formula of the intersection over union is IoU = TP / (TP + FN + FP), where TP represents the number of correctly detected targets, FN represents the number of missed targets, and FP represents the number of false detections.
[0030] In a preferred example of the present application: The calculation formula of the precision is Precision = TP / (TP + FP) = TP / all detections, where all detections is the predicted area of the algorithm's detection bounding box;
[0031] The calculation formula of the recall is Recall = TP / (TP + FN) = TP /
[0032] allground trusts, where allground trusts is the actual area of the actual annotation bounding box.
[0033] In a preferred example of the present application: The calculation formula of the average precision is AP =
[0034]
[0035] where r represents the recall rate, ρ(r) is the precision value of the recall rate r, and ρ interp (r n+1 ) is the maximum precision value in ρ(r) when the recall rate is greater than or equal to r.
[0036] In summary, the present application includes at least one of the following beneficial technical effects:
[0037] 1. In OREPA, a linear scaling layer is used to replace the traditional non - linear layer (such as the batch normalization BN layer) during training. This approach not only maintains optimization diversity but also enhances the representation ability. A complex structure is used during the training phase to improve performance, and during the inference phase, these complex structures are compressed into a single linear layer through equivalent transformation, thus not introducing additional inference time costs.
[0038] 2. Grouped convolution decomposes the convolution operation into multiple groups. Each group only needs to process a part of the input feature map, thus significantly reducing the computational complexity. Since each group only corresponds to a part of the input and output channels, grouped convolution can greatly reduce the number of model parameters, help prevent overfitting, and reduce memory requirements. Grouped convolution also enables each group to learn different aspects of the input features, which helps maintain feature diversity and avoid information loss. The independent grouping feature of grouped convolution makes it easier to perform parallel computing on multi - core processors, further improving the computational efficiency. By using different - sized convolution kernels in different groups, grouped convolution helps the network learn multi - scale features simultaneously.
[0039] 3. The core idea of the ECA attention mechanism is to capture the dependencies between channels without significantly increasing computational complexity. It achieves this through a simple per - channel learning strategy based on local cross - channel interactions, which significantly improves the recognition efficiency.
[0040] 4. In the prediction stage of the improved YOLOv9 detection model of this application, a small - object detection layer with a resolution of 160×160 is added. This additional detection layer is specifically used to capture small objects at a long distance. It has a larger receptive field and can more effectively identify and locate these small - sized objects. By integrating the original three detection layers and the newly added small - object detection layer, the model now has four detection layers with different receptive fields, thus being able to more comprehensively cover objects of various sizes. This multi - scale detection strategy not only improves the detection accuracy of small objects but also enhances the model's adaptability and scalability to scales. In practical applications, this means that the model can more accurately detect small objects at a long distance and can better handle the overlap problem between objects, thus comprehensively improving the effect of object detection. Brief Description of the Drawings
[0041] Figure 1 It is the overall flowchart of a method for detecting and identifying rail surface defects based on the improved YOLOv9 in an embodiment of this application.
[0042] Figure 2 It is the schematic diagram of the OREPA network structure in an embodiment of this application.
[0043] Figure 3 It is a schematic diagram of the modified network structure of RepNCSP in the embodiments of this application;
[0044] Figure 4 It is a schematic diagram of the convolution operation of the SCDown module in the embodiments of this application, where k is the size of the convolution kernel, s is the stride, and p is the grayscale padding;
[0045] Figure 5 It is a schematic diagram of the network structure of the ECA attention mechanism in the embodiments of this application. Detailed implementation manners
[0046] The following further elaborates on this application in conjunction with the accompanying drawings.
[0047] Refer to Figure 1 , in one embodiment, this application discloses an orbit surface defect detection and recognition method based on improved YOLOv9, which specifically includes the following steps:
[0048] S1: Collect high-definition images of the orbit surface through the China Railway High-Speed Comprehensive Inspection Train, and obtain an initial data set after quality screening;
[0049] Specifically, use the China Railway High-Speed Comprehensive Inspection Train to collect orbit surface defects, capture high-definition image data, and then conduct strict quality screening and preprocessing on the obtained images to remove those with poor image quality, such as blurred, overexposed, or jittery images.
[0050] S2: Use the LabelImg tool to label the initial data set, construct a labeled data set including "normal" and "defect" categories, and divide it into a training set, a validation set, and a test set according to a preset ratio of (8:1:1);
[0051] Specifically, use the labeling tool LabelImg to label the images collected by the China High-Speed Railway Comprehensive Inspection Train, and label the orbit surface conditions in these images as "normal" (as negative samples) or "defect" (as positive samples). After completing the labeling task, integrate the labeled data and the original images into a complete data set. Given the limited number of actual defect samples collected, to ensure the balance of positive and negative samples in the data set and accumulate enough defect samples for model training, we first screened the data set and removed some redundant negative samples. Subsequently, we divided the data set into a training set, a validation set, and a test set according to a ratio of 8:1:1.
[0052] S3: Construct an improved YOLOv9 detection model, train and optimize the improved YOLOv9 detection model, and expand the training set through data augmentation techniques;
[0053] Specifically, considering both the detection speed and detection accuracy, YOLOv9 is selected as the basic model, and it is mainly improved in the following aspects: 1. Improve RepNCSP in the RepNCSPELAN4 module of yolov9 by introducing the online reparameterized convolution OREPA; 2. Improve the downsampling module of yolov9 using SCDown; 3. Optimize SPPELAN by using grouped convolution and ECA attention mechanism; 4. Add a small object detection layer to increase the detection of small objects, and finally obtain an orbital surface defect detection model based on YOLOv9. On this basis, in order to further expand the training data, we implemented data augmentation techniques on the training set and validation set, including geometric transformations (such as rotation, flipping, cropping) and color adjustments (such as adjusting hue, saturation, and brightness), to increase the diversity of the data and the generalization ability of the model.
[0054] S4: Use the test set to evaluate the model performance and output precision, recall, mean average precision, and intersection over union data.
[0055] Specifically, after the model training is completed, in order to evaluate the performance of the model, we need to evaluate the images in the test set. At this time, multiple metrics can be used to measure the accuracy of the model, including precision, recall, and mean average precision (mAP). These metrics can comprehensively reflect the performance of the model in the object detection task. Intersection over union (IoU) is another important evaluation metric, which measures the overlap degree between the prediction box and the ground truth box. Through the comprehensive evaluation of metrics such as precision, recall, mAP, and IoU, we can comprehensively understand the performance of the model on the test set, further optimize the model performance, and improve its accuracy and reliability in practical applications.
[0056] In one embodiment, the process of constructing the improved YOLOv9 detection model in step S3 specifically includes the steps:
[0057] S31: Replace the original RepNCSP module with the online reparameterized convolution OREPA structure in the RepNCSPELAN4 module;
[0058] S32: Reconstruct the network downsampling process using the SCDown downsampling module;
[0059] S33: Integrate grouped convolution and ECA attention mechanism in the SPPELAN module;
[0060] S34: Add a small object detection layer with a resolution of 160×160 during the prediction stage of the model to capture small objects at a long distance.
[0061] In this embodiment, Online Convolutional Re-parameterization (OREPA) is an optimization technique for deep learning model training, aiming to reduce training costs and complexity while improving the performance of the model. Spatial-channel decoupled downsampling (SCDown) achieves efficient downsampling by decoupling space and channels. Group Convolution is an effective convolution operation optimization technique in deep learning, which reduces the computational amount and model parameters while maintaining the performance of the network. The Efficient Channel Attention (ECA) mechanism is a lightweight channel attention module used in computer vision tasks, which aims to improve the performance of the network with less computational cost. In the object detection task, objects of different sizes in the image exhibit different characteristics.
[0062] Specifically, objects closer to the camera occupy more pixels in the image. Therefore, the model can extract richer and more detailed feature information from these objects, and the noise is relatively low, which helps to improve the detection accuracy. However, for objects far from the camera, their pixel representation in the image is less, resulting in limited available feature information and a higher noise level, which will have an adverse impact on the accuracy of the detection algorithm. To address this challenge, the YOLOv9 model adopts a multi-scale feature map strategy to detect objects of different sizes. In the default configuration of the model, three detection layers are set, targeting objects with pixel sizes of 32×32, 16×16, and 8×8 respectively. These detection layers are obtained by downsampling the original 640×640 pixel image by 8, 16, and 32 times, forming feature maps with pixel sizes of 20×20, 40×40, and 80×80. The improved YOLOv9 detection model of this application adds a small object detection layer with a resolution of 160×160 during the prediction stage of the model. This additional detection layer is specifically used to capture small objects at a long distance. It has a larger receptive field and can more effectively identify and locate these small-sized objects. By integrating the original three detection layers and the newly added small object detection layer, the model now has four detection layers with different receptive fields, thus being able to more comprehensively cover objects of various sizes. This multi-scale detection strategy not only improves the detection accuracy of small objects but also enhances the adaptability and scalability of the model to scales. In practical applications, this means that the model can more accurately detect small objects at a long distance and can better handle the overlapping problem between objects, thus comprehensively improving the effect of object detection.
[0063] Refer to Figure 2 and Figure 3 In one embodiment, the online reparameterized convolution OREPA structure includes:
[0064] Adopt a composite structure containing multiple parallel convolutional branches during the training phase;
[0065] Generate an equivalent single convolutional layer through parameter fusion during the inference phase.
[0066] In this embodiment, OREPA includes two main phases. 1. Introduce a special linear scaling layer to optimize the performance of the online block: In OREPA, use a linear scaling layer to replace the traditional non-linear layer during training (such as the batch normalization BN layer). This not only maintains optimization diversity but also enhances the representation ability. 2. Reduce the training overhead by compressing complex training-time modules into a single convolution: Use complex structures during the training phase to improve performance, and during the inference phase, compress these complex structures into a single linear layer through equivalent transformation, thus not introducing additional inference time costs.
[0067] Refer to Figure 4 In one embodiment, the SCDown downsampling module includes:
[0068] Adjust the number of channels through 1×1 pointwise convolution;
[0069] Perform downsampling through a 3×3 depthwise separable convolutional layer with a stride of 2.
[0070] In this embodiment, the first convolutional operation mainly realizes the conversion of channels, that is, adjusts the number of channels through 1x1 pointwise convolution. The second convolutional operation performs spatial downsampling through 3x3 depth convolution, which halves the size of the feature map. Such a design reduces the computational cost while maximizing the retention of information. The design of SCDown enables the model to maintain high performance while reducing the consumption of computing resources. This decoupling process of space and channels helps the model to more efficiently capture and retain key features when dealing with targets of different scales, and is one of the key factors for reducing computational complexity and resource consumption.
[0071] Refer to Figure 5 In one embodiment, step S33 specifically includes:
[0072] S331: Divide the input feature map and convolutional kernel of the SPPELAN module into multiple groups for implementation, and each group performs convolutional operations independently;
[0073] S332: Perform global average pooling on the input feature map to obtain a channel descriptor x with a dimension of (C, 1, 1), where C is the number of channels;
[0074] S333: Use a one-dimensional convolution of size k to capture local cross-channel information;
[0075] S334: The output of the one-dimensional convolution passes through a Sigmoid function to generate a weight between 0 and 1 for adjusting the importance of each channel;
[0076] S335: Place the channel descriptor x and the weight w of the one-dimensional convolution into the preset formula y = σ(w * x), and calculate to generate the attention weight y, where * represents the convolution operation, σ is the Sigmoid activation function, and the dimension of y is also (C, 1, 1);
[0077] S336: Multiply the attention weight by the original input feature map to obtain the final output feature map.
[0078] In this embodiment, in the standard convolution operation, the convolution kernel convolves with each channel of the input feature map, and then the results of all convolutions are accumulated to form a channel of the output feature map. In grouped convolution, the input feature map and the convolution kernel are divided into several groups, and the convolution operations within each group are independent. For example, if the input feature map has C channels, the output feature map has N channels, and the convolution is divided into G groups, then each group will only contain C / G input channels and N / G output channels. Grouped convolution decomposes the convolution operation into multiple small groups, and each small group only needs to process a part of the input feature map, thus significantly reducing the computational complexity. Since each group only corresponds to a part of the input and output channels, grouped convolution can greatly reduce the number of model parameters, help prevent overfitting, and reduce memory requirements. Grouped convolution also enables each group to learn different aspects of the input features, which helps maintain the diversity of features and avoid information loss. The independent grouping feature of grouped convolution makes it easier to perform parallel computing on multi-core processors, further improving the computational efficiency. By using different-sized convolution kernels in different groups, grouped convolution helps the network learn multi-scale features simultaneously. By using different-sized convolution kernels in different groups, grouped convolution helps the network learn multi-scale features simultaneously. The core idea of the ECA attention mechanism is to capture the dependencies between channels without significantly increasing the computational complexity. It achieves this through a simple per-channel learning strategy based on local cross-channel interactions, rather than global average pooling like SENet (Squeeze-and-Excitation Networks).
[0079] Specifically, it is achieved by dividing the input feature map and the convolution kernel into multiple groups, and each group performs convolution operations independently; global average pooling is performed on the input feature map to obtain a channel descriptor with a dimension of (C, 1, 1), where C is the number of channels; then, a one-dimensional convolution (1D Conv) with a size of k is used to capture local cross-channel information, and the size k of the convolution kernel is a learnable parameter or can be preset in advance; the output of the one-dimensional convolution passes through a Sigmoid function to generate a weight between 0 and 1, which is used to adjust the importance of each channel; the formula of the ECA module can be expressed as y = σ(w * x); the final output feature map is obtained by multiplying the attention weight by the original input feature map.
[0080] In one embodiment, the intersection over union is used to measure the overlap degree between the predicted bounding box and the ground truth bounding box, and the calculation formula of the intersection over union is IoU = TP / (TP + FN + FP), where TP represents the number of correctly detected targets, FN represents the number of missed targets, and FP represents the number of false detections.
[0081] In one embodiment, the calculation formula of the precision is Precision = TP / (TP +
[0082] FP) = TP / alldetections, where alldetections is the predicted area of the algorithm's detected bounding box;
[0083] The calculation formula of the recall is Recall = TP / (TP + FN) = TP / allground trusts, where allground trusts is the actual area of the actual annotation bounding box.
[0084] In one embodiment, the calculation formula of the average precision is
[0085] where r represents the recall, ρ(r) is the precision value of the recall r, and ρ interp (r n+1 ) is the maximum precision value in ρ(r) when the recall is greater than or equal to r.
[0086] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. An orbit surface defect detection and recognition method based on improved YOLOv9, characterized in that, Including the steps: S1: Collect high-definition images of the track surface by the CRH comprehensive inspection train, and obtain the initial data set after quality screening; S2: Use the LabelImg tool to annotate the initial data set, construct an annotated data set including "normal" and "defect" categories, and divide it into a training set, a validation set, and a test set according to a preset ratio (8:1:1); S3: Construct an improved YOLOv9 detection model, train and optimize the improved YOLOv9 detection model, and expand the training set through data augmentation techniques; S4: Use the test set to evaluate the model performance, and output precision, recall, average precision, and intersection over union data.
2. The method for detecting and identifying rail surface defects based on improved YOLOv9 according to claim 1, characterized in that The process of constructing the improved YOLOv9 detection model in step S3 specifically includes the steps: S31: In the RepNCSPELAN4 module, use the online reparameterized convolution OREPA structure to replace the original RepNCSP module; S32: Use the SCDown downsampling module to reconstruct the network downsampling process; S33: In the SPPELAN module, fuse grouped convolution and ECA attention mechanism; S34: Add a small target detection layer with a resolution of 160×160 in the prediction stage of the model to capture small targets at a long distance.
3. The method for detecting and identifying rail surface defects based on improved YOLOv9 according to claim 2, wherein, The online reparameterized convolution OREPA structure includes: Adopt a composite structure containing multiple parallel convolution branches in the training stage; Generate an equivalent single convolution layer through parameter fusion in the inference stage.
4. A method for detecting and identifying rail surface defects based on improved YOLOv9 according to claim 2, characterized in that, The SCDown downsampling module includes: Adjust the number of channels through 1×1 pointwise convolution; Realize downsampling through a 3×3 depthwise separable convolution layer with a stride of 2.
5. The method for detecting and identifying rail surface defects based on improved YOLOv9 according to claim 2, characterized in that, Step S33 specifically includes: S331: Divide the input feature map and convolution kernel of the SPPELAN module into multiple groups to implement, and each group performs convolution operations independently; S332: Perform global average pooling on the input feature map to obtain a channel descriptor x with a dimension of (C, 1, 1), where C is the number of channels; S333: Use a one-dimensional convolution of size k to capture local cross-channel information; S334: The output of the one-dimensional convolution passes through a Sigmoid function to generate a weight between 0 and 1 for adjusting the importance of each channel; S335: Place the channel descriptor x and the weight w of the one-dimensional convolution in the preset formula y = σ(w * x), calculate and generate the attention weight y, where * represents the convolution operation, σ is the Sigmoid activation function, and the dimension of y is also (C, 1, 1); S336: Multiply the attention weight by the original input feature map to obtain the final output feature map.
6. According to the method for detecting and identifying track surface defects based on improved YOLOv9 described in claim 1, the intersection over union is used to measure the overlap degree between the predicted box and the ground truth box, and is characterized in that: The calculation formula of the intersection over union is IoU = TP / (TP + FN + FP), where TP represents the number of correctly detected targets, FN represents the number of missed targets, and FP represents the number of false detections.
7. A method for detecting and identifying rail surface defects based on improved YOLOv9 according to claim 6, characterized in that: The formula for calculating the precision is Precision = TP / (TP + FP) = TP / all detections, where all detections is the predicted area of the algorithm detection box; The formula for calculating the recall is Recall = TP / (TP + FN) = TP / all ground trusts, where all ground trusts is the actual area of the actual annotation box.
8. A method for detecting and identifying rail surface defects based on improved YOLOv9 according to claim 1, characterized in that: The calculation formula for the average precision rate is where r represents the recall rate, ρ(r) is the precision value of the recall rate r, and ρ interp (r n+1 ) is the maximum precision value among the corresponding precision values ρ(r) when the recall rate is greater than or equal to r.
Citation Information
Cited By
Method for detecting structural member problem based on YOLOv3 model
CN121708020A
A method for detecting structural component problems based on the YOLOv3 model
CN121708020B