Wheel set tread detection method and system based on improved YOLOv7 model

By improving the YOLOv7 model, GSConv and small-objective enhancement modules are used for feature fusion, and the improved EIoU loss function is used to solve the problems of low tread defect detection efficiency and low small-objective detection accuracy of high-speed train wheels, achieving more efficient and accurate detection effects.

CN120088206APending Publication Date: 2025-06-03ZHENGZHOU UNIVERSITY OF LIGHT INDUSTRY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510120579.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-25
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

In the prior art, high-speed train wheels have low efficiency in detecting tread defects, and small target defects are easily ignored, resulting in low detection accuracy.

Method used

The improved YOLOv7 model is adopted, and the normal convolution is replaced by using GSConv in the neck network and a small-objective enhancement module is added to perform multi-scale feature fusion and channel and spatial attention processing, while the improved EIoU loss function is used.

Benefits of technology

The accuracy and efficiency of wheel-to-wheel tread defect detection are improved, especially in small object detection, and the model's adaptability to complex scenarios is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088206A_ABST
    Figure CN120088206A_ABST
Patent Text Reader

Abstract

The invention discloses a wheel set tread detection method and system based on an improved YOLOv7 model, relates to the technical field of train wheel set fault detection, and is used for solving the problems that in the prior art, high-speed train wheel set tread defect detection efficiency is low, and small target defects are prone to being ignored. The technical key points of the invention comprise: providing a detection model based on an improved Yolov7 algorithm, replacing common convolution by using GSConv in a neck network, and reducing model calculation parameters by using lightweight convolution; a small target enhancement module is added in the neck network, multi-scale feature fusion is carried out on feature maps of different sizes in the backbone network through a multi-scale sub-module, and the fused feature maps and the feature maps before fusion are integrated; the integrated feature map is input into a channel attention sub-module and a space attention sub-module for processing, so that the problem that a small target detection result is not ideal is solved; and the loss function adopts an improved EIoU loss function, so that efficient detection is realized, and the damage classification accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of train wheel set fault detection, and specifically relates to a wheel set tread detection method and system based on an improved YOLOv7 model. Background Technique

[0002] The train wheel set is an important moving and stress-bearing component of high-speed trains. Due to the long-term operation of train wheel sets under harsh conditions for a long time, faults are likely to occur, and its common manifestations are tread peeling, abrasion, and pits. In the statistics of train accidents, those caused by wheel set faults account for 88%, so the efficient detection of wheel set tread defects is a necessary guarantee for the safe operation of trains.

[0003] Traditional methods require manual measurement of the tread one by one, and the information is not reasonably utilized, so efficient detection cannot be achieved. In the field of object detection, machine vision technology is often used for cumbersome detection work to reduce labor waste. However, when traditional machine vision faces the influence from complex scenes, due to serious interference problems such as noise and weather, its accuracy is reduced and the model stability is poor, and its performance advantages are difficult to reflect. The superior characteristics of convolutional neural networks have made them the focus of research in various fields. Among them, the Faster-RCNN series models and the YOLO series models are the most common, which represent high precision and high efficiency respectively. The YOLO series algorithms are widely used in various fields due to their high efficiency.

[0004] At present, object detection algorithms have developed rapidly, but the lack of small object detection accuracy is still a key problem to be solved urgently. In real complex scenes, small objects exist in large numbers. Compared with regular-sized objects, small objects have low pixels, small proportions, are easy to overlap, and are difficult to distinguish; at the same time, changes in illumination, shooting angles, etc. have a more significant impact on small objects, and there is often a lack of sufficient information to distinguish them during the detection process. Summary of the Invention

[0005] In view of the above problems, the present invention proposes a wheel set tread detection method and system based on an improved YOLOv7 model to solve the problems of low detection efficiency for high-speed train wheel set tread defects and easy neglect of small object defects in the prior art.

[0006] According to one aspect of the present invention, a wheel set tread detection method based on an improved YOLOv7 model is proposed, and the method includes:

[0007] Obtain an image data set including defective and non-defective train wheel sets;

[0008] Preprocess the image data set;

[0009] Use the preprocessed image data set to train a wheel set tread detection model based on an improved YOLOv7 model;

[0010] Use the trained wheel tread detection model to detect the image to be detected and obtain the detection result.

[0011] Further, the defects of the train wheelset include wheel tread peeling defects, wheel tread pit defects, and wheel tread bruise defects.

[0012] Further, the preprocessing includes: data augmentation, data enhancement, and class label annotation.

[0013] Further, the data augmentation includes data augmentation by copy-pasting, or training a generative adversarial network using the image dataset, and generating images corresponding to the defects through the generative adversarial network for dataset augmentation; the data enhancement includes: Gaussian blur, affine transformation, brightness transformation, downsampling pixel transformation, flipping transformation, Mixup, and Mosaic for stitching and mixing to enhance the data.

[0014] Further, the improvements of the YOLOv7 model include: using GSConv to replace ordinary convolution in the neck network; adding a small target enhancement module in the neck network, which is used to perform multi-scale feature fusion on feature maps of different sizes in the backbone network through a multi-scale sub-module, and integrating the fused feature map and the feature map before fusion; inputting the integrated feature map into a channel attention sub-module and a spatial attention sub-module for processing; and using an improved EIoU loss function for the loss function of the detection head.

[0015] Further, the expression of the improved EIoU loss function is:

[0016]

[0017] In the formula, x, y represent the coordinates of the center point of the predicted bounding box, x gt and y gt represent the coordinates of the center point of the ground truth bounding box, W g and H g represent the width and height of the minimum bounding box; IoU represents the intersection over union; λ 1 , λ 2 , λ 3 represent hyperparameters that adjust the influence of the center distance and the width and height differences; C d , C w , C h are constants used for normalization respectively; D c represents the distance between the center points of the predicted bounding box and the ground truth bounding box; w diff represents the width difference between the predicted bounding box and the ground truth bounding box; h diff represents the height difference between the predicted bounding box and the ground truth bounding box.

[0018] According to another aspect of the present invention, a wheel tread detection system based on an improved YOLOv7 model is proposed. The system includes:

[0019] A data acquisition module configured to acquire an image data set including defective and non-defective train wheel sets;

[0020] A preprocessing module configured to preprocess the image data set;

[0021] A detection model training module configured to train a wheel tread detection model based on an improved YOLOv7 model using the preprocessed image data set;

[0022] A wheel tread defect detection module configured to detect a to-be-detected image using the trained wheel tread detection model and obtain a detection result.

[0023] Further, the defective train wheel sets in the data acquisition module include wheel tread peeling defects, wheel tread pit defects, and wheel tread abrasion defects.

[0024] Further, the preprocessing in the preprocessing module includes: data augmentation, data enhancement, and class label annotation; the data augmentation includes data augmentation by copy-pasting, or training a generative adversarial network using the image data set, and generating images corresponding to defects through the generative adversarial network for data set augmentation; the data enhancement includes: Gaussian blur, affine transformation, brightness transformation, downsampling pixel transformation, flipping transformation, Mixup, and Mosaic for splicing and mixing to enhance data.

[0025] Further, the improvements in the YOLOv7 model in the detection model training module include: using GSConv to replace ordinary convolution in the neck network; adding a small object enhancement module in the neck network, where the small object enhancement module is used to perform multi-scale feature fusion on feature maps of different sizes in the backbone network, integrate the fused feature maps and the feature maps before fusion, and input the integrated feature maps into a channel attention sub-module and a spatial attention sub-module for processing; the loss function of the detection head uses an improved EIoU loss function; the expression of the improved EIoU loss function is:

[0026]

[0027] In the formula, x,y represent the center point coordinates of the prediction box, x gt , y gt represent the center point coordinates of the ground truth box, W g , H g represent the width and height of the minimum bounding box; IoU represents the intersection over union; λ1 , λ 2 , λ 3 represent hyperparameters that regulate the influence of the center distance and the width-height difference; C d , C w , C h are constants used for normalization respectively; D c represents the distance between the centers of the predicted bounding box and the ground truth bounding box; w diff represents the width difference between the predicted bounding box and the ground truth bounding box; h diff represents the height difference between the predicted bounding box and the ground truth bounding box.

[0028] The beneficial technical effects of the present invention are as follows:

[0029] The present invention proposes a wheel tread detection method and system based on an improved YOLOv7 model, in which an improved model based on the yolov7 algorithm is proposed. On the basis of the original YOLOv7 model, GSConv is used to replace ordinary convolution in the neck network, and lightweight convolution is used to reduce the model's computational parameters; a small target enhancement module is added to the neck network. The small target enhancement module is used to perform multi-scale feature fusion on feature maps of different sizes in the backbone network through multi-scale sub-modules, and integrate the fused feature maps and the feature maps before fusion; the integrated feature maps are input into the channel attention sub-module and the spatial attention sub-module for processing, thus solving problems such as unsatisfactory small target detection results; the loss function of the detection head adopts an improved EIoU loss function to achieve efficient detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] By referring to the accompanying drawings and reading the following detailed description, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown in an exemplary rather than restrictive manner, wherein:

[0031] Figure 1 is a flowchart of a wheel tread detection method based on an improved YOLOv7 model according to an embodiment of the present invention.

[0032] Figure 2 are example diagrams of three types of defects (pits, scratches, and spalls) on the wheel tread in an embodiment of the present invention.

[0033] Figure 3 is an example diagram of copying and pasting to expand the dataset in an embodiment of the present invention.

[0034] Figure 4 is an example diagram of a wheel tread with defects generated using stylegan3 in an embodiment of the present invention.

[0035] Figure 5It is an example diagram of the annotation boxes for various defects in the embodiments of the present invention; among them, blue corresponds to pits, green corresponds to scratches, and yellow corresponds to delamination.

[0036] Figure 6 It is the structural diagram of the improved YOLOv7 model in the embodiments of the present invention.

[0037] Figure 7 It is the structural schematic diagram of the small target enhancement module in the embodiments of the present invention.

[0038] Figure 8 It is the mAP example diagram of the improved YOLOv7 model in the embodiments of the present invention.

[0039] Figure 9 It is the result example diagram of using different models to identify pits in the embodiments of the present invention; among them, (a) corresponds to the improved YOLOv7 model with a confidence level of 0.94; (b) corresponds to the original YOLOv7 model with a confidence level of 0.86; (c) corresponds to the YOLOv5 model with a confidence level of 0.77.

[0040] Figure 10 It is the result example diagram of using different models to identify scratches in the embodiments of the present invention; among them, (a) corresponds to the improved YOLOv7 model with a confidence level of 0.95; (b) corresponds to the original YOLOv7 model with a confidence level of 0.88; (c) corresponds to the YOLOv5 model with a confidence level of 0.79.

[0041] Figure 11 It is the result example diagram of using different models to identify delamination in the embodiments of the present invention; among them, (a) corresponds to the improved YOLOv7 model with a confidence level of 0.97; (b) corresponds to the original YOLOv7 model with a confidence level of 0.89; (c) corresponds to the YOLOv5 model with a confidence level of 0.81.

[0042] Figure 12 It is the structural schematic diagram of a wheel tread detection system based on an improved YOLOv7 model described in the embodiments of the present invention. Detailed implementation manners

[0043] Next, the principles and spirit of the present invention will be described with reference to several exemplary embodiments. It should be understood that these embodiments are given only to enable those skilled in the art to better understand and then implement the present invention, and do not limit the scope of the present invention in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to be able to fully convey the scope of the present disclosure to those skilled in the art.

[0044] Those skilled in the art know that the embodiments of the present invention can be implemented as a system, device, equipment, method, or computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software. In this article, it should be understood that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.

[0045] An embodiment of the present invention proposes a method for detecting the tread of a wheel set based on an improved YOLOv7 model, as Figure 1 shown, the method includes:

[0046] S1. Obtain an image dataset including defective and non-defective train wheel sets; the defective train wheel sets include defects such as tread peeling of the wheel set, pit defects on the tread of the wheel set, and abrasion defects on the tread of the wheel set;

[0047] S2. Preprocess the image dataset;

[0048] S3. Use the preprocessed image dataset to train a wheel set tread detection model based on the improved YOLOv7 model;

[0049] S4. Use the trained wheel set tread detection model to detect the image to be detected and obtain the detection result.

[0050] The method starts from S1. In S1, an image dataset including defective and non-defective train wheel sets is obtained.

[0051] According to the embodiments of the present invention, the representation ability of the deep model highly depends on the diversity of the training data. However, there is currently no complete public dataset of wheel set damage targets. Therefore, when using a wheel set damage detection method based on deep learning, it is necessary to make the dataset by oneself.

[0052] First, various types of defective images are obtained by field shooting with a portable handheld wheel set damage acquisition device, but the number of images obtained is limited and the distribution of each defect type is uneven. The collected data images are three types of defects, namely tread peeling, pits, and abrasions of the wheel set tread, as Figure 2 shown.

[0053] Then S2 is executed. In S2, the image dataset is preprocessed.

[0054] According to the embodiments of the present invention, due to the insufficient content of the dataset, traditional data augmentation and copy-paste are used simultaneously to expand the dataset before training. An example of the copy-paste result is as Figure 3As shown, copy-pasting does not increase the training cost or inference time when augmenting the dataset, and it can be conveniently and effectively integrated into various code libraries. A dataset consisting of 654 images is obtained.

[0055] Subsequently, this dataset is used to train the Generative Adversarial Network - stylegan3 network. Through stylegan3, in response to the problem of dataset imbalance, it selectively generates images corresponding to defects for dataset augmentation, thereby obtaining a dataset composed of 1200 images, as Figure 4 shown.

[0056] The defects are labeled using LabelImg software. Subsequently, the training set, validation set, and test set are divided in the ratio of 8:1:1. Then, data is enhanced through splicing and mixing using Gaussian blur, affine transformation, brightness transformation, downsampling pixel transformation, flipping transformation, Mixup, and Mosaic to further augment the training set to 2160 images. Figure 5 It shows that there are obvious size differences in the defects.

[0057] Then, S3 is executed. In S3, the wheel tread detection model based on the improved YOLOv7 model is trained using the preprocessed image dataset.

[0058] According to the embodiments of the present invention, Yolov7 is a commonly used detection network, and its network structure consists of a Head network, a Backbone network, and a Neck network. Yolov7 has 3 commonly used models: Yolov5, Yolov7x, and Yolov7tiny. Among them, Yolov7tiny is the smallest model, which features fast detection speed and relatively low accuracy, and is mainly applied to scenarios with high speed requirements or limited computing resources. Yolov7 is suitable for scenarios that require high detection accuracy. The present invention uses Yolov7 as the basic model and adjusts some parameters to improve the model performance.

[0059] The present invention has improved the Neck layer of Yolov7. First, on the premise of ensuring detection accuracy, GSConv is introduced to replace Conv, reducing the network volume and computational complexity, and realizing the lightweight design of the model. In addition, since the size of the pits in the defects to be detected is small, a Small Target Enhancement Module (STE module) is added to perform feature fusion on the feature maps (p3, p4, p5) of three different scales in the Backbone layer to capture multi-scale features, so as to be able to more accurately capture pits of different sizes and shapes. The improved Yolov7 network structure is as Figure 6 shown.

[0060] Regarding the problem of low detection accuracy for small target defects, the added Small Target Enhancement Module is as Figure 7As shown, it includes a multi-scale sub-module, a channel attention sub-module, and a spatial attention sub-module. The main function of the small object enhancement module is to fuse features of different scales from the Backbone layer through 3D convolution and pooling in the multi-scale sub-module, especially to provide better feature representation for the detection of small objects; at the same time, the channel attention sub-module and the spatial attention sub-module are introduced for the multi-scale features to make full use of cross-channel information, achieving performance improvement without increasing the network complexity. While retaining key information, noise and irrelevant information are suppressed.

[0061] 1) Channel attention sub-module: It aims to emphasize which channels contain more important information. Its basic idea is to analyze each channel of the input feature map to find more informative channels and enhance their weights. The calculation steps of the channel attention sub-module are as follows:

[0062] a. Average pooling and max pooling: Perform average pooling and max pooling operations on the input feature map F ∈ R C×H×W (where C is the number of channels, H is the height, and W is the width) to obtain two 1×1×C descriptors:

[0063] (Average pooling result)

[0064] (Max pooling result)

[0065] These two operations compress the feature map in the spatial dimension, only retaining the information in the channel dimension.

[0066] b. Shared multi-layer perceptron (MLP): Perform shared multi-layer perceptron processing on the pooled features, usually two layers. First, reduce the dimension of the pooled 1x1xC features to 1×1×C / r (where r is the reduction ratio, a hyperparameter) through a fully connected layer, and then restore it to 1×1×C through another fully connected layer.

[0067] For the features after average pooling:

[0068] For the features after max pooling:

[0069] c. Merge and activation: Add the features processed by the two MLPs and obtain the final channel attention map through a Sigmoid activation function:

[0070] c = σ(M avg (F) + M max (F))

[0071] The obtained M c(F) is a 1×1×C tensor, and each element represents the attention weight of the corresponding channel.

[0072] 2) Spatial attention sub-module: This module focuses on finding important regions in the spatial dimension. The calculation steps of the spatial attention sub-module are as follows:

[0073] a. Channel pooling: For the input feature map F∈R C×H×W Perform max pooling and average pooling respectively in the channel dimension to obtain two feature maps of HxWx1:

[0074] (Average pooling result)

[0075] (Max pooling result)

[0076] Here, the feature map is compressed in the channel dimension, only retaining the spatial information.

[0077] b. Concatenation and convolution: Concatenate these two pooling results in the channel dimension to obtain a feature map of H×W×2; then perform a convolution operation through a 7×7 convolution kernel to reduce the number of channels to 1.

[0078]

[0079] The finally obtained M s (F) is a spatial attention map of H×W×1.

[0080] First, pass the input feature map through the channel attention sub-module to obtain the weighted feature map (where represents element-wise multiplication); then input F' into the spatial attention sub-module to obtain the finally weighted feature map

[0081] Furthermore, the present invention adopts an improved EIoU loss function for the loss function of the detection head.

[0082] The calculation complexity of the CIoU loss function is relatively high, which may increase the training time. Moreover, due to its sensitivity to the aspect ratio, in some cases, it may overemphasize the aspect ratio matching and ignore other factors, such as the overall positional relationship of the target. The EIoU loss function is selected for the above problems. The EIoU (Enhanced Intersection over Union) loss function is mainly used to evaluate the difference between the predicted bounding box and the ground truth bounding box in object detection, and its formula is as follows:

[0083] Calculate the center point distance D c : The center of the prediction box is (xp , y p ), the center of the ground truth box is (x g , y g ), then

[0084]

[0085] Calculate the difference in aspect ratio: width difference w diff and height difference h diff . There are various calculation methods. A common one is the simple difference, that is, w diff = |w p - w g |, h diff = |h p - h g |, and it can also be in the form of a ratio difference, etc.

[0086] The general form of the EIoU loss function is:

[0087]

[0088] where λ 1 , λ 2 and λ 3 are hyperparameters that adjust the influence of the center distance and the aspect ratio difference. C d , C w , C h are constants for normalization respectively, which are related to the size of the objects in the dataset, etc.

[0089] EIoU can improve the detection accuracy. It comprehensively considers various geometric information of the bounding boxes, including the overlapping area, the center point distance, and the aspect ratio difference, enabling the model to more accurately learn the position and shape information of the objects, thus improving the accuracy of object detection, especially performing better when dealing with complex situations such as object deformation and occlusion; its convergence speed is faster. Compared with some traditional loss functions, the aspect ratio loss of EIoU directly minimizes the difference between the width and height of the target box and the predicted box, and can more quickly guide the model to learn the appropriate bounding box parameters, accelerating the convergence speed of the model.

[0090] Introduce the distance attention mechanism of the Wise Intersection over Union (Wise - IoU) into the EIoU loss function to improve it:

[0091]

[0092] In the above formula, the superscript * indicates that this part is separated from the computational graph. The purpose of separation is to prevent the gradient generated by this part from affecting the convergence of the model. In deep learning, the computational graph calculates the gradient according to the chain rule. If not separated, the gradient calculation of the entire loss function may be affected due to abnormal gradients in this part (such as vanishing gradients or exploding gradients), resulting in unstable model training. After separating it, during the backpropagation process, the gradient calculation will ignore this part, making the model training more stable and conducive to convergence.

[0093] Then the improved EIoU loss function is as follows:

[0094]

[0095] In the formula, x and y represent the coordinates of the center point of the predicted bounding box, x gt , y gt represent the coordinates of the center point of the ground truth bounding box, W g , H g represent the width and height of the minimum enclosing box; IoU represents the intersection over union; λ 1 , λ 2 , λ 2 represent hyperparameters that adjust the influence of the center distance and width-height difference; C d , C w , C h are constants used for normalization respectively; D c represents the distance between the center points of the predicted bounding box and the ground truth bounding box; w diff represents the width difference between the predicted bounding box and the ground truth bounding box; h diff represents the height difference between the predicted bounding box and the ground truth bounding box.

[0096] Here, R WIoU plays the role of amplifying the IoU of ordinary-quality anchor boxes. When the anchor box coincides well with the target box, it will reduce the attention to the center point distance, thereby reducing the dominant position of high-quality anchor boxes in training, making the model pay more attention to ordinary-quality anchor boxes, and helping to improve the overall performance.

[0097] Then, execute S4. In S4, use the trained wheel tread detection model to detect the image to be detected and obtain the detection results; the detection results include: no defect, wheel tread peeling defect, wheel tread pit defect, and wheel tread scratch defect.

[0098] Further, verify the technical effect of the present invention through experiments.

[0099] The hardware and software environments used in the experiment are shown in Table 1.

[0100] Table 1 Experimental environment configuration

[0101]

[0102]

[0103] In the experiment, the SGD (Stochastic Gradient Descent) method was used to optimize the learning rate, and the number of training rounds was determined by comparing the loss functions of the training set and the validation set. Table 2 lists the parameters for training the network.

[0104] Table 2 Network training function

[0105]

[0106] To verify the superior performance of the improved Yolov7 model, the experiment measured mAP, FPS, model volume, etc. Some commonly used precision (P), recall (R), and mean average precision (mAP) metrics were selected to evaluate the model performance, and the formulas are defined as follows:

[0107]

[0108] In the formula, P and R represent precision and recall. The formulas are respectively:

[0109]

[0110] The analysis of the comparative experiment results is as follows.

[0111] To verify the detection performance of the improved model, a comparison was made with mainstream detection algorithms, and the results are shown in Table 3 and Figure 8 as shown. From Table 3 and Figure 8 it can be seen that the mAP of the improved model is higher than that of other algorithms. It is 1.6 percentage points higher than the original YOLOv7 model, and the model size is reduced by 73.91MB; it is 10.7 percentage points higher than the YOLOv5 model, and the model size is reduced by 94.69MB; it is 48.63 percentage points higher than SSD, and the model size is reduced by 122.11MB; it is 37.97 percentage points higher than Faster-Rcnn, and the model size is reduced by 154.91MB. Compared with other models, it has better detection accuracy and significantly reduces the model size while slightly improving the FPS; generally speaking, the performance of the improved model is better than that of other algorithms.

[0112] Table 3 Comparison of each detection model

[0113]

[0114] As shown in Table 4, five different loss functions were used on the original model of yolov7. Among them, the improved EIoU proposed in the present invention has the best effect. Compared with the original model of yolov7 (CIoU), mAP@0.5 is increased by 1.8% and mAP@0.5:0.95 is increased by 1.44%. The EIoU loss function improves the effect of small object bounding boxes.

[0115] As shown in Table 5, using the GSConv module in the neck network reduced the model size by 84 MB, while increasing the FPS to 87.1. However, the mAP@0.5 for small object detection (pit) decreased by 10.2%. When only the STE module was introduced, it increased the mAP for small object detection by 4.4%, but the model size increased by 14 MB, the detection speed decreased, and for the remaining object sizes, the mAP decreased by 2.6%. The mAP@0.5 for peeling decreased by 1.3%, and the mAP@0.5 for abrasion decreased by 1.3%. When both the STE module and GSConv were used, compared with yolov7 (EIoU), the model size decreased by 73.91 MB, the mAP@0.5 (all levels) increased by 1.6%, the pit mAP increased by 4.2%, the peeling mAP increased by 0.1%, the abrasion mAP increased by 0.3%, and the FPS increased by 1.9.

[0116] Table 4 Comparison results of different IoUs

[0117]

[0118] Table 5 Ablation experiment results

[0119]

[0120]

[0121] To verify the robustness of the model, the publicly available RSDDs dataset (railway track pit defects) with similar small object defects was selected for testing. Figure 9 It is an example diagram of the results of using different models to identify pits. Figure 10 It is an example diagram of the results of using different models to identify abrasions. Figure 11 It is an example diagram of the results of using different models to identify peeling.

[0122] In summary, to address the problems of being unable to quickly and accurately detect wheel tread defects and being unable to be applied to real complex scenarios, the present invention improves the YOLOv7 algorithm: adopting GSConv lightweight convolution in its neck network to reduce the model size, and at the same time adding a small object enhancement (STE) module to solve the problem of low pixels and difficulty in distinguishing small objects, and improving the classification accuracy of damage feature recognition.

[0123] The dataset is also augmented by various methods (including the generative adversarial network StyleGAN3, Copy-paste, etc.) so that the model can identify more defect features and improve the generalization ability of the detection model. The YOLOv7-STE model proposed in the present invention has the highest classification accuracy compared with traditional models. Compared with the original YOLOv7 model, the overall recognition accuracy is increased by 1.6 percentage points, the model size is reduced by 73.91 MB, and the accuracy of small target detection (for pits) is increased by 4.2 percentage points.

[0124] Another embodiment of the present invention proposes a wheel tread detection system based on an improved YOLOv7 model, as Figure 12 shown. The system includes:

[0125] A data acquisition module 1210 configured to acquire an image dataset including defective and non-defective train wheel sets;

[0126] A preprocessing module 1220 configured to preprocess the image dataset;

[0127] A detection model training module 1230 configured to train a wheel tread detection model based on an improved YOLOv7 model using the preprocessed image dataset;

[0128] A wheel tread defect detection module 1240 configured to detect a to-be-detected image using the trained wheel tread detection model and obtain a detection result.

[0129] In this embodiment, preferably, the defective train wheel sets in the data acquisition module 1210 include wheel tread peeling defects, wheel tread pit defects, and wheel tread abrasion defects.

[0130] In this embodiment, preferably, the preprocessing in the preprocessing module 1220 includes: data augmentation, data enhancement, and class label annotation; the data augmentation includes data augmentation using copy-paste, or training a generative adversarial network using the image dataset, and generating images corresponding to defects through the generative adversarial network for dataset augmentation; the data enhancement includes: Gaussian blur, affine transformation, brightness transformation, downsampling transformation, flipping transformation, Mixup, and Mosaic for splicing and mixing to enhance data.

[0131] In this embodiment, preferably, the improvements of the YOLOv7 model in the detection model training module 1230 include: using GSConv to replace ordinary convolution in the neck network; adding a small target enhancement module in the neck network, which is used to perform multi-scale feature fusion on feature maps of different sizes in the backbone network, integrate the fused feature maps and the feature maps before fusion, and input the integrated feature maps into the channel attention sub-module and the spatial attention sub-module for processing; the loss function of the detection head adopts an improved EIoU loss function; the expression of the improved EIoU loss function is:

[0132]

[0133] In the formula, x, y represent the coordinates of the center point of the predicted bounding box, x gt , y gt represent the coordinates of the center point of the ground truth bounding box, W g , H g represent the width and height of the minimum bounding box; IoU represents the intersection over union; λ 1 , λ 2 , λ 3 represent hyperparameters for adjusting the influence of the center distance and the width and height differences; C d , C w , C h are constants for normalization respectively; D c represents the distance between the center points of the predicted bounding box and the ground truth bounding box; w diff represents the width difference between the predicted bounding box and the ground truth bounding box; h diff represents the height difference between the predicted bounding box and the ground truth bounding box.

[0134] It should be noted that the functions of the wheel set tread detection system based on the improved YOLOv7 model in this embodiment can be illustrated by the aforementioned wheel set tread detection method based on the improved YOLOv7 model. For parts not detailed in the system embodiment, please refer to the above method embodiment.

[0135] It should be noted that although several units, modules or sub-modules are mentioned in the above detailed description, this division is only exemplary and not mandatory. In fact, according to the embodiments of the present invention, the features and functions of the two or more modules described above can be embodied in one module. Conversely, the features and functions of one module described above can be further divided and embodied by multiple modules.

[0136] Moreover, although the operations of the method of the present invention are described in a specific order in the drawings, this is not a requirement or implication that these operations must be performed in that specific order, or that all of the shown operations must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step and performed, and / or one step may be decomposed into multiple steps and performed.

[0137] Although the spirit and principles of the present invention have been described with reference to several specific embodiments, it should be understood that the present invention is not limited to the specific embodiments disclosed, and the division of each aspect does not mean that the features in these aspects cannot be combined for benefit. This division is only for convenience of expression. The present invention is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.

Claims

1. A wheelset tread detection method based on an improved YOLOv7 model, characterized in that: include: Obtaining an image dataset containing defective and non-defective train wheelsets; Preprocessing the image data set; The preprocessed image dataset is used to train a wheelset tread detection model based on the improved YOLOv7 model; The trained wheelset tread detection model is used to detect the image to be detected and obtain the detection result.

2. The wheelset tread detection method based on the improved YOLOv7 model according to claim 1, characterized in that: The train wheelset defects include wheelset tread peeling defects, wheelset tread pit defects, and wheelset tread abrasion defects.

3. The wheelset tread detection method based on the improved YOLOv7 model according to claim 1, characterized in that: The preprocessing includes: data expansion, data enhancement, and category labeling.

4. The wheelset tread detection method based on the improved YOLOv7 model according to claim 3 is characterized in that: The data expansion includes using copy and paste to expand the data, or using the image data set to train a generative adversarial network, and generating images of corresponding defects through the generative adversarial network to expand the data set; the data enhancement includes: Gaussian blur, affine transformation, brightness transformation, pixel drop transformation, flip transformation, Mixup, and Mosaic to splice and mix enhanced data.

5. The wheelset tread detection method based on the improved YOLOv7 model according to claim 1, characterized in that: The improvements of the YOLOv7 model include: using GSConv to replace ordinary convolution in the neck network; adding a small target enhancement module in the neck network, and the small target enhancement module is used to perform multi-scale feature fusion on feature maps of different sizes in the backbone network through a multi-scale sub-module, and integrate the fused feature map with the feature map before and after fusion; inputting the integrated feature map into the channel attention sub-module and the spatial attention sub-module for processing; the loss function of the detection head adopts an improved EIoU loss function.

6. The wheelset tread detection method based on the improved YOLOv7 model according to claim 5, characterized in that: The expression of the improved EIoU loss function is: In the formula, x,y represents the coordinates of the center point of the prediction box, x gt ,y gt Represents the coordinates of the center point of the real box, W g , H g represents the width and height of the minimum bounding box; IoU represents the intersection over union ratio; λ1, λ2, and λ3 represent the hyperparameters that adjust the center distance and the difference in width and height; C d , C w , C h are constants used for normalization; D c Represents the distance between the center point of the predicted box and the real box; w diff Indicates the width difference between the predicted box and the real box; h diff Indicates the height difference between the predicted box and the true box.

7. The wheel set tread detection system based on the improved YOLOv7 model is characterized by: include: A data acquisition module configured to acquire an image data set containing defective and non-defective train wheelsets; A preprocessing module, configured to preprocess the image data set; A detection model training module, configured to train a wheelset tread detection model based on an improved YOLOv7 model using the preprocessed image dataset; The wheelset tread defect detection module is configured to use the trained wheelset tread detection model to detect the image to be detected and obtain the detection result.

8. The wheelset tread detection system based on the improved YOLOv7 model according to claim 7, characterized in that: The train wheelset defects in the data acquisition module include wheelset tread peeling defects, wheelset tread pit defects, and wheelset tread abrasion defects.

9. The wheelset tread detection system based on the improved YOLOv7 model according to claim 7, characterized in that: The preprocessing in the preprocessing module includes: data expansion, data enhancement, and category labeling; the data expansion includes data expansion by copying and pasting, or using the image data set to train a generative adversarial network, and generating images of corresponding defects through the generative adversarial network to expand the data set; the data enhancement includes: Gaussian blur, affine transformation, brightness transformation, pixel drop transformation, flip transformation, Mixup, and Mosaic to splice and mix enhanced data.

10. The wheelset tread detection system based on the improved YOLOv7 model according to claim 7, characterized in that: The improvements of the YOLOv7 model in the detection model training module include: using GSConv to replace ordinary convolution in the neck network; adding a small target enhancement module in the neck network, the small target enhancement module is used to perform multi-scale feature fusion on feature maps of different sizes in the backbone network, and integrate the fused feature map with the feature map before fusion, and input the integrated feature map into the channel attention submodule and the spatial attention submodule for processing; the loss function of the detection head adopts the improved EIoU loss function; the expression of the improved EIoU loss function is: In the formula, x,y represents the coordinates of the center point of the prediction box, x gt ,y gt Represents the coordinates of the center point of the real box, W g , H g represents the width and height of the minimum bounding box; IoU represents the intersection over union ratio; λ1, λ2, and λ3 represent the hyperparameters that adjust the center distance and the difference in width and height; C d , C w , C h are constants used for normalization; D c Represents the distance between the center point of the predicted box and the real box; w diff Indicates the width difference between the predicted box and the real box; h diff Indicates the height difference between the predicted box and the true box.