Bullet train wheel tread anomaly detection method and device
The modified Yolov5-seg model with coordinate attention and spatial adaptive fusion enhances wheel tread abnormality detection, addressing inefficiencies in existing methods by improving precision and reducing manual intervention, ensuring timely fault identification and enhancing train safety.
Patent Information
- Application Number
- CN202510474799.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The existing EMU wheel tread abnormality detection methods have problems such as high false alarm rate, high false alarm rate, difficulty in precise fault location and segmentation, and require a lot of manual intervention, especially in small sizes or complex backgrounds.
The improved yolov5-seg model is adopted, combining the coordinate attention mechanism and the spatial adaptive fusion module (ASFF), abnormal detection is performed on the wheel tread image, fault area identification is performed through the object detection segmentation model, and unsupervised anomaly detection algorithm and manual review are combined to reduce manual intervention.
It improves the accuracy and efficiency of fault detection, reduces the intensity of manual labor, and can promptly detect and locate wheel tread faults, ensuring the safe operation of the EMU.
Smart Images

Figure CN120318199A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of image processing, and particularly to a method and device for detecting abnormal conditions of the tread surface of a high-speed train wheel. Background Art
[0002] The wheels of high-speed trains are an important part of the running gear of high-speed trains and are crucial for ensuring the safe operation of trains. The tread surface of the wheel is the part that contacts the rail. Common faults such as abrasions, spalls, and indentations not only cause large impact forces and strong vibrations on the wheel but may also affect the safe operation of high-speed trains, shorten the service life of the wheels and other components, extend the braking distance, and even damage the rail and the track. These faults seriously endanger the safety and stability of railway transportation.
[0003] Traditional methods for inspecting the tread surface of high-speed train wheels usually rely on manual inspection. After the high-speed train enters the maintenance station, the staff uses relevant equipment to check the condition of the wheels. With the increasing demand for railway transportation and the rapid growth in the number of high-speed trains, manual inspection faces great pressure, with low efficiency and long time consumption. Especially when the faults on the tread surface of the wheel are relatively subtle or difficult to detect, the defects of traditional methods are more obvious, easily leading to missed or misdetected inspections, affecting the overall maintenance efficiency and the safe operation of the vehicle. In recent years, the rapid development of artificial intelligence technology, especially the progress of deep learning and image processing technology, has brought new opportunities for detecting abnormal conditions of the tread surface of high-speed train wheels. The intelligent detection method based on image processing can automatically detect the tread surface image of the wheel in a short time, quickly identify potential abnormal areas, and feed back the fault information to the maintenance personnel, thus greatly improving the detection efficiency and accuracy and reducing the pressure of manual inspection. However, the existing image processing-based abnormal detection methods still have some deficiencies, mainly reflected in the following aspects.
[0004] The existing abnormal detection methods have a high false alarm rate and missed alarm rate during the processing process. Especially when dealing with fault areas with small sizes or complex backgrounds, the accuracy and efficiency are often insufficient, affecting the actual application effect. Most existing methods are difficult to simultaneously achieve accurate fault location and fine-grained segmentation of the fault area. Especially in the case of low image resolution or complex backgrounds, the segmentation effect is not good, resulting in the inability to accurately judge the type and severity of the fault. Although some existing methods can improve the efficiency of abnormal detection through unsupervised learning or preliminary screening algorithms, due to the inability to automatically label the fault area, a large amount of manual intervention is still required for labeling and confirmation, increasing the workload and labor cost.
[0005] Therefore, the method for detecting abnormal conditions of the tread surface of high-speed train wheels based on deep learning, especially the improved object detection and image segmentation technology, can effectively improve the accuracy and efficiency of fault detection, reduce manual intervention, and improve the intelligent level, and has become an effective solution to solve the above problems. Summary of the Invention
[0006] An embodiment of the present application provides a method and device for detecting abnormal conditions on the tread of a high - speed train wheel. The technical solution is as follows:
[0007] On the one hand, a method for detecting abnormal conditions on the tread of a high - speed train wheel is provided. The method is used for a full - range dynamic image detection system for high - speed trains entering the depot. The detection system includes a tread acquisition module. The method includes:
[0008] Collect the tread images of the wheel set of the high - speed train through the tread acquisition module;
[0009] Sort out and annotate the collected tread images of the wheel set;
[0010] Perform abnormal detection on the tread images of the wheel set through a target detection and segmentation model;
[0011] For the detected target fault area information, filter out small areas by area and promptly return the position where the fault occurs to the platform.
[0012] Optionally, the full - range dynamic image detection system for high - speed trains entering the depot is equipped with 4 tread acquisition modules. Each tread acquisition module includes left - and right - hand groups of acquisition modules. Each group of acquisition modules includes 5 area array cameras, and each area array camera uses an oblique upward shooting angle;
[0013] Each group of acquisition modules is respectively connected to a trigger magnet cylinder. The tread acquisition module is communicatively connected to a train - approaching magnet cylinder, and the train - approaching magnet cylinder is connected to a converter for information calculation.
[0014] Optionally, the step of collecting the tread images of the wheel set of the high - speed train through the tread acquisition module includes:
[0015] In response to the train - approaching magnet cylinder first receiving a train - approaching signal, record the time of the train - passing signal through the converter, where the train - approaching signal is the signal generated by the corresponding sensor of the train - approaching magnet cylinder when the high - speed train rolls over the rail;
[0016] In response to the wheel passing over each trigger magnet cylinder afterwards, calculate the train - passing speed through the converter according to the actual distance between the current trigger magnet cylinder and the train - approaching magnet cylinder;
[0017] Calculate the delay time of the area array cameras in different acquisition modules, so as to control the acquisition trigger time of different cameras;
[0018] In response to the tread acquisition module receiving the train - approaching signal, start the corresponding acquisition cameras through the tread acquisition module to sequentially capture the tread images of the wheel set of the high - speed train;
[0019] Transmit the collected tread images of the wheel set of the high - speed train to the server.
[0020] Optionally, the sorting and annotation processing of the collected wheel tread images includes:
[0021] Initial screening of suspected fault pictures from the wheel tread images through an unsupervised anomaly detection algorithm;
[0022] Manually review the suspected fault pictures and mark the fault areas pixel by pixel.
[0023] Optionally, before performing anomaly detection on the wheel tread images through the target detection and segmentation model, building a target detection and segmentation model is also included;
[0024] The building of the target detection and segmentation model includes:
[0025] Modify the original Proto module in the yolov5-seg model, where the network structure of the Yolov5-seg model consists of four parts: an input end, a BackBone network, a Neck structure, and a Head prediction layer;
[0026] Add a coordinate attention module and a spatial adaptive fusion (ASFF) module on the basis of the yolov5-seg model to obtain the target detection and segmentation model.
[0027] Optionally, the anomaly detection of the wheel tread images through the target detection and segmentation model includes:
[0028] Preprocess the wheel tread images through the input end;
[0029] Extract useful features from the preprocessed wheel tread images through the BackBone network to obtain corresponding feature maps;
[0030] Fuse the feature maps through the Neck structure, enhance the spatial position information of the feature maps at different scales through the coordinate attention module, and transfer the enhanced multi-scale feature maps to the Head prediction layer;
[0031] Adaptive selection of feature maps at different levels for weighted fusion through the spatial adaptive fusion ASFF module to obtain an adaptively fused multi-scale feature map. The adaptive selection criterion is to ensure the simultaneous processing of feature maps of different sizes, where the ASFF module is added in front of the Head prediction layer;
[0032] Perform target detection on the multi-scale feature maps through the Head prediction layer to obtain target fault area information including the boundary box of the fault area, fault category information, and segmented images.
[0033] Optionally, adaptively selecting feature maps of different levels through the ASFF module for weighted fusion to obtain an adaptively fused multi-scale feature map includes:
[0034] The ASFF module takes feature maps of different sizes as inputs, performs upsampling on the small-size feature maps at different ratios, and keeps the spatial features of the large-size feature maps unchanged with the same resolution.
[0035] Fuse the feature maps from different layers and normalize them through softmax, multiply and add them with the feature maps of the same resolution respectively to obtain an adaptively fused multi-scale feature map.
[0036] Optionally, enhancing the spatial position information of feature maps of different scales through the coordinate attention module and passing the enhanced multi-scale feature maps to the Head prediction layer includes:
[0037] Pool and encode the input feature maps along the X direction and Y direction respectively through the coordinate attention module.
[0038] Concatenate and fuse the encoded information in the X direction and Y direction, and decompose it into two independent tensors along the spatial dimension.
[0039] Perform convolutional transformation on the two tensors to adjust their number of channels until they match the number of channels of the input feature map.
[0040] Fuse the attention information in the X direction and Y direction with the original input feature map.
[0041] Pass the fused multi-scale feature maps to the Head prediction layer for object detection.
[0042] Optionally, performing object detection on the multi-scale feature maps through the Head prediction layer to obtain target fault region information including the boundary box of the fault region, fault category information, and segmented image includes:
[0043] Process the multi-scale feature maps through the Head prediction layer, including convolutional layers, pooling layers, and fully connected layers.
[0044] Add a segmentation branch Proto to the Yolov5-seg model, restore the image to the input size through a bottom-up decoding process, and finally generate the segmentation result of the image.
[0045] In the Yolov5-seg model, transfer the feature maps of different resolutions extracted at the Neck end to the object detection head, and the object detection head is used to optimize object box, class regression, and box position regression tasks, and finally output the boundary box, class information, and pixel-level segmentation result of the target fault region.
[0046] By referring to the bottom-up decoding structure in PaNet through the Proto module, the feature size is gradually restored from 80×80 to a resolution of 640×640. Through the concatenation of CBL and upsampling pooling operations, the input features are gradually restored to a size of 640×640 for pixel-level segmentation of target regions of different sizes;
[0047] After being processed by the target detection head and the segmentation part, the target fault region information including the bounding box, class information, and pixel-level segmentation image of the target fault region is obtained.
[0048] On the other hand, a device for detecting abnormal tread of EMU wheels is provided. The device is used for the all-round dynamic image detection system of EMU entering the depot. The detection system includes a tread acquisition module. The device includes:
[0049] An image acquisition module for acquiring the tread image of the EMU wheel pair through the tread acquisition module;
[0050] An image processing module for sorting and annotating the acquired tread image of the wheel pair;
[0051] An abnormal detection module for detecting abnormalities in the tread image of the wheel pair through a target detection and segmentation model;
[0052] A fault detection module for filtering out small regions by area from the detected target fault region information and timely returning the position where the fault occurs to the platform.
[0053] On the other hand, a server is provided. The server includes a processor and a memory; the memory stores at least one instruction, and the at least one instruction is used to be executed by the processor to implement the method for detecting abnormal tread of EMU wheels as described in the above aspect.
[0054] On the other hand, a computer-readable storage medium is provided. The storage medium stores at least one instruction, and the at least one instruction is used to be executed by a processor to implement the method for detecting abnormal tread of EMU wheels as described in the above aspect.
[0055] On the other hand, a computer program product is further provided. The computer program product stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the method for detecting abnormal tread of EMU wheels as described in the above aspect.
[0056] This application discloses a method for detecting abnormal conditions on the tread surface of high-speed train wheels. The method for detecting abnormal conditions on the tread surface of high-speed train wheels based on deep learning image processing technology can effectively improve the detection accuracy and efficiency of wheel tread faults by introducing an improved yolov5-seg model. By combining the coordinate attention mechanism and the spatially adaptive feature fusion module (ASFF), the model performs excellently in multi-scale feature fusion and target localization accuracy, especially being able to accurately identify fault regions of different sizes and having strong real-time performance. Compared with traditional manual detection methods, this method has obvious advantages in non-destructive detection, automated detection, and fault information feedback. It not only significantly improves the detection efficiency, reduces the manual labor intensity, but also can promptly detect and locate wheel tread faults, ensuring the safe operation of high-speed trains. Brief Description of the Drawings
[0057] Figure 1 Shows the layout diagram of the detection system provided by an exemplary embodiment of this application;
[0058] Figure 2 Shows the flow chart of the method for detecting abnormal conditions on the tread surface of high-speed train wheels provided by an exemplary embodiment of this application;
[0059] Figure 3 Shows the schematic diagram of the fault picture;
[0060] Figure 4 Shows the picture of the normal tread surface;
[0061] Figure 5 Is the fault picture detected as abnormal by the VAND algorithm;
[0062] Figure 6 Is the schematic diagram of the Backbone network structure;
[0063] Figure 7 Is the schematic diagram of pooling encoding in the X and Y directions;
[0064] Figure 8 Is the internal processing flow chart of the ASFF module;
[0065] Figure 9 Is the schematic diagram of the bottom-up decoding structure in PaNet;
[0066] Figure 10 Is the flow schematic diagram of the abnormal detection system for the tread surface image used on-site in the high-speed train depot;
[0067] Figure 11 Is the schematic diagram of the tread fault area. Detailed Embodiment
[0068] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe the embodiments of this application in detail with reference to the drawings.
[0069] As used herein, "a plurality of" means two or more. "And / or" describes the relationship between associated objects, indicating three possible relationships. For example, A and / or B can represent three cases: A exists alone, both A and B exist simultaneously, and B exists alone. The character " / " generally indicates an "or" relationship between the associated objects before and after.
[0070] Please refer to Figure 2 , which shows a flowchart of a method for detecting abnormal treads of EMU wheels provided by an exemplary embodiment of the present application. This method is used for the all-round dynamic image detection system for EMUs entering the depot. The detection system includes a tread acquisition module. The method includes:
[0071] Step 201: Acquire the tread image of the EMU wheel pair through the tread acquisition module.
[0072] As Figure 1 shown, first, an explanation is made in combination with the layout of the detection system. The all-round dynamic image detection system for EMUs entering the depot is equipped with 4 tread acquisition modules. Each tread acquisition module includes left and right groups of acquisition modules. Each group of acquisition modules includes 5 area array cameras, and each area array camera uses an oblique upward shooting angle.
[0073] Each group of acquisition modules is respectively connected to a trigger magnet cylinder. The tread acquisition module is communicatively connected to a train-receiving magnet cylinder, and the train-receiving magnet cylinder is connected to a converter for information calculation.
[0074] Thus, in a possible implementation manner, step 201 includes the following contents.
[0075] Content 1: In response to the train-receiving magnet cylinder receiving the train-receiving signal for the first time, record the time of the passing train signal through the converter, where the train-receiving signal is the signal generated by the corresponding sensor of the train-receiving magnet cylinder when the EMU rolls onto the rail.
[0076] Content 2: In response to the wheel passing over each trigger magnet cylinder afterwards, calculate the passing train speed through the converter according to the actual distance between the current trigger magnet cylinder and the train-receiving magnet cylinder.
[0077] Content 3: Calculate the delay time of the area array cameras in different acquisition modules, so as to control the acquisition trigger time of different cameras.
[0078] Content 4: In response to the tread acquisition module receiving the train-receiving signal, start the corresponding acquisition camera through the tread acquisition module to sequentially capture the tread images of the EMU wheel pair;
[0079] Content 5: Transmit the captured tread images of the EMU wheel pair to the server.
[0080] In a possible implementation, after the train passes through the acquisition area, the 360-degree dynamic image detection system for the multiple unit train entering the depot will send a transmission instruction to the transmission program, and the transmission program will start to transmit the pictures in the acquisition machine to the server. The acquisition machine is an acquisition management machine connected to the area array camera.
[0081] Combined with the attached Figure 1 Specifically, a total of 4 tread acquisition modules (31, 32, 33, 34) are laid on site. Each module consists of two groups of acquisition modules on the left and right. Each module consists of 5 area array cameras with an inclined upward shooting angle. After the trigger magnet cylinder 1 receives the trigger signal, it will start the acquisition and shooting of 5 cameras on the left side of each of the 31st and 33rd modules. After the trigger magnet cylinder 2 receives the trigger signal, it will start the acquisition and shooting of 5 cameras on the right side of each of the 31st and 33rd modules. And so on, the trigger magnet cylinder 3 and the trigger magnet cylinder 4 respectively control the acquisition work of the cameras in the 32nd and 34th modules. After the train has passed, the system will send a transmission instruction to start the transmission program to transmit the acquired pictures from the acquisition machine to the data processing server.
[0082] Step 202: Sort out and label the acquired wheel pair tread images.
[0083] In a possible implementation, step 202 includes the following content.
[0084] Content 1: Preliminary screen out suspected fault pictures from the wheel pair tread images through an unsupervised anomaly detection algorithm.
[0085] Content 2: Manually review the suspected fault pictures and label the fault areas pixel by pixel.
[0086] Regarding why the process of manual review is added, it needs to be specifically explained here. Since the number of tread pictures (i.e., wheel pair tread images) acquired for each multiple unit train is large (1280 pictures for 8-car formation, 2560 pictures for 16-car formation), and the camera shooting angle is fixed, if each picture is judged one by one, the time cost and labor cost are relatively high. We first use the unsupervised anomaly detection algorithm to preliminarily screen out suspected fault pictures, and then manually review the preliminarily screened fault pictures and label the fault areas pixel by pixel.
[0087] In an example, Figure 3 shows a schematic diagram of a fault picture, Figure 4 shows a normal tread picture, Figure 5Fault pictures detected by the VAND algorithm (the red area is the abnormal area). Taking 6 trips of pictures reviewed by the EMU operation depot team from May 19th to May 23rd, 2024 as the initial dataset, we selected common fault pictures among them (such as indentations, foreign objects, etc. The collection of major fault samples such as peeling and abrasion is relatively scarce and is being continuously collected). The unsupervised anomaly detection algorithm VAND (Visual Anomaly and Novelty Detection, VAND) is the algorithm we use to initially screen for tread anomalies. We selected 2927 normal pictures from the initial dataset as the training set, 732 normal and 85 abnormal pictures as the validation set, and the remaining 3840 pictures as the test set. The tread anomaly pictures detected by VAND are as Figure 5 shown. 81 anomalies were detected in the test set (after review, 15 were positive reports and 66 were false alarms, with no missed reports). It can be seen that the unsupervised anomaly detection algorithm can detect most anomalies, but it cannot give the exact size of the fault area (such as area, etc.), and the number of false alarms is relatively large, so it is not suitable for direct use in on-site detection. Therefore, we use the unsupervised anomaly detection as the initial screening tool, and professional personnel annotate the fault areas of the selected abnormal pictures pixel by pixel as the golden standard for subsequent analysis.
[0088] Step 203, perform anomaly detection on the wheel pair tread image through the object detection and segmentation model.
[0089] In a possible implementation manner, step 203 includes the following contents.
[0090] Content 1. Preprocess the wheel pair tread image through the input end.
[0091] The input end is responsible for image preprocessing. It uses Mosaic data augmentation to splice new images by randomly scaling, cropping, and arranging the input images, simulating various perspectives that may appear in the collected images (including imaging distance, imaging angle, etc.), enriching the diversity of the original dataset, and also helping to improve the detection performance of small targets, thereby improving the robustness of the model.
[0092] Content 2. Extract useful features from the preprocessed wheel pair tread image through the BackBone network to obtain the corresponding feature map.
[0093] The Backbone network is the backbone network used to extract image features. Its main function is to convert the original input image into a multi-layer feature map for subsequent object detection tasks. As Figure 6As shown in the figure, first, the preprocessed image is subjected to feature extraction through a convolution with a kernel of 6×6, a stride of 2, and an expansion of 2. This not only reduces the resolution of the original image (from 640×640 to 320×320), but also distributes the image information among different channels, ensuring that the calculation speed can be improved without losing image information. Similar to common backbones, yolov5-seg adopts the basic convolutional neural network CBL module, as well as the C3 module containing a residual structure and the spatial pyramid pooling module SPPF.
[0094] Content three: The feature maps are fused through the Neck structure, and the spatial position information of the feature maps at different scales is enhanced through the coordinate attention module, and the enhanced multi-scale feature maps are passed to the Head prediction layer.
[0095] The Neck structure ( Figure 6 the middle two columns) can perform multi-scale feature fusion on the feature maps and pass these features to the Head prediction layer.
[0096] As Figure 7 shown, for the input features through the coordinate attention module, pooling encoding along the X direction and the Y direction is respectively used;
[0097] The encoding information in the two directions is concatenated and fused, and decomposed into two independent tensors along the spatial dimension;
[0098] Convolutional transformation is performed to transform the number of channels of the two tensors into the same dimension as the number of channels of the input features;
[0099] The attention in the X direction and the Y direction is fused with the features of the original input. Therefore, we add the coordinate attention mechanism to the feature maps with a smaller resolution (20×20, 40×40) in front of the Head prediction layer, while fusing the relationship between channels and position information, enabling the model to more accurately locate and identify the target fault area.
[0100] Content four: Through the ASFF module, different levels of feature maps are adaptively selected for weighted fusion to obtain an adaptively fused multi-scale feature map. The criterion for adaptive selection is to ensure the simultaneous processing of feature maps of different sizes, and the ASFF module is added in front of the Head prediction layer.
[0101] Among them, the ASFF module takes feature maps of different sizes as inputs, performs upsampling on the small-sized feature maps at different ratios respectively, and keeps the spatial features of the large-sized feature maps unchanged with the same resolution; the feature maps from different layers are fused and normalized through softmax, and multiplied and added with the features with the same resolution respectively to obtain an adaptively fused multi-scale feature map.
[0102] In one example, as Figure 8 shown, the ASFF module takes features of different sizes as input. In our experiment, the features of Figure 3 the 17th, 21st, and 25th layers are used as input. The semantic features of small sizes are upsampled at different ratios respectively, while the spatial features of large sizes remain unchanged and the resolution is kept the same. Subsequently, the feature information from different layers is fused and normalized through softmax, and multiplied and added with the features of the same resolution respectively. As the neural network parameters are continuously iteratively optimized, this module acts as an adaptive weight feature fusion, which can not only locate faults from a larger field of view but also retain the fine-grained information of the faults, thus improving the accuracy of segmentation and location.
[0103] Content Five: The multi-scale feature map is subjected to object detection through the Head prediction layer to obtain the target fault region information including the fault region bounding box, fault category information, and the segmented image.
[0104] The Head prediction layer is Figure 6 the model structure in the last column of which can perform final predictions on feature maps of different sizes (20×20, 40×40, 80×80). Among them, the feature maps with smaller sizes (20×20, 40×40) undergo multiple convolutional feature extractions, with richer semantic information and stronger non-linear fitting ability, which is more helpful for improving the accuracy of classification tasks and the regression of large target bounding boxes in detection tasks. While the feature maps with larger sizes (80×80) represent richer image spatial information and are suitable for fine-grained pixel-level segmentation tasks.
[0105] First of all, it should be noted that before performing anomaly detection on the wheel tread image through the object detection and segmentation model, building the object detection and segmentation model is also included. The building of the object detection and segmentation model includes:
[0106] Content One: Modify the original Proto module in the yolov5-seg model. The network structure of the Yolov5-seg model consists of four parts: the input end, the BackBone network, the Neck structure, and the Head prediction layer.
[0107] In a possible implementation manner, 4 convolutional and upsampling pooling operations are concatenated after the original output features to expand the original feature map from 80×80 to 640×640 in size, facilitating the implementation of fine-grained segmentation tasks.
[0108] Content Two: Add a coordinate attention module and an ASFF module on the basis of the yolov5-seg model to obtain the object detection and segmentation model.
[0109] Among them, Yolov5 is the mainstream object detection framework in current industrial detection, which is characterized by fast speed, small model size, high accuracy, and easy to use. Yolov5-seg is an image segmentation model based on Yolov5. A segmentation head is added on the basis of Yolov5, enabling the model to perform object detection and segmentation simultaneously. In a possible implementation manner, a coordinate attention module and a spatial adaptive network ASFF are added on the basis of Yolov5-seg, and the original Proto module is modified, which can improve the detection performance of the model for small target fault areas. Figure 6 (The schematic diagram of the improved Yolov5-seg model framework is shown). The Yolov5-seg network structure mainly consists of four parts: an input end, a BackBone network, a Neck structure, and a Head prediction layer.
[0110] Step 204: Filter out small regions from the detected target fault area information according to the area, and return the position where the fault occurs to the platform in a timely manner.
[0111] The Head end is used to process features of different feature sizes (20×20, 40×40, 80×80), including convolutional layers, pooling layers, and fully connected layers, etc. In the modified Yolov5-seg model, a segmentation branch Proto is added. Referring to the decoding part of PaNet, the bottom-up decoding process of features is realized, and finally the image is restored to the input size to generate the segmentation result of the image. The object detection part still follows the original idea in Yolov5-seg. The features of different resolutions extracted by the Neck end are sent into the object detection head to optimize tasks such as object bounding box, category, and box position regression. Finally, the object detection bounding box and category information of the image and the pixel-level segmentation result are returned.
[0112] As Figure 9 shown in the schematic diagram of the bottom-up decoding structure in PaNet, the model structure Proto of the segmentation part draws on the bottom-up decoding structure in PaNet. For the feature information enhanced by ASFF, its feature size is 80×80. In the 80×80 size of the original image 640×640, it is relatively difficult to optimize the segmentation of some small-area fault regions. Therefore, through cascaded and upsampling pooling operations, the input features are gradually restored from 80×80 to 640×640 in size, and pixel-level segmentation of target regions of different sizes is realized from a larger field of view.
[0113] In addition, the object detection and segmentation model provided by this application also includes subsequent model training and iterative update of parameters.
[0114] The Yolov5-seg model uses stochastic gradient descent as the optimizer to iteratively optimize the model parameters. It is trained for 70 rounds on a machine with Windows 11 and a 3090 graphics card, and the model parameters with the best validation set metrics are saved. The Yolov5-seg loss function includes the loss functions for object detection and segmentation. The loss function Loss for the object detection part det is the same as the loss function in the original yolov5, including the regression loss l of the bounding box CIoU , the object confidence loss l obj and the class loss l cls .
[0115]
[0116] where K, S, and B are the feature maps output for different classes, as well as the total number of pixels and the number of anchor boxes respectively. a box , a obj , a cls represents the weights of the three loss functions l CIoU , l obj , l cls , with default values of 0.05, 0.3, 0.7. l CIoU optimizes the parameters by calculating the intersection over union between the target box and the predicted box, while the confidence loss and the class loss both use the binary cross-entropy function for optimization. indicates whether the j-th anchor box of the i-th central pixel in the k-th feature map is a positive sample, 1 if it is and 0 if it is not. is used to balance the weights of the output feature maps for each scale (20×20, 40×40, 80×80), with a default value of [4.0, 1.0, 0.4].
[0117] The loss function Loss for the image segmentation part seg includes two parts, the binary cross-entropy loss function loss BC E and the Dice loss function loss Dice . Before the predicted image is fed into the loss function, it is first normalized between (0, 1) by the Sigmoid function, and then the probability distribution difference between the predicted result and the annotation standard is statistically calculated. The Dice loss function can focus on the accurate classification of the foreground and background by the model. The combination of the two loss functions can better improve the performance of the model in terms of both the segmentation boundary and the discrimination of foreground and background classification. As shown in the following formula, y i represents the i-th pixel value in the predicted result, represents the i-th pixel value in the annotated image.
[0118]
[0119] The trained model saves the.pt file, which is suitable for the Python environment and requires the original model definition to load. However, most recognition machines do not have the Python environment configured, making it impossible to achieve cross-platform deployment and operation. Moreover, in a hardware environment with a GPU, the conversion between the two does not result in a decrease in the inference speed and detection performance. Therefore, we convert the.pt file into a TorchScript file so that the model can be used for inference in an environment without relying on Python.
[0120] The specific process of the tread image anomaly detection system used on-site in the motor car depot is as follows Figure 10 : When the motor car passes through the acquisition device, the acquisition device takes pictures and transmits the pictures back to the processing server. After the transmission is completed, the scheduling program sends the current train information in the database and the location where the acquired images are stored to the recognition program. When the recognition program receives the task vehicle information, it starts to traverse the pictures in the target folder. Each picture undergoes object detection and segmentation by the yolov5-seg model (where the confidence level of the tread fault in object detection is 0.5. If it exceeds 0.5, it is considered a tread fault; if it is less than 0.5, it is not). When a fault is detected by object detection, we will count the pixel attributes (such as area) of the fault area within the current fault box, and then multiply the pixel attributes by the pixel conversion ratio of the current acquired image (actually taking 0.15 mm / pixel) to obtain the tread fault area ( Figure 11 ). After traversing and processing a whole train (for an 8-car formation, the normal number of tread pictures acquired for a whole 8-car formation is 1280, and the processing speed is approximately 0.1 s per picture), the fault information of the whole train and its corresponding wheel positions are returned to the platform. The platform is responsible for inserting the data into the database and filtering the data according to the alarm threshold of the tread (the area sizes for different vehicle types and different fault identifications are different), and finally displaying it on the platform.
[0121] The embodiment of the present application also provides a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement the method for detecting anomalies in the tread of a motor car wheel provided in the above various embodiments.
[0122] Optionally, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), solid-state drive (SSD, Solid State Drives), or optical disc, etc. Among them, the random access memory may include resistive random access memory (ReRAM, Resistance RandomAccess Memory) and dynamic random access memory (DRAM, Dynamic Random Access Memory).
[0123] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.
[0124] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk, an optical disc, or the like.
[0125] The above are only alternative embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for detecting abnormal conditions of the tread surface of a high-speed train wheel, characterized in that, The method is used for the all-round dynamic image detection system of EMU entering the depot. The detection system includes a tread acquisition module. The method includes: Collect the tread images of the EMU wheel set through the tread acquisition module; Sort out and label the collected tread images of the wheel set; Perform anomaly detection on the tread images of the wheel set through the target detection and segmentation model; Filter out small areas according to the area of the detected target fault area information and return the position where the fault occurs to the platform in time.
2. The method according to claim 1, wherein The all-round dynamic image detection system of EMU entering the depot is equipped with 4 tread acquisition modules. Each tread acquisition module includes left and right groups of acquisition modules. Each group of acquisition modules includes 5 area array cameras, and each area array camera uses an oblique upward shooting angle; Each group of acquisition modules is respectively connected with a trigger magnetic cylinder. The tread acquisition module is communicatively connected with a receiving vehicle magnetic cylinder, and the receiving vehicle magnetic cylinder is connected with a converter for information calculation.
3. The method according to claim 2, wherein The step of collecting the tread images of the EMU wheel set through the tread acquisition module includes: In response to the receiving vehicle magnetic cylinder receiving the receiving vehicle signal for the first time, record the time of the passing vehicle signal through the converter, where the receiving vehicle signal is the signal generated by the corresponding sensor of the receiving vehicle magnetic cylinder when the EMU rolls over the rail; In response to the wheel pressing over each trigger magnetic cylinder afterwards, calculate the passing vehicle speed through the converter according to the actual distance between the current trigger magnetic cylinder and the receiving vehicle magnetic cylinder; Calculate the delay time of the area array cameras in different acquisition modules, so as to control the acquisition trigger time of different cameras; In response to the tread acquisition module receiving the receiving vehicle signal, start the corresponding acquisition camera through the tread acquisition module to sequentially capture the tread images of the EMU wheel set; Transmit the collected tread images of the EMU wheel set to the server.
4. The method according to claim 1, wherein The step of sorting out and labeling the collected tread images of the wheel set includes: Pre-screen suspected fault pictures from the tread images of the wheel set through an unsupervised anomaly detection algorithm; Manually review the suspected fault pictures and label the fault areas pixel by pixel.
5. The method according to claim 1, characterized in that Before performing anomaly detection on the tread images of the wheel set through the target detection and segmentation model, a target detection and segmentation model is also built; Building the target detection and segmentation model includes: Modify the original Proto module in the yolov5-seg model. The network structure of the Yolov5-seg model includes four parts: an input end, a BackBone network, a Neck structure, and a Head prediction layer; Add a coordinate attention module and a spatial adaptive fusion (ASFF) module on the basis of the yolov5-seg model to obtain the target detection and segmentation model.
6. The method according to claim 5, wherein The step of performing anomaly detection on the tread images of the wheel set through the target detection and segmentation model includes: Preprocess the tread images of the wheel set through the input end; Extract useful features from the preprocessed tread images of the wheel set through the BackBone network to obtain corresponding feature maps; Fuse the feature map through the Neck structure, enhance the spatial location information of feature maps at different scales through the coordinate attention module, and transfer the enhanced multi-scale feature maps to the Head prediction layer; Adaptive fusion of multi-scale feature maps is achieved by the Spatial Adaptive Fusion ASFF module adaptively selecting feature maps at different levels for weighted fusion. The criterion for adaptive selection is to ensure simultaneous processing of feature maps of different sizes. The ASFF module is added in front of the Head prediction layer; Perform object detection on the multi-scale feature maps through the Head prediction layer to obtain target fault region information including the bounding box of the fault region, fault category information, and segmented image.
7. The method according to claim 6, wherein The adaptive fusion of multi-scale feature maps by adaptively selecting feature maps at different levels for weighted fusion through the ASFF module includes: The ASFF module takes feature maps of different sizes as inputs, performs upsampling on small-size feature maps at different ratios, and keeps the spatial features of large-size feature maps unchanged with the same resolution; Fuse the feature maps from different layers and normalize them through softmax, multiply and add them with the feature maps of the same resolution respectively to obtain the adaptively fused multi-scale feature maps.
8. The method according to claim 6, characterized in that, The enhancement of the spatial location information of feature maps at different scales through the coordinate attention module and the transfer of the enhanced multi-scale feature maps to the Head prediction layer includes: Perform pooling encoding on the input feature maps along the X direction and Y direction respectively through the coordinate attention module; Concatenate and fuse the encoding information in the X direction and Y direction, and decompose it into two independent tensors along the spatial dimension; Perform convolutional transformation on the two tensors to adjust their number of channels until they match the number of channels of the input feature map; Fuse the attention information in the X direction and Y direction with the original input feature map; Transfer the fused multi-scale feature maps to the Head prediction layer for object detection.
9. The method according to claim 6, characterized in that, The object detection of the multi-scale feature maps through the Head prediction layer to obtain target fault region information including the bounding box of the fault region, fault category information, and segmented image includes: Process the multi-scale feature maps through the Head prediction layer, including convolutional layers, pooling layers, and fully connected layers; Add a segmentation branch Proto to the Yolov5-seg model, restore the image to the input size through a bottom-up decoding process, and finally generate the segmentation result of the image; In the Yolov5-seg model, transfer the feature maps of different resolutions extracted at the Neck end to the object detection head, which is used to optimize object box, class regression, and box position regression tasks, and finally output the bounding box, class information, and pixel-level segmentation result of the target fault region; By referring to the bottom-up decoding structure in PaNet through the Proto module, the feature size is gradually restored from 80×80 to a resolution of 640×640. Through the concatenation of CBL and upsampling pooling operations, the input features are gradually restored to a size of 640×640 for pixel-level segmentation of target regions of different sizes; After being processed by the target detection head and the segmentation part, the target fault area information including the bounding box of the target fault area, class information, and the pixel-level segmentation image is obtained.
10. An abnormal detection device for the tread of a bullet train wheel, characterized in that, The device is used for the all-round dynamic image detection system of the EMU entering the depot. The detection system includes a tread acquisition module. The device includes: An image acquisition module, which is used to acquire the tread image of the EMU wheel pair through the tread acquisition module; An image processing module, which is used to sort out and annotate the acquired tread image of the wheel pair; An anomaly detection module, which is used to perform anomaly detection on the tread image of the wheel pair through the target detection and segmentation model; A fault detection module, which is used to filter out small areas by area for the detected target fault area information and timely return the position where the fault occurs to the platform.
Citation Information
Patent Citations
Vehicle tread scratch fault detection method
CN112712552A
TensorRT accelerated yolov5s-seg-based infrared image instance segmentation and low-light image background fusion algorithm
CN118505723A