Method and system for detecting steel rail defects based on improved RetinaNet network

Through the improved RetinaNet network, the CBAM attention module and EfficientNetB0 network are used to solve the problems of inaccurate and low efficiency of rail defect detection in the prior art, and efficient and accurate detection in complex environments are achieved, which reduces maintenance costs and extends the service life of rails.

CN120107755APending Publication Date: 2025-06-06BEIJING XIAOMINGZHI IRON TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510372689.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the detection of rail defects, the existing technology has the situation of low damage detection rate, high false alarm rate, manual intervention, high labor intensity and easy to miss rail damage in the detection of rail defects. Especially in environments with complex backgrounds and large lighting changes, it is difficult to achieve accurate and efficient detection.

Method used

The improved RetinaNet network is adopted, and the improved RetinaNet network is formed by replacing the SE module in the EfficientNetB0 network with the CBAM attention module and replacing the backbone network ResNet in the RetinaNet network with the improved EfficientNetB0 network. Through the combination of pre-training and feature pyramid network, this network can more effectively extract features in rail images and perform defect detection.

Benefits of technology

It realizes efficient and accurate detection of rail defects in complex backgrounds and large lighting changes, reduces manual intervention and false alarm rates, improves detection efficiency and accuracy, reduces maintenance costs, and extends the service life of rails.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107755A_ABST
    Figure CN120107755A_ABST
Patent Text Reader

Abstract

The invention provides a method for detecting steel rail defects based on an improved RetinaNet network, and the method comprises the steps: improving an EfficientNetB0 network and a backbone network ResNet in the RetinaNet network, and forming the improved RetinaNet network; pre-training the improved RetinaNet network by using the constructed training set, verification set and test set to obtain a pre-trained improved RetinaNet network; acquiring an image of the steel rail in an actual detection area, inputting the image into the pre-trained improved RetinaNet network for defect detection and outputting a defect detection result; and comparing and analyzing a defect detection result with a defect standard. The invention further provides a system for detecting the steel rail defects based on the improved RetinaNet network. According to the invention, real-time detection of steel rail defects is realized, the maintenance cost of a railway is reduced, and the accident rate is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image target detection, and in particular to a method and system for detecting rail defects based on an improved RetinaNet network. Background Art

[0002] At present, my country has 18 railway bureaus and more than 180 railway sections. Railway is a high-speed means of transportation. Rails, as the infrastructure supporting train operation, play a vital role in ensuring the safety of train operation.

[0003] Detecting rail defects can help ensure the safety of railway transportation. Damage or defects in rails can cause vibration and noise when the train is running, which not only affects the normal operation of the train, but also increases maintenance costs and reduces transportation efficiency. Therefore, timely detection and repair of rail defects can reduce maintenance costs and improve transportation efficiency.

[0004] With the development of computer vision, deep learning and other technologies, it has become possible to detect rail defects with the help of image processing and intelligent algorithms. These advanced technologies can improve the accuracy and efficiency of detection. The application of big data and Internet of Things technologies provides more data support for rail defect detection, such as data collected by sensors, images taken by drones, etc. Using this data for analysis and detection can better identify rail problems. Early detection and repair of rail defects can avoid accidents caused by them, and help implement preventive maintenance and extend the service life of rails.

[0005] As the hardware equipment of large-scale rail flaw detection vehicles is gradually improved, the rail damage defect detection technology based on deep learning technology still needs to be further improved, because it has problems such as low damage detection rate, high false alarm rate, common need for manual intervention, high labor intensity and easy to miss rail damage.

[0006] When the rail background occupies most of the area and there are many background pixels, how to accurately and efficiently detect various types of rail damage defects so as to carry out effective maintenance of the rails later becomes an urgent problem to be solved.

[0007] Chinese patent application 201910903588.4 discloses a rail defect detection method based on image reconstruction and block threshold segmentation, which performs a median filtering preprocessing step on the rail image to be detected to obtain an enhanced image; reconstructs the enhanced image to obtain a reconstructed image; subtracts the enhanced image from the reconstructed image to obtain a differential image; divides the differential image into image blocks of the same size, where the image block size is 80*80; calculates the average grayscale difference between classes of all image blocks; performs threshold segmentation processing on image blocks with an average grayscale difference between classes greater than 50, and determines other image blocks as pure background blocks, and sets the pixel values ​​of pure background blocks to "0", performs foreground and background area judgment and threshold segmentation on each image block to obtain a binary image; performs an open operation on the binary image to finally obtain a detection result image. However, it has the following shortcomings: (1) High computational complexity: Image reconstruction and block threshold segmentation algorithms usually require a large number of pixel-level calculations and image processing operations, which may lead to high computational complexity of the algorithm, especially for large-size or high-resolution images, which may increase the running time and resource consumption of the algorithm; (2) Difficult parameter selection: Image reconstruction and block threshold segmentation algorithms usually involve the selection of some parameters, such as filter type, threshold segmentation method and block size. The selection of appropriate parameters is crucial to the performance of the algorithm, but this usually requires adjustment by experienced professionals, and different optimal parameters may be generated due to changes in image characteristics and environmental conditions; (3) Over-reliance on column consistency characteristics: If the steel (3) Loss of local information: The block threshold segmentation algorithm divides the image into multiple blocks and performs threshold processing on each block, which may lead to the loss of some local information, especially for edge or small-sized defects, which may be misjudged or ignored due to block processing; (4) Sensitivity to illumination changes: Even if the block threshold segmentation algorithm is used, it may still be affected by uneven illumination, especially in outdoor or open environments, where illumination conditions may change significantly, resulting in inaccurate threshold selection, which in turn affects the effect of defect detection.

[0008] Chinese patent application 202111120010.5 discloses a rail defect detection method and system based on a deep residual shrinkage network, which integrates image sharpening, image rotation, image random cropping and image mosaic processing methods, uses a soft thresholding function sf(inputM,t) to reduce noise on the damaged image data after cleaning, and uses a soft attention mechanism to calculate the global descriptive features between features, obtain more nonlinear features and complex correlation features between channels, and better extract the correlation between nonlinear features and feature channels, so as to better and more accurately classify damage defects. However, the rail defect detection method and system based on the deep residual shrinkage network failed to take into account the characteristics of high-resolution images, resulting in a relatively simple algorithm design, a relatively complex detection process, and low detection efficiency.

[0009] Railway lines are usually located in outdoor environments and are affected by factors such as weather and light. The amount of data generated by the railway system is huge, and rail defect detection requires high accuracy, because even small defects can lead to accidents, so algorithms and technologies must be able to accurately identify various types of defects. Railway transportation needs to ensure the continuity and stability of train operation, so rail defect detection needs to be completed in real time or near real time so that repair measures can be taken in time. Moreover, rail defect detection requires a lot of manpower and material costs. How to reduce costs and improve efficiency while ensuring quality is a challenge, and the detection equipment needs to be frequently maintained and calibrated to ensure its normal operation and accuracy, which is also a problem that needs to be solved. Summary of the invention

[0010] In order to solve the deficiencies in the prior art, the purpose of the present invention is to provide a method and system for detecting rail defects based on an improved RetinaNet network to improve the above problems. To achieve the above purpose, the technical solution adopted by the present invention is as follows: In a first aspect, the present invention provides a method for detecting rail defects based on an improved RetinaNet network, comprising: The SE module in the EfficientNetB0 network is replaced with the CBAM attention module to improve the EfficientNetB0 network to obtain an improved EfficientNetB0 network, and the backbone network ResNet in the RetinaNet network is replaced with the improved EfficientNetB0 network to improve the RetinaNet network to form an improved RetinaNet network; Constructing a training set, a validation set and a test set for pre-training the improved RetinaNet network, and pre-training the improved RetinaNet network using the training set, the validation set and the test set to obtain a pre-trained improved RetinaNet network, wherein the training set, the validation set and the test set include images of rails to be detected for defects; Determine an actual detection area, collect an image of the rail in the actual detection area, input the collected image of the rail in the actual detection area into the pre-trained improved RetinaNet network for defect detection, and the pre-trained improved RetinaNet network outputs a defect detection result of the image of the rail in the actual detection area; The defect detection results of the rail images in the actual inspection area are compared and analyzed with the defect standards, and defects exceeding the defect standards are marked and warned.

[0011] Preferably, the improved RetinaNet network includes the improved EfficientNetB0 network, the FPN feature pyramid network and the classification regression head, wherein: The improved EfficientNetB0 network extracts feature maps from the images of the rails with defects to be detected in the training set, the validation set, and the test set, and sequentially generates a first output feature map, a second output feature map, a third output feature map, a fourth output feature map, a fifth output feature map, a sixth output feature map, and a seventh output feature map; The FPN feature pyramid network fuses all feature maps: the seventh output feature map is subjected to 256 1*1 convolution kernel operations to generate the topmost fused feature map of the FPN feature pyramid network, which is recorded as the first fused feature map; the first fused feature map is upsampled with a step size of 2, and then added to the result obtained by the 256 1*1 convolution kernel operations on the sixth output feature map to generate a second fused feature map; the second fused feature map is upsampled with a step size of 2, and then added to the result obtained by the 256 1*1 convolution kernel operations on the fifth output feature map to generate a third fused feature map; the third fused feature map is upsampled with a step size of 2, and then added to the result obtained by the 256 1*1 convolution kernel operations on the fourth output feature map to generate a fourth fused feature map; the first fused feature map is subjected to a maximum pooling operation with a step size of 2 to generate a fifth fused feature map; The classification regression head marks the anchor box positions for all the obtained fusion feature maps and predicts the target types circled by the anchor boxes.

[0012] Preferably, constructing a training set, a validation set and a test set for pre-training the improved RetinaNet network, and pre-training the improved RetinaNet network using the training set, the validation set and the test set to obtain a pre-trained improved RetinaNet network, comprises: By collecting image data of the steel rail to be inspected for defects, an image data set of the steel rail to be inspected for defects is obtained; Dividing the rail image dataset into a training set, a validation set, and a test set; Preprocessing the images of each rail to be detected in the training set, the validation set, and the test set to obtain enhanced images of the rail to be detected; The improved RetinaNet network is pre-trained using the training set, the validation set and the test set to obtain a pre-trained improved RetinaNet network.

[0013] Preferably, the preprocessing of the images of each rail to be detected in the training set to obtain an enhanced image of the rail to be detected includes: The threshold is determined by the maximum inter-class variance method; An improved histogram equalization algorithm is used to preprocess the collected rail images. The gradient information calculated from different windows is weighted averaged to obtain feature representation. For high-frequency areas greater than the threshold, the values ​​of all pixel feature points are increased by 1 to highlight the high-frequency features. For high-frequency areas less than the threshold, the values ​​of all pixel feature points remain unchanged. The features obtained after weighted averaging are normalized to obtain the channel image to be processed, which is used as the input feature of the training model.

[0014] Preferably, the determining of the actual detection area, collecting an image of the rail in the actual detection area, inputting the collected image of the rail in the actual detection area into the pre-trained improved RetinaNet network for defect detection, and the pre-trained improved RetinaNet network outputting a defect detection result of the image of the rail in the actual detection area, comprises: According to the detection requirements, an actual detection area is determined, and an image of the rail in the actual detection area is collected, wherein the actual detection area is located in an image collection area of ​​the image data set to be detected; Inputting the collected image of the rail in the actual detection area into the pre-trained improved RetinaNet network for defect detection, and the pre-trained improved RetinaNet network performs backbone feature extraction, feature fusion and classification regression processing on the image of the rail; Output the defect detection result of the image of the rail in the actual detection area by the pre-trained improved RetinaNet network.

[0015] In a second aspect, the present invention also provides a system for detecting rail defects based on an improved RetinaNet network, comprising an improvement module, a pre-training module, a detection module, and an analysis module, wherein: The improvement module is used to replace the SE module in the EfficientNetB0 network with the CBAM attention module to improve the EfficientNetB0 network to obtain an improved EfficientNetB0 network, and to replace the backbone network ResNet in the RetinaNet network with the improved EfficientNetB0 network to improve the RetinaNet network to form an improved RetinaNet network; The pre-training module is used to construct a training set, a validation set and a test set for pre-training the improved RetinaNet network, and to pre-train the improved RetinaNet network using the training set, the validation set and the test set to obtain a pre-trained improved RetinaNet network; The detection module is used to determine the actual detection area, collect images of the rails in the actual detection area, input the collected images of the rails in the actual detection area into the pre-trained improved RetinaNet network for defect detection, and the pre-trained improved RetinaNet network outputs defect detection results for the images of the rails in the actual detection area; The analysis module is used to compare and analyze the defect detection results of the rail image in the actual detection area with the defect standard, and to mark and warn the defects that exceed the defect standard.

[0016] Preferably, the pre-training module includes a first acquisition unit, a division unit, an image pre-processing unit, and a training unit, wherein: The first acquisition unit is used to acquire image data of the steel rail to be inspected for defects by acquiring image data of the steel rail to be inspected for defects; The division unit is used to divide the rail image data set into a training set, a verification set and a test set; The image preprocessing unit is used to preprocess the images of each rail to be detected in the training set, the verification set, and the test set to obtain enhanced images of the rail to be detected; The training unit is used to pre-train the improved RetinaNet network using the training set, the validation set and the test set to obtain a pre-trained improved RetinaNet network.

[0017] Preferably, the detection module includes a second acquisition unit, a processing unit and an output unit, wherein: The second acquisition unit is used to determine the actual detection area according to the detection requirements, and acquire the image of the rail in the actual detection area, wherein the actual detection area is located in the image acquisition area of ​​the image data set to be detected; The processing unit is used to input the collected image of the rail in the actual detection area into the pre-trained improved RetinaNet network for defect detection, and the pre-trained improved RetinaNet network performs backbone feature extraction, feature fusion and classification regression processing on the image of the rail; The output unit is used to output the defect detection result of the image of the rail in the actual detection area by the pre-trained improved RetinaNet network.

[0018] In a third aspect, the present invention further provides an electronic device, comprising a memory and a processor, wherein the memory stores a program running on the processor, and when the processor runs the program, the steps of the above-mentioned method for detecting rail defects based on the improved RetinaNet network are executed.

[0019] In a fourth aspect, the present invention further provides a computer-readable storage medium having computer instructions stored thereon, which, when executed, execute the steps of the above-mentioned method for detecting rail defects based on the improved RetinaNet network.

[0020] The method and system for detecting rail defects based on the improved RetinaNet network of the present invention have the following beneficial effects compared with the prior art: (1) Through accurate rail defect detection, rail defects can be discovered and repaired in a timely manner before the problem seriously affects train operation, which can extend the service life of rails and related facilities and avoid large-scale maintenance and replacement costs caused by untimely discovery; (2) By establishing a regular rail inspection and maintenance mechanism, preventive maintenance can be implemented to avoid losses and accidents caused by unexpected problems, reduce train delays and transportation interruptions caused by rail defects, and improve the efficiency and reliability of railway transportation; (3) Ensure the safety of the railway transportation system by timely discovering and repairing rail defects and prevent accidents and failures caused by rail problems; standardize and systematize the shooting conditions and shooting specifications for rail defect detection to form industry standards; (4) The present invention can replace manual operation, reduce labor costs and improve efficiency; (5) The present invention detects all rail defects in real time and focuses on the maintenance of defective rails to reduce railway maintenance costs and accident rates; (6) The present invention takes into account the multi-scale feature fusion of high-resolution images in its design, which can effectively capture target information at different scales, is conducive to the detection of rail defects of different sizes, simplifies the detection process and improves the detection efficiency, and has strong generalization ability.

[0021] Other features and advantages of the present invention will be set forth in the description which follows, and in part will be apparent from the description, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 Schematic diagram of rail top surface defects; Figure 2 A flow chart of a method for detecting rail defects based on an improved RetinaNet network according to the present invention; Figure 3 This is a schematic diagram of the EfficientNetB0 network structure; Figure 4 It is a schematic diagram of the MBConv module structure; Figure 5 It is a structural diagram of the SE module; Figure 6 Schematic diagram of the CBAM attention module structure; Figure 7 This is a schematic diagram of the RetinaNet network structure; Figure 8 A schematic diagram of the improved RetinaNet network structure according to the present invention; Fig. 9 Schematic diagram of the predicted box and the true box; Fig.10 The figure is a schematic diagram of a system for detecting rail defects based on an improved RetinaNet network according to the present invention. DETAILED DESCRIPTION

[0023] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not intended to limit the scope of protection of the present invention.

[0024] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.

[0025] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0026] It should be noted that the modifications of "one" and "plurality" mentioned in the present invention are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more". "Plurality" should be understood as two or more.

[0027] During the rail inspection process, it is found that since the rails are in an outdoor environment, they are easily affected by uncontrollable factors such as weather, light, and background clutter, and a large part of the rail top surface is a defect-free area. The number of these pixels is very large, such as Figure 1 shown.

[0028] It is necessary to screen the low-frequency part of the image (such as the railway background, rail top surface background, etc.) and identify the defect type in real time based on the high-resolution image collected by the turnout comprehensive inspection trolley. Example 1

[0029] The present invention provides a method for detecting rail defects based on an improved RetinaNet network. The embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Figure 2 For a flow chart of a method for detecting rail defects based on an improved RetinaNet network according to the present invention, see Figure 2 The method for detecting rail defects based on the improved RetinaNet network of the present invention includes steps S100, S200, S300, and S400.

[0030] Step S100: The SE module in the EfficientNetB0 network is replaced with the CBAM attention module to improve the EfficientNetB0 network to obtain an improved EfficientNetB0 network, and the backbone network ResNet in the RetinaNet network is replaced with the improved EfficientNetB0 network to improve the RetinaNet network to form an improved RetinaNet network.

[0031] The EfficientNetB0 network includes convolutional layers, pooling layers, fully connected layers, and multiple MBConv modules, such as Figure 3 As shown. Figure 3 It can be seen that the main component of the EfficientNetB0 network is the lightweight flip bottleneck convolution MBConv module. The MBConv module structure is as follows Figure 4 As shown. Figure 4 It can be seen that the feature matrix input to MBconv will be upgraded after a 1*1 convolution layer, BatchNormalization processing and swish function activation, and then pass through a k*k depth separable convolution layer (DWConv), and then Batch Normalization processing and swish function activation again, and then pass through an SE module. A 1*1 convolution layer is connected after the SE module. After receiving the input of the SE module, the convolution layer convolves it and performs Batch Normalization processing, and finally passes through the drop out layer for a certain proportion of inactivation, and finally outputs the processed feature matrix. The structure of the SE module is as follows: Figure 5 As shown in the figure, the SE module is a channel attention mechanism that mainly obtains the importance of each channel through automatic learning, and then assigns weights to each channel according to the importance, so that the neural network focuses on channels with large weights and learns useful features.

[0032] It can be understood that the present invention replaces the SE module with the CBAM attention module to improve the EfficientNetB0 network. The structure diagram of the CBAM attention module is as follows: Figure 6 The CBAM module combines the Channel Attention Module and the Spatial Attention Module to enhance the convolutional neural network's ability to focus on images. The channel attention module passes the input feature map through two parallel MaxPool layers and AvgPool layers, changes the feature map from C*H*W to C*1*1, and then extracts features through the Share MLP module, and obtains two activated results through the ReLU activation function. The two output results are added element by element, and then a sigmoid activation function is used to obtain the output of the channel attention module, and then this output is multiplied by the original image to change back to the size of C*H*W. Global average pooling and global maximum pooling are used to obtain global statistics for each channel respectively (SENet only uses global average pooling), and the weights of the channels are learned through two fully connected layers. Then, the two processed results are added, and the weights are normalized to between 0 and 1 using the Sigmoid function, and each channel is scaled. Finally, the scaled channel features are multiplied with the original features to produce features with enhanced channel importance.

[0033] The spatial attention module obtains two 1*H*W feature maps by max pooling and average pooling from the output of the channel attention module, and then concatenates the two feature maps through the Concat operation, converts them into a 1-channel feature map through a 7*7 convolution (experiments show that 7*7 is better than 3*3), and then obtains the feature map of the spatial attention module through a sigmoid. Finally, the output result is multiplied by the original image to return to the size of C*H*W, and the maximum and average pooling are used to obtain the maximum and average values ​​of each spatial position. Specifically, since multiple channels are generated after convolution, finally, weights are applied to each spatial position on the feature map to produce features with enhanced spatial importance.

[0034] The SE module is replaced with the CBAM attention module to improve the EfficientNetB0 network, and the improved EfficientNetB0 network is obtained. Then, the improved EfficientNetB0 network is used to improve the backbone network of the RetinaNet network.

[0035] The RetinaNet network structure is as follows Figure 7 As shown, it includes a backbone network, an FPN feature pyramid network and a classification regression head, wherein the backbone network can adopt ResNet50 or ResNet101, and ResNet50 or ResNet101 are collectively referred to as ResNet here.

[0036] The improvement of the RetinaNet network in the present invention is to replace ResNet with the EfficientNetB0 network to obtain an improved RetinaNet network, such as Figure 8 shown.

[0037] The resolution of the 2D camera installed on the turnout comprehensive inspection vehicle for collecting rail images is high resolution, so the improved EfiicientNetB0 network is used to replace ResNet as the backbone network for feature extraction. EfficientNetB0 can smoothly control the depth, width and input resolution of the model to form a combination relationship. In addition, the three networks of EfficientNet-B0, ResNet-50 and DenseNet-169 have the highest accuracy in the ImageNet dataset and the least number of parameters, so the amount of calculation is the least. The number of parameters of EfficientNetB0 is 4.9 times lower than that of ResNet-50, and the top-1 accuracy has increased by 1.1%. Therefore, EfficientNetB0 is used for defect analysis of rail images compared to RetinaNet's original backbone network ResNet-50.

[0038] Improve RetinaNet's backbone network ResNet, replace it with the EfficientNet network, and introduce the CBAM attention mechanism to improve the EfficientNetB0 network to extract damaged channel and spatial features. The CBAM module combines channel attention and spatial attention, and can model both channel and spatial dimensions at the same time. This means that CBAM can more comprehensively capture the relationship between channels in the feature map and the relationship between spatial positions, thereby improving the expressiveness of feature representation. The CBAM module can achieve multi-level feature representation learning by stacking multiple sub-modules. This multi-level feature representation can better capture feature information at different levels, which helps to improve network performance.

[0039] The improved RetinaNet network includes an improved EfficientNetB0 network, an FPN feature pyramid network and a classification regression head, wherein: The improved EfficientNetB0 network includes a convolution module, a first MBConv module, a second MBConv module, a third MBConv module, a fourth MBConv module, a fifth MBConv module, a sixth MBConv module, and a seventh MBConv module, wherein: The convolution module is a common convolution layer with a kernel size of 3x3 and a stride of 2; The first MBConv module has a multiplication factor of 1, and the convolution kernel size used by its Depthwise Conv is 3x3. The first MBConv module repeats the MBConv structure once; The second MBConv module has a multiplication factor of 6, and the convolution kernel size used by its Depthwise Conv is 3x3. The second MBConv module repeats the MBConv structure twice; The third MBConv module has a multiplication factor of 6, and the convolution kernel size used by its Depthwise Conv is 5x5. The third MBConv module repeats the MBConv structure twice; The fourth MBConv module has a multiplication factor of 6, and the convolution kernel size used by its Depthwise Conv is 3x3. The fourth MBConv module repeats the MBConv structure 3 times; The fifth MBConv module has a multiplication factor of 6, and the convolution kernel size used by its Depthwise Conv is 5x5. The fifth MBConv module repeats the MBConv structure 3 times; The sixth MBConv module has a multiplication factor of 6, and the convolution kernel size used by its Depthwise Conv is 5x5. The sixth MBConv module repeats the MBConv structure 4 times; The seventh MBConv module has a multiplication factor of 6, and the convolution kernel size used by its Depthwise Conv is 3x3. The seventh MBConv module repeats the MBConv structure once.

[0040] It should be noted that the multiplication factor 1 or 6 means that the first 1x1 convolution layer in MBConv will expand the channels of the input feature matrix by 1 or 6 times.

[0041] The input image is processed by the convolution module and the first MBConv module in sequence, and the first output feature map L1 is output. The first output feature map is processed by the second MBConv module, and the second output feature map L2 is output. The second output feature map L2 is processed by the third MBConv module, and the third output feature map L3 is output. The third output feature map L3 is processed by the fourth MBConv module, and the fourth output feature map L4 is output by the fifth MBConv module, and the fifth output feature map L5 is processed by the sixth MBConv module, and the sixth output feature map L6 is output. The seventh output feature map L7 is processed by the seventh MBConv module.

[0042] The FPN feature pyramid network is an improvement on the previous RetinaNet network. It uses the 5-layer feature map extracted by the backbone network as the input of the FPN feature pyramid network. The FPN feature pyramid network transforms the scale of the upper feature map from top to bottom to construct a new feature map. The new feature map needs to keep the same scale as the feature map of the lower layer to ensure that the feature maps can be fused together. In the length and width directions, the upsampling method is used to pull the width and height of the lower feature map to the same size; in the depth direction, a 1×1 convolution is used to compress the depth of the upper feature map to the same depth as the lower feature map. The fusion of multi-scale feature maps simultaneously extracts the deep and shallow information of the feature map, increasing the receptive field.

[0043] The top-down process of the FPN feature pyramid network is performed by upsampling, and the lateral connection is to merge the upsampling result with the feature map of the same size generated from the bottom up. After the fusion, a 3*3 convolution kernel is used to convolve each fusion result to eliminate the aliasing effect of upsampling. Specifically, The seventh output feature map L7 is processed by 256 1*1 convolution kernels to generate the top-level feature of the FPN feature pyramid network, which is recorded as P5. P5 is upsampled with a step size of 2, and then added to the result obtained by the 256 1*1 convolution kernel operations on the sixth output feature map L6 to generate a feature map, which is recorded as P4; P4 is upsampled with a step size of 2, and then added to the result obtained by the 256 1*1 convolution kernel operations on the fifth output feature map L5 to generate a feature map, which is recorded as P3; P3 is upsampled with a step size of 2, and then added to the result obtained by the 256 1*1 convolution kernel operations on the fourth output feature map L4 to generate a feature map, which is recorded as P2; Perform a maximum pooling operation with a step size of 2 on P5 to generate a feature map, which is recorded as P6.

[0044] After obtaining the P2, P3, P4, P5 and P6 layer feature maps with complete semantics, the two-way detection head of the classification regression head is used to predict the anchor box position and the target type enclosed by the anchor box. The input feature maps of the two branches are both H×W×256, where The classification branch network uses 5 3×3 convolution operations to obtain a H×W×1 feature map and counts the classification scores of each sampling point; The regression branch network also uses five 3×3 convolution operations to obtain a H×W×4 feature map, and counts the four-dimensional vector of the position information of each sampling point.

[0045] Specifically, the classification prediction of the classification branch network: RetinaNet belongs to a single-stage target detection method. The generation of anchor boxes does not depend on the characteristics of the image itself, so a large number of negative samples that do not contain the target to be detected will be generated, making the entire training process inefficient. The use of the Focal loss function can greatly reduce the number of negative sample prediction boxes and improve the training speed. The Focal loss function is: FL = -α(1-p) γ log(p) Among them, FL represents the Focal loss function, α represents the weight parameter between categories, -log(p) is the initial cross entropy loss function, (1-p) γ represents the easy / difficult sample adjustment factor, γ is the focusing parameter, and p can be expressed as: The value range of w is 0~1, which indicates the probability that the model predicts that the image belongs to the foreground. The value of y is 1 or -1, which represents the foreground and background, respectively.

[0046] Regression prediction of regression branch network: The present invention uses CIOU function to replace the original IOU function, and the original IOU function is:

[0047] Among them, A represents the predicted box, B represents the real box, A∩B represents the intersection of the predicted box and the real box, and A∪B represents the union of the predicted box and the real box; If the two boxes do not intersect, it cannot reflect the distance between the two boxes, nor can it accurately reflect the degree of overlap between the two boxes.

[0048] The CIOU function can be expressed as: Among them, it can be understood that d is the distance between the center point of the predicted box and the real box, and c is the diagonal distance of the minimum circumscribed rectangle, such as Fig. 9 As shown, the lighter blue area represents the predicted box, the darker blue area represents the true box, and α is expressed as: in, Represents the correction factor, which is used to further adjust the loss function, taking into account the shape and direction of the target box. It can be expressed as: in, and are the width and height of the target box and the predicted box respectively.

[0049] The CIOU function introduces a minimum bounding box based on the IOU feature to solve the problem that the loss is equal to 0 when the detection box and the real box do not overlap. It describes the size of the overlapping area between the predicted box and the real box, increases the loss of the detection box scale, and increases the loss of length and width, so that the predicted box will be more consistent with the real box. The CIOU function can also measure the difference in the aspect ratio and the distance between the center points of the two to achieve real-time adjustment of network weights.

[0050] The CIOU function is used to measure the difference between the predicted box and the true box, and the model training parameters are updated in time to effectively improve the model detection accuracy. The improved EfficientNetB0 network effectively improves the accuracy by 1.2% and the average real-time detection time for each image is increased by 5ms.

[0051] Step S200: constructing a training set, a validation set and a test set for pre-training the improved RetinaNet network, and pre-training the improved RetinaNet network using the training set, the validation set and the test set to obtain a pre-trained improved RetinaNet network.

[0052] It can be understood that step S200 includes step S201, step S202, step S203 and step S204. Step S201: acquiring image data of the steel rail with defects to be detected by collecting image data of the steel rail with defects to be detected.

[0053] It should be noted that the present invention utilizes an improved RetinaNet network to detect rail defects, so an image data set of the rails to be detected is obtained by collecting image data of the rails to be detected, wherein one or more sections of the existing railway line can be selected when collecting image data of the rails to be detected, and image data can be collected using a BASLER line scanning 2D camera installed on a turnout comprehensive inspection trolley to obtain images of the rails to be detected, and multiple images of the rails to be detected constitute an image data set of the rails to be detected.

[0054] Furthermore, in order to take into account the effect of illumination on the image of the rail to be inspected for defects, the image acquisition time may include 9:30-11:30 am, 3:30-5:30 pm, and 11:00 pm-3:00 am.

[0055] Step S202: Divide the rail image dataset into a training set, a validation set and a test set.

[0056] It can be understood that before dividing the image dataset of the rail into a training set, a validation set, and a test set, the dat file of the image in the image dataset of the rail to be detected needs to be converted into a jpg file with a unified pixel of 1024×1024, and then the image dataset of the rail to be detected is divided according to the training set: validation set: test set, and the training set: validation set: test set division ratio can be 8:1:1.

[0057] Step S203: preprocessing the images of each rail to be detected in the training set to obtain an enhanced rail image to be detected, specifically, comprising the following steps: The threshold T is determined by the maximum between-class variance method (OSTU method); The improved histogram equalization algorithm is used to preprocess the collected rail images: first, a 3x3 window is used to calculate the gradient information around the center point, and the point value below the center point is subtracted from the point value above the center point, and the point value to the right of the center point is subtracted from the point value to the left of the center point. If their difference is greater than a threshold T, it means that there is a gradient at this center point, and this gradient information is recorded as G1. 1x3 and 3x1 windows are used to calculate the gradient information in the horizontal and vertical directions respectively, and the point value above the center point is subtracted from the point value below the center point, and the point value to the left of the center point is subtracted from the point value to the right of the center point. If their difference is greater than a threshold T, it means that there is a gradient at this center point, and this gradient information is recorded as G2 and G3 respectively; The gradient information calculated from different windows is weighted averaged to obtain a more balanced feature representation: a weight can be assigned to the gradient information calculated for each window, and then their weighted sum is used to obtain the integrated features, where for high-frequency areas greater than the threshold T, the values ​​of all pixel feature points are increased by 1 to highlight the high-frequency features, and for high-frequency areas less than the threshold T, the values ​​of all pixel feature points remain unchanged.

[0058] The features obtained after adjusting the pixel feature points are normalized to obtain a 224 pixel × 224 pixel × 3 channel image to be processed, which is used as the input feature of the training model.

[0059] The improved histogram equalization method can effectively improve the significance of defect features acquired by high-resolution cameras, avoid the influence of railway background noise, and thus improve the recall rate of rail top surface defect detection.

[0060] Step S204: pre-training the improved RetinaNet network using the training set, the validation set and the test set to obtain a pre-trained improved RetinaNet network.

[0061] It is understandable that the method of pre-training the improved RetinaNet network using the training set, the validation set and the test set can adopt the commonly used method in the art, which will not be described in detail here.

[0062] Step S300: determine the actual inspection area, collect images of the rails in the actual inspection area, input the collected images of the rails in the actual inspection area into the pre-trained improved RetinaNet network for defect detection, and the pre-trained improved RetinaNet network outputs the defect detection results of the images of the rails in the actual inspection area.

[0063] It can be understood that step S300 includes step S301, step S302 and step S303, wherein: Step S301: determining an actual detection area according to detection requirements, and collecting images of rails in the actual detection area, wherein the actual detection area is located in an image collection area of ​​an image data set to be detected; Step S302: inputting the collected image of the rail in the actual detection area into the pre-trained improved RetinaNet network for defect detection, wherein the pre-trained improved RetinaNet network performs backbone feature extraction, feature fusion, and classification regression processing on the image of the rail; Step S303: outputting the defect detection result of the image of the rail in the actual detection area by the pre-trained improved RetinaNet network.

[0064] Step S400: comparing and analyzing the defect detection results of the rail image in the actual detection area with the defect standard, and marking and warning the defects exceeding the defect standard.

[0065] It is understandable that the defect standard can be formulated based on my country's railway "Rail Damage Classification" (TB / T1778-2010) standard. Example 2

[0066] like Fig.10 As shown, this embodiment provides a system for detecting rail defects based on an improved RetinaNet network, see Fig.10 The system includes an improvement module 901, a pre-training module 902, a detection module 903, and an analysis module 904, wherein: Improvement module 901: used to replace the SE module in the EfficientNetB0 network with the CBAM attention module to improve the EfficientNetB0 network, thereby obtaining an improved EfficientNetB0 network, and to replace the backbone network ResNet in the RetinaNet network with the improved EfficientNetB0 network to improve the RetinaNet network, thereby forming an improved RetinaNet network; Pre-training module 902: used to construct a training set, a validation set and a test set for pre-training the improved RetinaNet network, and pre-train the improved RetinaNet network using the training set, the validation set and the test set to obtain a pre-trained improved RetinaNet network; Detection module 903: used to determine the actual detection area, collect images of the rails in the actual detection area, input the collected images of the rails in the actual detection area into the pre-trained improved RetinaNet network for defect detection, and the pre-trained improved RetinaNet network outputs defect detection results for the images of the rails in the actual detection area; Analysis module 904: used to compare and analyze the defect detection results of the rail images in the actual detection area with the defect standards, and mark and warn the defects that exceed the defect standards.

[0067] Specifically, the pre-training module 902 includes a first acquisition unit 9021, a division unit 9022, an image pre-processing unit 9023, and a training unit 9024, wherein: The first acquisition unit 9021 is used to acquire image data of the steel rail to be detected for defects by acquiring image data of the steel rail to be detected for defects; Division unit 9022: used to divide the rail image data set into a training set, a validation set and a test set.

[0068] Image preprocessing unit 9023: used to preprocess the images of each rail to be detected in the training set to obtain an enhanced image of the rail to be detected; Training unit 9024: used to pre-train the improved RetinaNet network using the training set, the validation set and the test set to obtain a pre-trained improved RetinaNet network.

[0069] Specifically, the detection module 903 includes a second acquisition unit 9031, a processing unit 9032 and an output unit 9033, wherein: The second acquisition unit 9031 is used to determine the actual detection area according to the detection requirements, and to acquire the image of the rail in the actual detection area, wherein the actual detection area is located in the image acquisition area of ​​the image data set to be detected; Processing unit 9032: used for inputting the collected images of the rails in the actual detection area into the pre-trained improved RetinaNet network for defect detection, wherein the pre-trained improved RetinaNet network performs trunk feature extraction, feature fusion and classification regression processing on the images of the rails; Output unit 9033: used to output the defect detection result of the image of the rail in the actual detection area by the pre-trained improved RetinaNet network.

[0070] It should be noted that, regarding the system in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here. Example 3

[0071] The embodiment of the present invention further provides an electronic device, including a memory and a processor, wherein the memory stores a program running on the processor, and the processor executes the steps of the method for detecting rail defects based on the improved RetinaNet network when running the program. The method for detecting rail defects based on the improved RetinaNet network is described in the above section and will not be described in detail. Example 4

[0072] An embodiment of the present invention further provides a computer-readable storage medium having computer instructions stored thereon. When the computer instructions are executed, the steps of the above-mentioned method for detecting rail defects based on the improved RetinaNet network are executed. The method for detecting rail defects based on the improved RetinaNet network is described in the aforementioned part and will not be described in detail.

[0073] Those skilled in the art can understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention is described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions recorded in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for detecting rail defects based on an improved RetinaNet network, comprising: The SE module in the EfficientNetB0 network is replaced with the CBAM attention module to improve the EfficientNetB0 network to obtain an improved EfficientNetB0 network, and the backbone network ResNet in the RetinaNet network is replaced with the improved EfficientNetB0 network to improve the RetinaNet network to form an improved RetinaNet network; Constructing a training set, a validation set and a test set for pre-training the improved RetinaNet network, and pre-training the improved RetinaNet network using the training set, the validation set and the test set to obtain a pre-trained improved RetinaNet network, wherein the training set, the validation set and the test set include images of rails to be detected for defects; Determine an actual detection area, collect an image of the rail in the actual detection area, input the collected image of the rail in the actual detection area into the pre-trained improved RetinaNet network for defect detection, and the pre-trained improved RetinaNet network outputs a defect detection result of the image of the rail in the actual detection area; The defect detection results of the rail images in the actual inspection area are compared and analyzed with the defect standards, and defects exceeding the defect standards are marked and warned.

2. The method for detecting rail defects based on the improved RetinaNet network according to claim 1 is characterized in that: The improved RetinaNet network includes the improved EfficientNetB0 network, the FPN feature pyramid network and the classification regression head, wherein: The improved EfficientNetB0 network extracts feature maps from the images of the rails with defects to be detected in the training set, the validation set, and the test set, and sequentially generates a first output feature map, a second output feature map, a third output feature map, a fourth output feature map, a fifth output feature map, a sixth output feature map, and a seventh output feature map; The FPN feature pyramid network fuses all feature maps: the seventh output feature map is subjected to 256 1*1 convolution kernel operations to generate the topmost fused feature map of the FPN feature pyramid network, which is recorded as the first fused feature map; the first fused feature map is upsampled with a step size of 2, and then added to the result obtained by the 256 1*1 convolution kernel operations on the sixth output feature map to generate a second fused feature map; the second fused feature map is upsampled with a step size of 2, and then added to the result obtained by the 256 1*1 convolution kernel operations on the fifth output feature map to generate a third fused feature map; the third fused feature map is upsampled with a step size of 2, and then added to the result obtained by the 256 1*1 convolution kernel operations on the fourth output feature map to generate a fourth fused feature map; the first fused feature map is subjected to a maximum pooling operation with a step size of 2 to generate a fifth fused feature map; The classification regression head marks the anchor box positions for all the obtained fusion feature maps and predicts the target types circled by the anchor boxes.

3. The method for detecting rail defects based on the improved RetinaNet network according to claim 1 is characterized in that: The step of constructing a training set, a validation set, and a test set for pre-training the improved RetinaNet network, and pre-training the improved RetinaNet network using the training set, the validation set, and the test set to obtain a pre-trained improved RetinaNet network includes: By collecting image data of the steel rail to be inspected for defects, an image data set of the steel rail to be inspected for defects is obtained; Dividing the rail image dataset into a training set, a validation set, and a test set; Preprocessing the images of each rail to be detected in the training set, the validation set, and the test set to obtain enhanced images of the rail to be detected; The improved RetinaNet network is pre-trained using the training set, the validation set and the test set to obtain a pre-trained improved RetinaNet network.

4. The method for detecting rail defects based on the improved RetinaNet network according to claim 3 is characterized in that: The preprocessing of the images of each rail to be detected in the training set to obtain an enhanced rail image to be detected includes: The threshold is determined by the maximum inter-class variance method; An improved histogram equalization algorithm is used to preprocess the collected rail images. The gradient information calculated from different windows is weighted averaged to obtain feature representation. For high-frequency areas greater than the threshold, the values ​​of all pixel feature points are increased by 1 to highlight the high-frequency features. For high-frequency areas less than the threshold, the values ​​of all pixel feature points remain unchanged. The features obtained after weighted averaging are normalized to obtain the channel image to be processed, which is used as the input feature of the training model.

5. The method for detecting rail defects based on the improved RetinaNet network according to claim 1, characterized in that: The determining of the actual detection area, collecting an image of the rail in the actual detection area, inputting the collected image of the rail in the actual detection area into the pre-trained improved RetinaNet network for defect detection, and the pre-trained improved RetinaNet network outputting a defect detection result of the image of the rail in the actual detection area, comprises: According to the detection requirements, an actual detection area is determined, and an image of the rail in the actual detection area is collected, wherein the actual detection area is located in an image collection area of ​​the image data set to be detected; Inputting the collected image of the rail in the actual detection area into the pre-trained improved RetinaNet network for defect detection, and the pre-trained improved RetinaNet network performs backbone feature extraction, feature fusion and classification regression processing on the image of the rail; Output the defect detection result of the image of the rail in the actual detection area by the pre-trained improved RetinaNet network.

6. A system for detecting rail defects based on an improved RetinaNet network, comprising an improvement module, a pre-training module, a detection module, and an analysis module, characterized in that: The improvement module is used to replace the SE module in the EfficientNetB0 network with the CBAM attention module to improve the EfficientNetB0 network to obtain an improved EfficientNetB0 network, and to replace the backbone network ResNet in the RetinaNet network with the improved EfficientNetB0 network to improve the RetinaNet network to form an improved RetinaNet network; The pre-training module is used to construct a training set, a validation set and a test set for pre-training the improved RetinaNet network, and to pre-train the improved RetinaNet network using the training set, the validation set and the test set to obtain a pre-trained improved RetinaNet network; The detection module is used to determine the actual detection area, collect images of the rails in the actual detection area, input the collected images of the rails in the actual detection area into the pre-trained improved RetinaNet network for defect detection, and the pre-trained improved RetinaNet network outputs defect detection results for the images of the rails in the actual detection area; The analysis module is used to compare and analyze the defect detection results of the rail image in the actual detection area with the defect standard, and to mark and warn the defects that exceed the defect standard.

7. The system for detecting rail defects based on the improved RetinaNet network according to claim 6 is characterized in that: The pre-training module includes a first acquisition unit, a division unit, an image pre-processing unit, and a training unit, wherein: The first acquisition unit is used to acquire image data of the steel rail to be inspected for defects by acquiring image data of the steel rail to be inspected for defects; The division unit is used to divide the rail image data set into a training set, a verification set and a test set; The image preprocessing unit is used to preprocess the images of each rail to be detected in the training set, the verification set, and the test set to obtain enhanced images of the rail to be detected; The training unit is used to pre-train the improved RetinaNet network using the training set, the validation set and the test set to obtain a pre-trained improved RetinaNet network.

8. The system for detecting rail defects based on the improved RetinaNet network according to claim 6 is characterized in that: The detection module includes a second acquisition unit, a processing unit and an output unit, wherein: The second acquisition unit is used to determine the actual detection area according to the detection requirements, and acquire the image of the rail in the actual detection area, wherein the actual detection area is located in the image acquisition area of ​​the image data set to be detected; The processing unit is used to input the collected image of the rail in the actual detection area into the pre-trained improved RetinaNet network for defect detection, and the pre-trained improved RetinaNet network performs backbone feature extraction, feature fusion and classification regression processing on the image of the rail; The output unit is used to output the defect detection result of the image of the rail in the actual detection area by the pre-trained improved RetinaNet network.

9. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a program running on the processor, and the processor executes the steps of the method for detecting rail defects based on an improved RetinaNet network as described in any one of claims 1 to 5 when running the program.

10. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instructions are executed, the steps of the method for detecting rail defects based on the improved RetinaNet network described in any one of claims 1 to 5 are executed.

Citation Information

Patent Citations

  • Steel rail defect detection method based on image reconstruction and partitioning threshold segmentation

    CN110687123A

  • Steel rail defect detection method and system based on deep residual shrinkage network

    CN113888488A