Rail surface defect detection method based on improved UPerNet and connected domain analysis

Through improved UPerNet network and connectivity domain analysis technology, the problems of limited information and low detection efficiency in rail surface defect detection are solved, and robust segmentation and parameter calculation of complex defects are realized, and detection accuracy and maintenance efficiency are improved.

CN116977280BActive Publication Date: 2025-05-13LANZHOU JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310737035.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-21
Publication Date
2025-05-13
Estimated Expiration
2043-06-21

AI Technical Summary

Technical Problem

The prior art has problems such as limited information, low detection efficiency, difficulty in effectively collecting defect images, and difficulty in directly connecting the defect parameters in the detection of rail surface defects, especially when dealing with complex edge shapes and random, multi-scale discrete defects.

Method used

Using the improved UPerNet network, Swin-T is used as the backbone network for feature extraction, combined with connectivity domain analysis and cross-card synchronous batch normalization, and using the Lovász-hinge loss function to improve detection accuracy and robustness.

Benefits of technology

It realizes robust segmentation of complex rail defects, improves detection accuracy, and can intuitively calculate the specific parameters of the defect, such as length, width and area, to help railway workers carry out effective repairs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116977280B_ABST
    Figure CN116977280B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of rail surface defect detection, and specifically to a rail surface defect detection method based on improved UPerNet and connected domain analysis. The Swin‑T network with a Transformer architecture is used to extract defect features, which makes full use of the global information in the image and avoids the problem of inductive preference. The regular window self-attention in Swin‑T reduces the number of parameters of the model. Secondly, gradient optimization is performed using cross-card synchronous batch normalization, and Lovász‑hinge is used as the loss function, incorporating the use of dependencies between pixels. In the face of complex rail scenes, the improved method of the present invention can achieve robust segmentation of defects and solve the problem of low prediction accuracy of defect edges. Finally, on the basis of the rail surface defect semantic segmentation map, the connected domain analysis is used to distinguish different defect areas therein, and the actual length and actual area of ​​the defect are calculated, which helps railway maintenance personnel to intuitively understand the defect parameters and make corresponding repairs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of rail surface defect detection, and in particular to a rail surface defect detection method based on improved UPerNet and connected domain analysis. Background Art

[0002] In the wheel-rail system, the rail bears the vertical load of the vehicle and plays a guiding role. The rail head is subject to longitudinal and lateral forces, and the entire track system will also vibrate. These forces will cause surface defects of the rail. Surface defects of the rail will aggravate the vertical vibration of the rail and transmit it to the locomotive, increasing the risk of resonance between the elastic components and the vehicle. At the same time, it increases the torque of the rail and weakens the guiding role of the wheel, endangering the lateral stability of the vehicle. Therefore, it is necessary to detect rail surface defects in a timely manner.

[0003] Due to the long line, complex natural environment and high train density, manual inspection can no longer meet the needs. With the development of machine vision and deep learning, non-contact image-based defect detection methods have become a research hotspot. Traditional visual inspection methods use features such as texture and low-level grayscale to identify and locate target areas in unmarked rail samples.

[0004] At present, the image-based detection method of rail surface defects still has the following problems: 1) The information available for identification in rail images is limited, resulting in the inefficiency of methods based on texture and grayscale features; 2) Due to the timely maintenance of the engineering department, it is difficult to effectively collect a large number of rail defect images, which affects the detection effect of supervised deep learning; 3) The GTC-80 track inspection vehicle widely used in China can meet the image detection of fastener defects, but its rail damage detection data is different from traditional images and cannot be used directly; 4) Rail surface defects mainly have two forms of manifestation: corrugation defects and discrete defects. Among them, discrete defects appear randomly, arbitrarily, and multi-scale, without any repeatable characteristics, which increases the difficulty of defect detection; 5) The complex edge shape of rail surface defects poses a challenge to the semantic segmentation algorithm; 6) The detection results of deep learning methods are difficult to directly link with the specific parameters of the defects.

[0005] In the prior art, the UPerNet network is a semantic segmentation network based on the pyramid pooling module (PPM) and the image pyramid (FPN). It combines the four feature layers extracted by ResNet, fully integrates the multi-scale information output at different stages, and reduces the loss of boundary information compared with the PSPNet network that also uses the image pyramid structure. However, the use of CNN for feature extraction in the UPerNet network has certain inductive preference problems, and it is difficult to use global information; when a single card is trained and the batch size is relatively small, the use of BN will affect the convergence effect of the model; when the cross entropy loss function is used, the dependency between adjacent pixels will be ignored. Step 1 Although the CNN-based network performs well in a variety of image processing tasks, the locality of the convolution operation and the inductive bias of weight sharing bring about the limitation of long-distance dependency, making it unable to learn the information in the entire image and unable to complete the information interaction in the image beyond the range of the convolution kernel. In contrast, the self-attention mechanism in the Transformer can dynamically adjust the receptive field, enabling it to have the ability to model long-distance dependencies and avoid the inductive preference problem in CNN. When using classic batch normalization (BN) as the normalization method, the input is divided into multiple parts for each iteration in network training, and then forward and backward propagation and gradient solution are performed on different GPUs respectively. The next iteration starts after the gradients and parameters are merged. The principle is as follows Figure 5 As shown. Since each GPU processes data separately in forward and backward propagation, the BN operation is also completed separately in each GPU. The samples normalized by the actual BN are limited to each GPU, which is equivalent to reducing the batch_size. In deep learning, a larger batch_size makes the training process more stable. When using cross entropy as the segmentation loss function, if the number of foreground pixels is much smaller than the number of background pixels, its inter-class competition mechanism will make the background pixels dominate, causing the optimization direction of the model to be seriously biased towards the background, which is very unfavorable for the segmentation of small-sized rail surface defects. At the same time, the cross entropy loss is calculated pixel by pixel, which ignores the relationship between pixels. There is a strong dependency between pixels in the image, which carries important information about the target structure. The use of cross entropy loss may cause the segmentation results to contain fuzzy predictions. The current deep learning detection algorithm is mainly used to determine whether there are rail defects, and is not combined with the specific parameters of the defects. Summary of the invention

[0006] The present invention provides a rail surface defect detection method based on improved UPerNet and connected domain analysis to solve the defects existing in the background technology.

[0007] In order to achieve the above object, the present invention provides the following technical solutions:

[0008] The rail surface defect detection method based on improved UPerNet and connected domain analysis is characterized by comprising the following steps:

[0009] Step 1: Feature extraction network: Use Swin-T as the backbone network to extract richer global information and more discriminative features. By limiting the attention calculation to non-overlapping local windows, the computational complexity is linear with the image size, thereby reducing the amount of calculation. The network structure is as follows: Figure 2 shown.

[0010] Preferably, the entire network adopts a hierarchical design. First, the input H×W×3 image is split into (H / 4)×(W / 4) non-overlapping tiles through the tile splitting module. Each tile is flattened into a 48-dimensional token vector. The linear embedding layer maps the tensor of dimension (H / 4)×(W / 4)×48 to an arbitrary dimension C to obtain a linear mapping of dimension (H / 4)×(W / 4)×C. The mapped tokens are then sent to paired SwinTransformer modules, which keep the number of input and output tokens unchanged at (H / 4)×(W / 4). The first Swin Transformer module and the linear embedding layer are jointly designated as the first layer. After that, the tile merging layer and the Swin Transformer modules are combined into subsequent layers to generate feature descriptions of different scales. The tile merging layer is responsible for downsampling operations, reducing tile resolution and expanding the receptive field layer by layer. The SwinTransformer module is used for feature extraction. As the network deepens, each layer changes the dimension of the tensor, thus forming a hierarchical representation.

[0011] Specifically, the Swin Transformer module consists of layer normalization (LN), a multi-head self-attention module, short-circuit connections, and a multi-layer perceptron (MLP) with a GELU activation function. The structure of two cascaded Swin Transformer modules is as follows: Figure 3 shown.

[0012] Specifically, the layer normalization is the normalization of a single training data to all neurons in a certain layer.

[0013] Specifically, the design of the multi-head self-attention module embeds the self-attention mechanism into the window, which includes two types, namely regular window self-attention and sliding window self-attention.

[0014] Specifically, the regular window self-attention (W-MSA) divides the image into multiple windows, and calculates the self-attention in each window. Assuming that the image is divided into h×w tiles, each window has M×M tiles. The computational complexity of W-MSA is shown in formula (2). W-MSA reduces the computational complexity. Since the number of tiles contained in each window is much smaller than the number of tiles in the image, the computational complexity of W-MSA is linearly related to the image size.

[0015] Ω(W-MSA)=4hwC 2 +2M 2 hwC (2)

[0016] Specifically, the sliding window self-attention (SW-MSA) introduces a sliding window method to achieve cross-window connection, increase the information transfer between different windows, expand the receptive field, and improve the representation ability of the model. The calculation relationship implied in the two series-connected SwinTransformers is shown in formula (3), where and z l They represent the features output by the two types of window self-attention and the features output by the multi-layer perceptron respectively.

[0017]

[0018] Specifically, the short-circuit connection forms a path that skips some layers by directly connecting the input of the front layer in the neural network to the output of the back layer, so that the back layer can more easily learn the difference between the input and output of the front layer; the multilayer perceptron is the simplest deep network, and the multilayer perceptron can be used to solve the classification of linearly inseparable data.

[0019] Specifically, the GELU activation function can be compared with a variety of nonlinear characteristics. It can reduce operations with poor accuracy and help the model converge better, while providing smooth adjustment and maintaining gradient stability. It can effectively reduce the problem of gradient explosion and effectively eliminate the problem of gradient diffusion, which helps to improve the forward performance of the model. In addition, GELU can improve the accuracy and effectiveness of predictions, and its accuracy is higher than that of commonly used activation functions (such as ReLU).

[0020] Step 2: Cross-card synchronous batch normalization training; Cross-card synchronous batch normalization (SyncBN) obtains the global mean μ and variance σ during forward propagation 2 , the global gradient is calculated during back propagation, the principle is as follows Figure 4 shown.

[0021] Preferably, each GPU calculates ∑x i and ∑x i 2, and then perform synchronous summation to calculate the global variance and mean. The calculation can be completed with one synchronization, which can reduce the number of cross-card synchronizations from two to one, and improve the normalized synchronization calculation speed, as shown in equations (7) and (8), where x i is the sample point, and m is the total number of sample points of multiple graphics cards.

[0022]

[0023] Step 3: Loss function: Jaccard index, also known as IoU, takes into account the imbalance between background and target pixels when evaluating the prediction results. The corresponding loss function is shown in formula (9), where is the prediction vector of the model, y * is the label vector, and c represents the category.

[0024]

[0025] Preferably, since the continuous prediction results are not differentiable after discretization, the prediction results are required to be discrete when using Jaccard loss. Jaccard loss is a submodular function, and after Lovász expansion, its input space can be changed from discrete {0,1}p to continuous And the output value of the original function on {0,1}p remains equal.

[0026] Specifically, for the binary segmentation problem such as rail surface defect segmentation, the present invention uses the continuous and differentiable Lovász-hinge after smooth extension of Jaccard loss as the loss function to achieve the purpose of optimizing the model parameters; ΔJ c It can be rewritten as a function of a set of error predictions m, where the mathematical expression of the error prediction definition is shown in formula (6);

[0027]

[0028] ΔJ c The continuous linear distribution function after Lovász expansion is shown in formula (7);

[0029]

[0030] The error prediction m i It is replaced by the hinge loss of the input image x, as shown in formula (8), where F(x) is the network output;

[0031]

[0032] g(m) is derived from m, as shown in formula (9);

[0033]

[0034] In summary, the loss function is shown in formula (10);

[0035]

[0036] Step 4: Calculate defect parameters: Calculate the actual length, width and area of ​​the rail surface defects based on semantic segmentation, which helps railway maintenance personnel to intuitively understand the specific size of the defects and make corresponding welding repairs.

[0037] Preferably, when calculating specific parameters such as the length, width and area of ​​the defect, it is necessary to distinguish different rail surface defects in the semantic segmentation map, and use a method based on connected domain analysis to distinguish them. The connected domain is the part of the image that is of the same type and connected in pixels, and the pixel values ​​are required to be the same and connected. The connected domain in the image can be marked by the connected domain analysis technology; in order to prevent the influence of pixel value fluctuations on the extraction of different connected domains, the image is usually required to be binarized before the connected domain analysis. The judgment of the neighborhood is indispensable in the connected domain analysis. The criterion for judging the neighborhood in the image is the 8-neighborhood definition method, and its principle is as follows Figure 5 shown.

[0038] In addition to the horizontal and vertical directions under the 8-neighborhood definition method, pixels are also judged to be connected when they are adjacent in two diagonal directions. The coordinates of the 8 adjacent pixels of point P0(x,y) under the 8-neighborhood method are: P1(x-1,y), P2(x+1,y), P3(x,y-1), P4(x,y+1), P5(x-1,y-1), P6(x+1,y-1), P7(x-1,y+1), and P8(x+1,y+1).

[0039] Step 5: Perform a connected domain analysis on the rail surface defect semantic segmentation map obtained in step 4, mainly using the connected domain processing function cv2.connectedComponentsWithStats() in OpenCV-Python to mark the connected domains. This function can calculate the number of connected domains, the area of ​​each connected domain, and the length and width of its circumscribed rectangle. The process of defect parameter calculation based on connected domain analysis in the present invention is as follows: Figure 6 shown.

[0040] Preferably, first, the semantic segmentation map is subjected to median filtering for denoising; second, it is converted into a binary grayscale image; then, a binary image is obtained by threshold segmentation; finally, a connected domain analysis is performed; the number of connected domains in each image, the mark of each pixel on the image, and the statistical information of each mark, including the length, width and area of ​​each contour, are obtained according to the connected domain analysis, and different connected domains are filled with different colors using a color table for distinction. The results of the connected domain analysis are as follows: Figure 7The connected domain analysis provides the actual length, width and area of ​​the defect, which allows maintenance personnel to have an intuitive understanding of the defect parameters and perform repair welding in a targeted manner according to the importance of the task.

[0041] Basic principle:

[0042] Transformer is an encoding-decoding network structure proposed by VASWANI et al. that is completely based on the attention mechanism and point-to-point fully connected layers. It eliminates recursion and convolution, improves the disadvantage of slow RNN training, and performs well in machine translation tasks. The BERT and GPT models based on the Transformer architecture also show advanced performance in natural language processing. Most computer vision tasks use the encoder module of the original Transformer and regard it as a new type of feature extractor. The image classification network ViT proposed by DOSOVITSKIY et al. divides the image into fixed-size image blocks, uses the image blocks with position embedding added as the input of the Transformer encoder, and then performs classification processing through MLP, achieving excellent results compared with the most advanced convolutional networks. CARION et al.

[26] An end-to-end object detection network DETR combining CNN and Transformer architecture is proposed. Object detection is regarded as a prediction problem from image to set. The self-attention mechanism of Transformer is used to simulate all pairwise interactions between elements in the sequence, making it more suitable for the specific constraints of set prediction. CHEN et al. developed an image processing model IPT based on Transformer architecture to solve image processing problems, and achieved image super-resolution, restoration and denoising tasks on a large number of damaged images generated by ImageNet. Li Yaoqian et al. proposed a semi-supervised video segmentation framework, introducing Transformer at the bottleneck layer to further extract contextual information, and improving the accuracy of endoscopic video segmentation under limited annotated data. Although the Transformer structure has been widely used in various visual tasks and achieved good results in recent years due to its advantage of extracting global contextual features, it has not yet been widely used in rail surface defect detection tasks.

[0043] The attention mechanism originates from the study of human vision. Due to the bottleneck of information processing, humans selectively focus on part of all information while ignoring other visible information. The attention mechanism in deep learning is a bionic imitation of the human visual attention mechanism. It is essentially a resource allocation mechanism that redistributes weights according to the importance of objects.

[0044] In Transformer, multiple self-attention layers are connected to form multi-head self-attention (MSA) to improve network performance. The principle is as follows Figure 1As shown. The self-attention layer converts the input vector into a query vector q, a key vector k, and a value vector v and compresses them into three different matrices Q, K, and V. The attention weight is calculated by Q and K, and then applied to V to obtain the entire weight and output. The vector with a higher probability will receive additional attention in the lower layer. The self-attention function between different input vectors is shown in formula (11), where d k is the dimension corresponding to K.

[0045]

[0046] Multi-head attention generates multiple self-attentions in the network, which act on the features in parallel, and combines the results of these self-attentions to obtain the final output, which helps the network capture richer feature information. Its mathematical expression is shown in formula (12).

[0047] MultiHead(Q,K,V)=Concat(head1,...head h )W O (12)

[0048] Among them, head i is the result of each self-attention action, as shown in formula (13), W i Q , W i K , W i V are projection weight matrices.

[0049] head i =Attention(QW i Q ,KW i K ,VW i V ) (13).

[0050] The present invention improves the UPerNet network, and its overall structure is as follows: Figure 1 As shown in the figure. The encoder network uses Swin-T for feature extraction, and the decoder network uses PPM and FPN. After upsampling, the learned discriminable features are projected into the pixel space. The last set of feature maps in each stage of the feature extraction network is {S2, S3, S4, S5}, and the feature maps output by FPN are {P2, P3, P4, P5}, where P5 is also a feature map that directly follows PPM.

[0051] The present invention has the following beneficial effects: The present invention proposes an improved UPerNet semantic segmentation network for rail surface defect detection, and adopts a Swin-T network with a Transformer architecture for defect feature extraction, which makes full use of the global information in the image and avoids the problem of inductive preference. The regular window self-attention in Swin-T reduces the number of parameters of the model. Secondly, gradient optimization is performed using cross-card synchronous batch normalization, and Lovász-hinge is used as the loss function, which incorporates the use of dependencies between pixels. In the face of complex rail scenes, the improved method of the present invention can achieve robust segmentation of defects and solve the problem of low prediction accuracy of defect edges. Finally, on the basis of the semantic segmentation map of rail surface defects, the connected domain analysis is used to distinguish different defect areas, and the actual length and actual area of ​​the defects are calculated, which helps railway maintenance personnel to intuitively understand the defect parameters and make corresponding repairs. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solution in the example of the present invention, some drawings in the example of the present invention are briefly introduced below.

[0053] Figure 1 It is the overall structure diagram of the method of the present invention;

[0054] Figure 2 It is the network structure diagram of Swin-T in the present invention;

[0055] Figure 3 It is a diagram of two SwinTransformer modules connected in series in the present invention;

[0056] Figure 4 It is a working principle diagram of SyncBN in the present invention;

[0057] Figure 5 A schematic diagram of an 8-neighborhood definition method for adjacent pixels in the present invention;

[0058] Figure 6 It is a flow chart of connected domain analysis in the present invention;

[0059] Figure 7 This is an example diagram of the connected domain analysis results in the present invention;

[0060] Figure 8 These are some Rail-BY dataset images in Example 1 of the present invention;

[0061] Fig. 9 This is the RSDDs dataset image in Example 1 of the present invention

[0062] Fig.10 This is a comparison chart of the visualization results of the method of the present invention and other semantic segmentation methods. DETAILED DESCRIPTION

[0063] In order to more clearly illustrate the purpose, technical solutions and advantages of the present invention, the functions and advantages of each module are specifically explained below in conjunction with the accompanying drawings.

[0064] Embodiment 1

[0065] The present invention is experimentally verified on two datasets. The first dataset is collected on-site on the transportation railway line of Baiyin Nonferrous Metals Company. The main defect type is peeled blocks. Since the image information of the rail bottom, ballast, etc. in the background is invalid for detection, it is uniformly processed into a 160×370 pixel image containing only the rail surface area. After cropping, it contains 914 defective images and is named Rail-BY dataset. The second dataset is the RSDDs dataset produced by researchers from Beijing Jiaotong University. Due to the timely maintenance of the railway engineering department, there are fewer rail surface defect samples. The defect images are expanded by rotation, mirroring, filtering, noise addition and DCGAN. The expanded Rail-BY dataset contains 9570 defective images, and the RSDDs dataset contains 6950 defective images. Some samples are as follows: Figure 8 and Fig. 9 shown.

[0066] Pixel accuracy (PA), intersection over union (IoU), Dice coefficient and Precision are used to measure the accuracy of segmentation, and parameter count (Params) and floating point operations (FLOPs) are used to evaluate the scale and complexity of the algorithm. PA is the proportion of correctly classified pixels, IoU represents the degree of overlap between the predicted result and the true label, Dice coefficient calculates the similarity between the predicted result and the true label, Precision represents the proportion of pixels predicted as defects by the segmentation network that are actually defects. The larger the four parameters, the more accurate the segmentation of the defect area. Their calculations are shown in formulas (14), (15), (16) and (17), respectively, where X represents the predicted result and Y represents the true label. Params is the total parameter of model training, FLOPs is the number of floating point operations, and the smaller the Params and FLOPs, the lower the complexity of the model.

[0067]

[0068] The experimental equipment uses Intel Core i9-10920X CPU@3.50GHz, 128GB running memory, dual NVIDIAGeForce RTX 3090GPU, based on Python3.8 compilation environment, PyTorch1.8.1 framework. In the experiment, the iterations are set to 160000 times, and the AdamW optimizer is used, where β1=0.9, β2=0.999, the learning rate is 0.00006, batch_size=2, and weight_decay=0.01. The loss curves of the validation set on the Rail-BY and RSDDs datasets are shown in the figure. Fig. 9 As shown in Figure 1, the training loss is stable after 160,000 iterations, indicating that the training effect is ideal.

[0069] In order to verify the effectiveness of the Swin-T feature extraction network, cross-card synchronous batch normalization, and the introduction of the Lovász-hinge loss function for improving the UPerNet network, the segmentation results of different modules are compared on the two datasets. The comparison of the segmentation results of different modules is shown in Table 1. The backbone network of the baseline is ResNet50, the normalization method is BN, and the loss function is the cross entropy loss function.

[0070] Table 1 Comparison of segmentation results of different modules of improved UPerNet

[0071]

[0072]

[0073] From Table 1, we can see that when only Swin-T replaces ResNet50, PA is improved by 8.06%, IoU by 11.42%, Dice by 7.3%, and Precision by 9.52% on the Rail-BY dataset. PA is improved by 3.48%, IoU by 4.45%, Dice by 2.59%, and Precision by 1.7% on the RSDDs dataset. The accuracy of UPerNet segmentation is improved comprehensively, and the performance improvement effect on the original UPerNet is the most significant. In addition, the use of Swin-T for feature extraction reduces Params by 6.55M, which shows that the backbone network based on the Transformer architecture is much ahead of the CNN-based network in performance. The use of SyncBN normalization and the Lovász-hinge loss function mainly improves PA. Since the two are not improvements in the network structure, they have little impact on the network complexity.

[0074] The training result parameter comparison of the improved UPerNet network of the present invention and other mainstream semantic segmentation networks on the two datasets is shown in Table 2.

[0075] Table 2 Comparison of defect segmentation results of this method and other networks

[0076]

[0077] It can be seen that the improved method of the present invention performs very well on both datasets. Even compared with the excellent DeepLabV3+, on the Rail-BY dataset, PA is improved by 6.88%, IoU is improved by 6.83%, Dice is improved by 4.2%, and Precision is improved by 1.38%. On the RSDDs dataset, PA is improved by 1.13%, IoU is improved by 0.64%, and Dice is improved by 0.37%. Only Precision is slightly inferior with a decrease of 0.41%. The FPS of the original UPerNet network on Rail-BY is 20.1Hz, and the FPS on RSDDs is 20.2Hz. The FPS of the improved UPerNet on the two datasets is reduced by 2.2Hz and 2.3Hz respectively. Compared with other semantic segmentation networks, the detection speed is average. The disadvantage is that although the improved method of the present invention has a reduced number of parameters compared with the original UPerNet model, the number of parameters is still large compared with other mainstream semantic segmentation algorithms. The visualization results of the improved method of the present invention and other segmentation networks are compared. Fig.10 shown.

[0078] Embodiment 2

[0079] The rail head width of the rail photographed in the present invention is 73 mm, and the rail surface defect images used are all 160×370 pixels. The corresponding rail length can be calculated in proportion to be approximately 168.81 mm, and the corresponding actual area of ​​the rail region S entire_real About 12323.31mm 2 . Obtain the defect pixel length L in the semantic segmentation map defect_seg Total length of the picture L entire_seg , then the actual rail defect length is as shown in formula (18), where L entire_real The actual rail length corresponding to the semantic segmentation map is 168.81 mm

[0080] Obtain the defect area S in the semantic segmentation map detect_seg Total area S entire_seg The actual defect area of ​​the rail is the percentage of

[0081] S defect_real As shown in formula (18), where S entire_realThe actual area of ​​the rail area corresponding to the semantic segmentation map is 12323.31mm 2 .

[0082]

[0083] After analyzing the connected domain, the length, width and area of ​​the defect are multiplied by the corresponding coefficient to obtain its actual length and width, and the actual area of ​​the defect is calculated. Since the background pixels in the image also belong to a connected domain, the connected domain with an area greater than 30,000 pixels is determined as the background.

[0084] in conclusion

[0085] In view of the low efficiency of traditional machine vision methods and the random and complex defect shapes in the process of rail surface defect detection, this paper proposes an improved UPerNet algorithm to detect defects. Using Swin-T as the backbone network, the inductive preference and feature extraction locality problems of the CNN-based backbone network are avoided; secondly, the cross-card synchronous batch normalization method is used to optimize the influence of BN on the convergence of the semantic segmentation model of large video memory; finally, the cross-entropy loss ignores the problem of pixel dependencies, and Lovász-hinge is used as the loss function. The improved effect is verified by comparative experiments of different modules and different segmentation algorithms. The experimental results show that the detection accuracy on both data sets has been improved to a certain extent. Finally, based on the semantic segmentation of rail surface defects, the connected domain processing technology is used to divide different defects in the same picture, and then the actual length and area of ​​each defect area are calculated by calculating the number of pixels.

[0086] The above content introduces the basic principles, main features and advantages of the examples of the present invention. Relevant practitioners should understand that the present invention is not limited to the above examples, and the above embodiments and descriptions are only for illustrating the principles of the present invention. The present invention can be applied to any other field with optimization properties. The present invention may also have various changes and improvements, which all fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.

Claims

1. A rail surface defect detection method based on improved UPerNet and connected domain analysis is characterized in that: The steps include: Step 1: Feature extraction network; Using Swin-T as the backbone network, it extracts richer global information and more discriminative features, and reduces the amount of computation by limiting the attention calculation to non-overlapping local windows so that the computational complexity is linear with the image size; Step 2: Synchronize batch normalization training across cards; Cross-card synchronous batch normalization SyncBN obtains the global mean μ and variance σ during forward propagation 2 , calculate the global gradient during backpropagation; Step 3: Loss function: Jaccard index, also known as IoU, takes into account the imbalance between background and target pixels when evaluating the prediction results. The corresponding loss function is shown in formula (1): in is the prediction vector of the model, y* is the label vector, c represents the category; Step 4: Calculate defect parameters: Calculate the actual length, width and area of ​​the rail surface defects based on semantic segmentation, which helps railway maintenance personnel to intuitively understand the specific size of the defects and make corresponding welding repairs; Step 5: Perform connected domain analysis on the semantic segmentation map of rail surface defects obtained in step 4, mainly using the connected domain processing function cv2.connectedComponentsWithStats() in OpenCV-Python to mark the connected domains. This function can calculate the number of connected domains, the area of ​​each connected domain, and the length and width of its circumscribed rectangle.

2. The rail surface defect detection method based on improved UPerNet and connected domain analysis according to claim 1 is characterized in that: In step 1, the entire network adopts a hierarchical design. First, the input H×W×3 image is split into (H / 4)×(W / 4) non-overlapping tiles through the tile splitting module. Each tile is flattened into a 48-dimensional token vector. The linear embedding layer maps the tensor of dimension (H / 4)×(W / 4)×48 to an arbitrary dimension C, and obtains a linear mapping of dimension (H / 4)×(W / 4)×C. Then the mapped tokens are sent to the paired SwinTransformer modules, which keep the number of input and output tokens unchanged at (H / 4)×(W / 4). The first SwinTransformer module and the linear embedding layer are jointly designated as the first layer; then, the tile merging layer and the Swin Transformer modules are combined into subsequent layers to generate feature descriptions of different scales. The tile merging layer is responsible for downsampling operations, reducing tile resolution and expanding the receptive field layer by layer. The SwinTransformer module is used for feature extraction. As the network deepens, each layer changes the dimension of the tensor, thus forming a hierarchical representation.

3. The rail surface defect detection method based on improved UPerNet and connected domain analysis according to claim 2 is characterized in that: The Swin Transformer module consists of layer normalization, a multi-head self-attention module, short-circuit connections, and a multi-layer perceptron with a GELU activation function.

4. The rail surface defect detection method based on improved UPerNet and connected domain analysis according to claim 3 is characterized in that: The layer normalization is the normalization of a single training data to all neurons in a certain layer; The design of the multi-head self-attention module embeds the self-attention mechanism into the window, which includes two types: regular window self-attention and sliding window self-attention. The regular window self-attention W-MSA divides the image into multiple windows and calculates the self-attention in each window. Assuming that the image is divided into h×w blocks and each window has M×M blocks, the computational complexity of W-MSA is shown in formula (2): Ω(W-MSA)=4hwC 2 + 2M 2 hwC (2) W-MSA reduces the computational complexity; since the number of tiles contained in each window is much smaller than the number of tiles in the image, the computational complexity of W-MSA is linearly related to the image size; the sliding window self-attention SW-MSA introduces the sliding window method to achieve cross-window connection, increase the information transmission between different windows, expand the receptive field, and improve the representation ability of the model. The computational relationship implied in the two series-connected Swin Transformers is shown in formula (3): in and z l They represent the features output by the two types of window self-attention and the features output by the multi-layer perceptron respectively; The short-circuit connection forms a path that skips some layers by directly connecting the input of the front layer in the neural network to the output of the back layer, so that the back layer can more easily learn the difference between the input and output of the front layer; the multilayer perceptron is the simplest deep network, and the multilayer perceptron can be used to solve the classification of linearly inseparable data; The GELU activation function can be compared with a variety of nonlinear characteristics. It can reduce operations with poor accuracy and help the model converge better. At the same time, it provides smooth adjustment and maintains the stability of the gradient. It can effectively reduce the problem of gradient explosion and the problem of gradient diffusion, which helps to improve the forward performance of the model. In addition, GELU can improve the accuracy and effectiveness of predictions, and its accuracy is higher than that of the commonly used ReLU activation function.

5. The rail surface defect detection method based on improved UPerNet and connected domain analysis according to claim 1 is characterized in that: In step 2, each GPU calculates ∑x i and ∑x i 2 , and then perform synchronous summation to calculate the global variance and mean. The calculation can be completed in one synchronization, which can reduce the number of cross-card synchronizations from two to one, and improve the normalized synchronization calculation speed, as shown in formula (4) and formula (5). where x i is the sample point, and m is the total number of sample points of multiple graphics cards.

6. The rail surface defect detection method based on improved UPerNet and connected domain analysis according to claim 1 is characterized in that: In step 3, since the continuous prediction results are not differentiable after discretization, the prediction results are required to be discrete when using Jaccardloss. Jaccardloss is a submodular function. After Lovász expansion, its input space can be expanded from discrete {0,1} p Become continuous And the original function is in {0,1} p The output values ​​on remain equal.

7. The rail surface defect detection method based on improved UPerNet and connected domain analysis according to claim 6 is characterized in that: For the binary segmentation problem such as rail surface defect segmentation, the present invention uses the continuous and differentiable Lovász-hinge after smooth extension of Jaccard loss as the loss function to achieve the purpose of optimizing model parameters; ΔJ c It can be rewritten as a function of a set of error predictions m, where the mathematical expression of the error prediction definition is shown in formula (6); ΔJ c The continuous linear distribution function after Lovász expansion is shown in formula (7; The error prediction m i It is replaced by the hinge loss of the input image x, as shown in formula (8), where F(x) is the network output; g(m) is derived from m, as shown in formula (9); g πi (m)=ΔJ c ({π1,…,π i })-ΔJ c ({π1,...,π i-1 }) (9) In summary, the loss function is shown in formula (10); 8. The rail surface defect detection method based on improved UPerNet and connected domain analysis according to claim 1 is characterized in that: In step 4, when calculating specific parameters such as the length, width and area of ​​the defect, it is necessary to distinguish different rail surface defects in the semantic segmentation map, and use a method based on connected domain analysis to distinguish them; the connected domain is the part of the image that is of the same type and connected in pixels, and the pixel values ​​are required to be the same and connected. The connected domain in the image can be marked by the connected domain analysis technology; in order to prevent the influence of pixel value fluctuations on the extraction of different connected domains, the image usually needs to be binarized before the connected domain analysis. The judgment of the neighborhood is indispensable in the connected domain analysis, and the criterion for judging the neighborhood in the image is the 8-neighborhood definition method.

9. The rail surface defect detection method based on improved UPerNet and connected domain analysis according to claim 8 is characterized in that: In addition to the horizontal and vertical directions under the 8-neighborhood definition method, pixels are also judged to be connected when they are adjacent in two diagonal directions. The coordinates of the 8 adjacent pixels of point P0(x,y) under the 8-neighborhood method are: P1(x-1,y), P2(x+1,y), P3(x,y-1), P4(x,y+1), P5(x-1,y-1), P6(x+1,y-1), P7(x-1,y+1), and P8(x+1,y+1).

10. The rail surface defect detection method based on improved UPerNet and connected domain analysis according to claim 1 is characterized in that: In step five, first, the semantic segmentation map is denoised by median filtering; secondly, it is converted into a binary grayscale image; then, a binary image is obtained by threshold segmentation; finally, a connected domain analysis is performed; based on the connected domain analysis, the number of connected domains in each image, the mark of each pixel on the image, and the statistical information of each mark, including the length, width and area of ​​each contour, are obtained, and different connected domains are filled with different colors using a color table for easy distinction; the connected domain analysis provides the actual length, width and actual area of ​​the defect, which allows maintenance personnel to have an intuitive understanding of the defect parameters and perform targeted repair welding according to the importance of the task.