Steel rail surface defect detection method

Through the patch partitioning and multi-scale upsampling decoder of the Swin Transformer model, the efficiency and accuracy of rail surface defect detection are solved, efficient and accurate defect detection is achieved, and the safety and reliability of railway transportation is improved.

CN120259199APending Publication Date: 2025-07-04EAST CHINA UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510289186.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art has problems such as low efficiency and insufficient accuracy in the detection of rail surface defects, especially in complex backgrounds, which are difficult to effectively identify small or fuzzy defects.

Method used

Using the Swin Transformer-based rail surface defect detection method, the feature information is extracted and fused layer by layer through patch partitioning, linear embedding, Multi-Swin Transformer Block and multi-scale upsampling decoder to generate a segmentation mask to identify defects.

Benefits of technology

It improves the accuracy and robustness of rail surface defect detection, can detect potential problems earlier, reduce train delays and accidents, and improves the safety and reliability of railway transportation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259199A_ABST
    Figure CN120259199A_ABST
Patent Text Reader

Abstract

The invention discloses a steel rail surface defect detection method. A steel rail image is divided into image blocks with fixed sizes through a patch partitioning module, and the image blocks are mapped to a low-dimensional feature space through a linear embedding layer. Then, the low-dimensional feature vectors are sequentially input into a Multi-Swin Transformer Block (a shift window multi-head self-attention block) and a patch merging module, and semantic information is extracted layer by layer; and then, a multi-scale up-sampling decoder receives feature maps of different sizes, multi-head self-attention is utilized to enhance the spatial relationship and semantic understanding ability, the resolution is gradually recovered through a patch expansion module, and a feature fusion effect is integrated and optimized through a channel. And the integrated feature map is recovered to the original input size through a patch expansion module with a one-time multiplying power of 4, and a segmentation mask is generated through a convolutional layer. The method can accurately sense the defect degree, realizes efficient and accurate steel rail surface defect detection and positioning, is suitable for railway maintenance and safety monitoring, and ensures stable operation of facilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a method for detecting defects on the surface of steel rails. Background Art

[0002] As the core load-bearing element of the railway system, the integrity of the surface state of the track is directly related to the safety of train operation. An efficient and accurate track defect detection technology is an important support for realizing the safe operation of railways. Traditional track detection methods mainly include manual inspection and automated methods based on simple image processing techniques. Although these methods have played a role in the early maintenance of railway systems, with the increasing requirements for railway transportation safety and the complexity of the track environment, their efficiency, accuracy, and the ability to recognize complex defect patterns are facing more and more challenges. Therefore, introducing an automatic non-destructive detection system for the surface detection of steel rails and developing an efficient and accurate method for detecting track surface defects have profound theoretical and practical significance for improving the safety and reliability of railway transportation.

[0003] In recent years, with the rapid development of computer vision and deep learning technologies, significant progress has been made in the application of image-based steel rail surface defect detection methods from traditional image processing techniques to deep learning algorithms. First, the first method uses a deep convolutional neural network to detect railway surface defects, improving the detection accuracy and robustness through deep feature learning, but there are still challenges in processing images with highly complex backgrounds, and it is necessary to further improve the generalization ability of the model and reduce the false detection rate. The second method proposes a coarse-to-fine model (CFE) for railway surface defect detection, locates abnormal points through mean shift, and uses a fine extractor to filter these noises, but this model faces challenges in computational complexity and needs to be further optimized to improve the real-time detection ability. The third method uses the YOLOv3 deep learning network to detect railway surface defects, achieving efficient and accurate defect localization and recognition through a real-time object detection framework, but there is still room for improvement in the detection of some particularly small or blurred defects, and the detection ability can be further improved through model optimization or data augmentation in the future. Summary of the Invention

[0004] To solve the problems existing in the above background art, the present invention provides the following technical solution: A method for detecting defects on the surface of steel rails, comprising the following steps:

[0005] 1. A method for detecting defects on the surface of steel rails, characterized by comprising the following steps:

[0006] S1, input an image of the surface of a steel rail with defects, construct a dataset for detecting defects on the surface of the steel rail, and preprocess the dataset of the surface images of the steel rail;

[0007] S2. Split the preprocessed image into sub-image patches of a fixed size through a patch partitioning module;

[0008] S3. Input all the above sub-image patches into a Linear Embedding layer to convert the sub-image patches into embedding features to reduce the computational complexity;

[0009] S4. Input the embedded feature vectors in S3 into a Multi-Swin Transformer Block (shifted window multi-head self-attention block) to obtain a feature map with the same size and semantic information;

[0010] S5. Repeat the PatchMerging (patch merging module) operation and the Multi-Swin Transformer Block (shifted window multi-head self-attention block) operation 3 times in sequence for the feature map obtained in S4, and output 3 kinds of feature maps respectively;

[0011] S6. Perform multi-head self-attention operations on the 4 feature maps output in S4 and S5 respectively to obtain 4 new feature maps;

[0012] S7. Gradually restore the resolution of the feature maps through the PatchExpanding (patch expanding module) with an upsampling ratio of 2 for different numbers of times for the 4 new feature maps in S6, and obtain 4 new feature maps again;

[0013] S8. Integrate the 4 new feature maps in S7 along the channel dimension and pass through a patch expanding module with an upsampling ratio of 4 once to restore the feature map resolution to the input size, and finally obtain a segmentation mask through a convolutional layer once, and the construction of the model is completed;

[0014] S9. Select the rail surface defect images in the rail surface defect detection dataset, and input the rail surface defect images and the labeled ground truth maps into the constructed model for training to obtain a defect detection model;

[0015] S10. Input the rail surface defect images in the test set, and after being processed by steps S2 - S8, obtain the predicted defect results.

[0016] Further, S2 specifically includes the following steps:

[0017] S21. First, input the image and split the input RGB image of H×W×3 into N non-overlapping and equal-sized (p 2 ×3) sub-images through the patch partitioning module, where p is the set image size and is initially set to 4, and H and W are the height and width of the input image respectively;

[0018] S22. Flatten these sub-images in the channel direction, that is, each 4×4×3 sub-image obtained by segmentation is subsequently flattened into a one-dimensional vector of length 48 (i.e., 4 2 ×3), and then followed by a layer normalization operation: First, calculate the mean and variance The input is x = (x1, x2, …… x m ), where x i represents the i-th dimensional feature of the input vector, m is the number of features, and then perform normalization where ∈ is a very small constant used to prevent division by zero, μ is the mean, and σ 2 is the variance;

[0019] S23. Finally, scale and translate the output to where γ and β are learnable parameters used to scale and translate the normalized features At this time, the output y of the patch partition module i has a size of

[0020] Furthermore, the specific steps of S3 are as follows:

[0021] S31. Given an input Y, where Y contains n non-overlapping patches, each patch y i has a shape of n is the total number of patches generated from the original RGB image of size H×W×C through the patch partition module, and H, W, and C are the height, width, and number of channels of the original image respectively;

[0022] S32. Perform a linear embedding operation: Z = YW + b, where W is the weight matrix with a shape of 48×C, C is the new feature dimension, b is the bias vector with a dimension of C, Y is the output of S31 through the patch partition module, and Z is the output after linear embedding with a shape of representing the high-dimensional feature space for further mapping.

[0023] Furthermore, the specific steps of S4 are as follows:

[0024] S41. Input the embedded feature vector in S3 and define it as Z (l-2) ∈R d×1 , Z (l-2) is the output after mapping in S32, where (l - 2) is set to distinguish different stages of the Multi-Swin Transformer Block, and then perform a layer normalization operation: First, calculate the mean and variance and perform normalization processing where is the result of normalization, ∈ is a very small constant used to prevent division by zero, d is the number of channels of the summation feature, μ is the calculated mean, and σ 2 is the calculated variance, and Z i (l-2) represents the number of channels of the i-th feature in the (l - 2)-th layer. Finally, the scaled and translated output is where is the result of the scaled and translated output, γ and β are learnable parameters, and have the same dimension d as Z (l-2) . The next step is to perform the W-MSA operation (window multi-head self-attention) on . The feature map is first divided into multiple non-overlapping small windows of size M×M, where M is the size of the divided image, and the self-attention calculation within the window is performed to obtain the linearly projected feature map which is linearly transformed to obtain Q, K, V: Q = W q Z, K = W k Z, V = W v Z, where Q, K, and V represent the query, key, and value matrices respectively, and W q , W k , W v are learnable weight matrices; then the multi-head self-attention operation is performed: head i = Attention(Q i , K i , V i ), MultiHead(K, Q, V) = Concact(head1,..., head i )W O , where Q i , K i , V i represent the query, key, and value matrices of the i-th attention head respectively, i is the number of heads in the multi-head self-attention mechanism, W O is a learnable weight matrix, and the Attention operation is: where softmax represents the activation function, K T is the transpose of K, is the scaling factor, d k is the dimension of Q and K, and the output of this part is R (l-1) . The output R (l-1) of the W-MSA and Z (l-2) are subjected to a residual connection H (l-1) = R (l-1) + Z (l-2) , where H(l-1) is the result after the residual connection, and then perform layer normalization on H (l-1) and then perform MLP (Multi-Layer Perceptron) operation: First is the linear transformation a l = W l H (l-1) + b l , where a l is the activation value of the l-th layer, W l is the weight matrix of the l-th layer, and b l is the bias vector of the l-th layer. where σ is the activation function GELU, is the output after applying the activation function, and make a residual connection between this output and H (l-1) : Z (l-1) = R (l-1) + Z (l-2) , where Z (l-1) is the output of the first module among the four modules of the Multi-Swin Transformer Block;

[0025] S42: Similar to S41, the difference is that SW-MSA operation is adopted in this step. First is window partitioning and offset: where Z shifted is the result after window partitioning and offset, is the output after layer normalization, shift(,) is the window offset operation, offset is set alternately according to the layer, and the rest of the operations are the same as S41;

[0026] S43: Similar to S41, the difference is that WSW-MSA operation is adopted in this step, and the strategy of horizontal shift window is adopted. First is window partitioning and offset: where is the output after layer normalization, shift(,) is the window offset operation, offset is set alternately according to the layer, and the rest of the operations are the same as S41;

[0027] S44: Similar to S41, the difference is that HSW-MSA operation is adopted in this step, and the strategy of horizontal shift window is adopted. First is window partitioning and offset: where is the output after layer normalization, shift(,) is the window offset operation, offset is set alternately according to the layer, and the rest of the operations are the same as S41. The final output is the feature map S l , and the shape of S l is

[0028] Furthermore, S5 specifically includes the following steps:

[0029] S51: Take the feature map S output by S44 l as the input, and the shape of S l is Perform a Patch Merging operation on it. First, rearrange the input tensor S l : S l rearranged is the result after tensor rearrangement, and then adjust the dimension order: S l transpose = transpose(S l rearranged , (0, 2, 1, 3, 4)), where S l transpose is the result after adjusting the dimension order. Finally, readjust the shape: S l merged is the result after readjusting the shape. Transform the tensor S after readjusting the shape l merged through a linear layer: S' = S l merged W + b, where S' is the result after the linear layer transformation, W is the weight matrix with the shape of 4C×2C, and b is the bias vector with the shape of 2C. The shape of the output S' is Then input S' into the Multi-Swin Transformer Block module. The operation is the same as that in step S4, and the output is the feature map S l +1 , with the shape of

[0030] S52: Repeat the operation of S51 twice. The first output is the feature map S l+2 , with the shape of The second output is the feature map S l+3 , with the shape of

[0031] Furthermore, S6 specifically includes the following steps: Perform self-attention operations on the feature maps S l , S l+1 , S l+2 , S l+3 obtained in steps S4 and S5 respectively: Obtain Q, K, V through linear transformation: Q = W q S, K = W k S, V = W v S, where Q, K, V represent the query, key, and value matrices respectively, and W q , W k , Wv is a learnable weight matrix; then perform multi-head self-attention operation: head i = Attention(Q i , K i , V i ), MultiHead(K, Q, V) = Concact(head1,..., head i )W O , where Q i , K i , V i represent the query, key, and value matrices of the i-th attention head respectively, and i is the number of heads in the multi-head self-attention mechanism. W O is a learnable weight matrix, and the Attention operation is: where softmax represents the activation function, K T is the transpose of K, is the scaling factor, d k is the dimension of Q and K, and finally obtain the new feature maps K l , K l+1 , K l+2 , K l+3 .

[0032] Furthermore, S7 specifically includes the following steps: perform patch expansion operations (Patch Expanding) with an upsampling ratio of 2 for 0 times, 1 time, 2 times, and 3 times on the new feature maps K l , K l+1 , K l+2 , K l+3 respectively. Taking K l+1 as an example, first rearrange and reshape the input: where K l+1 reshaped is the result after rearrangement and reshaping. Then adjust the dimension order for upsampling: K l+1 transposed = transpose(K l +1 reshaped , (0, 2, 1, 3, 4)), K l+1 transposed is the result after adjusting the dimension order. Then reshape the adjusted tensor to increase the resolution: K l+1 expanded is the result after reshaping. Finally, apply a linear layer for feature transformation to obtain the new feature map F l+1 : F l+1 = Sl +1 expanded W + b, where F l+1 is the result of the linear layer output, W is the weight matrix of shape C×2C, b is the bias vector of shape 2C, K l 、K l+2 、K l+3 's outputs are the new feature maps F l 、F l+2 、F l+3 respectively, where the sizes of F l 、F l+1 、F l+2 、F l+3 are all

[0033] Furthermore, S8 specifically includes the following steps: Concatenate the F l 、F l+1 、F l+2 、F l+3 output by S7 along the channel dimension into a feature map with a size of , and the operation is: F = concat(F l , F l+1 , F l+2 , F l+3 , axis = 3), where F is the concatenated feature map, axis = 3 indicates concatenation along the channel dimension. Then, pass F through a Patch Expanding module with an upsampling ratio of 4 to restore the feature map resolution to the input size, and the operation is: F reshaped is the result after rearrangement and shape adjustment, F transposed = transpose(F reshaped , (0, 2, 1, 3, 4)), F transposed is the result after adjusting the dimension order, F expanded = reshape(F transposed , (H, H, C)), F expanded is the result after reshaping again. Finally, pass F expanded through the convolutional layer to obtain the segmentation mask.

[0034] The beneficial effects of the present invention are as follows: The designed network model uses the hierarchical structure of Swin Transformer and multi-head self-attention feature fusion for feature extraction to capture the complex patterns of rail surface defects; and uses an improved multi-scale upsampling decoder for defect detection. In this way, the model can not only extract sufficient representations for effective defect detection, but also handle the challenges brought by the high reflectivity and small gray-scale changes on the rail surface as well as the complexity of the rail background. This method of fusing multi-scale features can improve the accuracy and robustness of rail surface defect detection.

[0035] By designing a rail surface defect detection method based on Swin Transformer, the present invention utilizes the model's hierarchical and global modeling capabilities, providing a new direction for accurately identifying and classifying track surface defects and achieving efficient and accurate detection of railway track surface defects. In addition, the improved Swin Transformer model can achieve higher accuracy and efficiency in track surface defect detection, which means that potential track problems can be discovered earlier, so that maintenance measures can be taken in a timely manner to effectively avoid the occurrence of safety accidents. This not only reduces train delays and accidents caused by track problems, improves the overall safety and reliability of railway transportation, but also provides scientific and accurate data support for railway maintenance, reduces maintenance costs, and improves the efficiency of maintenance work.

[0036] In summary, by introducing and improving the Swin Transformer model, the present invention aims to solve the problem of railway track surface defect detection, which not only enriches the application research of deep learning in the field of image processing, but also provides effective technical support for improving the safety and reliability of railway transportation systems. Description of the Drawings

[0037] Figure 1 It is a flowchart of the rail surface defect detection task for a rail surface defect detection method of the present invention.

[0038] Figure 2 It is a schematic diagram of the overall network model for a rail surface defect detection method of the present invention.

[0039] Figure 3 It is a flowchart of the shifted window multi-head self-attention block for a rail surface defect detection method of the present invention.

[0040] Figure 4 It is a flowchart of the WSW-MSA and HSW-MSA modules for a rail surface defect detection method of the present invention. Detailed Embodiments

[0041] To better understand the above objects, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. In the following description, many specific details are set forth to facilitate a thorough understanding of the present invention. However, the present invention can also be implemented in other ways different from those described herein. Therefore, the present invention is not limited by the specific embodiments disclosed below. In the present invention, H and W refer to the height and width of an image, Multi-SwinTransformer refers to a shifted window multi-head self-attention transformer, Multi-Swin TransformerBlock refers to a shifted window multi-head self-attention block, W-MSA refers to window multi-head self-attention, SW-MSA refers to shifted window multi-head self-attention, WSW-MSA refers to horizontally shifted window multi-head self-attention, HSW-MSA refers to vertically shifted window multi-head self-attention, and PatchMerging refers to a patch merging module.

[0042] As Figures 1 to 4 shown, a method for detecting rail surface defects includes the following steps:

[0043] S1. Input a rail surface image, construct a rail surface defect detection dataset. The dataset adopted in the present invention is the RSDDs dataset, which includes two types of datasets: The first is the type-I RSDDs dataset captured from the fast lane, which contains 67 challenging images. The second is the type-II RSDDs dataset captured from ordinary / heavy transport tracks, which contains 128 challenging images. Each image in the two datasets contains at least one defect, and the background is complex and has a lot of noise. Perform preprocessing operations on the rail surface defect images. First, resize each image and its mask to 224×224 and perform 20 times of data augmentation, including random horizontal flipping, random vertical flipping, brightness change, mirror operation, adding Gaussian noise, adding salt noise. The total number of enhanced images is 3900. Take 70% of the total samples as the training set, a total of 2730 images, and 30% as the test set, a total of 1170 images. Install python3.9, pytorch1.11, cuda11.3, and create a.py test file;

[0044] S2. Input the rail surface defect image and split it into fixed-size sub-image blocks through a patch partition module (Patch Partition), with a size of H / 4×W / 4×48. The specific steps are as follows:

[0045] S21. The input image first passes through Figure 2 the patch partition module in 2 to split the input RGB image of H×W×3 into N non-overlapping and equal-sized (p

[0046] S22. Flatten these sub - images in the channel direction, that is, each 4×4×3 sub - image obtained by segmentation is subsequently flattened into a one - dimensional vector of length 48 (i.e., 4 2 ×3), and then followed by a layer normalization operation: First, calculate the mean and variance The input is x=(x1,x2,……x m ), where x i represents the i - th dimensional feature of the input vector, m is the number of features. Secondly, perform normalization where ∈ is a very small constant used to prevent division by zero, μ is the mean, and σ 2 is the variance;

[0047] S23. Finally, scale and translate the output to where γ and β are learnable parameters used to scale and translate the normalized features At this time, the output size of the patch partition module is This method divides the input image into sub - image blocks of a fixed size through Patch Partition, flattens them into one - dimensional vectors, and at the same time uses Layer Normalization to standardize the feature distribution, enhancing the stability and generalization ability of the model, and providing a standardized and efficient input for subsequent deep feature learning.

[0048] S3. Input all the above - mentioned sub - image blocks into the Figure 2 linear embedding layer in, convert the sub - image blocks into embedding features to reduce the computational complexity, and the size is Specifically, it includes the following steps:

[0049] S31. Given an input Y, where Y contains n non - overlapping patches, and each patch y i has a shape of n is the total number of patches generated from the original RGB image of size H×W×C through the patch partition module, and H, W, and C are the height, width, and number of channels of the original image respectively;

[0050] S32. Then perform a linear embedding operation: Z = YW + b, where W is the weight matrix with a shape of 48×C, C is the new feature dimension. b is the bias vector with a dimension of C, Y is the output of S31 through the patch partition module, Z is the output after linear embedding, and the shape of the output Z is Represents a high-dimensional feature space for further mapping. This method converts sub-image patches into embedded features through Linear Embedding, reducing computational complexity and enhancing feature representation ability, mapping them to a high-dimensional feature space, and providing an optimized input for the deep feature learning of the subsequent Transformer module.

[0051] S4. Input the embedded feature vectors in S3 into the Multi-Swin Transformer Block to obtain a feature map with unchanged size and semantic information, which specifically includes the following steps:

[0052] S41. As shown in Figure 3 the first column, take the output of the previous step as its input and define it as Z (l-2) ∈R d×1 , where (l - 2) is set to distinguish different stages of the Multi-Swin Transformer Block. Then perform layer normalization: First, calculate the mean and variance and perform normalization processing where ∈ is a very small constant used to prevent division by zero. Finally, scale and translate the output to where γ and β are learnable parameters with the same dimension d as Z (l-2) . The next step is to perform the W-MSA operation (window multi-head self-attention) on . The feature map is first divided into multiple non-overlapping small windows of size M×M, and self-attention calculation within the windows is performed to obtain the linearly projected feature map . Through linear transformation, obtain Q, K, V: Q = W q Z, K = W k Z, V = W v Z, where W q , W k , W v are learnable weight matrices; then perform multi-head self-attention operation: head i = Attention(Q i , K i , V i ), MultiHead(K, Q, V) = Concact(head1,..., head i )W O , where i is the number of heads in the multi-head self-attention mechanism, W O is a learnable weight matrix, and the Attention operation is: is the scaling factor, d k is the dimension of Q and K, and the output of this part is R (l-1) . Take the output R of W-MSA (l-1) and Z (l-2) to perform a residual connection H (l-1) = R (l-1) +Z (l-2) , then perform layer normalization on H (l -1) and then perform MLP (Multi-Layer Perceptron) operation: First is the linear transformation a l = W l H (l-1) +b l , where a l is the activation value of the l-th layer, W l is the weight matrix of the l-th layer, and b l is the bias vector of the l-th layer. where σ is the activation function GELU, is the output after applying the activation function, and take the output and H (l-1) to perform a residual connection: Z (l-1) = R (l-1) +Z (l-2) , where Z (l-1) is the output of the first module in the four modules of the Multi-Swin Transformer Block. This method normalizes the feature distribution through layer normalization, uses window multi-head self-attention (W-MSA) to extract local and global information, and combines residual connection and multi-layer perceptron (MLP) to further optimize the feature expression ability, improving the learning stability and generalization ability of the model.

[0053] S42. Similar to S41, such as Figure 3 the second column, the difference is that SW-MSA operation is adopted in this module. First is window partitioning and offset: where is the output after layer normalization, shift(,) is the window offset operation, and offset is set alternately according to the layer. Then calculate the multi-head self-attention: head i = Attention(Q i , K i , V i ), MultiHead(K,Q,V) = Concact(head1,..., head i )W O , where i is the number of heads in the multi-head self-attention mechanism, W Ois the learnable weight matrix, and the Attention operation is as follows: is the scaling factor, d k is the dimension of Q and K. The remaining operations are the same as those in S41.

[0054] S43 is similar to S41, as Figure 3 the third column, with the difference that the WSW-MSA operation is adopted in this module, and the strategy of horizontal shift window is adopted. First is the window division and offset: where is the output after layer normalization, shift(,) is the window offset operation, and offset is set alternately according to the layer. And the specific division strategy is as Figure 4 the first row, then calculate the multi-head self-attention: head i = Attention(Q i , K i , V i ), MultiHead(K, Q, V) = Concact(head1,..., head i )W O , where i is the number of heads in the multi-head self-attention mechanism, W O is the learnable weight matrix, and the Attention operation is as follows: is the scaling factor, d k is the dimension of Q and K. The remaining operations are the same as those in S41.

[0055] S44 is similar to S41, as Figure 3 the fourth column, with the difference that the HSW-MSA operation is adopted in this module, and the strategy of horizontal shift window is adopted. First is the window division and offset: where is the output after layer normalization, shift(,) is the window offset operation, and offset is set alternately according to the layer. And the specific division strategy is as Figure 4 the second row, then calculate the multi-head self-attention: head i = Attention(Q i , K i , V i ), MultiHead(K, Q, V) = Concact(head1,..., head i )W O , where i is the number of heads in the multi-head self-attention mechanism, W O is the learnable weight matrix, and the Attention operation is as follows: is the scaling factor, d k is the dimension of Q and K. The rest of the operations are the same as those in S1. The final output is the feature map S l , S l has the shape of

[0056] S5, such as Figure 2 In the shifted window multi-head self-attention transformer, the feature map obtained from S4 is sequentially repeated 3 times with the Patch Merging operation and the Multi-Swin Transformer Block (shifted window multi-head self-attention block) operation to generate multi-level feature representations and establish context relationships between pixels in the feature map, specifically including the following steps:

[0057] S51. Taking the feature map S output by S44 l as the input, S l has the shape of Performing the Patch Merging operation on it. First, rearrange the input tensor S l : Then adjust the dimension order: S l rearranged = transpose(S l rearranged , (0, 2, 1, 3, 4)), and finally readjust the shape: Transform the reshaped tensor S l merged through a linear layer: S' = S l merged W + b, where W is the weight matrix with the shape of 4C×2C, and b is the bias vector with the shape of 2C. The shape of the output S' is Then input S' into the Multi-Swin Transformer Block module. The operation is the same as that in step (4), and the output is the feature map S l+1 , with the shape of This method gradually reduces the resolution of the feature map and expands the channel dimension through Patch Merging to reduce the computational complexity while retaining key information. Subsequently, it uses linear transformation for feature mapping and inputs it into the Multi-Swin Transformer Block to extract deep features and improve the global modeling ability of the model.

[0058] S52. Repeat the operation of S51 twice. The first output is the feature map S l+2 , with the shape of The second output is the feature map S l+3 , with a shape of

[0059] S6. As in Figure 2 the multi-scale upsampling decoder part, perform multi-head self-attention operations on the 4 different-sized feature maps output by the four Multi-SwinTransformer Blocks obtained in the above steps, specifically including the following steps:

[0060] First, perform self-attention operations on the S l , S l+1 , S l+2 , S l+3 obtained in the above steps respectively, and obtain Q, K, and V through linear transformation: Q = W q S, K = W k S, V = W v S, where W q , W k , W v are learnable weight matrices; then perform multi-head self-attention operations: head i = Attention(Q i , K i , V i ), MultiHead(K, Q, V) = Concact(head1,..., head i )W O , where i is the number of heads in the multi-head self-attention mechanism, W O is a learnable weight matrix, and the Attention operation is: is the scaling factor, d k is the dimension of Q and K, and finally obtain the new feature maps K l , K l+1 , K l+2 , K l+3 . This method can improve the model's understanding ability of spatial relationships and semantic information, thereby enhancing the accuracy and robustness of rail defect detection.

[0061] S7. As in Figure 2 the multi-scale upsampling decoder part, gradually restore the resolution of the feature maps by passing the above 4 new feature maps through the Patch Expanding module with an upsampling ratio of 2 for different numbers of times, including the following steps: For the new feature maps K l , K l+1 , K l+2 , K l+3Perform patch expansion operations with an upsampling factor of 2 for 0, 1, 2, and 3 times respectively, taking K l+1 as an example. First, rearrange and reshape the input: where K l+1 reshaped is the result after rearrangement and reshaping. Then, adjust the dimension order for upsampling: K l+1 transposed = transpose(K l+1 reshaped , (0, 2, 1, 3, 4)), K l+1 transposed is the result after adjusting the dimension order. Next, reshape the adjusted tensor to increase the resolution: K l+1 expanded is the result after reshaping. Finally, apply a linear layer for feature transformation to obtain a new feature map F l+1 : F l+1 = Sl +1 expanded W + b, where F l+1 is the result output by the linear layer, W is a weight matrix with shape C×2C, and b is a bias vector with shape 2C. The outputs of K l , K l+2 , K l+3 are the new feature maps F l , F l+2 , F l+3 respectively, where the sizes of F l , F l+1 , F l+2 , F l+3 are all This method uses the Patch Expanding module to gradually upsample feature maps of different scales, and gradually restores the resolution of the feature maps through tensor rearrangement, dimension transformation, and linear mapping, ensuring the alignment of features at different scales, while enhancing the model's ability to recover spatial information, providing high-quality feature representations for subsequent feature fusion and accurate defect detection.

[0062] S8. For example, in the multi-scale upsampling decoder part of Figure 2 , integrate the above 4 new feature maps along the channel dimension and pass them through a patch expansion module with an upsampling factor of 4 once to restore the resolution of the feature map to the input size. Finally, obtain a segmentation mask through a convolutional layer once, and the construction of the model is completed, including the following steps: Combine F l , F l+1 , F l+2 , Fl+3 Concatenate along the channel dimension to form a feature map with a size of . The operation is: F = concat(F l , F l+1 , F l+2 , F l+3 , axis = 3), where axis = 3 indicates concatenation along the channel dimension. Then, pass F through a PatchExpanding module with an upsampling ratio of 4 to restore the resolution of the feature map to the input. The operation is: F transposed = transpose(F reshaped , (0, 2, 1, 3, 4)), F expanded = reshape(F transposed , (H, H, C)). Finally, pass F expanded through a convolutional layer to obtain a segmentation mask.

[0063] S9. Randomly select two thousand rail surface defect images from the rail surface defect detection dataset, and input the two thousand rail surface defect images and the labeled ground truth maps into the constructed model for training to obtain a defect detection model. During training, use the Adam optimizer, use the dropout layer to randomly suppress 10% of the neurons, initialize the learning rate to 5e -4 , and decay it by a factor of 0.9 after one-tenth of all epochs to obtain a quality assessment model; set the initial number of epochs to 200 and set the batch-size to 128.

[0064] S10. Input the rail surface defect images in the test set. After being processed through steps S2 - S8, obtain the predicted defect results.

[0065] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, replacement, or improvement made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for detecting surface defects of rail, characterized in that, It includes the following steps: S1. Input the surface image of the rail with defects, construct a rail surface defect detection dataset, and preprocess the rail surface image dataset; S2. Segment the preprocessed image into sub-image blocks of a fixed size through a patch partition module; S3. Input all the above sub-image blocks into a Linear Embedding layer to convert the sub-image blocks into embedding features to reduce the computational complexity; S4. Input the embedded feature vectors in S3 into a Multi-Swin Transformer Block to obtain a feature map with unchanged size and semantic information; S5. Repeat the PatchMerging operation and the Multi-Swin Transformer Block operation 3 times in sequence for the feature map obtained in S4, and output 3 types of feature maps respectively; S6. Perform multi-head self-attention operations on the 4 feature maps output in S4 and S5 respectively to obtain 4 new feature maps; S7. Gradually restore the resolution of the feature maps through the PatchExpanding module with an upsampling ratio of 2 for different numbers of times for the 4 new feature maps in S6, and obtain 4 new feature maps again; S8. Integrate the 4 new feature maps in S7 in the channel dimension and pass through a PatchExpanding module with an upsampling ratio of 4 to restore the feature map resolution to the input size, and finally obtain a segmentation mask through a convolutional layer to complete the construction of the model; S9. Select the rail surface defect images from the rail surface defect detection dataset, and input the rail surface defect images and the annotated ground truth maps into the constructed model for training to obtain a defect detection model; S10. Input the rail surface defect images of the test set, and after being processed through steps S2 - S8, obtain the predicted defect results.

2. The method for detecting surface defects of a steel rail according to claim 1, characterized in that, The specific steps of S2 are as follows: S21. First, the input image is split by the patch partition module into N non-overlapping and equal-sized (p 2 ×3) sub-images from the input RGB image of H×W×3, where p is the set image size and is initially set to 4, and H and W are the height and width of the input image respectively; S22. Flatten these sub-images in the channel direction, that is, each 4×4×3 sub-image obtained by segmentation is subsequently flattened into a one-dimensional vector with a length of 48 (i.e., 4 2 ×3), and then perform a layer normalization operation: First, calculate the mean and variance The input is x = (x1, x2, …… x m ), where x i represents the i-th dimensional feature of the input vector, m is the number of features, and then perform normalization where ∈ is a very small constant used to prevent division by zero, μ is the mean, and σ 2 is the variance; S23. Finally, scale and translate the output to where γ and β are learnable parameters used to scale and translate the normalized features At this time, the output y of the patch partition module i with size 3. The rail surface defect detection method according to claim 2, wherein The specific steps of S3 are as follows: S31. Given an input Y, where Y contains n non - overlapping patches, each patch y i has a shape of n is the total number of patches generated from the original RGB image of size H×W×C through the patch partition module, and H, W, and C are the height, width, and number of channels of the original image respectively; S32. Perform a Linear Embedding operation: Z = YW + b, where W is the weight matrix with a shape of 48×C, C is the new feature dimension, b is the bias vector with a dimension of C, Y is the output of S31 through the patch partition module, and Z is the output after linear embedding with a shape of a high-dimensional feature space representing further mapping.

4. A method for detecting surface defects of a rail according to claim 3, characterized in that, The specific steps of S4 are as follows: S41. Input the embedded feature vector in S3 and define it as Z (l-2) ∈R d×1 , Z (l-2) is the output after mapping by S32, where (l - 2) is set to distinguish different stages of the Multi - Swin Transformer Block. Then, perform layer normalization operation: First, calculate the mean and variance and perform normalization processing where is the result of the normalization processing, ∈ is a very small constant used to prevent division by zero, d is the number of channels of the summed features, μ is the calculated mean, and σ 2 is the calculated variance, and Z i (l-2) represents the number of channels of the i - th feature in the (l - 2) - th layer. Finally, scale and shift the output to be where is the result of the scaled and shifted output, and γ and β are learnable parameters with the same dimension d as Z (l-2) . The next step is to perform the W - MSA operation (window multi - head self - attention) on . The feature map is first divided into multiple non - overlapping small windows of size M×M, where M is the size of the divided image. Perform self - attention calculation within the window to obtain the linearly projected feature map and obtain Q, K, V through linear transformation: Q = W q Z, K = W k Z, V = W v Z, where Q, K, and V represent the query, key, and value matrices respectively, and W q , W k , W v are learnable weight matrices; then perform the multi - head self - attention operation: head i = Attention(Q i , K i , V i ), MultiHead(K, Q, V) = Concact(head1,..., head i )W O , where Q i , K i , V i represent the query, key, and value matrices of the i - th attention head respectively, i is the number of heads in the multi - head self - attention mechanism, W O is a learnable weight matrix, and the Attention operation is: where softmax represents the activation function, and K T is the transpose of K, is the scaling factor, d k is the dimension of Q and K, and the output of this part is R (l-1) . The output R (l-1) of W-MSA and Z (l-2) are subjected to a residual connection to obtain H (l-1) = R (l-1) + Z (l-2) , where H (l-1) is the result after the residual connection. Then, after layer normalization of H (l-1) , an MLP (Multi-Layer Perceptron) operation is performed: First, a linear transformation a l = W l H (l-1) + b l , where a l is the activation value of the l-th layer, W l is the weight matrix of the l-th layer, and b l is the bias vector of the l-th layer. where σ is the activation function GELU, is the output after applying the activation function. This output and H (l-1) are subjected to a residual connection: Z (l-1) = R (l-1) + Z (l-2) , where Z (l-1) is the output of the first module among the four modules of the Multi-Swin Transformer Block; S42: Similar to S41, the difference is that the SW-MSA operation is adopted in this step. First is window partitioning and offset: where Z shifted is the result after window partitioning offset, is the output after layer normalization. shift(,) is the window offset operation. offset is set alternately according to the layer. The rest of the operations are the same as those in S41; S43: Similar to S41, the difference is that the WSW-MSA operation is adopted in this step, and the strategy of horizontal shift window is adopted. First is window division and offset: where is the output after layer normalization, shift(,) is the window offset operation, offset is set alternately according to the layer, and the rest of the operations are the same as S41; S44: Similar to S41, the difference is that the HSW-MSA operation is adopted in this step, and the strategy of horizontal shift window is adopted. First is window division and offset: where is the output after layer normalization, shift(,) is the window offset operation, offset is set alternately according to the layer, and the rest of the operations are the same as S41. The final output is the feature map S l S l has the shape of 5. A method for detecting surface defects of a rail according to claim 4, characterized in that, The specific steps of step S5 are as follows: S51: Take the feature map S output by S44 l as input, and the shape of S l is Perform patch merging operation on it. First, rearrange the input tensor S l : S l rearranged is the result after tensor rearrangement. Then, adjust the dimension order: S l transpose = transpose(S l rearrangle , (0, 2, 1, 3, 4)), where S l transpose is the result after adjusting the dimension order. Finally, readjust the shape: S l merged is the result after readjusting the shape. Apply the reshaped tensor S l merged to a linear layer for transformation: S' = S l merged W + b, where S' is the result after the linear layer transformation, W is the weight matrix with the shape of 4C×2C, and b is the bias vector with the shape of 2C. The shape of the output S' is Next, input S' into the Multi-Swin Transformer Block module. The operation is the same as that in step S4, and the output is the feature map S l +1 , with the shape of S52: Repeat the operation of S51 twice. The first output is the feature map S l+2 , with the shape of The second output is the feature map S l+3 , with the shape of 6. The rail surface defect detection method according to claim 5, characterized in that, Step S6 specifically includes the following steps: Feature maps S l 、S l+1 、S l+2 、S l+3 obtained in steps S4 and S5 are respectively subjected to self-attention operations: Q, K, and V are obtained through linear transformation: Q = W q S, K = W k S, V = W v S, where Q, K, and V respectively represent the query, key, and value matrices, and W q 、W k 、W v are learnable weight matrices; then multi-head self-attention operations are performed: head i = Attention(Q i , K i , V i ), MultiHead(K, Q, V) = Concact(head1,..., head i )W O , where Q i , K i , V i respectively represent the query, key, and value matrices of the i-th attention head, i is the number of heads in the multi-head self-attention mechanism, W O is a learnable weight matrix, and the Attention operation is: where softmax represents the activation function, K T is the transpose of K, is the scaling factor, d k is the dimension of Q and K, and finally new feature maps K l 、K l+1 、K l+2 、K l+3 are obtained.

7. A method for detecting surface defects of a rail according to claim 6, characterized in that, Step S7 specifically includes the following steps: For the new feature maps K l , K l+1 , K l+2 , K l+3 , perform patch expanding operations with an upsampling factor of 2 0 times, 1 time, 2 times, and 3 times respectively. Taking K l+1 as an example, first rearrange and reshape the input: where K l+1 reshaped is the result after rearrangement and reshaping. Then adjust the dimension order for upsampling: K l+1 transposed = transpose(K l+1 reshaped , (0, 2, 1, 3, 4)), K l+1 transposed is the result after adjusting the dimension order. Next, reshape the adjusted tensor to increase the resolution: K l+1 expanded is the result after reshaping. Finally, apply a linear layer for feature transformation to obtain the new feature map F l+1 : F l+1 = S l+1 expandedW + b, where F l+1 is the result output by the linear layer, W is the weight matrix with a shape of C × 2C, b is the bias vector with a shape of 2C, and the outputs of K l , K l+2 , K l+3 are the new feature maps F l , F l+2 , F l+3 respectively, where the sizes of F l , F l+1 , F l+2 , F l+3 are all 8. A method for detecting surface defects of a steel rail according to claim 7, characterized in that, The specific steps of step S8 include the following steps: Concatenate F output by S7 l , F l+1 , F l+2 , F l+3 along the channel dimension to form a feature map with a size of . The operation is: F = concat(F l , F l+1 , F l+2 , F l+3 , axis = 3), where F is the concatenated feature map, and axis = 3 indicates concatenation along the channel dimension. Then, pass F through a Patch Expanding module with an upsampling ratio of 4 to restore the feature map resolution to the input size. The operation is: F reshaped is the result after rearrangement and shape adjustment, F transposed = transpose(F reshaped , (0, 2, 1, 3, 4)), F transposed is the result after adjusting the dimension order, F expanded = reshape(F transposed , (H, H, C)), F expanded is the result after reshaping again. Finally, pass F expanded through a convolutional layer to obtain a segmentation mask.

Citation Information

Cited By

  • Defect detection method and device, computer equipment and computer readable storage medium

    CN120931612A