Train track defect detection method and system based on improved ViDT

By improving the FFN layer structure and the use of LPA modules of the ViDT model, high accuracy and high recall rate of train track defect detection are achieved, and the problem of poor detection effect of existing methods in complex environments is solved, and detection efficiency and safety are improved.

CN120298987APending Publication Date: 2025-07-11BEIJING MASS TRANSIT RAILWAY OPERATION CORPORATION LIMITED
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510279904.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing CNN-based train track defect detection methods have shortcomings in recognition accuracy and recall, especially in complex environments, and traditional manual inspection methods are inefficient and unsafe.

Method used

Using the improved ViDT model, the FFN layer in the original ViDT model is changed to a parallel structure, and the multi-scale feature map is processed using the LPA module to achieve the characteristics of local information and global information. The multi-scale feature map is generated through the hierarchical Swin Transformer backbone and RAM module, and finally the target position and category are generated by the detection head module.

Benefits of technology

It significantly improves the recognition accuracy and recall rate of the model, especially in complex environments, and is better than other mainstream target detection methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298987A_ABST
    Figure CN120298987A_ABST
Patent Text Reader

Abstract

The invention discloses a train track defect detection method and system based on improved ViDT, and the method comprises the steps: inputting a data set containing a train track surface defect image into an improved ViDT model for training, inputting the collected train track surface defect image into the improved ViDT model after the model training is completed, and detecting the train track surface defect. Outputting a defect detection result; the method comprises the following steps: changing an FFN layer in a ViDT original model into a parallel structure, and processing an output multi-scale feature map between a backbone network and a neck by using an LPA module; on the basis of an end-to-end target detection ViDT model, train track defects are detected, and a parallel FFN layer structure is designed through a linear gating mechanism so as to reduce information flow interference between different inputs; and for a multi-scale feature map output by the trunk part, feature fusion of local information and global information is realized by using a local pyramid attention network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of train track defect detection, and in particular to a train track defect detection method and system based on improved ViDT. Background Art

[0002] At present, the scale of my country's rail transit industry is growing rapidly, and rail transit safety has become a very important issue. Due to the high-speed operation of trains, any small defect on the track may cause a major accident. The traditional manual inspection method is not only inaccurate and inefficient, but also difficult to ensure the safety of technical workers. Therefore, applying machine vision technology to the task of train track defect detection and replacing "human inspection" with "machine inspection" has important practical significance.

[0003] With the advancement of deep learning theory and image processing technology, some object detection algorithms based on convolutional neural networks (CNN) have been applied to the field of rail defect detection. Zhang Ming used the optimized Faster RCNN algorithm to locate and identify rail surface defects, effectively solving the problems of low accuracy and missed detection of small target defects; Feng et al. combined YOLO and feature pyramids to use multi-scale feature maps to achieve rail defect detection in complex environments; Hsieh et al. developed an automatic detection system for rail fastener types based on YOLOv3; Du Shaocong et al. added an attention mechanism to the YOLOv5 algorithm to introduce global dependencies for defect features, which can achieve accurate detection of rail surface defects under different environmental conditions. Due to the lack of global information capabilities of CNN, in some complex environments, object detection algorithms based entirely on CNN may not be able to realize their due potential.

[0004] In recent years, the Transformer model has performed well in capturing global information and long-distance dependencies. This method was originally used in the field of natural language processing. Vison Transformer et al. proposed to apply Transformer to the field of computer vision. Nicolas Carion et al. proposed an end-to-end object detection method based on Transformer, called DETR. Zhu et al. introduced a deformable attention module to accelerate the convergence of DETR by utilizing multi-scale features. Song et al. integrated DETR and Vison Transformer to construct an efficient object detector ViDT. However, existing methods need to be improved in terms of recognition accuracy and recall rate. Summary of the invention

[0005] The purpose of the present invention is to provide a train track defect detection method and system based on improved ViDT, so as to further improve the recognition accuracy of the model and the model performance in complex environments.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] On the one hand, the present invention provides an improved ViDT train track defect detection method, including the following steps:

[0008] Input a dataset containing train track surface defect images into the improved ViDT model for training. After the improved ViDT model is trained, input the collected train track surface defect images to be detected into the improved ViDT model to detect train track surface defects and output the defect detection results;

[0009] The improved ViDT model changes the FFN layer in the original ViDT model to a parallel structure to avoid information interference between PATCH TOKEN and DET TOKEN; meanwhile, an LPA module is used to process the output multi-scale feature maps between the backbone network and the neck to achieve feature fusion of local information and global information;

[0010] The PATCH TOKEN generates multi-scale feature maps via a hierarchical Swin Transformer backbone and an LPA module, and the DET TOKRN extracts fine-grained features through a RAM module. Then the two are input into a deformable Transformer decoder together, and finally the detection head module generates the target position and category.

[0011] In some embodiments, the specific method of changing the FFN module in the original ViDT model to a parallel structure is as follows:

[0012] First, split the PATCH TOKEN after passing through the normalization layer and residual connection into two linear projections. One of the projections obtains semantic information through a depthwise separable convolutional layer and is activated through a gating function;

[0013] Secondly, multiply the two linear projections element-wise and output through a linear layer; the DET TOKEN remains unchanged and adopts the FFN layer design in the original ViDT model;

[0014] Finally, splice and fuse the output of the PATCH TOKEN and the output of the DET TOKEN as the input of the next Swin Transformer layer.

[0015] In some embodiments, the structure of the LPA module is as follows:

[0016] The first pyramid layer of the LPA module is used to obtain the global attention map;

[0017] The other pyramid layers of the LPA module divide the feature map into smaller dimensions, obtain more accurate target features through the local attention mechanism, and merge them into the attention map of the current pyramid layer;

[0018] Finally, the LPA module fuses the attention maps of different pyramid layers with each other to form the final attention map.

[0019] In some embodiments, in each pyramid layer of the other pyramid layers, the feature map is first divided into multiple feature vectors, and the weight vector of each feature vector is obtained through the attention mechanism; then these weight vectors are merged to obtain the weight vector of the current layer, and multiplied by the original feature map to obtain the output of each layer.

[0020] In some embodiments, the attention mechanism includes a channel attention CA module and a spatial attention SA module. The spatial attention SA module uses spatial information to find the task-related parts in the features, and the channel attention CA module measures the importance of the features by assigning different weights to the channels.

[0021] In some embodiments, the channel attention CA module is:

[0022] A S =σ(ξ(Concat(P avg (X),P max )));

[0023] The spatial attention SA module is:

[0024] A C =σ(FC2ReLU(FC1Pool avg (X)));

[0025] In the formula, σ is the sigmond function; P max is the adaptive max pooling operation; P avg is the adaptive average pooling operation; ξ is a 7×7 convolution; FC1 and FC2 are two fully connected layers; Pool avg is the average pooling layer.

[0026] On the other hand, the present invention provides an improved ViDT train track defect detection system based on the above method, including:

[0027] Dataset module: It is a storage module for including high-speed railway track image and general transportation track image datasets, and inputs the image dataset into the improved ViDT module;

[0028] Training module: Adopts the Pytorch deep learning framework and is used to train the improved ViDT model;

[0029] Improved ViDT Module: It is used to detect whether there are defects in the input images of train track surface defects. Its architecture changes the FFN layer in the original ViDT model to a parallel structure to avoid information interference between PATCH TOKEN and DET TOKEN. At the same time, an LPA module is used between the backbone network and the neck to process the output multi-scale feature maps to achieve feature fusion of local information and global information;

[0030] The PATCH TOKEN generates multi-scale feature maps through a hierarchical Swin Transformer backbone and an LPA module. The DET TOKRN extracts fine-grained features through a RAM module, and then the two are input into a deformable Transformer decoder together. Finally, the detection head module generates the target position and category;

[0031] Image Acquisition Module: It is used to acquire images of train track surface defects and input these images of train track surface defects to be detected into the improved ViDT module;

[0032] Output Module: It is used to output the detection results of the improved ViDT module.

[0033] Compared with the prior art, the present invention has the following beneficial effects:

[0034] Based on the end-to-end object detection ViDT model, the present invention detects train track defects. In the Swin Transformer backbone part, through a linear gating mechanism, the present invention designs a parallel FFN layer structure to reduce the interference of information flow between different inputs. For the multi-scale feature maps output by the backbone part, the present invention uses a local pyramid attention network to achieve feature fusion of local information and global information.

[0035] Experiments have proved that the improved ViDT model of the present invention has achieved relatively ideal results on the extended RSDDs dataset. Compared with the original model, the recognition accuracy and recall rate of the improved model have been significantly improved. Compared with other mainstream object detection methods, the present invention has more accurate detection accuracy. Description of the Drawings

[0036] Figure 1 It is a schematic diagram of the architecture of the FFN module in the original ViDT model;

[0037] Figure 2 It is a schematic diagram of the architecture of the FFN module in the improved ViDT model of the present invention;

[0038] Figure 3 It is a schematic diagram of the architecture of the LPA module in the improved ViDT model of the present invention;

[0039] Figure 4 Schematic diagram of the overall architecture of the improved ViDT model of the present invention;

[0040] Figure 5 During the training process of the improved ViDT model of the present invention, it is a schematic diagram of the changes in the learning rate, loss function value, and AP@0.5 index on the validation set. Specific implementation manners

[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0042] Embodiment 1:

[0043] Please refer to Figures 1 - 5 , a train track defect detection method based on improved ViDT, input a dataset containing images of train track surface defects into the improved ViDT model for training. After the improved ViDT model is trained, input the collected images of train track surface defects to be detected into the improved ViDT model to detect the train track surface defects and output the defect detection results.

[0044] The original ViDT model is:

[0045] The original ViDT model consists of a Swin Transformer backbone network with a Reconfigured Attention Module (RAM) and a neck structure without an encoder. The Swin Transformer backbone network uses a hierarchical structure and constructs feature maps of different sizes based on 4 different stages. Except for the first stage, each stage will downsample the feature map of the previous stage through a Patch Merging layer, and the number of channels will double. After downsampling, the feature map will be further feature-extracted through a series of stacked Swin Transformer blocks. The Swin Transformer module uses two attention mechanisms in pairs: the W-MSA module and the SW-MSA module. The W-MSA module divides the image into multiple windows and performs self-attention calculations within each window; the SW-MSA module realizes cross-window attention calculations through window displacement.

[0046] When the original ViDT model takes input, it adds learnable detection blocks (DET TOKEN) to the sequence of image patches (PATCH TOKEN), which are responsible for representing and predicting the categories and locations of various objects in the image. The introduced RAM module decomposes the single global attention related to image patches and detection blocks into PATCH×PATCH attention, DET×DET attention, and DET×PATCH attention.

[0047] The PATCH×PATCH attention is as follows: Using the attention calculation method of Swin Transform, the W-MSA module performs local attention calculation on each window partition, and the SW-MSA module realizes the connection between the shifted window partition and the window of the previous layer to obtain global information.

[0048] The DET×DET attention is as follows: Since the number of detection blocks specifies the number of objects to be detected and there is no locality, global self-attention calculation is performed while maintaining its number.

[0049] The DET×PATCH attention is as follows: Cross-attention is adopted, and an embedding is generated for each detection block. For each image patch, the key values in the detection blocks are aggregated to represent the target object. Since the detection blocks specify different objects, it generates different object embeddings for different objects.

[0050] Since the RAM module directly extracts fine-grained features suitable for object detection, the original ViDT model can use a neck structure without an encoder. To utilize multi-scale feature maps, the original ViDT model integrates a multi-layer deformable Transformer decoder. The decoder receives the multi-scale image patches generated at each stage in the Swin Transformer backbone and the detection blocks generated at the last stage, and finally outputs the target regions and categories via the detection head.

[0051] As Figure 1Shown is the FFN module in the original ViDT model. After the attention calculation of the RAM module, the DETTOKEN and PATCH TOKEN obtained are first combined through a concatenation operation and fused with the input of the current Swin Transformer layer through residual connection, and then output to the normalization layer for processing. Subsequently, feature transformation is performed through a typical FFN module, which consists of two linear transformations and an activation function. Through the FFN module, the model can capture the complex relationships between input features and enhance the feature expression ability. However, in the original ViDT model, for the processing of DET TOKEN, the FFN layer usually focuses on optimizing tasks such as object class prediction and bounding box position regression; while for PATCH TOKEN, the FFN layer pays more attention to the spatial details, semantic information, and fine-grained representation of the image. Due to the different focuses of processing these two inputs, simply concatenating them and passing them to a single FFN layer may lead to confusion and interference in the information flow, thus affecting the detection effect of the model.

[0052] As Figure 2 Shown is the FFN module in the improved ViDT model. The DETTOKEN and PATCH TOKEN are processed separately through parallel FFN modules. For PATCH TOKEN, a convolutional linear gating unit is used to replace the original FFN layer. In this unit, the PATCH TOKEN after passing through the normalization layer and residual connection is first split into two linear projections. One of the projections obtains semantic information through a depthwise separable convolutional layer and is activated through a gating function. Finally, the two linear projections are multiplied element-wise and output through a linear layer. By introducing a gating mechanism, the network can dynamically select important features according to different inputs, selectively "turn on" or "turn off" certain features, thereby enhancing the representation ability of the network. For DET TOKEN, the FFN layer design in the original ViDT model is followed in this embodiment. Finally, the outputs of PATCH TOKEN and DET TOKEN are concatenated and fused as the input of the next Swin Transformer layer.

[0053] Through the design of parallel FFN layers, the improved ViDT model of the present invention can independently extract and optimize the features of the two inputs, achieve more targeted feature transformation, reduce the interference between features, and thus improve the overall performance of the model.

[0054] As Figure 3 Shown is the architecture of the LPA module in the improved ViDT model. The multi-scale feature maps generated by the Swin Transformer backbone are fused with global features and local features through a Local Pyramid Attention Module (LPA).

[0055] The LPA module uses a pyramid structure to learn multi-scale features. First, the first layer of the pyramid obtains the global attention map. Then, the other layers of the pyramid divide the feature map into smaller dimensions, obtain more accurate target features through the local attention mechanism, and merge them into the attention map of the current pyramid layer. Finally, the attention maps of different pyramid layers are fused with each other to form the final attention map.

[0056] In each pyramid layer P i first, the feature map is divided into multiple feature vectors, and the weight vector of each feature vector is obtained through the attention mechanism. Then these weight vectors are merged to obtain the weight vector of the current layer, and multiplied by the original feature map to obtain the output of each layer. The attention mechanism consists of two attention blocks: the CA block and the SA block.

[0057] Among them, the SA module represents spatial attention, which uses spatial information to find the task-related parts in the features, and is defined as:

[0058] A S =σ(ξ(Concat(P avg (X),P max ))) (1);

[0059] In the formula, σ is the sigmond function; P max is the adaptive max pooling operation; P avg is the adaptive average pooling operation; ξ is a 7×7 convolution.

[0060] The CA module represents channel attention, which measures the importance of features by assigning different weights to channels. The calculation process of channel attention is defined as:

[0061] A C =σ(FC2ReLU(FC1Pool avg (X))) (2);

[0062] In the formula, FC1 and FC2 are two fully connected layers; Pool avg is the average pooling layer.

[0063] Such as Figure 4The figure shows a schematic diagram of the improved ViDT model architecture. In the present invention, the FFN module of the original ViDT model is improved to a parallel structure design to avoid information interference between PATCH TOKEN and DET TOKEN. At the same time, an LPA module is used between the backbone network and the neck to process the output multi-scale feature maps to achieve feature fusion of local information and global information, focus on the target area while suppressing irrelevant information, and improve the detection effect of the model in complex backgrounds. In this model, the PATCHTOKEN generates multi-scale feature maps through a hierarchical Swin Transformer backbone and an LPA module, and the DET TOKRN extracts fine-grained features through a RAM module. The two are input into a deformable Transformer decoder together, and finally the detection head module generates the target position and category.

[0064] Finally, the improved ViDT model is used to detect the surface defects of the train track.

[0065] The dataset uses the augmented RSDDs dataset to train and evaluate the model. Each image contains at least one defect, and the background is complex and has a lot of noise. There are 1,117 training images and 319 test images in total, and the number of defect categories is 8.

[0066] In this embodiment, the Pytorch deep learning framework is adopted. Under the Ubuntu system environment, the improved ViDT model is trained using an NVIDIA RTX4090 GPU. During training, the backbone part uses swin nano, the optimizer uses AdamW, the batch size is set to 2, the number of training epochs is 200, and the initial learning rate is set to 1×10 -4 During the training process, the changes in the learning rate, loss function value, and the AP@0.5 metric on the validation set are as Figure 5 shown. As the number of training epochs increases, the loss function value of the model gradually decreases, the entire network tends to converge, and the fitting effect is better. At the same time, the average precision AP@0.5 of the model on the validation set also stabilizes at about 0.95, indicating that the detection effect of the model is better.

[0067] To verify the improvement effect, the present invention conducted ablation experiments, and the experimental results are shown in Table 1. As can be seen from Table 1, although the model with the LPA module is slightly inferior to the original model in terms of AP@0.5, other indicators are better than the original model. Especially in the AP and AR of medium-sized targets, they are improved by 2.3% and 1.9% respectively. By using the parallel FFN module, the recognition accuracy of the model has been significantly improved, with an increase of 8.7% and 9.0% in the AP@0.5 and AP@0.5:0.95 indicators respectively, which shows the rationality and superiority of the improved parallel FFN structure of the present invention. Finally, by comprehensively using the LPA module and the parallel FFN layer, the highest AP@0.5 and the average precision on medium-sized targets are achieved.

[0068] Table 1 Detection effect table of different combination strategies

[0069]

[0070] To evaluate the advantages of the improved ViDT model in the rail defect detection task, the currently mainstream object detection algorithms YOLOv5n, YOLOv8s, YOLOV10n, and Deformable DETR are selected as comparison models. Table 2 shows the performance comparison of different models on the expanded RSDDs dataset. From the results in Table 2, it can be seen that the model of the present invention has the best comprehensive performance.

[0071] Table 2 Performance comparison table of mainstream object detection models

[0072]

[0073] Specifically, the improved ViDT model achieved the highest AP@0.5, which is 10.3% higher than Deformable DETR and 3.9% higher than YOLOv8s. In terms of the AP@0.5:0.95 indicator, the model of the present invention also performed outstandingly, with an increase of 15.1% compared to Deformable DETR and 0.4% compared to YOLOv8s, second only to YOLOV10n. Generally speaking, the model of the present invention has significant advantages in the rail defect detection task and can meet the actual engineering needs.

[0074] Embodiment 2

[0075] An improved ViDT-based train rail defect detection system, comprising:

[0076] Dataset module: including a storage module for high-speed railway track image and ordinary transportation track image datasets, and inputting the image dataset into the improved ViDT module.

[0077] Training module: Using the Pytorch deep learning framework, under the Ubuntu system environment, the improved ViDT model is trained using the NVIDIA RTX4090 GPU. When the average precision AP@0.5 stabilizes at around 0.95, it indicates that the detection effect of the model is good, and the training can be stopped.

[0078] Improved ViDT module: Used to detect whether there are defects in the input train track surface defect images. Its architecture changes the FFN layer in the original ViDT model to a parallel structure to avoid information interference between PATCH TOKEN and DET TOKEN. At the same time, the LPA module is used to process the output multi-scale feature maps between the backbone network and the neck to achieve the feature fusion of local information and global information.

[0079] The PATCH TOKEN generates multi-scale feature maps through the hierarchical Swin Transformer backbone and the LPA module, and the DET TOKRN extracts fine-grained features through the RAM module. Then the two are input into the deformable Transformer decoder together, and finally the detection head module generates the target position and category.

[0080] Image acquisition module: Using image acquisition devices such as cameras and high-definition cameras to collect train track surface defect images, and input these train track surface defect images to be detected into the improved ViDT module.

[0081] Output module: Used to output the detection results of the improved ViDT module.

[0082] An improved ViDT-based train track defect detection system of the present invention can be installed in a computer device. The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, such as an improved ViDT-based train track defect detection program. Among them, the memory includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), magnetic memories, disks, optical discs, etc. The processor is the control core of the electronic device, connecting various components of the entire computer device through various interfaces and lines, and executing various functions of the computer device and processing data by running or executing programs or modules stored in the memory and calling data stored in the memory.

[0083] The module described in the present invention refers to a series of computer program segments that can be executed by the processor of a computer device and can complete fixed functions, and are stored in the memory of the computer device.

[0084] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.

Claims

1. An improved ViDT-based train track defect detection method, characterized in that, It includes the following steps: Input the dataset containing the images of train track surface defects into the improved ViDT model for training. After the improved ViDT model is trained completely, input the collected images of train track surface defects to be detected into the improved ViDT model to detect the train track surface defects and output the defect detection results; The improved ViDT model changes the FFN layer in the original ViDT model to a parallel structure to avoid information interference between the PATCH TOKEN and the DET TOKEN; at the same time, use the LPA module between the backbone network and the neck to process the output multi-scale feature maps to achieve the feature fusion of local information and global information; The PATCH TOKEN generates multi-scale feature maps through the hierarchical Swin Transformer backbone and the LPA module, and the DET TOKRN extracts fine-grained features through the RAM module, and then the two are input into the deformable Transformer decoder together, and finally the detection head module generates the target position and category.

2. The improved ViDT train track defect detection method according to claim 1, characterized in that, The specific method of changing the FFN module in the original ViDT model to a parallel structure is as follows: First, split the PATCH TOKEN after passing through the normalization layer and the residual connection into two linear projections. One of the projections obtains semantic information through the depthwise separable convolutional layer and is activated through the gating function; Second, multiply the two linear projections element-wise and output through the linear layer; the DET TOKEN remains unchanged and adopts the FFN layer design in the original ViDT model; Finally, splice and fuse the output of the PATCH TOKEN and the output of the DET TOKEN as the input of the next SwinTransformer layer.

3. An improved ViDT-based train track defect detection method according to claim 1, characterized in that, The structure of the LPA module is: The first pyramid layer of the LPA module is used to obtain the global attention map; The other pyramid layers of the LPA module divide the feature map, obtain the target features through the local attention mechanism, and merge them into the attention map of the current pyramid layer; Finally, the LPA module fuses the attention maps of different pyramid layers with each other to form the final attention map.

4. The improved ViDT-based train track defect detection method according to claim 3, wherein In each pyramid layer of the other pyramid layers, first divide the feature map into multiple feature vectors, and obtain the weight vector of each feature vector through the attention mechanism; then merge these weight vectors to obtain the weight vector of the current layer, and multiply it with the original feature map to get the output of each layer.

5. The method for detecting train track defects based on the improved ViDT according to claim 4, wherein, The attention mechanism includes a channel attention CA module and a spatial attention SA module.

6. The improved ViDT-based train track defect detection method according to claim 5, wherein, The channel attention CA module is: A S = σ(ξ(Concat(P avg (X),P max ))); The spatial attention SA module is: A C = σ(FC2ReLU(FC1Pool avg (X))); where σ is the sigmond function; P max is the adaptive max pooling operation; P avg is the adaptive average pooling operation; ξ is a 7×7 convolution; FC1 and FC2 are two fully connected layers; Pool avg is the average pooling layer.

7. An improved ViDT train track defect detection system that applies the method described in any one of claims 1-6, characterized in that, It includes: Dataset module: It is a storage module including the high-speed railway track image and ordinary transportation track image datasets, and inputs the image dataset into the improved ViDT module; Training module: Adopt the Pytorch deep learning framework to train the improved ViDT model; Improved ViDT Module: It is used to detect whether there are defects in the input images of train track surface defects. Its architecture changes the FFN layer in the original ViDT model to a parallel structure to avoid information interference between PATCH TOKEN and DET TOKEN. At the same time, an LPA module is used between the backbone network and the neck to process the output multi-scale feature maps to achieve the feature fusion of local information and global information; The PATCH TOKEN generates multi-scale feature maps through the hierarchical Swin Transformer backbone and the LPA module. The DET TOKRN extracts fine-grained features through the RAM module, and then the two are input into the deformable Transformer decoder together. Finally, the detection head module generates the target position and category; Image Acquisition Module: It is used to acquire the images of train track surface defects and input the images of train track surface defects into the improved ViDT module; Output Module: It is used to output the detection results of the improved ViDT module.

Citation Information

Cited By

  • High-resolution image-oriented small and micro wetland ground feature type identification method, apparatus and device, and medium

    CN121236627A