Track foreign matter intrusion detection method and related device
By integrating deep separable and deformable convolutions with a deformable agent attention mechanism, the method addresses the precision-speed imbalance in railway foreign object detection, ensuring real-time warning capabilities for high-speed trains.
Patent Information
- Application Number
- CN202510287246.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-07-15
AI Technical Summary
The existing track obstacle detection methods are difficult to balance detection accuracy and inference speed, especially in small-objective detection and embedded equipment deployment, which cannot meet the real-time early warning requirements under high-speed operation of trains.
The MSD module that can separate convolution and deformable convolution with the fusion depth and the DAA module of deformable agent attention mechanism are adopted to enhance feature extraction capabilities and reduce the amount of calculation, optimize attention weight distribution through sparse query, and improve small object detection accuracy and inference speed.
Real-time detection of foreign objects on the track under high-speed operation of the train, improve detection accuracy and inference speed, meet real-time early warning needs, and is suitable for embedded equipment deployment.
Smart Images

Figure CN120318150A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical field of foreign object intrusion detection, and in particular, to a method and related device for detecting foreign object intrusion on a track. Background Art
[0002] High-speed railways and subways are widely constructed globally and are favored as the main means of passenger and cargo transportation due to their cost-effectiveness, speed, and safety. Foreign object intrusion (FOI) within the track gauge area poses a continuous threat to the safe operation of trains. When a train encounters a dangerous obstacle during operation, it must brake urgently to avoid accidents. However, due to the high speed of the train, the obstacle must be identified in advance and effective protective measures must be taken. Relying on the observation of train drivers is vulnerable to natural environment and human factors, resulting in an extended reaction time.
[0003] With the rapid development of deep learning in the field of image processing, models based on convolutional neural networks (CNNs) and Transformers have shown significant advantages in track obstacle detection, and there are also related studies currently. Although the existing research has a certain detection ability, these methods have problems in foreign object detection where the detection accuracy and inference speed cannot achieve a good balance. For example, methods with high detection accuracy are slower in inference speed, and methods with fast inference speed are difficult to meet the requirements in terms of detection accuracy. In addition, for the detection of small foreign objects on the track, existing methods are difficult to achieve fast and accurate detection. Generally speaking, 1. The ability of multi-scale object detection is insufficient, especially in the detection of small objects. 2. The model has a large amount of computation and is difficult to be deployed on embedded devices. 3. The detection speed cannot meet the real-time warning requirements under the high-speed operation of trains. Summary of the Invention
[0004] The embodiments of the present invention provide a method and related device for detecting foreign object intrusion on a track for rail transit foreign object intrusion detection. By fusing the depthwise separable convolution and deformable convolution MSD module, the feature extraction ability is enhanced and the computation amount is reduced. At the same time, the deformable proxy attention mechanism DAA module is designed to optimize the attention weight distribution through sparse queries, effectively improving the detection accuracy and inference speed of small objects in complex scenarios, and being able to meet the real-time warning requirements under the high-speed operation of trains.
[0005] In a first aspect, the embodiments of the present invention provide a method for detecting foreign object intrusion on a track, including:
[0006] Obtaining a real-time railway track image through an on-vehicle camera;
[0007] Performing feature extraction on the railway track image to obtain image features;
[0008] Input the image features into the trained foreign object detection model for foreign object intrusion detection to obtain the railway track foreign object intrusion detection result. Among them, the foreign object detection model is improved based on the RT-DETR object detection model. The foreign object detection model includes an MSD module. The MSD module is a module that combines depthwise separable convolution (DWC) and deformable convolution (DCNv4). The foreign object detection model uses a deformable agent attention mechanism (DAA) module to replace the AIFI module in RT-DETR. The DAA module optimizes the attention weight distribution through sparse queries.
[0009] In some embodiments, in the MSD module, the depthwise separable convolution (DWC) is used to process the channel features of the first part and the second part. The channel features of the third part use the deformable convolution (DCNv4) to adaptively focus on the important feature regions in the railway track image by dynamically adjusting the shape and position of the convolution kernel. The channel features of the fourth part do not perform any operations. After concatenating the channel features of the first part, the second part, the third part, and the fourth part, a shuffle operation is performed to enable cross-fusion of information between all channels.
[0010] In some embodiments, the shuffle operation to enable cross-fusion of information between all channels is expressed as follows:
[0011] Outout=Shuffle(cat(DWC3(x0),DWC5(x1),DCN(x2),X3) (1)
[0012] Where DWC3 and DWC5 represent DWC with convolution kernel sizes of 3×3 and 5×5.
[0013] In some embodiments, the DAA module includes a query vector Q, a proxy vector A, a key vector K, and a value vector V. The proxy vector A serves as a proxy for the query vector Q, aggregates information from the key vector K and the value vector V, and then broadcasts the information back to the query vector Q.
[0014] In some embodiments, deformable points in the deformable attention mechanism are used in the DAA module. The DAA module includes a deformable attention module of an offset network. The DAA module generates offsets based on the query features as reference points, creates deformable points to determine the positions of the proxy vector A and the key vector K, so that the proxy vector A and the key vector K can notice the key regions of the railway track image.
[0015] In some embodiments, the DAA module generates offsets based on the query features as reference points, creates deformable points to determine the positions of the proxy vector A and the key vector K, including:
[0016] The image features obtain a uniform grid size through a downsampling factor;
[0017] Generate reference points through a uniform grid in the image space;
[0018] Obtain the reference point offset;
[0019] Linearly project the input features to generate a query vector Q;
[0020] Put the query vector Q into a dedicated lightweight network to generate an offset;
[0021] Obtain sampled features through spatial sampling after the deformable points;
[0022] Perform linear projection and pooling operations on the sampled features to obtain the positions of the proxy vector A and the key vector K.
[0023] In some embodiments, the input-output process of the DAA module includes:
[0024] Given an input image feature x ∈ R H×W×C , obtain a uniform grid size of H G = H / r, W G = W / r through a downsampling factor r, and generate reference points through a uniform grid in the image space for subsequent deformable convolution operations; then obtain the reference point offset, linearly project the input features to generate a query q = xW q , and put it into a dedicated lightweight network θ offset (·) to generate an offset θ offset (q). Obtain sampled features through spatial sampling after the deformable points, and perform linear projection and pooling operations to obtain keys and proxies, as shown in equations (2) and (3):
[0025]
[0026]
[0027] where and Agent Token represent the embeddings of the deformed keys and proxies; set the sampling function to bilinear interpolation, as shown in equation (4):
[0028]
[0029] where the function g(a,b) represents the interpolation weight function in the x and y directions, and (r x ,r y ) represents in z ∈ RH×W×C Indices at all positions above; after obtaining q and the agent token, the further inference of the DAA module is expressed as:
[0030]
[0031] where A ∈ R n×c is the agent token defined by formula (2), σ(·) represents the Softmax function, which is used for agent aggregation and agent broadcasting, and v = xW v is the value obtained by linear projection of the input features; specifically, first, the agent vector A is regarded as a query, and information is aggregated from all features through the σ(AK T )V operation to obtain V A , then the agent vector A is used as the key vector k and the query vector Q for Softmax calculation and then dot-product operation with V A as the value v, broadcasting the global information in the agent features back to each feature, and obtaining the final output O; by using a small number of agents A as the agents of q, collecting information from K and V and presenting it to Q, setting the number of agent vectors A to be much smaller than the number of query vectors Q, the final output of the DAA module is:
[0032]
[0033] where B1 and B2 are agent biases, combining agent biases to utilize position information and using depthwise separable convolution DWC to maintain feature diversity.
[0034] In a second aspect, an embodiment of the present invention further provides an orbital foreign object intrusion detection device, and the device includes:
[0035] An acquisition module, configured to acquire real-time railway track images through an on-vehicle camera;
[0036] An extraction module, configured to extract features from the railway track images to obtain image features;
[0037] A detection module, configured to input the image features into a trained foreign object detection model for foreign object intrusion detection to obtain an orbital foreign object intrusion detection result, where the foreign object detection model is improved based on the object detection model of RT-DETR, the foreign object detection model includes an MSD module, the MSD module is a module that combines depthwise separable convolution DWC and deformable convolution DCNv4, the foreign object detection model uses a deformable agent attention mechanism DAA module to replace the AIFI module in RT-DETR, and the DAA module optimizes the attention weight distribution through sparse queries.
[0038] In a third aspect, an embodiment of the present invention further provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the rail foreign object intrusion detection method described in the first aspect is implemented.
[0039] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium storing computer-executable instructions for executing the rail foreign object intrusion detection method described in the first aspect.
[0040] According to the rail foreign object intrusion detection method and related devices provided by the embodiments of the present invention, the rail foreign object intrusion detection method includes: obtaining a real-time railway track image through an on-vehicle camera; extracting features from the railway track image to obtain image features; inputting the image features into a trained foreign object detection model for foreign object intrusion detection to obtain a rail foreign object intrusion detection result, wherein the foreign object detection model is improved based on the RT-DETR object detection model, the foreign object detection model includes an MSD module, and the MSD module is a module that combines depthwise separable convolution (DWC) and deformable convolution (DCNv4). The foreign object detection model replaces the AIFI module in RT-DETR with a deformable proxy attention mechanism (DAA) module, and the DAA module optimizes the attention weight distribution through sparse queries. The MSD module proposed in the embodiment of the present invention combines depthwise separable convolution (DWC) and deformable convolution (DCNv4) to expand the receptive field and improve the multi-scale object detection ability while reducing the computational amount, thereby improving the real-time performance on embedded devices and the detection accuracy in the track environment. The embodiment of the present invention also proposes a deformable proxy attention (DAA) module. This module focuses on important regions through deformable points, enabling the extraction of more critical feature information, and using a small number of agents as queries to avoid redundancy between attention weights, thereby improving the inference speed of the model. Based on this, the embodiment of the present invention uses an on-vehicle camera to detect rail transit obstacles, enhances the feature extraction ability and reduces the computational amount through the MSD module, and at the same time designs the DAA module to optimize the attention weight distribution through sparse queries, effectively improving the detection accuracy and inference speed of small targets in complex scenarios, thereby enhancing the ability to identify foreign objects on the track and meeting the real-time warning requirements under high-speed train operation to ensure the safe operation of the train. Description of the Drawings
[0041] Figure 1 is a flowchart of the rail foreign object intrusion detection method provided by an embodiment of the present invention;
[0042] Figure 2It is a schematic diagram of the MSD module structure provided by an embodiment of the present invention;
[0043] Figure 3 They are comparison graphs of each convolution provided by an embodiment of the present invention;
[0044] Figure 4 It is a schematic diagram of the deformable proxy attention DAA module structure provided by an embodiment of the present invention;
[0045] Figure 5 They are the comparison of the original image provided by an embodiment of the present invention and the feature heat maps before and after using the DAA module;
[0046] Figure 6 It is a schematic diagram of the improved RT-DETR structure provided by an embodiment of the present invention;
[0047] Figure 7 It is a distribution diagram of the target height and width and the target center point position in the TAD dataset provided by an embodiment of the present invention;
[0048] Figure 8 They are the comparison of the true annotation box, the predicted image of the RT-DETR model, and the predicted image of the improved RTDETR model provided by an embodiment of the present invention;
[0049] Figure 9 It is a schematic diagram of the structure of the track foreign object intrusion detection device provided by an embodiment of the present invention;
[0050] Figure 10 It is a schematic diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0051] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0052] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the description, claims and the following drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0053] In the embodiments of the present invention, words such as "furthermore", "exemplarily", or "optionally" are used to represent examples, illustrations, or explanations, and should not be construed as being more preferred or having more advantages than other embodiments or design solutions. The use of words such as "furthermore", "exemplarily", or "optionally" aims to present related concepts in a specific manner.
[0054] First, several terms involved in the present invention are analyzed:
[0055] MSD module (Multi-Scale Deformable Module): A module that combines depthwise separable convolution (DWC) and deformable convolution (DCNv4) to enhance the multi-scale feature extraction ability.
[0056] DAA mechanism (Deformable Agent Attention Mechanism): A mechanism for optimizing the attention weight distribution, which improves the inference speed through sparse queries.
[0057] RT-DETR: A Transformer-based object detection model for real-time object detection tasks.
[0058] ResNet: (Residual Network) is a deep convolutional neural network that constructs a residual learning structure by introducing shortcut connections, effectively solving the problem of gradient disappearance in the training of deep networks, making it possible to train deep networks with more than a hundred layers, and achieving breakthrough performance in image recognition tasks.
[0059] Softmax: A normalized exponential function that converts a real-valued vector into a probability distribution. By performing an exponential operation on each element and dividing by the sum of the exponentials of all elements, all output values fall within the interval (0, 1) and the sum is 1. It is often used in the output layer of neural networks for multi-classification tasks to represent the class prediction probability.
[0060] For the convenience of further describing the working principle of the embodiments of the present invention, the following first gives an introduction to the related technical scenarios.
[0061] High-speed railways and subways are widely built globally and are favored as the main means of passenger and freight transportation due to their cost-effectiveness, speed, and safety. Foreign object intrusion (FOI) within the track clearance area poses a continuous threat to the safe operation of trains. When a train encounters a dangerous obstacle during operation, it must brake urgently to avoid accidents. However, due to the high speed of the train, obstacles must be identified in advance and effective protective measures must be taken. Relying on the observation of train drivers is easily affected by natural environments and human factors, resulting in an extended reaction time.
[0062] With the rapid development of deep learning in the field of image processing, models based on convolutional neural networks (CNNs) and Transformers have shown significant advantages in railway obstacle detection, and there are also related studies. Although the current studies have certain detection capabilities, these methods have problems in foreign object detection where the detection accuracy and inference speed cannot achieve a good balance. For example, methods with high detection accuracy are slow in inference speed, and methods with fast inference speed are difficult to meet the requirements in detection accuracy. In addition, for the detection of small foreign objects on the railway track, existing methods are difficult to achieve fast and accurate detection. Generally speaking, there are three problems: 1. Insufficient multi-scale object detection ability, especially in the detection of small objects. 2. The model has a large amount of calculation and is difficult to be deployed on embedded devices. 3. The detection speed cannot meet the real-time warning requirements under the high-speed operation of the train.
[0063] Based on this, the present invention provides a method and related device for detecting railway foreign object intrusion. The method for detecting railway foreign object intrusion includes: obtaining a real-time railway track image through an on-vehicle camera; extracting features from the railway track image to obtain image features; inputting the image features into a trained foreign object detection model for detecting railway foreign object intrusion to obtain a railway foreign object intrusion detection result. The foreign object detection model is improved based on the RT-DETR object detection model. The foreign object detection model includes an MSD module, and the MSD module is a module that combines depthwise separable convolution (DWC) and deformable convolution (DCNv4). The foreign object detection model uses a deformable agent attention mechanism (DAA) module to replace the AIFI module in RT-DETR, and the DAA module optimizes the attention weight distribution through sparse queries. The MSD module proposed in the embodiment of the present invention combines depthwise separable convolution (DWC) and deformable convolution (DCNv4) to expand the receptive field and improve the multi-scale object detection ability while reducing the amount of calculation, thereby improving the real-time performance on embedded devices and the detection accuracy in the railway environment. The embodiment of the present invention also proposes a deformable agent attention mechanism (DAA) module. This module focuses on important regions through deformable points, enabling the extraction of more critical feature information, and using a small number of agents as queries to avoid redundancy between attention weights, thereby improving the inference speed of the model. Based on this, the embodiment of the present invention uses an on-vehicle camera to detect railway traffic obstacles, enhances the feature extraction ability and reduces the amount of calculation through the MSD module, and at the same time designs the DAA module to optimize the attention weight distribution through sparse queries, effectively improving the detection accuracy and inference speed of small objects in complex scenarios, thereby enhancing the ability to identify foreign objects on the railway track, meeting the real-time warning requirements under the high-speed operation of the train, and ensuring the safe operation of the train.
[0064] The following further elaborates on the embodiments of the present invention in conjunction with the accompanying drawings.
[0065] As shown Figure 1 in Figure 1 Figure 1, it is a flowchart of a rail foreign object intrusion detection method provided by an embodiment of the present invention. The rail foreign object intrusion detection method may include, but is not limited to, steps S101 to S103.
[0066] Step S101: Obtain real-time railway track images through an on-vehicle camera;
[0067] Step S102: Extract features from the railway track images to obtain image features;
[0068] Step S103: Input the image features into a trained foreign object detection model for foreign object intrusion detection to obtain a rail foreign object intrusion detection result. Among them, the foreign object detection model is improved based on the RT-DETR object detection model. The foreign object detection model includes an MSD module. The MSD module is a module that combines depthwise separable convolution (DWC) and deformable convolution (DCNv4). The foreign object detection model uses a deformable agent attention mechanism (DAA) module to replace the AIFI module in RT-DETR. The DAA module optimizes the attention weight distribution through sparse queries.
[0069] It can be understood that the MSD module proposed in the embodiment of the present invention, by combining depthwise separable convolution (DWC) and deformable convolution (DCNv4), expands the receptive field and improves the multi-scale object detection ability while reducing the computational amount, thereby improving the real-time performance on embedded devices and the detection accuracy in the rail environment.
[0070] It can be understood that an embodiment of the present invention proposes a deformable agent attention mechanism (DAA) module. This module focuses on important regions through deformable points, enabling the extraction of more critical feature information, and using a small number of agents as queries to avoid redundancy between attention weights, thereby improving the inference speed of the model.
[0071] The embodiment of the present invention uses an on-vehicle camera to detect rail transit obstacles, enhances the feature extraction ability and reduces the computational amount through the MSD module, and at the same time designs the DAA module to optimize the attention weight distribution through sparse queries, effectively improving the detection accuracy and inference speed of small targets in complex scenarios, thereby enhancing the recognition ability of foreign objects on the track and meeting the real-time warning requirements under high-speed train operation to ensure the safe operation of the train.
[0072] It is understandable that obstacle detection in rail transit faces the dual challenges of coexistence of multi-scale targets (such as large obstacles in the near field and tiny debris in the far field) and complex background interference (rail textures, lighting changes). Due to the limitations of the fixed geometric structure and single receptive field of traditional convolutional neural networks (CNNs), it is difficult to balance multi-scale feature extraction at a low computational cost. Therefore, the present invention proposes an MSD module (Multi-Scale Deformable Module), which realizes the following objectives through the cascaded fusion of depthwise separable convolution (DWC) and deformable convolution (DCNv4):
[0073] 1. Expand the effective receptive field: adaptively cover the feature regions of different-scale targets;
[0074] 2. Reduce computational redundancy: reduce parameters and computational volume through channel separation and sparse sampling;
[0075] 3. Enhance deformation modeling: accurately capture the features of irregular targets such as curved tracks and occluded foreign objects.
[0076] It is understandable that, as Figure 2 shown, the image features are input into the module, and the channels are evenly divided into four parts. The MSD combines depthwise separable convolution (DWC) and deformable convolution (DCNV4). According to the effective receptive field theory, increasing the convolution kernel size is more effective in expanding the feature receptive field of the model. Therefore, in the first two parts of the present invention, different-sized convolution kernels are used for the channel features to benefit the detection of multi-scale targets, and DWC can reduce the computational volume while maintaining the feature extraction ability to improve efficiency compared with ordinary convolution. Since there is some redundancy in the previous feature extraction, the present invention does not perform any operations on the features of the fourth part of the channels, thereby reducing the computational volume while maintaining the accuracy. For the feature channels of the third part, the present invention uses deformable convolution (DCN) for feature extraction because for complex images with deformation or non-rigid changes, it is difficult for a fixed convolution kernel to adapt to the changes in the image content, while deformable convolution (DCNV4) adaptively focuses on important feature regions in the image by dynamically adjusting the shape and position of the convolution kernel, which helps to extract complex local features. After splicing the four parts, the present invention performs a shuffle operation to enable cross-fusion of information between all channels, thereby improving the expression ability of the model. The expression is as follows:
[0077] Output=Shuffle(cat(DWC3(x0),DWC5(x1),DCN(x2),x3) (1)
[0078] Wherein, DWC3 and DWC5 represent DWC with convolution kernel sizes of 3×3 and 5×5 respectively.
[0079] It should be noted that in order to verify the effectiveness of multi-scale in the backbone network in improving the detection accuracy in the track environment, the present invention compares the results of the proposed MSD module under different combinations of convolutional kernels, including the ResNet structure, convolutional kernels of 3×3 and 5×5, convolutional kernels of 3×3 and 7×7, convolutional kernels of 5×5 and 7×7, and convolutional kernels of 7×7 and 9×9. The accuracies of these are 86.2%, 87.3%, 86.9%, 86.5% and 86.7% respectively to verify the effectiveness of the parallel multi-scale convolutional kernels. As Figure 3 shown, compared with the single-scale convolutional kernel model, the multi-scale convolutional kernel model achieves higher detection accuracy. Specifically, when using the 3×3 convolutional kernel and the 5×5 convolutional kernel in parallel structure, the model achieves the highest detection accuracy, reaching an mAP of 87.30%. In addition, the detection accuracy of the parallel structure in all experiments is higher than that of using the ResNet structure in the original backbone network. All convolutions use the more lightweight DWC, which proves that parallel convolution can better capture multi-scale features in feature extraction. At the same time, this also proves that there is feature redundancy in the feature extraction stage of the image, and part of the features can be extracted without affecting the final detection accuracy.
[0080] It can be understood that the global self-attention mechanism of the traditional Transformer has problems of high computational complexity and redundant weight interference, and it is difficult to achieve real-time inference on embedded devices. Aiming at the characteristics of sparse target spatial distribution (foreign objects usually only occupy a local area of the image) in the rail transit obstacle detection scenario, the present invention proposes a DAA mechanism (Deformable AgentAttention), and achieves a balance between efficiency and accuracy through the following innovations:
[0081] 1. Agent sparsification: Replace all pixels with a small number of learnable agents as query vectors to reduce the attention calculation amount;
[0082] 2. Dynamic deformation focusing: Predict the deformation offset of the agent to make it adaptively focus on the potential foreign object area;
[0083] 3. Hierarchical feature aggregation: Cross-scale fuse the agent attention weights to enhance the robustness of small target detection.
[0084] It can be understood that due to the high computational complexity of the standard Transformer and its strong dependence on computing resources, it is difficult to apply it to complex rail transit environments. To solve this problem, the present invention combines the advantages of AagentAttention and Deformable Attention to propose a Deformable Agent Attention (DAA). Figure 4Shows a schematic diagram of Deformable Agent Attention (DAA). In the traditional attention triple (Q, K, V), agent attention introduces a set of additional agent vectors A, defining a new four - tuple attention mechanism (Q, A, K, V). Among them, the agent vector A first acts as a proxy for the query vector Q, aggregates information from K and V, and then broadcasts the information back to Q. Since the number of agent vectors can be designed to be much smaller than the number of query vectors, agent attention can model global information at a lower computational cost. Agent attention combines the powerful global modeling ability of Softmax attention and the high - efficiency computational ability of linear attention, inherits their advantages, and enjoys low computational complexity and high model expressiveness. To enable this module to pay more attention to relevant regions and capture more useful feature information, deformable points in DeformableAttention are used in the DAA module. Through a deformable attention module containing an offset network, offsets are generated based on the query features as reference points, thereby creating deformable points to determine the positions of agent A and key K. This enables agent A and key K to fully notice more critical regions, enhancing the original self - attention module with higher flexibility and efficiency, and capturing more useful information features.
[0085] It can be understood that, as Figure 4 shown, given an input feature x ∈ R H×W×C , a uniform grid size of H G = H / r, W G = W / r is obtained through a downsampling factor r, and reference points are generated through the uniform grid in the image space for subsequent deformable convolution operations. Subsequently, the offset of the reference point is obtained. The input feature is linearly projected to generate a query q = xW q , which is put into a dedicated lightweight network θ offset (·) to generate an offset θ offset (q). After the deformable points, sampled features are obtained through spatial sampling, and linear projection and pooling operations are performed to obtain keys and agents, as shown in equations (2) and (3):
[0086]
[0087]
[0088] where and Agent Token represent the embeddings of the deformed keys and agents. The present invention sets the sampling function as bilinear interpolation, as shown in equation (4):
[0089]
[0090] Among them, the function g(a, b) represents the interpolation weight function in the x and y directions, and (r x , r y ) represents the indices of all positions where z ∈ R H×W×C . After obtaining q and the agent token, the present invention further infers that DAA can be expressed as:
[0091]
[0092] where a ∈ R n×c is the agent token defined by formula (2), σ(·) represents the Softmax function, which is used for agent aggregation and agent broadcasting, and v = xW v is the value obtained by linear projection of the input features. Specifically, the present invention first regards the agent A as a query, and aggregates information from all features through the σ(AK T )V operation to obtain V A . Then, the present invention takes A as the key k, performs a Softmax calculation with Q, and then performs a dot product operation with V A as the value v, broadcasts the global information in the agent features back to each feature, and obtains the final output O. In this way, the calculation of the similarity between Q and K is avoided, and the information exchange between each query-key is realized through the agent vector. By using a small number of agents A as the agents of q, collecting information from K and V and presenting it to Q, the present invention sets the number of A to be much smaller than the number of Q, so as to achieve global modeling ability with linear computational complexity. The final output of the DAA module is:
[0093]
[0094] where B1 and B2 are the agent biases Agent Bias, and the Agent Bias is combined to better utilize the position information and maintain feature diversity by using the Depthwise Convolution (DWC) module.
[0095] To prove the superiority of the proposed Deformable Agent Attention (DAA) of the present invention in obstacle intrusion detection in complex environments, the present invention selects the current leading attention mechanisms for comparative experiments. All attention mechanisms replace the AIFI module in the RT-DETR model for experiments on the TAD dataset, and the benchmark module is the AIFI module in RT-DETR. In addition, to reduce the test error, the FPS values are all averaged after testing 5 times. The results are shown in Table 1.
[0096] Table 1 Comparative experimental results of different attention mechanisms
[0097]
[0098] It can be seen from Table 1 that compared with the AIFI module in the basic model RT-DETR, the Deformable Agent Attention module proposed in the present invention achieves a significant improvement in mAP, F1 score and FPS while slightly increasing the model size and the number of parameters. Among them, mAP50 and mAP50:95 are increased by 1.2% and 1.3% respectively, and at the same time, the FPS is increased by 10.5 frames per second. Jingguang cascade Group Attention has an improvement of 12.8 frames per second in detection speed, but its accuracy is only slightly improved by 0.2% mAP. Both Deformable Attention and Agent Attention have improvements in accuracy and detection speed, but there is still room for improvement. The Transformer-based Blind-Spot Network (TBSN) module only improves the accuracy by 0.2% while increasing the computational cost (1.7 GFLOPs) and the number of parameters (1.87M), which is not friendly to the deployment on airborne devices. In addition, from the perspective of the F1 score, the attention module proposed in the present invention performs the best among these modules, and the F1 score reaches 86.45, indicating that the model finds a better balance between precision and recall.
[0099] To further illustrate the attention ability of the Deformable Agent Attention proposed in the present invention on important features, the present invention uses Grad-CAM to visualize the regions concerned by RT-DETR and the Deformable Agent Attention, so as to compare the feature heat maps before and after using this attention mechanism. Figure 5 (a-d) are the original images, Figure 5 (e-h) represent the feature heat maps of the original RT-DETR, Figure 5 (i-l) represent the feature heat maps with DAA added. The darker the color, the higher the degree of attention of the model to the features. It can be seen from the first column and the second column that the model of the present invention has stronger perception ability in complex track environments such as darker light and smaller targets, and more accurate detection ability for some floating objects and pedestrians, showing stronger global modeling ability. In addition, in the third column and the fourth column, the model proposed in the present invention reduces the excessive attention to interfering elements, and by paying attention to more necessary regions, the model of the present invention has stronger target recognition ability. The visualization results show that even in complex track environments, the model of the present invention has powerful global modeling ability and stable anti-interference ability.
[0100] It can be understood that, as Figure 6As shown, the present invention improves the algorithm based on the RT-DETR model to achieve efficient and accurate detection of track obstacles. The images in the dataset will be processed through the MSD_Block to achieve more efficient multi-scale feature extraction, and depthwise separable convolution and feature redundancy reduction operations are adopted to significantly reduce the computational amount. Subsequently, the low-resolution feature maps are fed into the DAA module to obtain richer and more accurate semantic information, and a small number of proxies are used as queries to avoid redundancy between attention weights, thereby accelerating the inference speed of the model. By replacing the AIFI module in RTDETR with the DAA module, the candidate Agents and k are focused on important regions, avoiding the high computational cost and high model complexity caused by the standard multi-head self-attention mechanism applying unified global attention to the image. The representation process is as follows:
[0101] F5 = DAA(P5) (7)
[0102] where DAA represents deformable proxy attention, and P5 is the higher-level feature layer P5 with richer semantic information from the Figure 6 backbone network.
[0103] It can be understood that to meet the requirements of the rail transit obstacle detection task for data diversity and scene coverage, the present invention constructs a railway dataset including a training set, a validation set, and a test set, and the dataset is sourced from infraDataset, Rail Dataset, and RailFOD23.
[0104] Table 2 Data in the TAD dataset
[0105]
[0106] To enable the model to perform real-time obstacle detection in front of the train, the present invention eliminates the images with non-train-driver front views in the three datasets to enhance the generalization ability of the model in the traffic scene in front of the train. In addition, some detections of people are not labeled in the original dataset, and some switches, crossings, etc. are not labeled. The present invention corrects the annotation information of the integrated dataset and names it the Train Assisted Dataset (TAD). Figure 7 shows the target height-width distribution in the TAD dataset ( Figure 7 (a)) and the target center point position distribution ( Figure 7(b)), from which it can be seen that most of the obstacles in the railway environment in this dataset are composed of small targets, which poses a certain challenge to the model to achieve high-precision real-time detection. The detection scenarios include targets near the tracks during the day and at night. The detection objects are divided into eight categories: level crossings, switches, signal lights, people, bird nests, plastic bags, floating objects, and balloons. The purpose of detecting level crossings, switches, and signal lights is to remind drivers to pay more attention to safety when driving on such sections. Obstacles such as bird nests, plastic bags, floating objects, and balloons pose a serious threat to the driver's vision and pantographs, transmission lines, etc. The description of each type of object in the dataset is shown in Table 2. There are a total of 11,957 images, which are randomly divided into a training set, a validation set, and a test set according to the ratios of 75%, 10%, and 15%.
[0107] It can be understood that to evaluate the performance of the proposed model under the same experimental conditions, the present invention uses the following indicators to compare the detection model: mean average precision (mAP), F1-score, giga floating-point operations per second (GFLOPs), frames per second (FPS), and the number of model parameters (Params). As the harmonic mean of precision and recall, the F1-score is a comprehensive evaluation index that balances these two indicators. A higher F1-score indicates a better balance between precision and recall, and its calculation method is shown in formula (8).
[0108]
[0109] Average precision (AP) is used to measure the average highest precision at different recall rate levels for each obstacle category. mAP reflects the average detection precision of the model on all obstacle categories. The higher the value, the better the multi-category recognition performance. The specific calculation formulas are shown in equations (9) and (10):
[0110]
[0111]
[0112] where S represents the total number of categories.
[0113] FPS is used to quantify the real-time inference speed. The higher the value, the faster the processing ability. The calculation formula is shown in equation (11):
[0114]
[0115] In the formula, t avg represents the inference time consumption of a single image.
[0116] In addition, GFLOPs and Params are used to measure the computational complexity and model scale respectively. The lower the values of these two indicators, the lower the algorithm complexity, and the more conducive it is to be deployed on resource-constrained embedded devices.
[0117] To evaluate the detection ability and improvement effect of the algorithm, the present invention conducted comparative experiments and ablation experiments on a railway dataset. All experiments were compared and analyzed under the same software and hardware environment. The experiments were carried out under the Pytorch 2.0.1 deep learning framework on Ubuntu 22.04. The configuration of the workstation used was an Intel Core i9-12900K central processing unit and a 24GB NVIDIA GeForce RTX 3090 graphics card. The input image size was scaled to 640×640 when input into the model. During training, the batch size was set to 16, the epoch was set to 120, a momentum optimizer was adopted, the learning rate was set to 0.0001, the momentum was set to 0.9, and the weight decay value was set to 0.0001.
[0118] In the comparative experiments of each algorithm, the present invention compared the improved RT-DETR of the present invention with advanced object detection models such as YOLOv11 and YOLOv12. The experimental results are shown in Table 3. The improved RT-DETR achieved a detection accuracy of 87.9% mAP at a detection speed of 90 FPS. Due to the large backbone network, RT-DETR-34 and RT-DETR-50 have slow detection speeds, high computational complexity, large number of parameters, and large model volume, so they are not suitable for deployment on airborne devices for real-time detection. Although YOLOv11n is suitable for deployment on embedded devices, its detection accuracy is too low and it is not suitable for scenarios with high detection accuracy requirements such as railway environments. Although YOLOv11m and YOLOv12m have excellent detection speeds, their detection accuracies are 0.8% and 0.6% lower than the improved model of the present invention respectively. In addition, their computational amounts are 17.5 GFLOPs and 16.9 GFLOPs higher than the improved model of the present invention respectively, the number of parameters are 5.05M and 5.12M higher than the improved model of the present invention respectively, and the model volumes are 9.4MB and 9.6MB larger than the improved model of the present invention respectively. This shows that the improved model of the present invention is more easily deployable on embedded devices such as airborne devices.
[0119] Table 3 Experimental results of comparison of each model
[0120]
[0121] Figure 8 Shows the detection results of the improved model in different complex railway environments. Figure 8 (a)-(l) show the comparison of detection results between RT-DETR and the improved model in different track environments, Figure 8 (a)-(d) are the original images, Figure 8 (e)-(h) are the detection results of RT-DETR,Figure 8 (i)-(l) are the improved model detection results. From these images, it can be clearly seen that the improved model of the present invention exceeds RT-DETR in terms of detection accuracy, missed detection, and false detection. Specifically, Figure 8 (e) shows that the confidence of the detection result of RT-DETR in the night environment is not high, and this situation is also reflected in Figure 8 (g) and Figure 8 (h). Figure 8 (f) shows that the crossing is missed and the roadside sign is wrongly detected as a person. In contrast, the improved model can perform correct detection, as shown in Figure 7 (i)-(l).
[0122] For the detection of foreign object intrusion in the track environment, the present invention can well balance the model detection accuracy and detection speed compared with the prior art, and the model size is significantly reduced compared with the basic model RT-DETR, and it can be well applied to embedded vehicle-mounted devices. In the dataset proposed by the present invention, the improved model achieves 87.9% mAP and a detection speed of 90 frames per second on the railway dataset.
[0123] In addition, as shown in Figure 9 , an embodiment of the present invention also discloses a track foreign object intrusion detection device, which includes:
[0124] An acquisition module 110, configured to acquire real-time railway track images through an in-vehicle camera;
[0125] An extraction module 120, configured to extract features from the railway track images to obtain image features;
[0126] A detection module 130, configured to input the image features into a trained foreign object detection model for foreign object intrusion detection to obtain a track foreign object intrusion detection result, where the foreign object detection model is improved based on the object detection model of RT-DETR, the foreign object detection model includes an MSD module, the MSD module is a module that combines depthwise separable convolution DWC and deformable convolution DCNv4, the foreign object detection model uses a deformable proxy attention mechanism DAA to replace the AIFI module in RT-DETR, and DAA optimizes the attention weight distribution through sparse queries.
[0127] The track foreign object intrusion detection device of the embodiment of the present invention is used to execute the track foreign object intrusion detection method in the above embodiment, and its specific processing process is the same as that of the track foreign object intrusion detection method in the above embodiment, and will not be elaborated here one by one.
[0128] In addition, as shown in Figure 10As shown in the figure, an embodiment of the present invention also discloses an electronic device, which includes: at least one processor 210; at least one memory 220 for storing at least one program; when the at least one program is executed by the at least one processor 210, the rail foreign object intrusion detection method in any of the previous embodiments is implemented.
[0129] In addition, an embodiment of the present invention also discloses a computer-readable storage medium, in which computer-executable instructions are stored, and the computer-executable instructions are used to execute the rail foreign object intrusion detection method in any of the previous embodiments.
[0130] The system architecture and application scenarios described in the embodiments of the present invention are for more clearly explaining the technical solutions of the embodiments of the present invention, and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. Those skilled in the art know that with the evolution of the system architecture and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present invention are equally applicable to similar technical problems.
[0131] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.
[0132] In the hardware implementation, the division between the functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component can have multiple functions, or a function or step can be executed by several physical components in cooperation. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, digital versatile disk (DVD), or other optical disk storage, magnetic cassette, tape, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, communication media typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0133] As used in this specification, the terms "component", "module", "system", etc. are used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable, an execution thread, a program, or a computer. By way of illustration, both an application running on a computing device and the computing device can be components. One or more components can reside within a process or execution thread, and a component can be located on one computer or distributed between two or more computers. In addition, these components can execute from various computer-readable media having various data structures stored thereon. A component can, for example, communicate by way of signals with other components through a local or remote process according to one or more data packets (e.g., data from two components interacting with each other from a local system, a distributed system, or a network, such as the Internet interacting with other systems by way of signals).
Claims
1. An orbital foreign object intrusion detection method, comprising: Obtaining a real-time railway track image through an on-vehicle camera; Performing feature extraction on the railway track image to obtain image features; Inputting the image features into a trained foreign object detection model for foreign object intrusion detection to obtain an orbital foreign object intrusion detection result, wherein the foreign object detection model is improved based on an RT-DETR object detection model, the foreign object detection model includes an MSD module, the MSD module is a module combining depthwise separable convolution DWC and deformable convolution DCNv4, the foreign object detection model uses a deformable proxy attention mechanism DAA module to replace the AIFI module in RT-DETR, and the DAA module optimizes the attention weight distribution through sparse queries.
2. The method according to claim 1, wherein In the MSD module, the depthwise separable convolution DWC is used to process the channel features of the first part and the second part, the channel features of the third part use the deformable convolution DCNv4 to adaptively focus on important feature regions in the railway track image by dynamically adjusting the shape and position of the convolution kernel, the channel features of the fourth part do not perform any operations, and after splicing the channel features of the first part, the second part, the third part, and the fourth part, a shuffle operation is performed to enable cross-fusion of information between all channels.
3. The method according to claim 2, wherein The cross-fusion of information between all channels by the shuffle operation is expressed as follows: Output = Shuffle(cat(DWC3(x0), DWC5(x1), DCN(x2), x3)) (1) Where DWC3 and DWC5 represent DWC with convolution kernel sizes of 3×3 and 5×5.
4. The method according to claim 1, wherein The DAA module includes a query vector Q, a proxy vector A, a key vector K, and a value vector V. The proxy vector A serves as a proxy for the query vector Q, aggregates information from the key vector K and the value vector V, and then broadcasts the information back to the query vector Q.
5. The method according to claim 4, characterized in that, The DAA module uses deformable points in the deformable attention mechanism. The DAA module includes a deformable attention module of an offset network. The DAA module generates an offset based on the query feature as a reference point, creates deformable points to determine the positions of the proxy vector A and the key vector K, so that the proxy vector A and the key vector K can notice the key areas of the railway track image.
6. The method according to claim 5, characterized in that, The DAA module generates an offset based on the query feature as a reference point and creates deformable points to determine the positions of the proxy vector A and the key vector K, including: The image features obtain a uniform grid size through a downsampling factor; Generating reference points through a uniform grid in the image space; Obtaining a reference point offset; Linearly projecting the input features to generate a query vector Q; Putting the query vector Q into a dedicated lightweight network to generate an offset; Obtaining a sampled feature through spatial sampling after the deformable points; Performing linear projection and pooling operations on the sampled feature to obtain the positions of the proxy vector A and the key vector K.
7. The method according to claim 1, wherein The input-output process of the DAA module includes: Given the input image feature x∈R H×W×C , obtain H G = H / r, W G = W / r as the uniform grid size of the image space, and generate reference points through the uniform grid in the image space for subsequent deformable convolution operations; subsequently, obtain the reference point offset, linearly project the input feature to generate the query q = xW q , and put it into a dedicated lightweight network θ offset (·) to generate the offset θ offset (q), obtain sampled features through spatial sampling after the deformable points, and perform linear projection and pooling operations to obtain the key and proxy, as shown in equations (2) and (3): Among them, and Agent Token represent the embeddings of the deformed key and the agent; set the sampling function to bilinear interpolation, as shown in Equation (4): Among them, the function g(a, b) represents the interpolation weight function in the x and y directions, and (r x , r y ) represents the indices of all positions where z ∈ R H×W×C . After obtaining q and the agent token, the further inference of the DAA module is expressed as: where \(A\in R\) n×c is the proxy token defined by formula (2), \(\sigma(\cdot)\) represents the Softmax function, which is used for proxy aggregation and proxy broadcasting, and \(v = xW\) v is the value obtained by linear projection of the input features; specifically, first, the proxy vector \(A\) is regarded as a query, and information is aggregated from all features through the \(\sigma(A K\) T ) \(V\) operation to obtain \(V\) A , then the proxy vector \(A\) is used as the key vector \(k\) and the query vector \(Q\) for Softmax calculation and then dot - product operation with \(V\) A as the value \(v\) to broadcast the global information in the proxy features back to each feature and obtain the final output \(O\); by using a small number of proxies \(A\) to act as agents for \(q\), collecting information from \(K\) and \(V\) and presenting it to \(Q\), setting the number of proxy vectors \(A\) to be much smaller than the number of query vectors \(Q\), the final output of the DAA module is: Among them, B1 and B2 are proxy biases. The proxy biases are combined to utilize the position information and the depthwise separable convolution DWC is used to maintain feature diversity.
8. An orbital foreign object intrusion detection device, characterized in that The device includes: An acquisition module, configured to acquire real-time railway track images through an on-vehicle camera; An extraction module, configured to extract features from the railway track images to obtain image features; A detection module, configured to input the image features into a trained foreign object detection model for foreign object intrusion detection to obtain a track foreign object intrusion detection result. Among them, the foreign object detection model is improved based on the RT-DETR object detection model. The foreign object detection model includes an MSD module, and the MSD module is a module that combines the depthwise separable convolution DWC and the deformable convolution DCNv4. The foreign object detection model uses a deformable proxy attention mechanism DAA module to replace the AIFI module in RT-DETR, and the DAA module optimizes the attention weight distribution through sparse queries.
9. An electronic device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the method for detecting track foreign object intrusion according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing computer-executable instructions for executing the method for detecting track foreign object intrusion according to any one of claims 1 to 7.