Foreign object intrusion detection method and system for railway perimeter protection

Through multimodal data spatiotemporal calibration and dynamic fusion technology, the problems of poor environmental adaptability and low efficiency of multi-sensor data fusion in traditional railway perimeter protection have been solved, high-precision foreign object detection in complex environments has been achieved, and the railway perimeter security protection capability has been improved.

CN120598946BActive Publication Date: 2025-10-10SICHUAN NATIONAL INNOVATION VISION UHD VIDEO TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511095121.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-10-10
Estimated Expiration
2045-08-06

AI Technical Summary

Technical Problem

In traditional railway perimeter protection systems, single sensors have poor environmental adaptability, multimodal data fusion is difficult, and there is a contradiction between computational efficiency and accuracy, making it difficult to achieve efficient foreign object detection in complex environments.

Method used

Through the spatiotemporal calibration of multimodal data, dynamic fusion weight allocation and deformable feature alignment technology, the spatiotemporal calibration parameters are used to unify multimodal data, and the deformable convolution offset field is generated in combination with the cross-modal semantic association matrix. The feature space mapping relationship is adaptively adjusted to achieve efficient fusion of multimodal data.

Benefits of technology

It can achieve foreign object detection with sub-meter positioning accuracy under complex weather and terrain conditions, reduce deployment and maintenance costs, improve the detection rate of tiny foreign objects, reduce the false alarm rate, and provide all-weather, high-precision intelligent railway perimeter protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120598946B_ABST
    Figure CN120598946B_ABST
Patent Text Reader

Abstract

The application provides a foreign matter intrusion detection method and system for railway perimeter protection, relates to the technical field of railway perimeter protection, and first unifies the space-time reference of a millimeter wave radar, a visible light video and a fiber vibration signal based on a space-time calibration parameter, eliminates the physical space deviation between heterogeneous sensor data, and ensures the accurate alignment of multi-modal features under complex weather and terrain conditions;Secondly, the real-time reliability of each mode is dynamically evaluated through a data quality index, a deformable convolution offset field is generated combined with a cross-modal semantic correlation matrix, the feature space mapping relationship is adaptively adjusted, and the interference of video blur, radar noise and vibration false alarm is effectively suppressed;Finally, based on a lightweight model, the dynamically fused joint features are efficiently inferred, sub-meter level positioning accuracy of foreign matter detection is realized on an edge computing device, existing railway monitoring facilities are upgraded, maintenance cost is reduced, and an intelligent solution is provided for railway perimeter safety protection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of railway perimeter protection, and in particular to a foreign body intrusion detection method and system for railway perimeter protection. Background Art

[0002] Railway perimeter protection systems require real-time detection of foreign objects within the track area, such as fallen rocks, metal fragments, and animals, to prevent train collisions. Traditional methods primarily rely on single sensors, such as visible light cameras or millimeter-wave radars. These methods suffer from bottlenecks such as poor environmental adaptability, difficulty integrating multimodal data, and a conflict between computational efficiency and accuracy. Visible light cameras experience significant image quality degradation in rain, fog, and low light conditions, resulting in high rates of missed detections. Millimeter-wave radars struggle to identify non-metallic foreign objects and have insufficient spatial resolution. Fiber optic vibration sensors are susceptible to mechanical noise, resulting in high false alarm rates. Furthermore, differences in the spatiotemporal coordinate systems of different sensors, such as between 3D radar point clouds and 2D video pixel coordinates, lead to significant data alignment errors. Traditional fixed-weight fusion methods are unable to dynamically adapt to complex environmental changes. Existing deep learning models are parameter-intensive, making them difficult to run in real time on edge devices. Lightweight models also struggle to process the high-dimensional features of heterogeneous, multimodal data. To address these challenges, a foreign object detection method is urgently needed that integrates multimodal data, dynamically optimizes fusion strategies, and adapts to complex railway scenarios.

[0003] Therefore, it is necessary to provide a foreign body intrusion detection method and system for railway perimeter protection to solve the above technical problems. Summary of the Invention

[0004] In order to solve the above technical problems, the present invention provides a foreign object intrusion detection method and system for railway perimeter protection, which achieves the beneficial effects of fusing multimodal data, dynamically optimizing fusion strategies and adapting to complex railway scenarios.

[0005] The present invention provides a foreign body intrusion detection method for railway perimeter protection, comprising:

[0006] S1: Synchronously collect multimodal raw data of the railway perimeter, calculate spatiotemporal calibration parameters, calibrate the multimodal raw data using the spatiotemporal calibration parameters, and extract multimodal features of the calibrated multimodal raw data;

[0007] S2: Calculate the data quality index and cross-modal semantic association matrix of the calibrated multimodal original data based on the multimodal features, and calculate the initial fusion weight matrix of the calibrated multimodal original data based on the data quality index;

[0008] S3: Generate deformable convolution offset field parameters based on the cross-modal semantic association matrix, adjust the spatial correspondence of multimodal features based on the deformable convolution offset field parameters, and obtain a multimodal spatial alignment feature map;

[0009] S4: Perform weighted fusion on the multimodal spatially aligned feature maps based on the initial fusion weight matrix and the deformable convolution offset field parameters to obtain a joint feature map;

[0010] S5: Input the joint feature map into the pre-trained foreign body intrusion detection model to obtain the foreign body detection result.

[0011] Preferably, in step S1, the multimodal raw data includes millimeter wave radar point cloud, visible light video and optical fiber vibration signal.

[0012] Preferably, in step S1, the spatiotemporal calibration parameters include lens distortion correction parameters and perspective transformation matrix of the visible light video, sensor coordinate conversion matrix, time synchronization deviation compensation value and vibration signal propagation delay.

[0013] Preferably, in step S1, the distortion of the video frame image of the visible light video is eliminated by lens distortion correction parameters, the perspective transformation of the video frame image of the visible light video is performed based on the perspective transformation matrix to generate a front view, and a unified railway physical coordinate system is established with the center line of the rail as the reference as the reference coordinate system for multimodal space alignment.

[0014] Preferably, in step S2, the data quality index includes a radar modal quality index, a video modal quality index and a vibration modal quality index, wherein the radar modal quality index is calculated by the signal-to-noise ratio and spatial coverage density of the millimeter-wave radar point cloud, the video modal quality index is calculated by the Laplace gradient modulus and inter-frame optical flow stability of the visible light video, and the vibration modal quality index is calculated by the energy proportion and frequency domain entropy value of the optical fiber vibration signal in a preset target frequency band.

[0015] Preferably, in step S2, generating the cross-modal semantic association matrix includes the following steps:

[0016] Apply the cosine similarity algorithm to the reflection intensity features of the millimeter-wave radar point cloud and the texture features of the video frame of the visible light video to generate the radar and video correlation matrix;

[0017] Mutual information calculation is applied to the spatial density distribution of millimeter-wave radar point cloud and the spectral entropy value of optical fiber vibration signal to generate radar and vibration correlation matrix;

[0018] The spatiotemporal feature map of the visible light video is extracted, and the optical fiber vibration signal is converted into a spatiotemporal energy distribution map through frequency domain and spatiotemporal conversion. The spatiotemporal feature map and spatiotemporal energy distribution map are processed through a multi-head attention mechanism to generate a spatiotemporal correlation matrix between the video and vibration.

[0019] Preferably, in step S3, generating the deformable convolution offset field parameters includes:

[0020] The cross-modal semantic association matrix is ​​input into a multi-level convolutional neural network, and the initial offset is output. The initial offset is iteratively optimized through the deformable convolution layer to obtain the deformable convolution offset field parameters.

[0021] Preferably, in step S4, obtaining the joint feature map includes the following steps:

[0022] The spatial confidence score of each spatial position in the multimodal spatial alignment feature map is calculated based on the deformable convolution offset field parameters. The fusion weights of visible light video, millimeter wave radar point cloud, and optical fiber vibration signal in the initial fusion weight matrix are adjusted based on the preset spatial confidence score threshold.

[0023] The fusion weights of visible light video, millimeter wave radar point cloud and optical fiber vibration signal in the adjusted initial fusion weight matrix are normalized to obtain the final fusion weight matrix;

[0024] Based on the final fusion weight matrix, the multimodal spatially aligned feature maps are weightedly fused to generate a joint feature map.

[0025] Preferably, the step of calculating the spatial confidence score includes:

[0026] Based on the deformable convolution offset field parameters, the two-dimensional offset of each spatial position in the multimodal spatial alignment feature map is obtained;

[0027] Based on the two-dimensional offset, the Euclidean norm of the offset is calculated for each spatial position;

[0028] Based on the Euclidean norm of the offset, the spatial confidence score is calculated through a preset nonlinear mapping function.

[0029] The present invention also provides a foreign body intrusion detection system for railway perimeter protection, which is applied to a foreign body intrusion detection method for railway perimeter protection, comprising:

[0030] The data acquisition and calibration module is used to synchronously collect the railway perimeter multimodal raw data, calculate the spatiotemporal calibration parameters, calibrate the multimodal raw data using the spatiotemporal calibration parameters, and extract the multimodal features of the calibrated multimodal raw data;

[0031] A cross-modal dynamic weight generation module is used to calculate the data quality index and cross-modal semantic association matrix of the calibrated multimodal raw data based on the multimodal features, and calculate the initial fusion weight matrix of the calibrated multimodal raw data based on the data quality index;

[0032] The semantically constrained feature alignment module is used to generate deformable convolution offset field parameters based on the cross-modal semantic association matrix, adjust the spatial correspondence of multimodal features based on the deformable convolution offset field parameters, and obtain a multimodal spatial alignment feature map;

[0033] Multimodal data fusion module, which is used to perform weighted fusion of multimodal spatial alignment feature maps based on the initial fusion weight matrix and deformable convolution offset field parameters to obtain a joint feature map;

[0034] The foreign body detection and recognition module is used to input the joint feature map into the pre-trained foreign body intrusion detection model to obtain the foreign body detection results.

[0035] Compared with related technologies, the foreign body intrusion detection method and system for railway perimeter protection provided by the present invention have the following beneficial effects:

[0036] The present invention solves the core problems of poor environmental adaptability and low efficiency of multi-sensor data fusion in traditional railway perimeter protection through multi-modal data spatiotemporal calibration, dynamic fusion weight allocation and deformable feature alignment technology. First, based on the spatiotemporal calibration parameters, the spatiotemporal benchmarks of millimeter-wave radar, visible light video and optical fiber vibration signals are unified to eliminate the physical space deviation between heterogeneous sensor data and ensure the precise alignment of multi-modal features under complex weather and terrain conditions. Second, the real-time reliability of each modality is dynamically evaluated through data quality indicators, and a deformable convolution offset field is generated by combining the cross-modal semantic association matrix. Adaptively adjust the feature space mapping relationship to effectively suppress the interference of video blur, radar noise and vibration false alarms; finally, based on the lightweight model, efficient inference is performed on the dynamically fused joint features to achieve foreign object detection with sub-meter positioning accuracy on the edge computing device, while being compatible with the upgrade of existing railway monitoring facilities to reduce deployment and maintenance costs. The solution of the present invention shows strong robustness in complex scenarios such as rain and fog, night, and curves, which not only improves the detection rate of tiny foreign objects, but also greatly reduces the false alarms caused by environmental noise, providing an all-weather, high-precision, low-latency intelligent solution for railway perimeter security protection. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 This is a flow chart of a foreign body intrusion detection method for railway perimeter protection according to the present invention;

[0038] Figure 2 This is a module structure diagram of a foreign object intrusion detection system for railway perimeter protection according to the present invention. DETAILED DESCRIPTION

[0039] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It will be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all of the structures. Furthermore, the embodiments of the present invention and the features of the embodiments may be combined with each other unless there is a conflict.

[0040] It should also be noted that, for ease of description, only portions relevant to the present invention are shown in the accompanying drawings, rather than all of the contents. Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the various operations (or steps) as being processed sequentially, many of the operations can be performed in parallel, concurrently, or simultaneously. In addition, the order of the various operations can be rearranged. The process can be terminated when its operations are completed, but may also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0041] Example 1

[0042] A foreign body intrusion detection method for railway perimeter protection, in the specific implementation process, such as Figure 1 Shown, including:

[0043] Step S1: synchronously collect multimodal raw data of the railway perimeter, calculate spatiotemporal calibration parameters, calibrate the multimodal raw data using the spatiotemporal calibration parameters, and extract multimodal features of the calibrated multimodal raw data.

[0044] Specifically, in step S1, the multimodal raw data includes millimeter wave radar point cloud, visible light video and optical fiber vibration signal.

[0045] Specifically, in step S1 , the spatiotemporal calibration parameters include lens distortion correction parameters and perspective transformation matrix of the visible light video, sensor coordinate conversion matrix, time synchronization deviation compensation value and vibration signal propagation delay.

[0046] Specifically, in step S1, the distortion of the video frame image of the visible light video is eliminated by using the lens distortion correction parameters, the perspective transformation of the video frame image of the visible light video is performed based on the perspective transformation matrix to generate a front view, and a unified railway physical coordinate system is established with the center line of the rail as the reference coordinate system for multimodal space alignment.

[0047] During the implementation process, a heterogeneous sensor network is used to synchronously collect multimodal raw data from the railway perimeter, including millimeter-wave radar point clouds, visible light video, and fiber-optic vibration signals. For example, the millimeter-wave radar uses frequency-modulated continuous wave (FMCW) technology, operating in the 76-81 GHz frequency band, with a 120° horizontal field of view, a 30° vertical field of view, and a scanning frequency of 20 Hz. It outputs 3D point cloud data containing distance, azimuth, pitch angle, and reflection intensity. Visible light video is captured using a global shutter industrial camera with a resolution of 1920×1080@30fps, and a wide dynamic range sensor is used to ensure image quality in backlit scenes or tunnel entrances. The fiber-optic vibration signal is based on phase-sensitive optical time-domain reflectometry technology. The sensing fiber laid along the track monitors vibration events in the 0.1-2 kHz frequency band at a 10 kHz sampling rate. Sub-microsecond time synchronization between sensors is achieved using the IEEE 1588 PTP protocol. The master clock source is a GPS-disciplined atomic clock. Simultaneously, a hardware trigger signal is generated by an FPGA to ensure strict alignment of the millimeter-wave radar scanning period, video frame exposure, and fiber-optic vibration sampling window. The calculation of spatiotemporal calibration parameters includes five core components. First, the lens distortion correction parameters for visible light video are obtained through the checkerboard calibration method to obtain the intrinsic parameter matrix and distortion coefficients, and the undistort function of OpenCV is used to eliminate radial and tangential distortion. Second, the perspective transformation matrix is ​​calculated based on the camera installation parameters, including pitch angle and height, to map the oblique-view video frames to the plane view of the railway track. Third, the sensor coordinate transformation matrix is ​​solved by corner point matching on the calibration plate and least-squares optimization. The radar point cloud is transformed into the railway physical coordinate system using the rigid transformation matrix, with a calibration residual of ≤0.1m. The video pixel coordinates are mapped to the same coordinate system using the perspective transformation matrix and a preset scale factor. Fourth, the time synchronization deviation compensation value is calculated based on the master-slave clock offset using the Sync and Delay_Req messages of the PTP protocol. The FPGA hardware trigger mode is switched when the network delay jitter is ≥50ms. Fifth, the vibration signal propagation delay is calculated based on the distance of the vibration event location from the sensor node based on the speed of sound wave propagation in the optical fiber. The spatiotemporal consistency calibration is achieved by superimposing the propagation delay. The calibrated multimodal data is further subjected to feature extraction. The millimeter-wave radar point cloud is transformed into coordinates to extract reflection intensity features and spatial density features. The visible light video is used to extract texture features using the HOG descriptor and calculate the inter-frame motion consistency score using the Farneback optical flow algorithm. The optical fiber vibration signal is converted to the frequency domain using STFT and the energy proportion and spectral entropy of the preset target frequency band are extracted.Finally, a unified railway physical coordinate system is constructed with the centerline of the rail as the X-axis (extension direction), the direction perpendicular to the rail as the Y-axis, and the ground height Z=0. For example, the track gauge is measured by a total station to fit the centerline equation y=0±δ (δ≤5mm), and the mileage post K0+000 is used as the origin. Benchmark points are set every 50m along the rail to realize multi-sensor joint calibration, ensuring that the radar point cloud position error is ≤0.1m, the video front view pixel resolution is 0.05m / pixel, and the vibration signal positioning error is ≤0.3m after calibration, providing a high-precision spatiotemporal alignment basis for subsequent multimodal fusion and detection.

[0048] Step S2: Calculate the data quality index and cross-modal semantic association matrix of the calibrated multimodal original data based on the multimodal features, and calculate the initial fusion weight matrix of the calibrated multimodal original data based on the data quality index.

[0049] Specifically, in step S2, the data quality indicators include radar modal quality indicators, video modal quality indicators and vibration modal quality indicators, among which the radar modal quality indicator is calculated by the signal-to-noise ratio and spatial coverage density of the millimeter-wave radar point cloud, the video modal quality indicator is calculated by the Laplace gradient modulus and inter-frame optical flow stability of the visible light video, and the vibration modal quality indicator is calculated by the energy proportion and frequency domain entropy value of the optical fiber vibration signal in the preset target frequency band.

[0050] Specifically, in step S2, the generation of the cross-modal semantic association matrix includes the following steps:

[0051] Apply the cosine similarity algorithm to the reflection intensity features of the millimeter-wave radar point cloud and the texture features of the video frame of the visible light video to generate the radar and video correlation matrix;

[0052] Mutual information calculation is applied to the spatial density distribution of millimeter-wave radar point cloud and the spectral entropy value of optical fiber vibration signal to generate radar and vibration correlation matrix;

[0053] The spatiotemporal feature map of the visible light video is extracted, and the optical fiber vibration signal is converted into a spatiotemporal energy distribution map through frequency domain and spatiotemporal conversion. The spatiotemporal feature map and spatiotemporal energy distribution map are processed through a multi-head attention mechanism to generate a spatiotemporal correlation matrix between the video and vibration.

[0054] During the specific implementation process, first, the data quality indicators of millimeter-wave radar point cloud, visible light video and optical fiber vibration signal are calculated respectively based on the calibrated multimodal features. Among them, the radar modal quality index is calculated jointly by the signal-to-noise ratio and spatial coverage density of the millimeter-wave radar point cloud. The signal-to-noise ratio is defined as the ratio of the mean value of the effective target reflection intensity to the standard deviation of the noise. The spatial coverage density is the ratio of the number of effective point clouds in the preset local grid to the theoretical maximum. The video modal quality index is evaluated by the Laplace gradient modulus and inter-frame optical flow stability of the visible light video. The Laplace gradient modulus is calculated by the sum of the absolute values ​​of the second-order derivatives of the image grayscale, and the optical flow stability is quantified by the standard deviation of the optical flow vector between two consecutive frames. The optical fiber vibration modal quality index is calculated by extracting the short-time Fourier transform energy proportion and frequency domain entropy value of the optical fiber vibration signal in the preset target frequency band. Subsequently, a cross-modal semantic association matrix is ​​generated based on multimodal features. The radar-video association matrix calculates the matching degree of reflection intensity features and texture features through the cosine similarity algorithm. That is, the reflection intensity vector of the millimeter-wave radar point cloud at each spatial position is normalized with the video HOG feature vector of the visible light video and then the inner product is calculated to generate the radar-video association matrix. The radar-vibration association matrix calculates the statistical correlation between the spatial density distribution of the point cloud and the spectral entropy of the optical fiber vibration signal through the mutual information. Specifically, the histogram statistics joint probability distribution and marginal distribution are used to calculate the radar-vibration association matrix. The video-vibration spatiotemporal association matrix is ​​finally output by fusing the spatiotemporal feature map of the visible light video and the spatiotemporal energy distribution of the optical fiber vibration signal through the multi-head attention mechanism. After the cross-modal association matrix is ​​constructed, the initial fusion weight matrix is ​​dynamically allocated based on data quality indicators: the initial weights of radar, video, and vibration modes are determined by the normalized quality indicators respectively. This process quantitatively evaluates the real-time reliability of each modality and combines the complementarity of cross-modal semantic associations to provide a robust weight allocation basis for subsequent multimodal alignment and fusion, effectively suppressing low-quality modal interference and improving detection robustness in complex scenarios.

[0055] Step S3: Based on the cross-modal semantic association matrix, generate deformable convolution offset field parameters, adjust the spatial correspondence of multimodal features based on the deformable convolution offset field parameters, and obtain a multimodal spatial alignment feature map.

[0056] Specifically, in step S3, the generation of deformable convolution offset field parameters includes:

[0057] The cross-modal semantic association matrix is ​​input into a multi-level convolutional neural network, and the initial offset is output. The initial offset is iteratively optimized through the deformable convolution layer to obtain the deformable convolution offset field parameters.

[0058] In the specific implementation process, the cross-modal semantic association matrix is ​​first input as input to a multi-level convolutional neural network for initial offset prediction. For example, the multi-level convolutional neural network consists of three levels of convolutional modules. The first level is a 3×3 convolutional layer with 64 output channels, a stride of 1, and an activation function called ReLU, which is used to extract local cross-modal association features. The second level is a 1×1 convolutional layer with 128 output channels, a stride of 1, and an activation function called LeakyReLU with a negative slope of 0.2, which is used for feature channel dimensionality increase and nonlinear enhancement. The third level is a 3×3 transposed convolutional layer with 2 output channels, corresponding to horizontal and vertical offsets, a stride of 1, no activation function, and directly outputs the initial offset field. After the initial offset field is generated, it is iteratively optimized through the deformable convolution layer. The initial offset field is spliced ​​with the cross-modal semantic association matrix and then input into the deformable convolution layer. For example, the deformable convolution layer uses a deformable convolution kernel with a kernel size of 3×3, an expansion rate of 2, and a filling mode of reflect. The input features are adaptively sampled through the deformable convolution operation, and the residual between the output feature and the target alignment position is calculated. The residual term is optimized through the mean square error loss function. The number of iterations is set to 5, and the offset is updated each iteration to finally obtain the optimized deformable convolution offset field parameters. The geometric meaning is the distance that the multimodal features of each spatial position need to be adjusted in the horizontal and vertical directions to achieve precise alignment. Subsequently, the spatial correspondence of multimodal features is dynamically adjusted based on the deformable convolution offset field parameters: For the millimeter-wave radar point cloud feature map, bilinear interpolation is used to resample it to the target position along the direction indicated by the offset field to generate an aligned radar feature map. For the visible light video feature map, the deformable RoI Align method is used to dynamically adjust the position of the region of interest based on the optimized deformable convolution offset field parameters to extract the aligned video feature map. For the optical fiber vibration signal feature map, the vibration features are mapped to a unified coordinate system using a spatiotemporal transformation matrix combined with the vibration propagation delay compensation value. The optimized deformable convolution offset field parameters are then superimposed and fine-tuned to obtain the vibration alignment feature map. Finally, by pixel-level alignment of the spatially aligned multimodal feature maps in a unified railway physical coordinate system, sensor perspective differences and installation deviations are eliminated, providing a high-precision spatial alignment foundation for subsequent dynamic fusion and foreign object detection.

[0059] Step S4: Perform weighted fusion on the multimodal spatial alignment feature maps based on the initial fusion weight matrix and the deformable convolution offset field parameters to obtain a joint feature map.

[0060] Specifically, in step S4, obtaining the joint feature map includes the following steps:

[0061] The spatial confidence score of each spatial position in the multimodal spatial alignment feature map is calculated based on the deformable convolution offset field parameters. The fusion weights of visible light video, millimeter wave radar point cloud, and optical fiber vibration signal in the initial fusion weight matrix are adjusted based on the preset spatial confidence score threshold.

[0062] The fusion weights of visible light video, millimeter wave radar point cloud and optical fiber vibration signal in the adjusted initial fusion weight matrix are normalized to obtain the final fusion weight matrix;

[0063] Based on the final fusion weight matrix, the multimodal spatially aligned feature maps are weightedly fused to generate a joint feature map.

[0064] Specifically, the calculation steps of the spatial confidence score include:

[0065] Based on the deformable convolution offset field parameters, the two-dimensional offset of each spatial position in the multimodal spatial alignment feature map is obtained;

[0066] Based on the two-dimensional offset, the Euclidean norm of the offset is calculated for each spatial position;

[0067] Based on the Euclidean norm of the offset, the spatial confidence score is calculated through a preset nonlinear mapping function.

[0068] In the implementation process, first, based on the deformable convolution offset field parameters, the spatial confidence score of each spatial position in the multi-modal spatial alignment feature map is calculated, which is realized by the following steps: the horizontal and vertical direction offsets are extracted from the deformable convolution offset field parameters, the Euclidean norm is calculated, and the offset Euclidean norm is converted into a spatial confidence score through a preset nonlinear mapping function, the score range is (0, 1], wherein when the geometric length of the offset, i.e. the offset Euclidean norm, is zero, the spatial confidence score is 1, indicating perfect alignment, and the larger the offset Euclidean norm is, the lower the spatial confidence score is, i.e. the lower the alignment reliability is. Subsequently, based on the preset spatial confidence score threshold, for example, the preset spatial confidence score threshold is 0.6, the initial fusion weight matrix is dynamically adjusted: when the spatial confidence score is less than 0.6, the visible light video fusion weight is attenuated to 30% of the original value, the millimeter wave radar point cloud fusion weight is enhanced to 150% of the original value, and the fiber vibration signal fusion weight is further constrained according to the radar point cloud density, specifically, if the radar point cloud density is less than 20 points / m2, it is forced to be zero, otherwise it is scaled by a preset ratio; when the spatial confidence score is greater than or equal to 0.6, the initial fusion weight is retained. The adjusted weight matrix needs to be normalized to obtain the final fusion weight matrix, ensuring that the sum of the fusion weights at any position is 1, avoiding feature amplitude distortion. Finally, the multi-modal spatial alignment feature map is weighted and fused based on the normalized weight matrix, i.e. the final fusion weight matrix, to generate a joint feature map. This process maximizes the complementary advantages of multi-modal through a spatial confidence driven weight adjustment mechanism in complex scenes, meeting the real-time protection needs of railway perimeter.

[0069] Step S5: inputting the joint feature map into a pre-trained foreign object intrusion detection model to obtain a foreign object detection result.

[0070] In the implementation process, the joint feature map is efficiently inferred by the pre-trained lightweight foreign object intrusion detection model, and the foreign object category, position and threat level are output. For example, the pre-trained lightweight foreign object intrusion detection model adopts a hybrid architecture of Transformer and depth separable convolution, the backbone network is based on an improved YOLOv7-tiny, the calculation amount is reduced by replacing the standard convolution, and 4 head Transformer modules are introduced to enhance the global feature modeling capability, and the input size is 640x640x3. During inference, the joint feature map is fused with the detection head to output a prediction tensor, and the repeated frame is filtered out by non-maximum suppression, and the result with a confidence of greater than or equal to 0.6 is retained, the threat score is integrated with the foreign object category weight, target size and motion speed, and finally the JSON format detection information is output, and the foreign object intrusion detection is completed.

[0071] The working principle of the foreign object intrusion detection method for railway perimeter protection provided by the application is as follows:

[0072] First, millimeter-wave radar point clouds, visible light video, and fiber-optic vibration signals are synchronously collected. Using spatiotemporal calibration parameters, including lens distortion correction parameters, perspective transformation matrices, sensor coordinate conversion matrices, time synchronization bias compensation values, and vibration signal propagation delays, the multimodal data are mapped to a unified railway physical coordinate system, eliminating the spatiotemporal bias introduced by sensor heterogeneity. The reliability of each modality is then evaluated based on data quality indicators. The complementarity between modalities is quantified using a cross-modal semantic association matrix to generate an initial fusion weight matrix. Furthermore, a deformable convolutional network is used to analyze cross-modal correlation features. Deformable convolution offset field parameters are iteratively optimized to dynamically adjust the spatial alignment of multimodal features and obtain a multimodal spatially aligned feature map. The fusion weights are adaptively adjusted using spatial confidence scores to suppress interference from low-quality modalities. Finally, the normalized weighted fused joint feature map is input into a lightweight detection model. A multi-scale feature pyramid and threat scoring mechanism are used to output the foreign object category, sub-meter positioning coordinates, and risk level, completing foreign object intrusion detection for railway perimeter protection.

[0073] Example 2

[0074] A foreign body intrusion detection system for railway perimeter protection, in the specific implementation process, such as Figure 2 Shown, including:

[0075] The data acquisition and calibration module 100 is used to synchronously acquire multimodal raw data of the railway perimeter, calculate spatiotemporal calibration parameters, calibrate the multimodal raw data using the spatiotemporal calibration parameters, and extract multimodal features of the calibrated multimodal raw data;

[0076] A cross-modal dynamic weight generation module 200 is configured to calculate a data quality index and a cross-modal semantic association matrix of the calibrated multimodal original data based on the multimodal features, and to calculate an initial fusion weight matrix of the calibrated multimodal original data based on the data quality index;

[0077] A semantically constrained feature alignment module 300 is configured to generate deformable convolution offset field parameters based on a cross-modal semantic association matrix, adjust the spatial correspondence of multimodal features based on the deformable convolution offset field parameters, and obtain a multimodal spatial alignment feature map;

[0078] A multimodal data fusion module 400 is configured to perform weighted fusion on the multimodal spatially aligned feature maps based on an initial fusion weight matrix and deformable convolution offset field parameters to obtain a joint feature map;

[0079] The foreign body detection and identification module 500 is used to input the joint feature map into the pre-trained foreign body intrusion detection model to obtain the foreign body detection result.

[0080] The working principle of the foreign matter intrusion detection system for railway perimeter protection provided by the application is as follows:

[0081] The data acquisition and calibration module 100 synchronously acquires the millimeter wave radar point cloud, visible light video and fiber vibration signal, eliminates the space-time deviation of the multi-modal data through the space-time calibration parameter, and maps to a unified coordinate system with the track center line as the reference; the cross-modal dynamic weight generation module 200 evaluates the reliability of each mode based on the data quality index, and generates an initial fusion weight matrix in combination with a cross-modal semantic correlation matrix; the semantic constraint feature alignment module 300 iteratively optimizes the offset field parameter through a multi-level convolutional neural network and a deformable convolution, dynamically corrects the space mapping error of the multi-modal features, and generates a spatially aligned multi-modal spatial alignment feature map; the multi-modal data fusion module 400 dynamically adjusts the initial fusion weight matrix based on the spatial confidence score to obtain a final fusion weight matrix, weights and fuses the multi-modal spatial alignment feature map based on the final fusion weight matrix, and obtains a joint feature map; finally, the foreign matter detection and recognition module 500 inputs the joint feature map into a pre-trained lightweight foreign matter intrusion detection model, outputs the foreign matter category, position and risk level through a multi-scale feature pyramid and a threat scoring mechanism, and outputs the complete foreign matter intrusion detection and result.

[0082] The present application is described with reference to flowcharts and / or block diagrams according to the methods, devices (systems) and computer program products of the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The device that implements the functions specified in one flow or multiple flows and / or blocks. Figure 1 The device that implements the functions specified in one flow or multiple flows and / or blocks.

[0083] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program. The program can be stored in a computer-readable storage medium, including a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, magnetic disk storage, or magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.

Claims

1. A foreign body intrusion detection method for railway perimeter protection, characterized in that: The foreign body intrusion detection method comprises the following steps: S1: Synchronously collect multimodal raw data of the railway perimeter, calculate spatiotemporal calibration parameters, calibrate the multimodal raw data using the spatiotemporal calibration parameters, and extract multimodal features of the calibrated multimodal raw data; S2: Calculate the data quality index and cross-modal semantic association matrix of the calibrated multimodal original data based on the multimodal features, and calculate the initial fusion weight matrix of the calibrated multimodal original data based on the data quality index; S3: Generate deformable convolution offset field parameters based on the cross-modal semantic association matrix, adjust the spatial correspondence of multimodal features based on the deformable convolution offset field parameters, and obtain a multimodal spatial alignment feature map; S4: Perform weighted fusion on the multimodal spatially aligned feature maps based on the initial fusion weight matrix and the deformable convolution offset field parameters to obtain a joint feature map; S5: Input the joint feature map into the pre-trained foreign body intrusion detection model to obtain the foreign body detection result.

2. A foreign body intrusion detection method for railway perimeter protection according to claim 1, characterized in that: In step S1, the multimodal raw data includes millimeter wave radar point cloud, visible light video and optical fiber vibration signal.

3. The method for detecting foreign body intrusion for railway perimeter protection according to claim 2, characterized in that: In step S1 , the spatiotemporal calibration parameters include lens distortion correction parameters and perspective transformation matrix of the visible light video, sensor coordinate conversion matrix, time synchronization deviation compensation value and vibration signal propagation delay.

4. A foreign body intrusion detection method for railway perimeter protection according to claim 3, characterized in that: In step S1, the distortion of the video frame image of the visible light video is eliminated by using the lens distortion correction parameters, the perspective transformation of the video frame image of the visible light video is performed based on the perspective transformation matrix to generate a front view, and a unified railway physical coordinate system is established with the center line of the rail as the reference coordinate system for multimodal space alignment.

5. The foreign body intrusion detection method for railway perimeter protection according to claim 4 is characterized in that: In step S2, the data quality indicators include radar modal quality indicators, video modal quality indicators and vibration modal quality indicators, among which the radar modal quality indicator is calculated by the signal-to-noise ratio and spatial coverage density of the millimeter-wave radar point cloud, the video modal quality indicator is calculated by the Laplace gradient modulus and inter-frame optical flow stability of the visible light video, and the vibration modal quality indicator is calculated by the energy proportion and frequency domain entropy value of the optical fiber vibration signal in the preset target frequency band.

6. A foreign body intrusion detection method for railway perimeter protection according to claim 5, characterized in that: In step S2, the generation of the cross-modal semantic association matrix includes the following steps: Apply the cosine similarity algorithm to the reflection intensity features of the millimeter-wave radar point cloud and the texture features of the video frame of the visible light video to generate the radar and video correlation matrix; Mutual information calculation is applied to the spatial density distribution of millimeter-wave radar point cloud and the spectral entropy value of optical fiber vibration signal to generate radar and vibration correlation matrix; The spatiotemporal feature map of the visible light video is extracted, and the optical fiber vibration signal is converted into a spatiotemporal energy distribution map through frequency domain and spatiotemporal conversion. The spatiotemporal feature map and spatiotemporal energy distribution map are processed through a multi-head attention mechanism to generate a spatiotemporal correlation matrix between the video and vibration.

7. A foreign body intrusion detection method for railway perimeter protection according to claim 6, characterized in that: In step S3, the generation of deformable convolution offset field parameters includes: The cross-modal semantic association matrix is ​​input into a multi-level convolutional neural network, and the initial offset is output. The initial offset is iteratively optimized through the deformable convolution layer to obtain the deformable convolution offset field parameters.

8. The method for detecting foreign body intrusion for railway perimeter protection according to claim 7, characterized in that: In step S4, obtaining the joint feature map includes the following steps: The spatial confidence score of each spatial position in the multimodal spatial alignment feature map is calculated based on the deformable convolution offset field parameters. The fusion weights of visible light video, millimeter wave radar point cloud, and optical fiber vibration signal in the initial fusion weight matrix are adjusted based on the preset spatial confidence score threshold. The fusion weights of visible light video, millimeter wave radar point cloud and optical fiber vibration signal in the adjusted initial fusion weight matrix are normalized to obtain the final fusion weight matrix; Based on the final fusion weight matrix, the multimodal spatially aligned feature maps are weightedly fused to generate a joint feature map.

9. The method for detecting foreign body intrusion for railway perimeter protection according to claim 8, characterized in that: The steps for calculating the spatial confidence score include: Based on the deformable convolution offset field parameters, the two-dimensional offset of each spatial position in the multimodal spatial alignment feature map is obtained; Based on the two-dimensional offset, the Euclidean norm of the offset is calculated for each spatial position; Based on the Euclidean norm of the offset, the spatial confidence score is calculated through a preset nonlinear mapping function.

10. A foreign body intrusion detection system for railway perimeter protection, characterized in that: A foreign body intrusion detection method for railway perimeter protection according to any one of claims 1 to 9, wherein the foreign body intrusion detection system comprises: The data acquisition and calibration module is used to synchronously collect the railway perimeter multimodal raw data, calculate the spatiotemporal calibration parameters, calibrate the multimodal raw data using the spatiotemporal calibration parameters, and extract the multimodal features of the calibrated multimodal raw data; A cross-modal dynamic weight generation module is used to calculate the data quality index and cross-modal semantic association matrix of the calibrated multimodal raw data based on the multimodal features, and calculate the initial fusion weight matrix of the calibrated multimodal raw data based on the data quality index; The semantically constrained feature alignment module is used to generate deformable convolution offset field parameters based on the cross-modal semantic association matrix, adjust the spatial correspondence of multimodal features based on the deformable convolution offset field parameters, and obtain a multimodal spatial alignment feature map; Multimodal data fusion module, which is used to perform weighted fusion of multimodal spatial alignment feature maps based on the initial fusion weight matrix and deformable convolution offset field parameters to obtain a joint feature map; The foreign body detection and recognition module is used to input the joint feature map into the pre-trained foreign body intrusion detection model to obtain the foreign body detection results.

Citation Information

Patent Citations

  • Faster R-CNN-based railway abnormal intrusion behavior detection method

    CN113534276A

  • Railway perimeter intrusion early warning method

    CN115083088A

Cited By

  • Railway perimeter intrusion detection system and method based on multi-modal fusion

    CN121938094A

  • Railway video quality self-diagnosis system based on multi-domain self-calibration

    CN122510797A