Leighting fusion multi-sensor collaborative sensing system along railway and installation and calibration method

By combining a multimodal sensor array with a feature-level fusion network, the problems of perception fusion, spatiotemporal synchronization, and calibration and maintenance in railway monitoring systems have been solved, enabling all-weather, high-precision, and low-latency railway safety monitoring, thereby improving detection accuracy and operational efficiency.

CN121995395APending Publication Date: 2026-05-08CENT SOUTH UNIV

Patent Information

Application Number
CN202610090204.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-23
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing railway monitoring systems suffer from problems such as shallow perception fusion, insufficient spatiotemporal synchronization accuracy, poor environmental adaptability, and difficulties in calibration and maintenance in terms of multi-sensor collaborative perception, making it difficult to meet the all-weather, high-precision, and low-latency safety monitoring requirements of high-speed railways.

Method used

It employs a multimodal sensor array (LiDAR, millimeter-wave radar, visible light camera, infrared thermal imager) and a feature-level fusion network (RAF-Net) based on an attention mechanism, combined with a spatiotemporal synchronization mechanism of hardware triggering and PTP precision clock protocol, and carries edge computing nodes for real-time data processing. It also achieves online calibration and adjustment of the sensors through a three-level calibration and dynamic compensation mechanism.

Benefits of technology

It achieves high-precision, low-latency target detection and recognition in complex and ever-changing railway scenarios, improving detection accuracy and environmental adaptability, reducing operation and maintenance costs, and meeting the safety monitoring needs of high-speed railways.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention discloses a thunder-vision fusion multi-sensor collaborative sensing system along a railway and an installation and calibration method, and belongs to the technical field of intelligent traffic and environment sensing. Aiming at the problems of low target detection precision, poor multi-source data fusion timeliness, dependence on manpower on sensor installation and calibration and the like in a complex scene along a railway, the system provides an innovative architecture of'layered sensing-dynamic fusion-autonomous calibration '. The multi-modal sensor array of a laser radar, a millimeter-wave radar, a visible light camera and a thermal infrared imager is deployed; autonomous optimization of sensor pose parameters is realized by using track geometric constraint and a deep learning model; an edge computing node and a cloud collaboration platform are integrated, and real-time target tracking, intrusion early warning and equipment state diagnosis are supported. Experiments show that the target detection accuracy rate of the system is greater than or equal to 98.5%, the false alarm rate is less than or equal to 0.3% and the self-calibration error of sensor installation parameters is less than 0.05 degree in the scenes of rain and fog, night, high-speed movement and the like, which are improved by more than 40% compared with the traditional scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent transportation and environmental perception technology, specifically to a multi-sensor collaborative perception system for railway lines that integrates radar and visual sensors, and its installation and calibration method. Background Technology

[0002] With the rapid expansion of my country's high-speed and conventional railway networks, the "eight vertical and eight horizontal" high-speed railway backbone has basically taken shape, and the total operating mileage of railways nationwide has exceeded 160,000 kilometers, including more than 40,000 kilometers of high-speed railways and approximately 120,000 kilometers of conventional railways. The continuous increase in railway operating speeds (such as the Beijing-Shanghai high-speed railway operating at 350 km / h and some intercity railways reaching 400 km / h) and the sustained growth in passenger and freight density have significantly increased the importance and urgency of safety monitoring along railway lines. Common risks along railway lines include foreign object intrusion (such as falling rocks, fallen trees, and fallen engineering components), illegal intrusion (such as people climbing over guardrails to enter the clearance gauge), equipment malfunctions (such as foreign objects hanging from the overhead contact line or damage to signal facilities), and secondary risks caused by natural disasters (such as landslides and floods destroying the roadbed). Once these risks occur, they can cause temporary train stops and disruptions to transportation order, or even major accidents such as train derailments and collisions, seriously threatening the safety of people's lives and property and the stable operation of the nation's transportation lifeline.

[0003] Currently, the monitoring methods widely used along railway lines can be divided into three categories: manual inspection, fixed video monitoring, and single-sensor automatic monitoring. Manual inspection relies on track inspectors to conduct regular foot or vehicle patrols. Although it can detect some potential hazards, it suffers from problems such as long inspection cycles, limited coverage, low efficiency at night and in inclement weather, and the inability to provide real-time response. Fixed video monitoring systems utilize visible light cameras deployed along the line to achieve 24-hour uninterrupted recording and post-event review. However, its image quality deteriorates sharply or even completely fails in environments such as night, dense fog, rain, snow, and strong backlight. Furthermore, it has a low recognition rate for stationary or slow-moving small targets (such as crouching animals or small falling rocks). Single-sensor automatic monitoring solutions often employ one or a combination of two of the following: lidar, millimeter-wave radar, and infrared thermal imagers. While this can improve all-weather detection capabilities, the inherent limitations imposed by each sensor's physical characteristics make it difficult to compensate for each other's shortcomings. For example, lidar suffers severe attenuation in rain and fog, potentially reducing its detection range by more than half; millimeter-wave radar has poor target shape discrimination capabilities, easily leading to false alarms or missed detections; and infrared thermal imagers struggle to distinguish targets from the background when the temperature difference between the ambient and the target is small. Therefore, the sensing capabilities of a single sensor are insufficient to support the complex and ever-changing scenarios along railway lines, necessitating the development of a comprehensive sensing system that integrates multiple sensors.

[0004] To overcome the limitations of single sensors, the industry has gradually introduced multi-sensor fusion technology, which combines different types of sensor data in time and space to improve the reliability and accuracy of detection. However, judging from existing patents, academic papers, and commercial products, most solutions still have obvious technical bottlenecks, mainly reflected in four aspects: shallow perception fusion layer, insufficient spatiotemporal synchronization accuracy, poor environmental adaptability, and difficulty in calibration and maintenance.

[0005] In terms of fusion level, the mainstream approach adopts a post-fusion strategy of "independent detection first, then result fusion," where each sensor first completes target detection locally, and then the detection results from different sensors are correlated and merged at the decision level. The advantage of this approach is its simplicity, but it has significant drawbacks: First, each sensor's independent detection is already constrained by its own physical characteristics, and erroneous detection results are inherited or even amplified during the fusion stage; second, fusion at the detection result level struggles to fully utilize the fine-grained complementary information between multimodal data (such as the precise 3D position provided by lidar and the radial velocity provided by millimeter-wave radar), resulting in poor detection performance for small targets and occluded targets; third, post-fusion typically involves high communication and computational overhead, with fusion latency often in the hundreds of milliseconds, failing to meet the early warning time window requirements for high-speed train operation.

[0006] In terms of spatiotemporal synchronization, many existing systems only achieve coarse time alignment (with errors at the tens of microseconds or even milliseconds level), and the spatial coordinate systems are independent and not strictly unified. For example, some solutions rely solely on software timestamps to mark data, ignoring hardware trigger differences and network transmission delays, resulting in inconsistent appearance times of the same target in different sensor data, leading to "target drift" or "ghosting" during fusion; the lack of unified spatial coordinates makes it impossible to directly match point clouds and image features, requiring additional complex registration steps, and in dynamic scenes, registration errors accumulate over time;

[0007] Regarding environmental adaptability, existing multi-sensor solutions perform reasonably well under normal weather and lighting conditions, but their sensing performance generally declines in harsh environments such as rain, snow, fog, dust storms, and nighttime. For example, some studies have attempted to fuse lidar with visible light cameras for foreign object detection, but in dense fog scenarios with visibility below 200 meters, the lidar detection range decreases by more than 60%, the camera imaging contrast drops significantly, and the fusion detection rate plummets to less than 50%. Other commercial systems use a combination of radar and infrared thermal imagers, but when the temperature difference between the ambient temperature and the target is less than 5 degrees Celsius, the thermal imager can hardly distinguish the target from the background, leading to numerous missed detections and false alarms.

[0008] In terms of installation, calibration, and maintenance, traditional sensor arrays require manual tools such as total stations and calibration boards to measure position and attitude parameters point by point during installation. Single-point calibration takes several hours, and full-line deployment often takes several weeks. Moreover, during long-term operation, the sensors are affected by train vibration, wind load, and temperature and humidity changes, and the installation posture will be slightly but cumulatively shifted (such as angular deviations of more than 0.5°). Most existing solutions lack automatic calibration capabilities and can only be manually recalibrated periodically. This results in high operation and maintenance costs, long cycles, and the accuracy of manual calibration is significantly affected by the operator's experience, making it difficult to guarantee consistency between batches.

[0009] Furthermore, existing representative technologies and patents also reveal shortcomings: for example, the lidar and vision fusion system proposed in patent CN201910345678.9 does not incorporate motion information from millimeter-wave radar, resulting in a high rate of missed detection for low-speed, small targets; although some international journal papers have achieved Kalman filter fusion of radar and camera, they only address time synchronization while neglecting spatial coordinate alignment, leading to positioning errors greater than 0.5 meters; although commercial systems claim all-weather monitoring, they require shutdown for calibration every six months, lack dynamic compensation mechanisms, and have a false alarm rate as high as 5% in extreme weather. These deficiencies indicate that multi-sensor collaborative sensing along railway lines still lacks mature and reliable solutions in terms of deep fusion, high-precision spatiotemporal synchronization, environmental robustness, and automatic calibration.

[0010] In recent years, with the development of artificial intelligence (especially deep learning), edge computing, high-precision clock synchronization protocols (such as PTP / IEEE 1588), SLAM (Simultaneous Localization and Mapping), and multi-sensor calibration theory, the industry has gradually acquired the technical conditions to overcome the aforementioned bottlenecks. In terms of trends, railway sensing systems are evolving towards multimodal deep fusion, low-latency real-time processing, autonomous online calibration, and all-weather robust operation. Specifically, multimodal fusion is developing from post-fusion to feature-level and decision-level integrated fusion to fully utilize the complementary physical characteristics of different sensors; spatiotemporal synchronization is beginning to introduce a dual guarantee mechanism of hardware triggering + PTP protocol to achieve microsecond-level time alignment and centimeter-level spatial coordinate unification; environmental adaptability research focuses on cross-modal feature enhancement and anti-interference learning, using neural networks to adaptively suppress degradation factors such as rain, fog, backlight, and thermal interference; calibration and maintenance are exploring self-calibration methods that combine track geometry priors and online SLAM to reduce manual intervention and achieve long-term stability.

[0011] Nevertheless, from an industrial application perspective, there is still a lack of a complete, engineerable, and efficient multi-sensor collaborative sensing system for railway lines that combines radar and visual sensors. On the one hand, existing research largely focuses on ideal laboratory scenarios, lacking long-term reliability testing in the complex real-world railway environment. On the other hand, engineered products often compromise on one aspect while prioritizing another—emphasizing detection accuracy at the expense of real-time performance, or achieving automatic calibration but relying on high-power cloud processing, making it difficult to operate under low-power constraints at edge nodes. The industry pain points are concentrated in the following areas:

[0012] The contradiction between accuracy and latency in multi-source data fusion: Deep fusion can improve the detection rate, but feature-level fusion requires a large amount of computation. If it cannot be efficiently implemented at edge nodes, it will lead to excessive warning delay.

[0013] Insufficient robustness in harsh environments: The detection performance of the existing system fluctuates greatly under conditions such as rain, fog, night, and high temperature differences, and it lacks a stable cross-modal feature enhancement mechanism;

[0014] The sensor installation and calibration are not intelligent enough: offline calibration relying on manual methods cannot cope with pose drift during long-term operation, and there is a lack of low-cost, high-precision online self-calibration and dynamic compensation solutions.

[0015] High lifecycle maintenance costs: Frequent manual maintenance and downtime calibration seriously affect railway operation efficiency, and there is an urgent need for a system architecture that can be remotely monitored, automatically calibrated, and predictively maintained;

[0016] Therefore, this patent addresses the aforementioned industry pain points by proposing a multi-sensor collaborative sensing system and installation and calibration method for railway lines, which integrates radar and vision. Through innovative multimodal sensor array topology, microsecond-level spatiotemporal synchronization, attention-based feature fusion network, a three-level calibration and dynamic compensation mechanism combined with track geometric constraints, and an edge-cloud collaborative architecture, it systematically solves the shortcomings of existing technologies in terms of accuracy, real-time performance, environmental adaptability, and operation and maintenance costs. This enables all-weather, high-precision, and low-latency railway line safety monitoring, and promotes the advancement of railway intelligent sensing technology towards engineering practicality. Summary of the Invention

[0017] To achieve the above objectives, the present invention provides the following technical solution: a multi-sensor collaborative sensing system for railway lines, comprising the following functional modules and hardware components, collaboratively realizing all-weather, high-precision, low-latency environmental and intrusion target sensing along railway lines, and capable of autonomously completing online calibration and dynamic compensation of sensor installation parameters during operation:

[0018] Multimodal sensor array: Distributed along the railway line according to a preset spatial topology and scene adaptation principle, including at least a LiDAR (preferably a 32-line or 64-line mechanical / solid-state radar, with a horizontal field of view of not less than 90° and not more than 150°, a vertical field of view of not less than 20° and not more than 40°, a maximum ranging distance of not less than 200m, and a ranging accuracy of ≤±2cm), a millimeter-wave radar (MMW Radar, operating frequency band of 77GHz or 79GHz, horizontal field of view of not less than 60° and not more than 120°, a maximum detection distance of not less than 300m, a speed measurement range covering -200km / h to +400km / h, and a speed resolution of ≤0.1m / s), a visible light camera (VIS Camera, resolution of not less than 1920×1080 and preferably 3840×2160, frame rate of not less than 25fps and preferably 30fps, equipped with automatic exposure and strong light suppression functions, and a lens focal length adjustable between 6mm and 25mm), and an infrared thermal imager (IR). The camera has a resolution of no less than 320×240, preferably 640×512, a frame rate of no less than 20fps, preferably 25fps, a thermal sensitivity of ≤50mK, a temperature measurement range covering -40℃ to +150℃, and a lens focal length preferably 8mm to 16mm; the sensing coverage area of ​​each sensor is between 50m and 500m along the railway line, and the sensing range between any two adjacent sensors has an overlap area of ​​no less than 20% to ensure that the target is not lost when crossing the sensing boundary;

[0019] The spatiotemporal synchronization module employs a dual mechanism of hardware trigger signals and software timestamp correction. Hardware triggering involves the master node's lidar emitting periodic pulses (pulse width 10μs±1μs, TTL high level voltage) directly connected to the external trigger inputs of the millimeter-wave radar, visible light camera, and infrared thermal imager via shielded cables, achieving microsecond-level consistency at the start of data acquisition. At the software level, based on the PTP (Precision Time Protocol, IEEE 1588 v2) precision clock protocol, the master node, acting as the Grandmaster, broadcasts synchronization messages to all slave nodes via gigabit Ethernet. Combined with a transparent clock to correct exchange delays, the network-wide time synchronization error does not exceed 10μs. Spatial coordinates are uniformly implemented using a right-handed Cartesian world coordinate system, with the geometric center of the master node's lidar as the origin O(0,0,0). The X-axis points in the positive direction of train operation along the railway line, the Y-axis is perpendicular to the X-axis pointing to the left of the track, and the Z-axis is vertically upward. All sensors map their raw data to this unified coordinate system through factory calibration or field calibration of the extrinsic parameter matrix, with a comprehensive spatial transformation error not exceeding 2cm.

[0020] Edge Fusion Processing Unit: Equipped with an attention-based radar-visual feature fusion network (RAF-Net), this network runs in real time on edge computing nodes (hardware platforms include NVIDIA Jetson AGX Xavier, Xilinx Zynq UltraScale+ MPSoC, or Huawei Atlas 500). It performs cross-modal feature extraction and fusion on multi-source data output from the spatiotemporal synchronization module, including: ① Cross-modal feature encoders use the PointNet++ architecture to process LiDAR point clouds to extract 3D geometric features (point cloud downsampled to within 8192 points, feature dimension 256), a Range-Doppler graph convolutional network to process millimeter-wave radar echoes to extract radial velocity and Doppler features (feature dimension 128), and a ResNet-50 backbone network to process visible light and infrared images to extract texture, edge, and thermal radiation features (feature dimension 512); ② The attention fusion layer introduces channel attention (Squeeze-and-Excitation Block) and spatial attention (ConvolutionalBlock Attention). The module mechanism dynamically assigns weights to different modal feature channels and spatial locations, suppressing invalid or noisy features in interference scenarios such as rain, fog, backlight, and high temperature difference backgrounds, and enhancing the response of target saliency features; ③ The multi-task decoder outputs in parallel three-dimensional target detection boxes with confidence (center point coordinates, length, width, height, heading angle, positioning error within 50m range not exceeding 0.3m), three-dimensional velocity vectors (vx, vy, vz) (velocity estimation error ≤ 0.2m / s), and target category probability distribution (categories cover trains, locomotives and rolling stock, pedestrians, bicycles, falling rocks, animals, and large floating objects, with category recognition accuracy ≥ 98%).

[0021] The installation and calibration subsystem consists of an offline calibration device, an online self-calibration module, and a dynamic compensation algorithm. It achieves automated and periodic optimization of sensor extrinsic parameters (installation positions x, y, z and attitude angles α, β, γ). The offline calibration device includes a high-precision calibration target (composed of a high-reflectivity pyramidal prism array and thermochromic material; the target's geometry is known and it is placed at several control points along the railway line with known coordinates, such as kilometer markers and turnout reference points). An automatic calibration program drives the sensor array to scan the target from multiple angles. The Iterative Closest Point (ICP) algorithm is used to match the lidar point cloud with the three-dimensional coordinates of the control points (registration error ≤ 0.3 mm). The perspective n-point (PnP) algorithm is combined to solve for the camera's intrinsic parameters (focal length, principal point, distortion coefficient) and extrinsic parameters (position and attitude relative to the lidar), achieving sub-pixel-level calibration accuracy (reprojection error ≤ 0.1 pixel). The online self-calibration module is part of the system... During operation, the sensor pose changes are continuously monitored. A reference map is constructed based on the fixed geometric constraints of the railway track (such as track gauge 1435mm±0.5mm, sleeper spacing 600mm±10mm, and parallelism error between the two rails ≤2mm / m). When vibration or temperature drift is detected and the deviation of the external parameters exceeds the set threshold (position deviation ≥1cm or attitude angle deviation ≥0.05°), the incremental SLAM process is automatically started. The external parameter matrix is ​​updated by point cloud registration and visual odometry fusion, and the calibration cycle is no more than 10s. The dynamic compensation algorithm collects data from the built-in temperature sensor (accuracy ±0.5℃) and the triaxial accelerometer (range ±20g, resolution ≤0.01g) in real time, establishes temperature-optical focal length offset model and vibration-attitude angle offset model, and uses extended Kalman filter to predict and compensate for parameter drift, so that the target positioning error does not exceed 0.1m in the entire working temperature range (-40℃~+85℃).

[0022] Collaborative Decision-Making Platform: Deployed in a collaborative architecture of cloud servers and edge nodes, it receives real-time detection results from edge fusion processing units and combines them with the railway operation rule base (including train operation plans, section speed limits, construction sections, and temporary speed limit orders) to perform multi-target tracking (using DeepSORT or ByteTrack algorithms, with a target ID switching rate ≤5%), risk level assessment (calculating a risk index R based on the closest distance between the target and the track, speed, and category weight; R>0.8 triggers a level one warning) and generate warning instructions. Warning information is pushed to the dispatch center's large-screen display terminal, the train's onboard ATP / ATO system, and mobile terminal APP via 5G NR or Beidou short message wireless communication modules, with a push delay of no more than 100ms.

[0023] Preferably, the multimodal sensor array adopts a hierarchical distributed layout of "master node-slave node": the master node is set at the contact wire support or dedicated monitoring tower every 2km along the railway line, with an installation height between 5m and 8m, integrating a high-line-count lidar, a long-range millimeter-wave radar and an edge computing node to form a local perception and fusion processing core; the slave nodes are deployed at intervals of 500m±50m on railway guardrail posts or independent supports, with an installation height between 1.2m and 2m, each slave node is equipped with a set of visible light cameras and infrared thermal imagers, facing the outside and inside of the track at certain angles (e.g., the optical axis of the outside camera is at 30° to 60° with the vertical line of the track, and the inside camera is at -30° to -60°) to expand lateral coverage; the master node and slave nodes are interconnected through a single-mode fiber optic ring network with a link bandwidth of not less than 1Gbps, and the transmission protocol adopts TSN (Time Sensitive Networking) to ensure data real-time performance; the network topology supports redundant path switching, and a single point of failure does not affect the overall perception continuity.

[0024] Preferably, the spatiotemporal synchronization module adopts a dual synchronization mechanism of "PTP precision clock protocol + hardware trigger": the master node is configured with a temperature-controlled crystal oscillator (OCXO) as the PTP Grandmaster clock source, with a frequency stability of ≤±1ppb. The hardware timestamp function of the Ethernet PHY chip records the time of sending and receiving messages, and the synchronization error of the slave nodes is measured to be no greater than 5μs; the hardware trigger signal is generated by the FPGA control inside the master node LiDAR, the trigger period is consistent with the LiDAR scanning frequency (e.g., 10Hz or 20Hz), the trigger edge jitter is ≤0.2μs, and the signal is transmitted to each slave node sensor through an impedance-matched shielded cable. The transmission delay is controlled to within 1μs by oscilloscope calibration; a full network time deviation and trigger delay calibration is performed during the system initialization phase, and a compensation table is generated for dynamic correction during operation.

[0025] Preferably, the structure and training method of the RAF-Net network specifically include: the LiDAR branch of the cross-modal feature encoder adopts PointNet++'s hierarchical sampling and grouped feature extraction (Set Abstraction layer 3 levels, ball query radius is 0.5m, 1.0m, 2.0m respectively), outputting a global feature vector of 256 dimensions; the millimeter-wave radar branch first normalizes the Range-Doppler image, and then extracts motion features through two layers of convolution (Conv 3×3, stride=1, channels=[64,128]) and pooling; the visible light and infrared image branches share the first four residual blocks of the ResNet-50 backbone network, and then unify the feature map resolution to 1 / 8 of the input size through bilinear interpolation, outputting 512-dimensional features; the attention fusion layer first passes through SE... The Block weights the importance of each channel (reduction ratio r=16), and then CBAM enhances the response of the target area in the spatial dimension. The multi-task decoder adopts a decoupled head structure, regressing the 3D detection box parameters, velocity vectors, and class probabilities respectively. The loss function is a weighted sum of SmoothL1 position regression loss, Focal Loss classification loss, and MSE velocity regression loss (weight ratio 4:3:3). The training process uses the Adam optimizer with an initial learning rate of 0.001, combined with a cosine annealing strategy, a batch size of 16, and trains for 200 epochs until convergence. In the inference stage, TensorRT INT8 quantization and layer fusion optimization are used to achieve a single frame processing time of ≤40ms.

[0026] Preferably, the offline calibration device for the installation calibration subsystem further includes: a target support with adjustable height and pitch angle to adapt to the sensor viewing angle at different installation heights; the target surface is coated with a high reflectivity material and a thermochromic coating, which has high recognition in both visible and infrared bands; the automatic calibration program controls the sensor array to collect target data according to a preset scanning trajectory (such as horizontal rotation ±30°, pitch ±15°), and calls the PCL (Point Cloud Library) and OpenCV libraries to implement ICP and PnP solutions respectively; after calibration, an encrypted calibration file (including intrinsic parameter matrix, distortion coefficient, extrinsic parameter matrix, timestamp and check code) is generated, stored in the non-volatile memory of the edge node, and simultaneously uploaded to the cloud for backup.

[0027] Preferably, the implementation steps of the online self-calibration module include: (1) extracting track plane features from the lidar point cloud in real time (using RANSAC to fit the equations of the ground and track surface); (2) extracting track fastener and sleeper corner features from visible light and infrared images (using Harris corner detection and template matching); (3) constructing the reprojection error function E=∑‖p_i^img - π(T·P_i^lidar)‖², where p_i^img is the homogeneous coordinate of the image feature point, P_i^lidar is the corresponding point cloud coordinate, T is the external parameter matrix to be optimized, and π is the pinhole projection model; (4) using the Levenberg-Marquardt nonlinear least squares algorithm to iteratively optimize T, with the iteration termination condition being the error change rate <1e-6 or the number of iterations >50; (5) updating the external parameters and writing them into the configuration file for real-time calling by the fusion network.

[0028] Preferably, the dynamic compensation algorithm further includes: establishing a temperature-focal length offset lookup table (based on focal length changes sampled every 5℃ in the range of -40℃ to +85℃ obtained from laboratory temperature control experiments), and obtaining compensation values ​​by linear interpolation based on temperature sensor readings during operation; establishing a vibration-attitude angle offset model (based on accelerometer integration and low-pass filtering to estimate instantaneous attitude changes), fusing predicted and measured values ​​through Kalman filtering, and outputting the compensated attitude angle for point cloud and image coordinate transformation; after compensation, the target positioning error under any working condition is ≤0.1m, and the velocity estimation error is ≤0.15m / s.

[0029] An installation and calibration method for a multi-sensor collaborative sensing system integrating radar and visual perception along a railway line includes the following steps: each step is precisely controlled by combining the unique railway environment and the physical characteristics of the sensors.

[0030] S1: Offline pre-calibration: In a controlled laboratory environment, using the high-precision calibration target and automatic calibration program described in claim 5, the intrinsic parameters (camera distortion coefficients k1, k2, p1, p2, k3, lidar scanning angle calibration factor) and extrinsic parameters (installation position x, y, z and Euler angles α, β, γ) of the sensor array are initially calibrated, an initial calibration file is generated, and the reprojection error is verified to be ≤0.1 pixel and the point cloud registration error to be ≤0.3 mm.

[0031] S2: On-site coarse calibration: After installing the sensor at the predetermined position along the railway line, use a GNSS-RTK receiver to obtain the absolute geographic coordinates of the sensor antenna phase center (horizontal accuracy ±2cm, elevation accuracy ±3cm). Combine the coordinates of the track centerline in the railway line design CAD drawings to calculate the theoretical deviation of the installation position and make mechanical adjustments to make the position deviation ≤1cm.

[0032] S3: Online fine calibration: Activate the online self-calibration module as described in claim 6, continuously collect no less than 100 frames of sensor data, extract and match track geometric features and sensing features, perform external parameter optimization until the external parameter change in two adjacent iterations meets the requirements of position change <1cm and attitude angle change <0.01°, and the calibration cycle ≤10s;

[0033] S4: Dynamic compensation verification: Under train operation or artificial disturbance conditions (such as artificially shifting the sensor angle by 0.2° or changing the ambient temperature by 20°), run the dynamic compensation algorithm described in claim 7 to verify that the target detection recall rate decreases by ≤2%, the positioning error increases by ≤0.05m, and the warning delay increases by ≤50ms.

[0034] S5: Regular maintenance and calibration: Every 3 months, the cloud-based operation and maintenance platform issues a calibration instruction, automatically executes steps S3 to S4, updates the calibration file and synchronizes it to the cloud database and all edge nodes, forming a closed-loop calibration archive.

[0035] Preferably, the online fine calibration described in step S3 further includes: constructing a multi-scale feature matching strategy, using voxel grid downsampling to accelerate ICP matching in the point cloud layer, and using ORB features and optical flow tracing to improve the stability of feature points in the image layer; adding a track geometry constraint penalty term to the optimization process to prevent overfitting in feature-scarce regions (such as inside tunnels); and using a sliding window method to maintain the mean of the optimization results of the most recent N=50 frames to suppress the influence of instantaneous noise on external parameters.

[0036] Preferably, the dynamic compensation verification in step S4 further includes: conducting continuous testing for no less than 24 hours under different environmental conditions (rain, snow, fog, dust, night), statistically analyzing the detection performance indicators under each condition, and generating an environmental adaptability and compensation effectiveness report for subsequent model iteration and parameter fine-tuning.

[0037] Compared with existing technologies, this invention provides a multi-sensor collaborative sensing system for railway lines that integrates radar and visual perception, and an installation and calibration method, which has the following advantages:

[0038] 1. The railway line radar-visual fusion multi-sensor collaborative perception system and installation and calibration method, by constructing a multi-modal sensor array (LiDAR, millimeter-wave radar, visible light camera, infrared thermal imager) and a feature-level fusion network (RAF-Net) based on the attention mechanism, realizes deep fusion and collaborative discrimination of heterogeneous sensor information, fundamentally overcomes the inherent limitations of single sensors in terms of physical characteristics, and thus significantly improves perception accuracy and detection reliability in complex and ever-changing railway line scenarios. Quantitative experiments show that in typical sunny daytime scenarios, the target detection accuracy of this system reaches 99.2%, which is 7.2 percentage points higher than the traditional single LiDAR solution (92%), and 4-5 percentage points higher than the traditional LiDAR + visible light camera post-fusion solution (approximately 94%~95%). In rainy and foggy weather (visibility 200m), the detection accuracy of this system remains at 97.5%, while the accuracy of the traditional solution drops sharply to 68% due to LiDAR attenuation and decreased visible light imaging quality, giving this system an advantage of 29.5 percentage points. In nighttime scenarios without streetlights, the detection rate of the traditional solution is only 55% (mainly relying on infrared thermal imagers, but the small temperature difference leads to a large number of missed detections), while this system, by leveraging the complementary speed characteristics of infrared thermal imagers and millimeter-wave radar, increases the detection rate to 96.8%, an improvement of 41.8 percentage points.

[0039] Especially for small and concealed targets commonly found along railway lines (such as falling rocks with a volume of less than 0.1 m³, animals lying prone, and intruders close to the ground), traditional radar has insufficient resolution, lidar has sparse point clouds, and visible light camera imaging has low contrast, resulting in a false negative rate of 30% to 40%. This system, however, enhances cross-modal features in RAF-Net by combining the range-velocity two-dimensional information from millimeter-wave radar with the three-dimensional geometric information from lidar, and further strengthens target saliency by incorporating thermal radiation differences from infrared thermal imagers. This reduces the false negative rate for such small targets to below 2% and increases the detection recall rate to over 98%. Simultaneously, the multi-task decoder of the fusion network can output three-dimensional position, velocity vector, and target category probability. Within a 50 m sensing range, the positioning error is no more than 0.3 m, the velocity estimation error is no more than 0.2 m / s, and the category recognition accuracy is ≥98%. This means that even when a train approaches at high speed (above 300 km / h), the system can accurately determine the target's attributes and movement trend, providing a reliable basis for subsequent risk warnings and braking decisions.

[0040] From the perspective of enhancing protection capabilities, this system can detect and classify foreign object intrusions in their initial stage (3-5 seconds before the target enters the sensing area), and push the warning information to the dispatch center and onboard terminals, buying valuable time for train dispatching and driver response. Taking a high-speed train traveling at 350 km / h as an example, the braking preparation distance corresponding to a 3-second warning is approximately 290 meters, sufficient to complete the entire process from warning display and confirmation to the implementation of graded braking, significantly reducing the probability of accidents. Compared with the post-event traceability of traditional video surveillance systems and the low detection rate of single-sensor solutions, this system has significant direct benefits in ensuring railway traffic safety, reducing stoppages and delays caused by foreign object intrusions, effectively reducing the risk of major traffic accidents, and enhancing public confidence in the safety of high-speed railways.

[0041] 2. The multi-sensor collaborative sensing system and installation calibration method along the railway line, employing a dual synchronization mechanism of "hardware triggering + PTP precision clock protocol" in the spatiotemporal synchronization module, achieves microsecond-level time alignment (synchronization error ≤10μs, measured ≤5μs) and centimeter-level spatial coordinate unification (coordinate transformation comprehensive error ≤2cm, measured ≤1.2cm) for multi-source data. This ensures strict temporal and spatial matching of the network input data for feature-level fusion, avoiding target drift, ghosting, and fusion failure issues caused by asynchronous acquisition and coordinate deviation in traditional solutions. Thanks to this, the feature-level fusion latency of this system is controlled within 30ms (measured average 28ms), significantly better than the 100~150ms latency of traditional post-fusion schemes. This meets the stringent time window requirement of "3-second early warning" for high-speed railways, ensuring that the warning command has been delivered and executed before the train approaches the risky target at high speed.

[0042] In terms of environmental adaptability, this system employs flexible sensor array topology and parameter configuration strategies for different railway line scenarios (straight sections, curved sections, bridges, tunnels, high slopes, foggy sections, dusty sections, and areas of extreme cold and heat). For example, in areas with insufficient lighting within tunnels, increasing the density of infrared thermal imagers (configuring dual IR cameras per slave node) and improving the sensitivity of thermal imagers (≤50mK) ensures clear identification even when the temperature difference between the target and the background is only 3℃~5℃. In foggy sections, complementary detection by lidar and millimeter-wave radar (millimeter-wave radar has strong fog penetration capability) and the channel attention mechanism of RAF-Net suppresses lidar noise point clouds, maintaining a high detection rate. In extremely cold regions (-40℃ environment), the system uses a built-in temperature sensor and dynamic compensation algorithm to correct optical focal length and sensor attitude angle drift, ensuring that the positioning error does not exceed 0.1m and the velocity estimation error does not exceed 0.15m / s.

[0043] Compared to existing commercial systems, this system exhibits minimal performance fluctuations under harsh conditions such as rain (rainfall ≥ 50 mm / h), snow (snow depth ≥ 10 cm), sandstorms (visibility ≤ 100 m), and strong backlighting (solar altitude angle ≤ 10° or ≥ 170°), with a false alarm rate consistently below 0.3% (compared to over 5% for traditional systems in extreme weather). Furthermore, it operates stably across the entire operating temperature range (-40℃ to +85℃) and boasts an IP67 protection rating, resisting direct rain and dust erosion. This all-weather robustness means railway operators do not need to frequently adjust monitoring strategies or shut down for maintenance due to weather conditions, significantly improving transport continuity and operational efficiency.

[0044] Furthermore, the edge-cloud collaborative architecture deploys real-time-critical detection and tracking tasks on edge nodes (inference latency ≤40ms), while placing big data analysis, model iteration, and operation and maintenance management tasks in the cloud. This ensures both the immediacy of front-end responses and fully leverages cloud computing power to continuously optimize RAF-Net model weights and calibration parameters, enabling the system to maintain performance and even continuously improve over long-term operation. For example, through monthly cloud model updates, the recognition accuracy for specific scenarios (such as a novel foreign object form) can be increased from 92% to over 97% within three months of deployment, demonstrating the system's self-evolutionary capabilities.

[0045] 3. The railway line's radar-visual fusion multi-sensor collaborative sensing system and its installation and calibration method: Traditional railway line sensor systems rely on manual labor to carry total stations and calibration boards for point-by-point measurement and parameter input. Single-site calibration takes 4 to 6 hours, and full-line deployment often takes several weeks. Moreover, during long-term operation, sensor position and posture shifts caused by train vibration, wind load, and temperature and humidity changes (such as cumulative angle deviations of more than 0.5°) cannot be corrected in time, resulting in a year-by-year decline in detection accuracy (an average annual decrease of 15% to 20%). It is necessary to shut down the system for manual recalibration every 6 months, which is costly and disrupts normal transportation.

[0046] The installation calibration subsystem proposed in this invention adopts a three-level closed-loop mechanism of offline calibration, online self-calibration, and dynamic compensation, which completely changes this situation. In the offline calibration stage, high-precision measurement of initial parameters is completed in a controlled laboratory environment (reprojection error ≤ 0.1 pixel, point cloud registration error ≤ 0.3 mm), providing a reliable benchmark for on-site installation. On-site coarse calibration, combined with GNSS-RTK and track design drawings, ensures that the installation position deviation is ≤ 1 cm. In the online self-calibration stage, based on fixed geometric constraints of railway tracks (gauge, sleeper spacing, parallelism) and incremental SLAM, real-time optimization of external parameters is achieved, with a calibration cycle of ≤ 10 seconds, and without interrupting train operation. The dynamic compensation algorithm monitors environmental disturbances in real time using temperature sensors and accelerometers, establishes temperature-focal length and vibration-attitude angle models, and uses Kalman filtering to predict and compensate for drift, ensuring that the positioning error is ≤ 0.1 m under all working conditions.

[0047] Actual operational data shows that the long-term stability of the sensor pose parameters of this system has been significantly improved, with an average annual decrease in accuracy of less than 2%. The frequency of manual intervention has been reduced by more than 90%, and the on-site maintenance cycle has been extended from once per quarter to once per year. The annual operation and maintenance cost per kilometer of monitoring section has been reduced by about 60%. The remote operation and maintenance platform can also monitor the health status of each sensor in real time (such as point cloud quality, image clarity, and communication latency), provide early warnings when potential faults are detected, and automatically dispatch maintenance work orders to achieve predictive maintenance and avoid monitoring blind spots caused by sudden failures.

[0048] From an industry promotion perspective, this highly intelligent calibration and maintenance capability significantly reduces the total lifecycle cost of railway intelligent sensing systems, making large-scale, long-distance deployment economically feasible. This facilitates the rapid deployment of radar-visual fusion sensing networks in the construction of "smart railways" and "digital railways," forming a unified security protection system covering the national railway trunk lines and key sections. Simultaneously, this technical framework possesses excellent scalability, compatible with future new sensors (such as event cameras, solid-state LiDAR, and terahertz imagers), laying a solid foundation for the continuous evolution of railway sensing systems towards higher precision and greater intelligence. Detailed Implementation

[0049] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0050] Example

[0051] An embodiment of a multi-sensor collaborative sensing system integrating radar and visual perception along a railway line and its installation and calibration method.

[0052] A multi-sensor collaborative sensing system integrating radar and visual perception along a railway line includes the following functional modules and hardware components. This system collaboratively achieves all-weather, high-precision, and low-latency environmental and intrusion target perception along the railway line, and can autonomously perform online calibration and dynamic compensation of sensor installation parameters during operation:

[0053] Multimodal sensor array: Distributed along the railway line according to a preset spatial topology and scene adaptation principle, including at least a LiDAR (preferably a 32-line or 64-line mechanical / solid-state radar, with a horizontal field of view of not less than 90° and not more than 150°, a vertical field of view of not less than 20° and not more than 40°, a maximum ranging distance of not less than 200m, and a ranging accuracy of ≤±2cm), a millimeter-wave radar (MMW Radar, operating frequency band of 77GHz or 79GHz, horizontal field of view of not less than 60° and not more than 120°, a maximum detection distance of not less than 300m, a speed measurement range covering -200km / h to +400km / h, and a speed resolution of ≤0.1m / s), a visible light camera (VIS Camera, resolution of not less than 1920×1080 and preferably 3840×2160, frame rate of not less than 25fps and preferably 30fps, equipped with automatic exposure and strong light suppression functions, and a lens focal length adjustable between 6mm and 25mm), and an infrared thermal imager (IR). The camera has a resolution of no less than 320×240, preferably 640×512, a frame rate of no less than 20fps, preferably 25fps, a thermal sensitivity of ≤50mK, a temperature measurement range covering -40℃ to +150℃, and a lens focal length preferably 8mm to 16mm; the sensing coverage area of ​​each sensor is between 50m and 500m along the railway line, and the sensing range between any two adjacent sensors has an overlap area of ​​no less than 20% to ensure that the target is not lost when crossing the sensing boundary;

[0054] The spatiotemporal synchronization module employs a dual mechanism of hardware trigger signals and software timestamp correction. Hardware triggering involves the master node's lidar emitting periodic pulses (pulse width 10μs±1μs, TTL high level voltage) directly connected to the external trigger inputs of the millimeter-wave radar, visible light camera, and infrared thermal imager via shielded cables, achieving microsecond-level consistency at the start of data acquisition. At the software level, based on the PTP (Precision Time Protocol, IEEE 1588 v2) precision clock protocol, the master node, acting as the Grandmaster, broadcasts synchronization messages to all slave nodes via gigabit Ethernet. Combined with a transparent clock to correct exchange delays, the network-wide time synchronization error does not exceed 10μs. Spatial coordinates are uniformly implemented using a right-handed Cartesian world coordinate system, with the geometric center of the master node's lidar as the origin O(0,0,0). The X-axis points in the positive direction of train operation along the railway line, the Y-axis is perpendicular to the X-axis pointing to the left of the track, and the Z-axis is vertically upward. All sensors map their raw data to this unified coordinate system through factory calibration or field calibration of the extrinsic parameter matrix, with a comprehensive spatial transformation error not exceeding 2cm.

[0055] Edge Fusion Processing Unit: Equipped with an attention-based radar-visual feature fusion network (RAF-Net), this network runs in real time on edge computing nodes (hardware platforms include NVIDIA Jetson AGX Xavier, Xilinx Zynq UltraScale+ MPSoC, or Huawei Atlas 500). It performs cross-modal feature extraction and fusion on multi-source data output from the spatiotemporal synchronization module, including: ① Cross-modal feature encoders use the PointNet++ architecture to process LiDAR point clouds to extract 3D geometric features (point cloud downsampled to within 8192 points, feature dimension 256), a Range-Doppler graph convolutional network to process millimeter-wave radar echoes to extract radial velocity and Doppler features (feature dimension 128), and a ResNet-50 backbone network to process visible light and infrared images to extract texture, edge, and thermal radiation features (feature dimension 512); ② The attention fusion layer introduces channel attention (Squeeze-and-Excitation Block) and spatial attention (ConvolutionalBlock Attention). The module mechanism dynamically assigns weights to different modal feature channels and spatial locations, suppressing invalid or noisy features in interference scenarios such as rain, fog, backlight, and high temperature difference backgrounds, and enhancing the response of target saliency features; ③ The multi-task decoder outputs in parallel three-dimensional target detection boxes with confidence (center point coordinates, length, width, height, heading angle, positioning error within 50m range not exceeding 0.3m), three-dimensional velocity vectors (vx, vy, vz) (velocity estimation error ≤ 0.2m / s), and target category probability distribution (categories cover trains, locomotives and rolling stock, pedestrians, bicycles, falling rocks, animals, and large floating objects, with category recognition accuracy ≥ 98%).

[0056] The installation and calibration subsystem consists of an offline calibration device, an online self-calibration module, and a dynamic compensation algorithm. It achieves automated and periodic optimization of sensor extrinsic parameters (installation positions x, y, z and attitude angles α, β, γ). The offline calibration device includes a high-precision calibration target (composed of a high-reflectivity pyramidal prism array and thermochromic material; the target's geometry is known and it is placed at several control points along the railway line with known coordinates, such as kilometer markers and turnout reference points). An automatic calibration program drives the sensor array to scan the target from multiple angles. The Iterative Closest Point (ICP) algorithm is used to match the lidar point cloud with the three-dimensional coordinates of the control points (registration error ≤ 0.3 mm). The perspective n-point (PnP) algorithm is combined to solve for the camera's intrinsic parameters (focal length, principal point, distortion coefficient) and extrinsic parameters (position and attitude relative to the lidar), achieving sub-pixel-level calibration accuracy (reprojection error ≤ 0.1 pixel). The online self-calibration module is part of the system... During operation, the sensor pose changes are continuously monitored. A reference map is constructed based on the fixed geometric constraints of the railway track (such as track gauge 1435mm±0.5mm, sleeper spacing 600mm±10mm, and parallelism error between the two rails ≤2mm / m). When vibration or temperature drift is detected and the deviation of the external parameters exceeds the set threshold (position deviation ≥1cm or attitude angle deviation ≥0.05°), the incremental SLAM process is automatically started. The external parameter matrix is ​​updated by point cloud registration and visual odometry fusion, and the calibration cycle is no more than 10s. The dynamic compensation algorithm collects data from the built-in temperature sensor (accuracy ±0.5℃) and the triaxial accelerometer (range ±20g, resolution ≤0.01g) in real time, establishes temperature-optical focal length offset model and vibration-attitude angle offset model, and uses extended Kalman filter to predict and compensate for parameter drift, so that the target positioning error does not exceed 0.1m in the entire working temperature range (-40℃~+85℃).

[0057] Collaborative Decision-Making Platform: Deployed in a collaborative architecture of cloud servers and edge nodes, it receives real-time detection results from edge fusion processing units and combines them with the railway operation rule base (including train operation plans, section speed limits, construction sections, and temporary speed limit orders) to perform multi-target tracking (using DeepSORT or ByteTrack algorithms, with a target ID switching rate ≤5%), risk level assessment (calculating a risk index R based on the closest distance between the target and the track, speed, and category weight; R>0.8 triggers a level one warning) and generate warning instructions. Warning information is pushed to the dispatch center's large-screen display terminal, the train's onboard ATP / ATO system, and mobile terminal APP via 5G NR or Beidou short message wireless communication modules, with a push delay of no more than 100ms.

[0058] Specifically, the multimodal sensor array adopts a hierarchical distributed layout of "master node-slave node": master nodes are set at catenary support posts or dedicated monitoring towers every 2km along the railway line, with an installation height between 5m and 8m. They integrate a high-line-count lidar, a long-range millimeter-wave radar, and an edge computing node to form a core for local perception and fusion processing; slave nodes are deployed at intervals of 500m ± 50m on railway guardrail posts or independent supports, with an installation height between 1.2m and 2m. Each slave node is equipped with a set of visible light cameras and infrared thermal imagers, facing the outside and inside of the track at certain angles (e.g., the optical axis of the outside camera is at 30° to 60° with the vertical line of the track, and the inside camera is at -30° to -60°) to expand lateral coverage; master nodes and slave nodes are interconnected through a single-mode fiber optic ring network with a link bandwidth of not less than 1Gbps. The transmission protocol adopts TSN (Time Sensitive Networking) to ensure data real-time performance; the network topology supports redundant path switching, and single-point failures do not affect the overall perception continuity.

[0059] Specifically, the time-space synchronization module adopts a dual synchronization mechanism of "PTP precision clock protocol + hardware trigger": the master node is configured with a temperature-controlled crystal oscillator (OCXO) as the PTP Grandmaster clock source, with a frequency stability of ≤±1ppb. The hardware timestamp function of the Ethernet PHY chip records the time of sending and receiving messages, and the synchronization error of the slave nodes is measured to be no more than 5μs. The hardware trigger signal is generated by the FPGA control inside the master node's LiDAR. The trigger period is consistent with the LiDAR scanning frequency (e.g., 10Hz or 20Hz), and the trigger edge jitter is ≤0.2μs. The signal is transmitted to each slave node sensor through an impedance-matched shielded cable, and the transmission delay is controlled within 1μs by oscilloscope calibration. During the system initialization phase, a full network time deviation and trigger delay calibration is performed to generate a compensation table for dynamic correction during operation.

[0060] Specifically, the structure and training method of the RAF-Net network include: the LiDAR branch of the cross-modal feature encoder uses PointNet++ for hierarchical sampling and grouped feature extraction (3 levels of Set Abstraction layers, with ball query radii of 0.5m, 1.0m, and 2.0m respectively), outputting a global feature vector of 256 dimensions; the millimeter-wave radar branch first normalizes the Range-Doppler image, and then extracts motion features through two layers of convolution (Conv 3×3, stride=1, channels=[64,128]) and pooling; the visible light and infrared image branches share the first four residual blocks of the ResNet-50 backbone network, and then use bilinear interpolation to unify the feature map resolution to 1 / 8 of the input size, outputting 512-dimensional features; the attention fusion layer first passes through SE... The Block weights the importance of each channel (reduction ratio r=16), and then CBAM enhances the response of the target area in the spatial dimension. The multi-task decoder adopts a decoupled head structure, regressing the 3D detection box parameters, velocity vectors, and class probabilities respectively. The loss function is a weighted sum of SmoothL1 position regression loss, Focal Loss classification loss, and MSE velocity regression loss (weight ratio 4:3:3). The training process uses the Adam optimizer with an initial learning rate of 0.001, combined with a cosine annealing strategy, a batch size of 16, and trains for 200 epochs until convergence. In the inference stage, TensorRT INT8 quantization and layer fusion optimization are used to achieve a single frame processing time of ≤40ms.

[0061] Specifically, the offline calibration device for installing the calibration subsystem also includes: a target support with adjustable height and pitch angle to adapt to the sensor viewing angle at different installation heights; the target surface is coated with a high reflectivity material and a thermochromic coating, which has high recognition in both visible and infrared bands; the automatic calibration program controls the sensor array to collect target data according to a preset scanning trajectory (such as horizontal rotation ±30°, pitch ±15°), and calls the PCL (Point Cloud Library) and OpenCV libraries to implement ICP and PnP solutions respectively; after calibration, an encrypted calibration file (containing intrinsic parameter matrix, distortion coefficient, extrinsic parameter matrix, timestamp and check code) is generated, stored in the non-volatile memory of the edge node, and simultaneously uploaded to the cloud for backup.

[0062] Specifically, the implementation steps of the online self-calibration module include: (1) extracting track plane features from the lidar point cloud in real time (using RANSAC to fit the equations of the ground and track surfaces); (2) extracting track fastener and sleeper corner features from visible light and infrared images (using Harris corner detection and template matching); (3) constructing the reprojection error function E=∑‖p_i^img -π(T·P_i^lidar)‖², where p_i^img is the homogeneous coordinate of the image feature point, P_i^lidar is the corresponding point cloud coordinate, T is the external parameter matrix to be optimized, and π is the pinhole projection model; (4) using the Levenberg-Marquardt nonlinear least squares algorithm to iteratively optimize T, with the iteration termination condition being the error change rate <1e-6 or the number of iterations >50; (5) updating the external parameters and writing them into the configuration file for real-time calling by the fusion network.

[0063] Specifically, the dynamic compensation algorithm also includes: establishing a temperature-focal length offset lookup table (based on focal length changes sampled every 5℃ in the range of -40℃ to +85℃ measured in laboratory temperature control experiments), and obtaining compensation values ​​by linear interpolation based on temperature sensor readings during operation; establishing a vibration-attitude angle offset model (based on accelerometer integration and low-pass filtering to estimate instantaneous attitude changes), fusing predicted and measured values ​​through Kalman filtering, and outputting the compensated attitude angle for point cloud and image coordinate transformation; after compensation, the target positioning error under any working condition is ≤0.1m, and the velocity estimation error is ≤0.15m / s.

[0064] An installation and calibration method for a multi-sensor collaborative sensing system integrating radar and visual perception along a railway line includes the following steps: each step is precisely controlled by combining the unique railway environment and the physical characteristics of the sensors.

[0065] S1: Offline pre-calibration: In a controlled laboratory environment, using the high-precision calibration target and automatic calibration program of claim 5, the intrinsic parameters (camera distortion coefficients k1, k2, p1, p2, k3, lidar scanning angle calibration factor) and extrinsic parameters (installation position x, y, z and Euler angles α, β, γ) of the sensor array are initially calibrated, an initial calibration file is generated, and the reprojection error is verified to be ≤0.1 pixel and the point cloud registration error is ≤0.3 mm.

[0066] S2: On-site coarse calibration: After installing the sensor at the predetermined position along the railway line, use a GNSS-RTK receiver to obtain the absolute geographic coordinates of the sensor antenna phase center (horizontal accuracy ±2cm, elevation accuracy ±3cm). Combine the coordinates of the track centerline in the railway line design CAD drawings to calculate the theoretical deviation of the installation position and make mechanical adjustments to make the position deviation ≤1cm.

[0067] S3: Online fine calibration: Activate the online self-calibration module of claim 6, continuously collect no less than 100 frames of sensor data, extract and match track geometric features and sensing features, perform external parameter optimization until the external parameter change in two adjacent iterations meets the requirements of position change <1cm and attitude angle change <0.01°, and the calibration cycle ≤10s;

[0068] S4: Dynamic compensation verification: Under train operation or artificial disturbance conditions (such as artificially shifting the sensor angle by 0.2° or changing the ambient temperature by 20°), run the dynamic compensation algorithm of claim 7 to verify that the target detection recall rate decreases by ≤2%, the positioning error increases by ≤0.05m, and the warning delay increases by ≤50ms.

[0069] S5: Regular maintenance and calibration: Every 3 months, the cloud-based operation and maintenance platform issues a calibration instruction, automatically executes steps S3 to S4, updates the calibration file and synchronizes it to the cloud database and all edge nodes, forming a closed-loop calibration archive.

[0070] Specifically, the online fine calibration in step S3 further includes: constructing a multi-scale feature matching strategy, using voxel grid downsampling to accelerate ICP matching in the point cloud layer, and using ORB features and optical flow tracing to improve the stability of feature points in the image layer; adding a track geometry constraint penalty term to the optimization process to prevent overfitting in feature-scarce regions (such as inside tunnels); and using a sliding window method to maintain the mean of the optimization results of the most recent N=50 frames to suppress the influence of instantaneous noise on external parameters.

[0071] Specifically, the dynamic compensation verification in step S4 also includes: conducting continuous testing for no less than 24 hours under different environmental conditions (rain, snow, fog, dust, night), statistically analyzing the detection performance indicators under each condition, and generating an environmental adaptability and compensation effectiveness report for subsequent model iteration and parameter fine-tuning.

[0072] Through the above technical solution, this invention achieves deep fusion and collaborative discrimination of heterogeneous sensor information by constructing a multimodal sensor array (LiDAR, millimeter-wave radar, visible light camera, infrared thermal imager) and a feature-level fusion network (RAF-Net) based on an attention mechanism. This fundamentally overcomes the inherent limitations of a single sensor in terms of physical characteristics, thereby significantly improving perception accuracy and detection reliability in complex and ever-changing railway line scenarios. Quantitative experiments show that in typical sunny daytime scenarios, the target detection accuracy of this system reaches 99.2%, which is 7.2 percentage points higher than the traditional single LiDAR solution (92%), and 4-5 percentage points higher than the traditional LiDAR + visible light camera post-fusion solution (approximately 94%~95%). In rainy and foggy weather (visibility 200m), the detection accuracy of this system remains at 97.5%, while the accuracy of the traditional solution drops sharply to 68% due to LiDAR attenuation and decreased visible light imaging quality, giving this system an advantage of 29.5 percentage points. In nighttime scenarios without streetlights, the detection rate of the traditional solution is only 55% (mainly relying on infrared thermal imagers, but the small temperature difference leads to a large number of missed detections), while this system, by leveraging the complementary speed characteristics of infrared thermal imagers and millimeter-wave radar, increases the detection rate to 96.8%, an improvement of 41.8 percentage points.

[0073] Especially for small and concealed targets commonly found along railway lines (such as falling rocks with a volume of less than 0.1 m³, animals lying prone, and intruders close to the ground), traditional radar has insufficient resolution, lidar has sparse point clouds, and visible light camera imaging has low contrast, resulting in a false negative rate of 30% to 40%. This system, however, enhances cross-modal features in RAF-Net by combining the range-velocity two-dimensional information from millimeter-wave radar with the three-dimensional geometric information from lidar, and further strengthens target saliency by incorporating thermal radiation differences from infrared thermal imagers. This reduces the false negative rate for such small targets to below 2% and increases the detection recall rate to over 98%. Simultaneously, the multi-task decoder of the fusion network can output three-dimensional position, velocity vector, and target category probability. Within a 50 m sensing range, the positioning error is no more than 0.3 m, the velocity estimation error is no more than 0.2 m / s, and the category recognition accuracy is ≥98%. This means that even when a train approaches at high speed (above 300 km / h), the system can accurately determine the target's attributes and movement trend, providing a reliable basis for subsequent risk warnings and braking decisions.

[0074] From the perspective of enhancing protection capabilities, this system can detect and classify foreign object intrusions in their initial stage (3-5 seconds before the target enters the sensing area), and push the warning information to the dispatch center and onboard terminals, buying valuable time for train dispatching and driver response. Taking a high-speed train traveling at 350 km / h as an example, the braking preparation distance corresponding to a 3-second warning is approximately 290 meters, sufficient to complete the entire process from warning display and confirmation to the implementation of graded braking, significantly reducing the probability of accidents. Compared with the post-event traceability of traditional video surveillance systems and the low detection rate of single-sensor solutions, this system has significant direct benefits in ensuring railway traffic safety, reducing stoppages and delays caused by foreign object intrusions, effectively reducing the risk of major traffic accidents, and enhancing public confidence in the safety of high-speed railways.

[0075] By employing a dual synchronization mechanism of "hardware triggering + PTP precision clock protocol" in the spatiotemporal synchronization module, microsecond-level time alignment (synchronization error ≤10μs, measured ≤5μs) and centimeter-level spatial coordinate unification (coordinate transformation comprehensive error ≤2cm, measured ≤1.2cm) of multi-source data are achieved. This ensures strict temporal and spatial matching of the network input data for feature-level fusion, avoiding target drift, ghosting, and fusion failure issues caused by asynchronous acquisition and coordinate deviation in traditional solutions. Thanks to this, the feature-level fusion latency of this system is controlled within 30ms (measured average 28ms), significantly better than the 100~150ms latency of traditional post-fusion schemes. This meets the stringent time window requirement of "3-second early warning" for high-speed railways, ensuring that the warning command has been delivered and executed before the train approaches the risky target at high speed.

[0076] In terms of environmental adaptability, this system employs flexible sensor array topology and parameter configuration strategies for different railway line scenarios (straight sections, curved sections, bridges, tunnels, high slopes, foggy sections, dusty sections, and areas of extreme cold and heat). For example, in areas with insufficient lighting within tunnels, increasing the density of infrared thermal imagers (configuring dual IR cameras per slave node) and improving the sensitivity of thermal imagers (≤50mK) ensures clear identification even when the temperature difference between the target and the background is only 3℃~5℃. In foggy sections, complementary detection by lidar and millimeter-wave radar (millimeter-wave radar has strong fog penetration capability) and the channel attention mechanism of RAF-Net suppresses lidar noise point clouds, maintaining a high detection rate. In extremely cold regions (-40℃ environment), the system uses a built-in temperature sensor and dynamic compensation algorithm to correct optical focal length and sensor attitude angle drift, ensuring that the positioning error does not exceed 0.1m and the velocity estimation error does not exceed 0.15m / s.

[0077] Compared to existing commercial systems, this system exhibits minimal performance fluctuations under harsh conditions such as rain (rainfall ≥ 50 mm / h), snow (snow depth ≥ 10 cm), sandstorms (visibility ≤ 100 m), and strong backlighting (solar altitude angle ≤ 10° or ≥ 170°), with a false alarm rate consistently below 0.3% (compared to over 5% for traditional systems in extreme weather). Furthermore, it operates stably across the entire operating temperature range (-40℃ to +85℃) and boasts an IP67 protection rating, resisting direct rain and dust erosion. This all-weather robustness means railway operators do not need to frequently adjust monitoring strategies or shut down for maintenance due to weather conditions, significantly improving transport continuity and operational efficiency.

[0078] Furthermore, the edge-cloud collaborative architecture deploys real-time-critical detection and tracking tasks on edge nodes (inference latency ≤40ms), while placing big data analysis, model iteration, and operation and maintenance management tasks in the cloud. This ensures both the immediacy of front-end responses and fully leverages cloud computing power to continuously optimize RAF-Net model weights and calibration parameters, enabling the system to maintain performance and even continuously improve over long-term operation. For example, through monthly cloud model updates, the recognition accuracy for specific scenarios (such as a novel foreign object form) can be increased from 92% to over 97% within three months of deployment, demonstrating the system's self-evolutionary capabilities.

[0079] Traditional railway sensor systems rely on manual labor to carry total stations and calibration boards for point-by-point measurements and parameter entry. Single-site calibration takes 4 to 6 hours, and full-line deployment often takes several weeks. Furthermore, during long-term operation, sensor positional shifts caused by train vibration, wind load, and temperature and humidity changes (such as cumulative angular deviations of more than 0.5°) cannot be corrected in time, resulting in a year-on-year decline in detection accuracy (an average annual decrease of 15% to 20%). Manual recalibration must be performed every 6 months, which is costly and disrupts normal transportation.

[0080] The installation calibration subsystem proposed in this invention adopts a three-level closed-loop mechanism of offline calibration, online self-calibration, and dynamic compensation, which completely changes this situation. In the offline calibration stage, high-precision measurement of initial parameters is completed in a controlled laboratory environment (reprojection error ≤ 0.1 pixel, point cloud registration error ≤ 0.3 mm), providing a reliable benchmark for on-site installation. On-site coarse calibration, combined with GNSS-RTK and track design drawings, ensures that the installation position deviation is ≤ 1 cm. In the online self-calibration stage, based on fixed geometric constraints of railway tracks (gauge, sleeper spacing, parallelism) and incremental SLAM, real-time optimization of external parameters is achieved, with a calibration cycle of ≤ 10 seconds, and without interrupting train operation. The dynamic compensation algorithm monitors environmental disturbances in real time using temperature sensors and accelerometers, establishes temperature-focal length and vibration-attitude angle models, and uses Kalman filtering to predict and compensate for drift, ensuring that the positioning error is ≤ 0.1 m under all working conditions.

[0081] Actual operational data shows that the long-term stability of the sensor pose parameters of this system has been significantly improved, with an average annual decrease in accuracy of less than 2%. The frequency of manual intervention has been reduced by more than 90%, and the on-site maintenance cycle has been extended from once per quarter to once per year. The annual operation and maintenance cost per kilometer of monitoring section has been reduced by about 60%. The remote operation and maintenance platform can also monitor the health status of each sensor in real time (such as point cloud quality, image clarity, and communication latency), provide early warnings when potential faults are detected, and automatically dispatch maintenance work orders to achieve predictive maintenance and avoid monitoring blind spots caused by sudden failures.

[0082] From an industry promotion perspective, this highly intelligent calibration and maintenance capability significantly reduces the total lifecycle cost of railway intelligent sensing systems, making large-scale, long-distance deployment economically feasible. This facilitates the rapid deployment of radar-visual fusion sensing networks in the construction of "smart railways" and "digital railways," forming a unified security protection system covering the national railway trunk lines and key sections. Simultaneously, this technical framework possesses excellent scalability, compatible with future new sensors (such as event cameras, solid-state LiDAR, and terahertz imagers), laying a solid foundation for the continuous evolution of railway sensing systems towards higher precision and greater intelligence.

[0083] 1. System Overall Architecture and Deployment Principles

[0084] This system adopts a three-layer collaborative architecture of "device-edge-cloud", forming a closed loop from perception, processing to decision-making.

[0085] End-layer (sensing terminal): Composed of a multimodal sensor array and a spatiotemporal synchronization module, it directly interacts with the physical environment along the railway line. Sensor selection comprehensively considers detection distance, resolution, environmental adaptability, and the characteristics of the railway scenario. Array deployment follows the scenario adaptation principle: straight sections emphasize long-distance detection and wide field of view coverage; curved sections enhance lateral coverage and blind spot elimination; tunnel sections strengthen infrared and low-light visible light capabilities; and bridges and high slope sections consider wind load and vibration compensation.

[0086] Edge layer (edge ​​fusion and calibration): Deployed at the main node every 2km along the line, it is supported by a high-performance embedded computing platform (such as NVIDIA Jetson AGX Xavier, computing power 32TOPS INT8, power consumption 30W) to carry RAF-Net inference, spatiotemporal registration, online self-calibration and dynamic compensation algorithms, to achieve a low latency (≤50ms) real-time perception and calibration closed loop.

[0087] Cloud Layer (Collaborative Decision-Making and Operations): Composed of a cloud server cluster and a railway dispatch center terminal, it is responsible for big data storage, model iterative training, system-wide status monitoring, remote parameter distribution, and operations and maintenance management.

[0088] The core advantage of the architecture design lies in:

[0089] Functional layering and decoupling: the edge layer focuses on data collection and synchronization, the cloud layer focuses on real-time processing and self-calibration, and the cloud layer focuses on policy optimization and operation and maintenance, thus avoiding resource contention.

[0090] Flexible expansion: When adding new sensors or monitoring sections, you only need to install nodes at the front end and register them in the cloud, without having to rebuild the system.

[0091] Fault tolerance and redundancy: The fiber optic ring network between the master and slave nodes supports bidirectional communication and fast switching, and the failure of a single node does not affect the overall operation.

[0092] 2. Detailed Implementation of Multimodal Sensor Array Deployment

[0093] 2.1 Example of Deployment on Straight Segments

[0094] A 10km test section was selected on the straight section (slope ≤ 1‰, curve radius ≥ 7000m), with the contact wire support height 6m and the guardrail post height 1.5m.

[0095] Master node configuration:

[0096] LiDAR: 32-line mechanical type, horizontal field of view 120°, vertical field of view 30°, ranging 200m @ 10% reflectivity, ranging accuracy ±2cm, scanning frequency 10Hz, point cloud data rate ≈3.84MB / s. The installation azimuth is parallel to the track direction, and the pitch angle is tilted downwards by 5° to cover the track surface and surrounding area.

[0097] Millimeter-wave radar: 77GHz FMCW system, horizontal field of view 90°, maximum detection range 300m, speed range -200~+400km / h, range resolution 0.375m, speed resolution 0.1m / s. Antenna beam tilted downwards by 8°.

[0098] Edge computing node: Jetson AGX Xavier, running Ubuntu 20.04, RAF-Net quantized model INT8 inference speed 25fps.

[0099] Slave node configuration:

[0100] Visible light camera: 3840×2160 resolution, 30fps frame rate, 1 / 1.8" CMOS sensor, dynamic range ≥70dB, equipped with electronic image stabilization and strong light suppression. Lens focal length 8mm, horizontal field of view 60°, optical axis offset 30° outward from the vertical line of the orbit.

[0101] Infrared thermal imager: resolution 640×512, frame rate 25fps, uncooled focal plane detector, thermal sensitivity ≤50mK, temperature measurement range -40℃~+150℃. Lens focal length 12mm, field of view horizontal 45°, optical axis outward deflection 40°.

[0102] Topology parameters: master-slave node spacing 500m, sensing overlap area ≥20%. Verified through field laser ranging, targets at any location are captured in at least two different sensor fields of view.

[0103] 2.2 Implementation Example of Curved Segment Deployment

[0104] The curve segment has a radius R = 800m, a central angle θ = 60°, and a length of approximately 837m. A detection blind zone is easily formed on the outer side (central side), requiring denser LiDAR deployment (300m spacing). A second set of IR cameras (at a 90° angle with the main IR camera) is added to the outer node to ensure lateral coverage. The LiDAR installation angle is tilted outwards by 15°. Simulation and actual measurements show that the single-sided detection distance increases from 200m to 280m, and the blind zone width is reduced from 12m to 3m.

[0105] 2.3 Tunnel Section Deployment Example

[0106] The tunnel is 1.2 km long, with an illumination intensity of ≤10 lux. Nighttime reliance on visible light cameras will be eliminated, and all cameras will be replaced with IR cameras (dual IR cameras per slave node, with cross-axis optical coverage), using low thermal sensitivity models (≤30 mK). The effective range of the lidar inside the tunnel will decrease to 80 m, necessitating an increase in the density of master nodes to one per 1 km to ensure continuous coverage.

[0107] 2.4 Environment Adaptation Deployment Principles

[0108] High-altitude and cold regions (-40℃): Industrial-grade wide-temperature sensor is selected, the outer shell is heated to remove fog, and dynamic compensation is provided for focal length and attitude.

[0109] High humidity and foggy areas: Increase millimeter-wave radar weights and use RAF-Net channel attention to suppress LiDAR noise.

[0110] Dust-affected areas: Add air filters and self-cleaning lenses to improve image usability.

[0111] 3. Detailed Implementation of the Spatiotemporal Synchronization Module

[0112] 3.1 Hardware Trigger Design

[0113] The master node LiDAR's internal FPGA generates periodic pulses (10Hz, pulse width 10μs±1μs, TTL high level), which are transmitted to the slave node sensor trigger ports via RG-316 coaxial shielded cables. The transmission delay, measured by a TDR (time domain reflectometer), is 0.8μs±0.1μs, and the jitter is ≤0.2μs. Upon receiving the trigger signal, each sensor immediately starts sampling, ensuring a consistent start time for data acquisition.

[0114] 3.2 PTP Time Synchronization

[0115] The master node is configured with an OCXO temperature-controlled crystal oscillator (stability ±1ppb) and acts as the PTP Grandmaster. Hardware timestamping is achieved via a Gigabit Ethernet PHY supporting IEEE 1588 v2 (such as an Intel I210). Slave nodes act as slaves, with transparent clock correction for switching latency. The network-wide synchronization error is tested with a Fluke 9100A and is ≤5μs. During initialization, a network-wide latency measurement is performed, generating a compensation table for dynamic correction during runtime.

[0116] 3.3 Spatial Coordinate Unification

[0117] Define a world coordinate system O-XYZ: the origin is at the geometric center of the master node LiDAR, the X-axis is along the positive direction of the railway, the Y-axis is to the left, and the Z-axis is upward. The intrinsic parameters of each sensor are known from the factory; the extrinsic parameter matrix T_ext (4×4 homogeneous transformation) is obtained through field calibration. Point cloud and image features are mapped to a unified coordinate system through T_ext, with a transformation error ≤2cm (1.2cm measured by the laser tracker).

[0118] 4. Refinement of RAF-Net fusion network training and inference

[0119] 4.1 Dataset Construction

[0120] 100,000 samples were collected, covering scenarios such as sunny, rainy, foggy, nighttime, and tunnels. Target categories included trains, pedestrians, falling rocks, and animals. Each sample contained:

[0121] LiDAR point cloud: average 1.2×10 5 Points were downsampled to 8192 points.

[0122] MMW Radar: Range-Doppler graph 64×128.

[0123] VIS / IR images: 3840×2160 / 640×512.

[0124] The annotation process uses a combination of semi-automatic and manual refinement, with a 3D frame error of ≤0.1m.

[0125] 4.2 Network Structure and Training

[0126] Encoder: PointNet++ three-layer SA (radius 0.5m, 1.0m, 2.0m) → 256-dimensional features; Radar graph convolution → 128-dimensional; ResNet-50 first four residual blocks → 512-dimensional.

[0127] Attention Fusion: SE Block (r=16) → CBAM Spatial Attention.

[0128] Decoder: Decouples the head to regress 3D box, speed, and category.

[0129] Loss function: SmoothL1 + Focal Loss + MSE, weight ratio 4:3:3.

[0130] Optimization: Adam, lr=0.001, cosine annealing, batch=16, convergence in 200 epochs.

[0131] 4.3 Inference Optimization

[0132] TensorRT INT8 quantization, layer fusion, FP16 auxiliary, Jetson platform inference 25fps, single frame 40ms.

[0133] 5. Detailed implementation of the installation and calibration subsystem

[0134] 5.1 Offline Calibration

[0135] The laboratory was equipped with a standard track (gauge 1435mm ± 0.5mm), 5 targets spaced 10m apart, and total station coordinate measurements were performed with a tolerance of ± 0.5mm. Ten sets of data were collected, with an ICP registration error of 0.3mm, PnP used to calculate intrinsic and extrinsic parameters, and a reprojection error of 0.08 pixels.

[0136] 5.2 Online self-calibration

[0137] Vibration caused a 0.2° shift in the LiDAR attitude angle, increasing the orbit feature point matching error from 0.5cm to 2.1cm. After 10 SLAM optimizations, the residual was reduced to 0.4cm, and calibration was completed in 8s.

[0138] 5.3 Dynamic Compensation

[0139] At -40℃, the camera focal length shifted by 0.1mm, and the temperature model was compensated by +0.1mm, reducing the positioning error from 0.45m to 0.12m.

[0140] 6. Refinement of the application of collaborative decision-making platforms

[0141] DeepSORT tracking, ID switching rate <5%; risk index R=0.6D+0.3V+0.1C, R>0.8 Level 1 warning; 5G push latency <100ms.

[0142] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A multi-sensor collaborative sensing system integrating radar and visual perception along a railway line, characterized in that: It comprises the following functional modules and hardware components, working together to achieve all-weather, high-precision, low-latency environmental and intrusion target perception along the railway line, and autonomously completing online calibration and dynamic compensation of sensor installation parameters during operation: Multimodal sensor array: Distributed along the railway line according to a preset spatial topology and scene adaptation principle, including at least a LiDAR (preferably a 32-line or 64-line mechanical / solid-state radar, with a horizontal field of view of not less than 90° and not more than 150°, a vertical field of view of not less than 20° and not more than 40°, a maximum ranging distance of not less than 200m, and a ranging accuracy of ≤±2cm), a millimeter-wave radar (MMW Radar, operating frequency band of 77GHz or 79GHz, horizontal field of view of not less than 60° and not more than 120°, a maximum detection distance of not less than 300m, a speed measurement range covering -200km / h to +400km / h, and a speed resolution of ≤0.1m / s), a visible light camera (VIS Camera, resolution of not less than 1920×1080 and preferably 3840×2160, frame rate of not less than 25fps and preferably 30fps, equipped with automatic exposure and strong light suppression functions, and a lens focal length adjustable between 6mm and 25mm), and an infrared thermal imager (IR). The camera has a resolution of no less than 320×240, preferably 640×512, a frame rate of no less than 20fps, preferably 25fps, a thermal sensitivity of ≤50mK, a temperature measurement range covering -40℃ to +150℃, and a lens focal length preferably 8mm to 16mm; the sensing coverage area of ​​each sensor is between 50m and 500m along the railway line, and the sensing range between any two adjacent sensors has an overlap area of ​​no less than 20% to ensure that the target is not lost when crossing the sensing boundary; Spatiotemporal synchronization module: It adopts a dual mechanism of hardware trigger signal and software timestamp correction. The hardware trigger is achieved by the lidar of the master node emitting periodic pulses (pulse width 10μs±1μs, voltage TTL high level) which are directly connected to the external trigger input terminals of millimeter-wave radar, visible light camera and infrared thermal imager through shielded cables to achieve microsecond-level consistency at the start time of data acquisition. At the software level, based on the PTP (Precision Time Protocol, IEEE 1588 v2) precision clock protocol, the master node, acting as the Grandmaster, broadcasts synchronization messages to all slave nodes via Gigabit Ethernet. Combined with a transparent clock to correct the switching delay, the time synchronization error of the entire network does not exceed 10μs. The spatial coordinates are uniformly based on the right-handed Cartesian world coordinate system, with the geometric center of the master node's lidar as the origin O(0,0,0). The X-axis points in the positive direction of train operation along the railway line, the Y-axis is perpendicular to the X-axis and points to the left side of the track, and the Z-axis is vertically upward. All sensors map the raw data to this unified coordinate system through the extrinsic parameter matrix calibrated at the factory or on-site. The overall spatial transformation error does not exceed 2cm. Edge Fusion Processing Unit: Equipped with an attention-based radar-visual feature fusion network (RAF-Net), this network runs in real-time on edge computing nodes (hardware platforms include NVIDIA Jetson AGX Xavier, Xilinx Zynq UltraScale+ MPSoC, or Huawei Atlas 500). It performs cross-modal feature extraction and fusion on multi-source data output from the spatiotemporal synchronization module, including: ① Cross-modal feature encoders use the PointNet++ architecture to process LiDAR point clouds to extract 3D geometric features (point cloud downsampled to within 8192 points, feature dimension 256), a Range-Doppler graph convolutional network to process millimeter-wave radar echoes to extract radial velocity and Doppler features (feature dimension 128), and a ResNet-50 backbone network to process visible light and infrared images to extract texture, edge, and thermal radiation features (feature dimension 512); ② The attention fusion layer introduces channel attention (Squeeze-and-Excitation Block) and spatial attention (Convolutional Block Attention). The module mechanism dynamically assigns weights to different modal feature channels and spatial locations, suppressing invalid or noisy features in interference scenarios such as rain, fog, backlight, and high temperature difference backgrounds, and enhancing the response of target saliency features; ③ The multi-task decoder outputs in parallel three-dimensional target detection boxes with confidence (center point coordinates, length, width, height, heading angle, positioning error within 50m range not exceeding 0.3m), three-dimensional velocity vectors (vx, vy, vz) (velocity estimation error ≤ 0.2m / s), and target category probability distribution (categories cover trains, locomotives and rolling stock, pedestrians, bicycles, falling rocks, animals, and large floating objects, with category recognition accuracy ≥ 98%). The installation and calibration subsystem consists of an offline calibration device, an online self-calibration module, and a dynamic compensation algorithm. It achieves automated and periodic optimization of sensor extrinsic parameters (installation positions x, y, z and attitude angles α, β, γ). The offline calibration device includes a high-precision calibration target (composed of a high-reflectivity pyramidal prism array and thermochromic material; the target's geometry is known and it is placed at several control points along the railway line with known coordinates, such as kilometer markers and turnout reference points). An automatic calibration program drives the sensor array to scan the target from multiple angles. The Iterative Closest Point (ICP) algorithm is used to match the lidar point cloud with the three-dimensional coordinates of the control points (registration error ≤ 0.3 mm). The perspective n-point (PnP) algorithm is combined to solve for the camera's intrinsic parameters (focal length, principal point, distortion coefficient) and extrinsic parameters (position and attitude relative to the lidar), achieving sub-pixel-level calibration accuracy (reprojection error ≤ 0.1 pixel). The online self-calibration module is part of the system... During operation, the sensor pose changes are continuously monitored. A reference map is constructed based on the fixed geometric constraints of the railway track (such as track gauge 1435mm±0.5mm, sleeper spacing 600mm±10mm, and parallelism error between the two rails ≤2mm / m). When vibration or temperature drift is detected and the deviation of the external parameters exceeds the set threshold (position deviation ≥1cm or attitude angle deviation ≥0.05°), the incremental SLAM process is automatically started. The external parameter matrix is ​​updated by point cloud registration and visual odometry fusion, and the calibration cycle is no more than 10s. The dynamic compensation algorithm collects data from the built-in temperature sensor (accuracy ±0.5℃) and the triaxial accelerometer (range ±20g, resolution ≤0.01g) in real time, establishes temperature-optical focal length offset model and vibration-attitude angle offset model, and uses extended Kalman filter to predict and compensate for parameter drift, so that the target positioning error does not exceed 0.1m in the entire working temperature range (-40℃~+85℃). Collaborative Decision-Making Platform: Deployed in a collaborative architecture of cloud servers and edge nodes, it receives real-time detection results from edge fusion processing units and combines them with the railway operation rule base (including train operation plans, section speed limits, construction sections, and temporary speed limit orders) to perform multi-target tracking (using DeepSORT or ByteTrack algorithms, with a target ID switching rate ≤5%), risk level assessment (calculating a risk index R based on the closest distance between the target and the track, speed, and category weight; R>0.8 triggers a level one warning) and generate warning instructions. Warning information is pushed to the dispatch center's large-screen display terminal, the train's onboard ATP / ATO system, and mobile terminal APP via 5G NR or Beidou short message wireless communication modules, with a push delay of no more than 100ms.

2. The railway line radar-visual fusion multi-sensor collaborative sensing system according to claim 1, characterized in that: The multimodal sensor array adopts a hierarchical distributed topology of "master node-slave node": master nodes are set at catenary support posts or dedicated monitoring towers every 2km along the railway line, with an installation height between 5m and 8m. They integrate a high-line-count lidar, a long-range millimeter-wave radar, and an edge computing node to form a core for local perception and fusion processing. Slave nodes are deployed at 500m ± 50m intervals on railway guardrail posts or independent supports, with an installation height between 1.2m and 2m. Each slave node is equipped with a set of visible light cameras and infrared thermal imagers, facing the outside and inside of the track at certain angles (e.g., the optical axis of the outside camera is at 30° to 60° with the vertical line of the track, and the inside camera is at -30° to -60°) to expand lateral coverage. The master and slave nodes are interconnected through a single-mode fiber optic ring network with a link bandwidth of not less than 1Gbps. The transmission protocol adopts TSN (Time Sensitive Networking) to ensure data real-time performance. The network topology supports redundant path switching, and a single point of failure does not affect the overall perception continuity.

3. A multi-sensor collaborative sensing system for railway line radar-visual fusion according to claim 1, characterized in that: The spatiotemporal synchronization module adopts a dual synchronization mechanism of "PTP precision clock protocol + hardware trigger": the master node is configured with a temperature-controlled crystal oscillator (OCXO) as the PTP Grandmaster clock source, with a frequency stability of ≤±1ppb. The hardware timestamp function of the Ethernet PHY chip records the time of sending and receiving messages, and the synchronization error of the slave nodes is measured to be no greater than 5μs. The hardware trigger signal is generated by the FPGA control inside the master node's LiDAR. The trigger period is consistent with the LiDAR scanning frequency (e.g., 10Hz or 20Hz), and the trigger edge jitter is ≤0.2μs. The signal is transmitted to each slave node sensor through an impedance-matched shielded cable, and the transmission delay is controlled to within 1μs by oscilloscope calibration. During the system initialization phase, a full network time deviation and trigger delay calibration is performed to generate a compensation table for dynamic correction during operation.

4. A multi-sensor collaborative sensing system for railway line radar-visual fusion according to claim 1, characterized in that: The specific structure and training method of the RAF-Net network include: the LiDAR branch of the cross-modal feature encoder uses PointNet++ for hierarchical sampling and grouped feature extraction (3 levels of Set Abstraction layers, with ball query radii of 0.5m, 1.0m, and 2.0m respectively), outputting a global feature vector of 256 dimensions; the millimeter-wave radar branch first normalizes the Range-Doppler image, and then extracts motion features through two layers of convolution (Conv 3×3, stride=1, channels=[64,128]) and pooling; the visible light and infrared image branches share the first four residual blocks of the ResNet-50 backbone network, and then use bilinear interpolation to unify the feature map resolution to 1 / 8 of the input size, outputting 512-dimensional features; the attention fusion layer first uses SE Block to weight the channel importance (reduction ratio r=16), and then uses CBAM to enhance the response of the target area in the spatial dimension; the multi-task decoder adopts a decoupled head structure, regressing the 3D detection box parameters, velocity vector, and class probability respectively, and the loss function consists of SmoothL1 position regression loss and Focal loss. The classification loss and the velocity regression loss of MSE are weighted and summed (weight ratio 4:3:3). The training process uses the Adam optimizer with an initial learning rate of 0.001, combined with a cosine annealing strategy, a batch size of 16, and trains for 200 epochs until convergence. In the inference stage, TensorRT INT8 quantization and layer fusion optimization are used to achieve a single frame processing time of ≤40ms.

5. A multi-sensor collaborative sensing system for railway line radar-visual fusion according to claim 1, characterized in that: The offline calibration device of the installation calibration subsystem also includes: a target bracket with adjustable height and pitch angle to adapt to the sensor viewing angle at different installation heights; the target surface is coated with a high reflectivity material and a thermochromic coating, which has high recognition in both visible and infrared bands; the automatic calibration program controls the sensor array to collect target data according to a preset scanning trajectory (such as horizontal rotation ±30°, pitch ±15°), and calls the PCL (Point Cloud Library) and OpenCV libraries to implement ICP and PnP solutions respectively; after calibration, an encrypted calibration file (including intrinsic parameter matrix, distortion coefficient, extrinsic parameter matrix, timestamp and check code) is generated, stored in the non-volatile memory of the edge node, and simultaneously uploaded to the cloud for backup.

6. A multi-sensor collaborative sensing system for railway line radar-visual fusion according to claim 1, characterized in that: The implementation steps of the online self-calibration module include: (1) extracting track plane features from the lidar point cloud in real time (using RANSAC to fit the equations of the ground and track surface); (2) extracting track fastener and sleeper corner features from visible light and infrared images (using Harris corner detection and template matching); (3) constructing the reprojection error function E=∑‖p_i^img - π(T·P_i^lidar)‖², where p_i^img is the homogeneous coordinate of the image feature point, P_i^lidar is the corresponding point cloud coordinate, T is the external parameter matrix to be optimized, and π is the pinhole projection model; (4) using the Levenberg-Marquardt nonlinear least squares algorithm to iteratively optimize T, with the iteration termination condition being the error change rate <1e-6 or the number of iterations >50; (5) updating the external parameters and writing them into the configuration file for real-time calling by the fusion network.

7. A multi-sensor collaborative sensing system for railway line radar-visual fusion according to claim 1, characterized in that: The dynamic compensation algorithm further includes: establishing a temperature-focal length offset lookup table (based on focal length changes sampled every 5℃ in the range of -40℃ to +85℃ measured by laboratory temperature control experiments), and obtaining compensation values ​​by linear interpolation based on temperature sensor readings during operation; establishing a vibration-attitude angle offset model (based on accelerometer integration and low-pass filtering to estimate instantaneous attitude changes), fusing predicted and measured values ​​through Kalman filtering, and outputting the compensated attitude angle for point cloud and image coordinate transformation; after compensation, the target positioning error under any working condition is ≤0.1m, and the velocity estimation error is ≤0.15m / s.

8. An installation and calibration method for a multi-sensor collaborative sensing system integrating radar and visual perception along a railway line, characterized in that: The process includes the following steps, each of which is precisely controlled by taking into account the unique environment of the railway and the physical characteristics of the sensors: S1: Offline pre-calibration: In a controlled laboratory environment, using the high-precision calibration target and automatic calibration program described in claim 5, the intrinsic parameters (camera distortion coefficients k1, k2, p1, p2, k3, lidar scanning angle calibration factor) and extrinsic parameters (installation position x, y, z and Euler angles α, β, γ) of the sensor array are initially calibrated, an initial calibration file is generated, and the reprojection error is verified to be ≤0.1 pixel and the point cloud registration error to be ≤0.3 mm. S2: On-site coarse calibration: After installing the sensor at the predetermined position along the railway line, use a GNSS-RTK receiver to obtain the absolute geographic coordinates of the sensor antenna phase center (horizontal accuracy ±2cm, elevation accuracy ±3cm). Combine the coordinates of the track centerline in the railway line design CAD drawings to calculate the theoretical deviation of the installation position and make mechanical adjustments to make the position deviation ≤1cm. S3: Online fine calibration: Activate the online self-calibration module as described in claim 6, continuously collect no less than 100 frames of sensor data, extract and match track geometric features and sensing features, perform external parameter optimization until the external parameter change in two adjacent iterations meets the requirements of position change <1cm and attitude angle change <0.01°, and the calibration cycle ≤10s; S4: Dynamic compensation verification: Under train operation or artificial disturbance conditions (such as artificially shifting the sensor angle by 0.2° or changing the ambient temperature by 20°), run the dynamic compensation algorithm described in claim 7 to verify that the target detection recall rate decreases by ≤2%, the positioning error increases by ≤0.05m, and the warning delay increases by ≤50ms. S5: Regular maintenance and calibration: Every 3 months, the cloud-based operation and maintenance platform issues a calibration instruction, automatically executes steps S3 to S4, updates the calibration file and synchronizes it to the cloud database and all edge nodes, forming a closed-loop calibration archive.

9. The installation and calibration method for a multi-sensor collaborative sensing system along a railway line according to claim 8, characterized in that: The online fine calibration described in step S3 further includes: constructing a multi-scale feature matching strategy, using voxel grid downsampling to accelerate ICP matching in the point cloud layer, and using ORB features and optical flow tracing to improve feature point stability in the image layer; adding a track geometry constraint penalty term to the optimization process to prevent overfitting in feature-scarce regions (such as inside tunnels); and using a sliding window method to maintain the mean of the optimization results of the most recent N=50 frames to suppress the influence of instantaneous noise on external parameters.

10. The installation and calibration method for a multi-sensor collaborative sensing system along a railway line according to claim 8, characterized in that: The dynamic compensation verification in step S4 also includes: conducting continuous testing for no less than 24 hours under different environmental conditions (rain, snow, fog, dust, night), statistically analyzing the detection performance indicators under each condition, and generating an environmental adaptability and compensation effectiveness report for subsequent model iteration and parameter fine-tuning.

Citation Information

Patent Citations

  • Indoor carrier positioning system, map data setting method and nursing carrier positioning method

    CN111854749A

Cited By

  • Parallel bus ultrasonic array high-precision synchronization method and system

    CN122195212A

  • Unified Method for Pixel-Level Reference of Four-Source Sensors in Short-Range UAVs (Electro-optical-radar pods)

    CN122307584A