Light rail vehicle safety monitoring system based on multi-sensor fusion
Through multi-sensor fusion technology, light rail environment images are captured in real time and point cloud data are generated. Combined with optical flow method and CNN deformation detection network, the problem of insufficient dynamic object recognition accuracy in the light rail vehicle monitoring system is solved, and high-precision dynamic object recognition and hierarchical early warning is achieved to ensure the safety of light rail vehicles.
Patent Information
- Application Number
- CN202510796262.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-16
AI Technical Summary
The existing light rail vehicle monitoring system lacks real-time response capabilities, making it difficult to effectively identify and distinguish dynamic objects, especially low-altitude threats such as drones, and the recognition accuracy is insufficient in complex environments and the false alarm rate is high.
Multi-sensor fusion technology is adopted to capture environmental images in real time and generate point cloud data through multi-modal perception modules. It combines optical flow method and CNN deformation detection network to calculate object deformation, uses Doppler compensation algorithm to obtain dynamic object distance, and makes multi-modal decisions through D-S evidence theory to achieve high-precision dynamic object recognition and hierarchical early warning.
It realizes high-precision identification and hierarchical decision-making of dynamic objects, reduces the false alarm rate, and ensures the safety of light rail vehicles, especially effectively monitoring threats such as drones in high-speed and complex environments.
Smart Images

Figure CN120339989A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of light rail vehicle safety monitoring, and specifically to a light rail vehicle safety monitoring system based on multi-sensor fusion. Background Art
[0002] As the core of urban public transportation, the safety of light rail is directly related to the lives of passengers and the operation order of the city. With the acceleration of the urbanization process, the passenger flow carried by the light rail system continues to grow, becoming an important link connecting various regions of the city. However, behind the efficient operation of the light rail lurk complex safety challenges. According to the analysis of historical accident cases, the potential safety hazards of light rail mainly stem from factors such as equipment failures, human operation errors, natural disasters, and external interferences. For example, in 2021, a drone flew at ultra-low altitude from the Fotuguan section to the Daping section of Chongqing Rail Transit Line 2 and collided with a high-speed train, resulting in an emergency braking of the train. Although no casualties were caused, the incident exposed the serious threat of unapproved drone flights (illegal flights) to rail operations. According to the "Regulations on the Operation Management of Urban Rail Transit", it is illegal to fly drones within 100 meters on both sides of the track. However, due to the lack of real-time monitoring means, similar incidents continue to occur despite repeated bans.
[0003] Existing technologies mostly rely on manual inspections and basic sensors, with problems such as limited coverage and reliance on empirical judgments, making it difficult to respond to sudden risks in real time. Traditional methods mostly use single sensors, lacking the fusion analysis of multi-modal data, resulting in insufficient recognition accuracy of dynamic objects and a high false alarm rate. In recent years, the frequent occurrence of unapproved drone flight incidents has exposed the weak real-time capture ability of existing monitoring systems for low-altitude dynamic targets and the lack of adaptability to complex environments. Summary of the Invention
[0004] To achieve the above objectives, the present invention is realized through the following technical solutions: A multi-modal perception module that captures images of the environment around the light rail in real time, demarcates visual marker areas, performs three-dimensional scanning on the visual marker areas, generates environmental point cloud data in real time, and uses the Doppler compensation algorithm to obtain the distances of dynamic objects. A data processing module that uses the optical flow method and the CNN deformation detection network to calculate the deformation of objects between adjacent frames, calculates the relative motion of objects based on the object deformation analysis, outputs the object contours based on the instance segmentation of MaskR-CNN, and at the same time uses the dynamic target clustering algorithm to perform multi-target tracking on the object contours, and combines the environmental point cloud data to output the three-dimensional coordinates and motion vectors of the objects. The fusion and analysis module performs time synchronization based on the PTP protocol, uses the checkerboard calibration method for laser vision joint calibration, and makes multi-modal decisions using the D-S evidence theory. It outputs a fusion decision by inputting evidence sources, where the evidence sources include the confidence of deformation analysis, the reliability of laser ranging, and the probability of the classification network. The control and response module obtains the fusion decision and conducts a hierarchical early warning mechanism and response.
[0005] Further, the process of delimiting the visual marker area is as follows: The median filtering and bilateral filtering combined algorithm is used to eliminate high-frequency noise in the environmental image. Adaptive histogram equalization is performed based on the Retinex theory. The lens distortion parameters are pre-calibrated using the calibration method, and the offset is corrected by online calibration. The adaptive threshold segmentation algorithm is used to delimit the visual marker area.
[0006] Further, the process of real-time generating environmental point cloud data is as follows: The LiDAR emits laser pulses simultaneously at different vertical angles through a multi-beam laser emitter. Each beam is spaced at a certain angle, and 360° coverage in the horizontal direction is achieved through a rotating mirror or electronic scanning method. Combining with the pose data of the IMU, the three-dimensional coordinates of each point are calculated.
[0007] Further, the process of obtaining the distance of dynamic objects is as follows: Within a single-frame scanning period, the motion trajectory is calculated by numerically integrating the data of the IMU and the wheel odometer, and after achieving microsecond-level time synchronization of the laser points and sensor data by combining with the PTP protocol, the pre-integration result of the IMU is distributed to each laser point by linear interpolation to compensate for motion distortion. Using the Doppler compensation algorithm, the relative velocity and distance of dynamic objects are calculated by analyzing the frequency offset of the laser echo signal. The reflected signal is collected at the receiving end, and the time difference between transmission and reception and the phase change of the signal are recorded. The echo signal is spectrally analyzed by fast Fourier transform to extract the frequency shift amount. The target velocity is calculated based on the frequency shift amount, and the Kalman filter is used to track the velocity change in real time.
[0008] Further, the process of calculating the deformation of objects between adjacent frames is as follows: S201: Input two consecutive light rail environmental images; analyze the displacement of all pixel points using the optical flow algorithm; output the optical flow field map. S202: Refine the optical flow field map obtained by the optical flow method. By inputting adjacent frames, the feature map is extracted through a pre-trained convolutional neural network. S203: Use a specific network structure to compare the feature differences between adjacent frames. The network outputs the deformation ratio of the key areas of the object in each frame, and combines with the optical flow result for dominant fusion.
[0009] Further, the process of calculating the relative motion of the object is as follows: For each detected object region, calculate the deformation ratio variance between adjacent frames, set a motion threshold, and make a motion judgment: If the deformation ratio variance ≤ the motion threshold, it is considered that the object is relatively stationary; If the deformation ratio variance > the motion threshold, it is marked as a dynamic object; At the same time, combine the motion vectors in the optical flow field to calculate the relative velocity and motion direction of the object.
[0010] Further, the process of outputting the object contour, as well as the 3D coordinates and motion vectors of the object is as follows: Input the environmental image and the dynamic object candidate area, and generate candidate boxes through the RPN module of Faster R-CNN; Based on the dynamic target clustering algorithm, input the object contour data of adjacent frames, extract the center point coordinates, velocity vectors and shape features of each object; perform clustering grouping, and use the DBSCAN algorithm to cluster the targets: according to the spatial distance and velocity similarity, divide the objects in consecutive frames into the same trajectory cluster; perform Kalman filtering on the clustering results to associate the trajectories and predict the target positions in the next frame; For the clustered target trajectories and the lidar point cloud data, calibrate the internal and external parameters of the camera and the lidar through a deep learning-based calibration algorithm, Match the 2D center point of the object contour with the 3D point cloud in the point cloud data through the perspective projection model; use the calibration parameters and the projection formula to calculate the 3D coordinates of the object; combine the time series of the point cloud data to calculate the three-dimensional velocity vector of the object, evaluate its motion direction and threat level; output a dynamic object list containing the category, 2D contour mask, 3D coordinates and motion vectors of each object, as well as the trajectory information of the object in consecutive frames.
[0011] Further, the process of the laser-vision joint calibration is as follows: Use the corner points of the checkerboard calibration board as 2D / 3D feature points, calculate the camera internal parameters through the Zhang Zhengyou calibration method, and determine the camera external parameters based on multi-view images; obtain the point cloud data by scanning the checkerboard with the lidar, and reverse-infer its installation pitch angle θ and height h; convert the point cloud to the camera coordinate system, and construct the rotation and translation matrix from the lidar to the camera; optimize the matrix parameters by minimizing the projection error between the point cloud and the image corner points; project the point cloud onto the image plane to verify the coincidence degree to achieve multi-modal spatial alignment.
[0012] Further, the process of outputting the fusion decision is as follows: By converting the CNN deformation confidence, laser ranging stability, and MaskR-CNN classification probability into a BPA function and calculating the conflict factor K, and using the Dempster rule to synthesize evidence to generate a comprehensive confidence level, hierarchical decision-making for high-threat, potential-threat, and safe objects is achieved.
[0013] Furthermore, the process of the hierarchical early warning mechanism and response is as follows: Dynamically adjust the threat threshold and optimize the mapping rule based on historical data; monitor the fusion decision result in real time and conduct regulation.
[0014] The light rail vehicle safety monitoring system based on multi-sensor fusion provided by the present invention has the following beneficial effects: (1) The present invention realizes the spatial alignment of LiDAR and camera through the checkerboard calibration method, combines the Doppler compensation algorithm to eliminate motion distortion, ensures that the calculation accuracy of the three-dimensional coordinates and motion vectors of dynamic objects, such as drones, reaches the centimeter level; uses the optical flow algorithm to analyze the pixel-level motion vector, combines the pre-trained CNN network to extract feature differences, accurately identifies the object deformation ratio, and distinguishes the static background from the dynamic threat; by converting the CNN deformation confidence, laser ranging stability, and MaskR-CNN classification probability into a BPA function, and calculating the conflict factor K, and using the Dempster rule to synthesize evidence, a comprehensive confidence level is generated; this method can effectively reduce the false alarm rate and achieve hierarchical decision-making for high-threat objects, such as drones approaching the track.
[0015] (2) The present invention optimizes the threat level division according to the environment to avoid misjudgment caused by sensor performance fluctuations; adopts a combined algorithm of median filtering and bilateral filtering to eliminate high-frequency noise, combines CLAHE adaptive histogram equalization to improve the image quality in low-visibility scenarios such as tunnels, rain and fog, and night; realizes microsecond-level time synchronization based on the PTP protocol, combines IMU for motion distortion compensation, and ensures the spatio-temporal consistency of point cloud data in high-speed scenarios; processes data in real time through an embedded system, such as an ARM processor, and combines a cloud computing platform to realize historical data learning and model iteration, and continuously optimize the threat recognition rule. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a schematic diagram of the system flow of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0018] Example 1 Please refer to Figure 1 , Example 1 of this application provides a light rail vehicle safety monitoring system based on multi-sensor fusion. The system includes: A multi-modal perception module that captures the surrounding environment image of the light rail in real time, delimits the visual marking area; performs three-dimensional scanning on the visual marking area, generates environmental point cloud data in real time, and uses the Doppler compensation algorithm to obtain the distance of dynamic objects; Capturing the surrounding environment image of the light rail in real time: Multiple groups of high-resolution industrial cameras are deployed at the front and flanks of the light rail, covering the forward track and the areas beside the tracks on both sides; Forward main camera: Equipped with optical image stabilization and a long focal length lens, the focal length ≥ 50mm, supports the HDR mode, and adapts to sudden changes in light inside and outside the tunnel; Lateral wide-angle camera: The horizontal field of view angle ≥ 120°, used to capture dynamic targets such as pedestrians and vehicles beside the track; Infrared auxiliary camera: Integrated with a non-cooled infrared detector, with a wavelength of 8 - 14μm, compensating for rain, fog, and low visibility at night; Among them, multi-camera μs-level time synchronization is achieved through PTP to ensure the spatio-temporal consistency of environmental data at the same moment, and the trigger frequency is dynamically adjusted according to the vehicle speed; A combined algorithm of median filtering and bilateral filtering is used to eliminate high-frequency noises such as raindrops and snowflakes; adaptive histogram equalization is performed based on the Retinex theory to solve the problems of strong backlighting at the tunnel entrance and low illuminance inside the tunnel. The lens distortion parameters are pre-calibrated using the calibration method, and the offset caused by mechanical vibration is corrected by combining online calibration; For example: When the light rail is running in the rain, the camera is full of small dot-like high-frequency noises. Median filtering is used to remove these randomly distributed small noise points. Median filtering effectively removes isolated noise points without affecting edge information by replacing each pixel value with the median of all pixel values in its neighborhood window. However, using median filtering alone may not be able to completely eliminate all types of noises, especially when the noise is relatively dense; and further apply bilateral filtering. Bilateral filtering not only considers the weight in terms of spatial distance but also the weight of pixel value differences, which means it can retain important edge features while smoothing the image. Even in the case of dense raindrops or snowflakes, it can ensure the clarity of the key object contours in the image, providing high-quality input for subsequent object detection and classification; Delimiting the visual marking area: The surrounding environment image of the light rail is captured in real time through a high-resolution camera, and an adaptive threshold segmentation algorithm is used to delimit the visual marking area, focusing on key areas such as along the track, signal equipment, and dynamic objects such as drones and pedestrians; Preprocess the environmental image using median filtering or Gaussian filtering to eliminate salt-and-pepper noise and Gaussian noise caused by vibration and light changes. For example, perform smoothing processing through a 5×5 Gaussian kernel to retain the edge information of the track along the line and the UAV target. To address the problem of uneven illumination during light rail operation, use histogram equalization or gray-scale stretching to enhance the image contrast and improve the robustness of subsequent threshold segmentation. For example, at the tunnel entrance, use the CLAHE algorithm, contrast-limited adaptive histogram equalization, to enhance the details in the dark area. Divide the image into local regions of N×N, such as 11×11 or 15×15 pixel blocks. Calculate the average gray value of the pixels in the neighborhood as the reference threshold through the adaptiveThreshold function in OpenCV. Use a Gaussian kernel (σ = 0.5) to perform weighted summation of the neighborhood pixels to reduce the edge blurring effect. Introduce a compensation factor C (usually taking values from 2 to 10) based on the reference threshold. The formula is: Formula 1: Formula 2: In the formula, is the dynamic threshold at position , which is used for tasks such as image segmentation or feature detection to distinguish the foreground and background; is the local average gray value at position , which is obtained by calculating the average value of the pixel gray values within a certain area around this point; the local average gray value reflects the image brightness characteristics near this position; is the compensation factor, which is a constant, usually with a value range of 2 to 10. It dynamically adjusts the threshold according to the scene to adapt to different lighting conditions. When the lighting is strong, increasing the value of C can suppress background interference. When the lighting is low, decreasing the value of C can avoid missing the target detection; is the Gaussian kernel function, where is the standard deviation, which is used to control the smoothing degree of the Gaussian filter. Gaussian filtering is a commonly used image smoothing technique that reduces image noise and retains edge information; is the convolution operator, indicating that the Gaussian kernel function is convolved with the image ; the result of the convolution is the result of smoothing the image; is the gray value of the original image at position ; Among them, C is dynamically adjusted according to the scene: when the lighting is strong, increase the value of C to suppress background interference, and when the lighting is low, decrease the value of C to avoid missing the target detection; Compare each pixel gray value with its corresponding threshold to generate a binary image: If the pixel value ≥ the threshold, mark it as the foreground, visually mark the candidate region, and assign a value of 255; If the pixel value < threshold, mark it as the background and assign 0; Perform a closing operation on the segmentation result, first dilate and then erode to fill holes, and eliminate small-area noise regions through connected component analysis, such as isolated points with an area < 50 pixels, and finally output the contour of the stable visual marking region; Perform a 3D scan on the visual marking region: Utilize the multi-beam scanning technology of LiDAR to generate high-precision environmental point cloud data in real time; through the motion distortion compensation algorithm, based on the incremental motion model of IMU / wheel odometer, eliminate the point cloud deformation error caused by the high-speed movement of the light rail; Multi-beam laser scanning and point cloud generation: LiDAR emits laser pulses simultaneously at different vertical angles, such as from -15° to +15°, through multi-beams, such as 16-line, 32-line, and 128-line laser emitters; for example, the Velodyne VLP-16 lidar covers 30° in the vertical direction, with a 2° interval between each beam; achieve 360° coverage in the horizontal direction through a rotating mirror or electronic scanning method, and the single-frame scanning time is about 0.1 second; after the laser beam hits an object and reflects, the receiver records the echo time difference and intensity information, and combines the pose data of IMU and wheel odometer to calculate the three-dimensional coordinates of each point; the original data is stored in LAS, XYZ, or PCD format after parsing, containing position, reflection intensity, and timestamp information; for example, binary data packets transmitted through the UDP protocol need to be decoded into a processable point cloud structure; Flow of the motion distortion compensation algorithm: Within a single-frame scanning period, utilize the IMU accelerometer and gyroscope data, combined with the wheel odometer pulse count, to calculate the incremental motion trajectory, position, attitude, and velocity through numerical integration; assign an accurate timestamp to each laser point based on the scanning angle or hardware trigger, and achieve microsecond-level synchronization with the IMU / wheel odometer data through the PTP protocol; use linear interpolation or spherical linear interpolation to distribute the IMU pre-integration result to each laser point according to the timestamp; Convert the local coordinate system of each laser point, i.e., the LiDAR coordinate system, to the global coordinate system; adopt factor graph optimization, such as the GTSAM framework, to jointly optimize the LiDAR point cloud matching residuals, IMU pre-integration constraints, and wheel odometer errors, and minimize the joint error of the global pose and motion distortion; Obtaining the distance of dynamic objects: Combined with the Doppler compensation algorithm, by analyzing the frequency offset of the laser echo signal, accurately calculate the relative velocity and distance of dynamic objects, with an accuracy of up to centimeter level; this algorithm can effectively distinguish static backgrounds and dynamic obstacles, and solve the problem of error accumulation of traditional ranging methods in high-speed scenarios; The receiving end collects the reflected signal and records the time difference between transmission and reception and the phase change of the signal. It should be noted that in a high-speed scenario, the laser echo signal may generate Doppler frequency shift due to the movement of the target, and the received signal is adjusted. The echo signal is subjected to spectral analysis through fast Fourier transform or wavelet transform to extract the frequency shift amount Δf. For example, the least squares algorithm for phase difference is used to model the phase difference of multiple subcarriers to solve the Doppler frequency shift. The target velocity v is calculated based on the frequency shift amount, and the Kalman filter or adaptive filter is used to track the velocity change in real time. For example, in a satellite communication system, the Doppler error is controlled within ±5 Hz through Kalman filtering.
[0019] The data processing module uses the optical flow method and the CNN deformation detection network to calculate the deformation of the object between adjacent frames. Based on the object shape analysis, the relative motion of the object is calculated. Based on the instance segmentation of MaskR-CNN, the object contour is output. At the same time, the dynamic target clustering algorithm is used to perform multi-target tracking on the object contour, and combined with the environmental point cloud data, the three-dimensional coordinates and motion vectors of the object are output. Calculate the deformation of the object between adjacent frames: Calculate the motion vector at the pixel level between consecutive image frames, that is, the moving direction and magnitude of each point between two frames. S201: Input two consecutive light rail environment images; use the optical flow algorithm to analyze the displacement of all pixel points; output the optical flow field map, which is represented as the two-dimensional motion vector of each pixel point, reflecting the motion change between adjacent frames. S202: Refine the optical flow field map obtained by the optical flow method to improve the ability to analyze non-linear deformation in complex scenarios; extract the feature map by inputting adjacent frames into a pre-trained convolutional neural network. S203: Use a specific network structure to compare the feature differences between adjacent frames; the network outputs the deformation ratio of the key area of the object in each frame, and combines the optical flow results to integrate the advantages of the two methods and improve the deformation detection accuracy. Advantages of the two methods: Combine the accurate pixel-level motion vector information provided by the optical flow method and the advanced feature difference analysis ability provided by the CNN to achieve the purpose of being able to accurately capture small displacement changes and effectively identify and quantify the non-linear deformation of the object, significantly improving the overall accuracy of deformation detection, especially in complex and changing actual application scenarios, especially during the monitoring of the surrounding environment of light rail vehicles. Calculation of the relative motion of the object: Determine whether the object is stationary or moving dynamically relative to the light rail; for each detected object area, calculate the variance of the deformation ratio between its adjacent frames and set a motion threshold (15%): If the variance of the deformation ratio ≤ 15%, the object is considered relatively stationary; If the variance of the deformation ratio > 15%, it is marked as a dynamic object; At the same time, combined with the motion vectors in the optical flow field, calculate the relative velocity and motion direction of the object to assist subsequent tracking; Motion threshold: Tending to set a relatively conservative threshold of 15% based on safety considerations, and adjust it in a timely manner according to on-site feedback and effect evaluation; Output object contour: Input the environmental image and the candidate area of the dynamic object, generate candidate boxes through the RPN module of Faster R-CNN, and the candidate boxes cover the areas that may contain objects, such as drones, pedestrians, etc.; Classification and mask generation: For each candidate box, use the classifier of Mask R-CNN to determine the object category, such as drones, vehicles; at the same time, generate a binary mask to accurately segment the object contour; perform morphological operations on the segmentation result, such as dilation and erosion, to remove noise and ensure the smoothness of the contour boundary; output the mask image and class label of each object; Multi-object tracking and 3D coordinate calculation: Based on the dynamic object clustering algorithm, input the object contour data (mask image) and optical flow motion vectors of adjacent frames; extract the center point coordinates, velocity vectors and shape features of each object; perform clustering grouping, and use the DBSCAN algorithm to cluster the targets: according to the spatial distance (Euclidean distance) and velocity similarity, divide the objects in consecutive frames into the same trajectory cluster and filter out noise points; Perform Kalman filtering on the clustering result to associate the trajectories and predict the target position in the next frame; Based on the dynamic object clustering algorithm, according to the spatial distance and velocity similarity of the objects, divide the objects in consecutive frames into different trajectory clusters. For each trajectory cluster, extract its center point coordinates, velocity vectors and shape features and other information as the initial state vector; use the Kalman filter to predict the state of each trajectory cluster; the Kalman filter is a recursive algorithm that uses the state estimate value of the previous moment and the observation value of the current moment to update the state estimate value of the current moment; when processing each frame, first use the state transition model to predict the position and velocity of the next state; for example, assuming that the target moves in a uniform straight line, the position of the next frame can be predicted based on the current position and velocity, compare the actually observed target position with the predicted value, and calculate the difference between the two; adjust the predicted value according to this difference to obtain a more accurate target position estimate; by continuously repeating the above prediction and update steps, the Kalman filter can not only provide the prediction of the target's future position, but also effectively reduce the position estimation error caused by sensor noise or external interference, significantly improving the stability and accuracy of target tracking; use the corrected state estimate value to predict the target's position in the next frame; 3D coordinate and motion vector calculation: The target trajectory after clustering and the lidar point cloud data are used to calibrate the internal and external parameters of the camera and lidar through a deep learning-based calibration algorithm to ensure the unity of the coordinate system. The 2D center point of the object contour is matched with the 3D point cloud in the point cloud data through a perspective projection model. Using the calibration parameters and projection formula, the 3D coordinates of the object are calculated. Combining the time series of the point cloud data, the three-dimensional velocity vector of the object is calculated to evaluate its motion direction and threat level. An output dynamic object list contains the category, 2D contour mask, 3D coordinates, and motion vector of each object, as well as the trajectory information of the object in consecutive frames.
[0020] The fusion and analysis module performs time synchronization based on the PTP protocol and uses the checkerboard calibration method for joint laser-vision calibration. The D-S evidence theory is adopted for multi-modal decision-making, and by inputting the evidence sources, the fusion decision is output. The evidence sources include the confidence of deformation analysis, the reliability of lidar ranging, and the probability of the classification network. Time synchronization based on the PTP protocol: Time synchronization based on the PTP protocol: The master clock sends a Sync message to the slave clock and records the transmission timestamp t1. After receiving the Sync message, the slave clock records the reception timestamp t2. The slave clock sends a Delay_Req message to the master clock, and the master clock records the reception timestamp t3 and sends a Delay_Resp message back to the slave clock, including t3. After receiving the Delay_Resp, the slave clock records the reception timestamp t4. And the average path delay and time offset are calculated. The slave clock corrects the local time according to the calculated Delay and Offset to synchronize the master and slave clocks to sub-microsecond accuracy. Joint laser-vision calibration: The conversion relationship between the lidar coordinate system and the camera coordinate system is established to achieve the spatial alignment of multi-modal data. A checkerboard calibration board is used, and the corner points of its black and white grids can be used as accurate 2D / 3D feature points. The corner point coordinates are extracted from the checkerboard image, and the camera internal parameters are calculated using Zhang's calibration method. The rotation matrix and translation vector of the camera relative to the world coordinate system are determined through multi-view images. By scanning the checkerboard calibration board with the lidar, the point cloud data is obtained, and the installation pitch angle θ and installation height h of the lidar are deduced from the point cloud data. The lidar point cloud data is converted to the camera coordinate system to obtain the rotation and translation matrix from the lidar to the camera. The values of the rotation and translation matrix from the lidar to the camera are optimized by minimizing the projection error between the point cloud data and the image corner points. The lidar point cloud data is projected onto the camera image plane through the calibration parameters to verify the coincidence degree. Output fusion decision: Confidence of deformation analysis: The CNN deformation detection network from the data processing module; representing the credibility m1 of the deformation ratio D of the object between adjacent frames. Reliability of Laser Ranging: Derived from lidar distance measurement; representing the stability of lidar point cloud data m2; Probability of Classification Network: Derived from MaskR-CNN instance segmentation; representing the prediction probability of object categories (such as drones, pedestrians) m3; Convert the confidence of each evidence source into a BPA function, calculate the conflict factor K between evidence sources, and synthesize evidence through Dempster's rule; generate a fusion decision based on the comprehensive trust degree: High confidence (>0.8): Marked as high-threat objects, such as drones approaching the orbit; Medium confidence (0.5 - 0.8): Marked as potential threats and trigger early warnings; Low confidence (<0.5): Determined as safe objects, such as static obstacles; Control and Response Module, obtain the fusion decision, and perform a hierarchical early warning mechanism and response; Obtain the fusion decision result, spatio-temporal alignment data, and environmental context information, and conduct threat level classification: Table 1: Threat Level Classification Table: Adjust the threat threshold in combination with environmental context information. For example, in rainy or snowy weather, the reliability of lidar decreases, and the m threshold needs to be increased; optimize the mapping rule through historical data learning. For example, if the false alarm rate of drones in certain areas is high, the threat level can be dynamically reduced; Through real-time monitoring, continuously receive the fusion decision result m, and compare it with the preset threshold; dynamically adjust the threshold according to the current environment (such as at night, strong light). For example, at night, the reliability of lidar decreases, and m needs to be increased to 0.8 to trigger a level-three early warning; Low Threat (Level-One Early Warning): Operation: Record logs, no manual intervention required.
[0021] Technical Implementation: Quickly write to the local database through an embedded system (such as an ARM processor); Medium Threat (Level-Two Early Warning): Operation: Acoustic and optical alarm + operator interface prompt; Technical Implementation: Acoustic and optical alarm: Control the buzzer and LED indicator through GPIO; Operator Interface: Display the real-time image and location of the threat target on the HMI (Human Machine Interface); High Threat (Level-Three Early Warning): Operation: Emergency braking + automatic intervention (such as laser interference, RF blocking); Technical Implementation: Emergency braking: Send an emergency stop instruction to the light rail control system via the CAN bus; Automatic intervention: Laser interference: Activate the high-energy laser module and implement strikes according to a preset strategy, such as three-step damage; RF blocking: Directionally interfere with the UAV communication link and force it to return or land.
[0022] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution.
[0023] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0024] As described above, the above are only specific implementation manners of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application.
Claims
1. A light rail vehicle safety monitoring system based on multi-sensor fusion, characterized in that, The system includes: A multi-modal perception module that captures images of the environment around the light rail in real time and demarcates the visual marker area; performs three-dimensional scanning on the visual marker area to generate environmental point cloud data in real time, and uses the Doppler compensation algorithm to obtain the distance of dynamic objects; A data processing module that uses the optical flow method and the CNN deformation detection network to calculate the deformation of objects between adjacent frames; calculates the relative motion of objects based on object shape analysis; outputs the object contour based on instance segmentation of MaskR-CNN; at the same time, uses the dynamic target clustering algorithm to perform multi-target tracking on the object contour, and combines the environmental point cloud data to output the three-dimensional coordinates and motion vectors of the object; A fusion and analysis module that performs time synchronization based on the PTP protocol and uses the checkerboard calibration method to perform joint calibration of laser vision; uses the D-S evidence theory for multi-modal decision-making, and outputs a fusion decision by inputting the evidence source; where the evidence source includes the confidence of deformation analysis, the reliability of laser ranging, and the probability of the classification network; A control and response module that obtains the fusion decision and performs a hierarchical warning mechanism and response.
2. The safety monitoring system for light rail vehicles based on multi-sensor fusion according to claim 1, wherein The process of demarcating the visual marker area is as follows: Adopt a combined algorithm of median filtering and bilateral filtering to eliminate high-frequency noise in the environmental image; perform adaptive histogram equalization based on the Retinex theory, apply the calibration method to pre-calibrate the lens distortion parameters, combine online calibration to correct the offset, and use the adaptive threshold segmentation algorithm to demarcate the visual marker area.
3. The safety monitoring system for light rail vehicles based on multi-sensor fusion according to claim 2, characterized in that, The process of generating environmental point cloud data in real time is as follows: LiDAR emits laser pulses simultaneously at different vertical angles through a multi-beam laser emitter; each beam is separated by a certain angle and covers 360° horizontally through a rotating mirror or electronic scanning method, and combines the pose data of the IMU to calculate the three-dimensional coordinates of each point.
4. The safety monitoring system for light rail vehicles based on multi-sensor fusion according to claim 3, characterized in that, The process of obtaining the distance of dynamic objects is as follows: Within a single-frame scanning period, calculate the motion trajectory through numerical integration by fusing the IMU and wheel odometer data, and after synchronizing the laser points and sensor data in time using the PTP protocol, linearly interpolate the IMU pre-integration results to each laser point; Use the Doppler compensation algorithm to calculate the relative speed and distance of dynamic objects by analyzing the frequency shift of the laser echo signal; collect the reflected signal at the receiving end, record the time difference between transmission and reception and the phase change of the signal; perform spectral analysis on the echo signal through fast Fourier transform to extract the frequency shift; calculate the target speed based on the frequency shift and use the Kalman filter to track the speed change in real time.
5. The safety monitoring system for light rail vehicles based on multi-sensor fusion according to claim 1, characterized in that, The process of calculating the deformation of objects between adjacent frames is as follows: S201: Input two consecutive light rail environment images; analyze the displacement of all pixel points using the optical flow algorithm; output the optical flow field map; S202: Refine the optical flow field map obtained by the optical flow method, input adjacent frames, and extract the feature map through a pre-trained convolutional neural network; S203: Use a specific network structure to compare the feature differences between adjacent frames; The network outputs the deformation ratio of the key areas of the object in each frame, and combines the optical flow results for dominant fusion.
6. The safety monitoring system for light rail vehicles based on multi-sensor fusion according to claim 5, characterized in that, The process of calculating the relative motion of objects is as follows: For each detected object region, calculate the variance of the deformation ratio between adjacent frames, set a motion threshold, and make a motion judgment: If the variance of the deformation ratio ≤ the motion threshold, the object is considered relatively stationary; If the variance of the deformation ratio > the motion threshold, it is marked as a dynamic object; At the same time, combine the motion vectors in the optical flow field to calculate the relative velocity and motion direction of the object.
7. The safety monitoring system for light rail vehicles based on multi-sensor fusion according to claim 6, characterized in that, The process of outputting the object contour, as well as the 3D coordinates and motion vectors of the object is as follows: Input the environmental image and the dynamic object candidate area, and generate candidate boxes through the RPN module of Faster R-CNN; Based on the dynamic object clustering algorithm, input the object contour data of adjacent frames, and extract the center point coordinates, velocity vectors and shape features of each object; Perform clustering grouping, and use the DBSCAN algorithm to cluster the targets: according to the spatial distance and velocity similarity, divide the objects in consecutive frames into the same trajectory cluster; Perform Kalman filter association on the clustering results to predict the target position in the next frame; For the target trajectory after clustering and the lidar point cloud data, calibrate the internal and external parameters of the camera and lidar through a deep learning-based calibration algorithm, Match the 2D center point of the object contour with the 3D point cloud in the point cloud data through the perspective projection model; use the calibration parameters and projection formula to calculate the 3D coordinates of the object; combine the time series of the point cloud data to calculate the three-dimensional velocity vector of the object, and evaluate its motion direction and threat level; Output a list of dynamic objects containing the category, 2D contour mask, 3D coordinates and motion vectors of each object, as well as the trajectory information of the objects in consecutive frames.
8. The safety monitoring system for light rail vehicles based on multi-sensor fusion according to claim 1, characterized in that, The process of the laser-vision joint calibration is as follows: Use the corner points of the checkerboard calibration board as 2D / 3D feature points, calculate the camera internal parameters through the Zhang Zhengyou calibration method, and determine the camera external parameters based on multi-view images; obtain the point cloud data by scanning the checkerboard with a lidar, and reverse calculate its installation pitch angle θ and height h; convert the point cloud to the camera coordinate system, and construct the rotation and translation matrix from the lidar to the camera; optimize the matrix parameters by minimizing the projection error between the point cloud and the image corner points; project the point cloud onto the image plane to verify the coincidence degree and perform multi-modal space alignment.
9. The safety monitoring system for light rail vehicles based on multi-sensor fusion according to claim 1, characterized in that, The process of the output fusion decision is as follows: Convert the CNN deformation confidence, lidar ranging stability and Mask R-CNN classification probability into a BPA function and calculate the conflict factor K, use the Dempster rule to synthesize evidence to generate a comprehensive trust degree, and make a hierarchical decision on high-threat, potential-threat and safe objects.
10. The safety monitoring system for light rail vehicles based on multi-sensor fusion according to claim 1, characterized in that, The process of the hierarchical early warning mechanism and response is as follows: Dynamically adjust the threat threshold and optimize the mapping rule based on historical data; monitor the fusion decision result in real time and perform regulation.
Citation Information
Patent Citations
Target situation fusion sensing method and system based on multiple sensors
CN110866887A
Traveling crane hoisting safety detection method and system based on multi-sensor calibration
CN119117939A
Unmanned subway obstacle intelligent detection system based on multi-mode AI sensor
CN119339354A
Mine vehicle safe driving detection system based on vision and radar fusion
CN120065228A
Method and apparatus for determining point cloud set corresponding to target object
WO2022088104A1
Cited By
Parking lot multi-dimensional security monitoring system using infrared sensing technology
CN120912989A
Tunnel illumination control method and system based on distributed sensing
CN122093994A
Disaster scene real scene situation multi-mode reconstruction method and system
CN122510468A
Disaster scene real scene situation multi-mode reconstruction method and system
CN122510468B