Rail transit vehicle obstacle detection method based on multimodal data fusion
Through the obstacle detection method of multimodal data fusion, the complementary advantages of visual images, millimeter-wave radar and ultrasonic radar are utilized to dynamically adjust the detection mode, which solves the problem of low detection accuracy of rail transit vehicles in harsh environments and improves the robustness and environmental adaptability of detection.
Patent Information
- Application Number
- CN202511052741.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-07-30
AI Technical Summary
Existing rail transit vehicle obstacle detection technology has low detection accuracy in severe weather or complex environments and lacks an adaptive adjustment mechanism, resulting in insufficient safety.
A multimodal data fusion method is adopted, combining visual images, millimeter-wave radar and ultrasonic radar. Through environmental perception classification, trust calculation and mode decision-making, the detection mode is dynamically adjusted to achieve the complementary advantages and adaptive switching of sensor data.
It improves the robustness and environmental adaptability of obstacle detection, reduces the impact of weather on detection results, and ensures the safety of rail transit.
Smart Images

Figure CN120561777B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of track detection, and in particular to a rail transit vehicle obstacle detection method based on multimodal data fusion. Background Art
[0002] In the field of rail transit vehicle operation safety, the reliability of obstacle detection technology directly impacts driving safety. Existing technologies generally use single-modal data, such as visual images, millimeter-wave radar, or ultrasonic radar, for obstacle detection. However, this detection method has significant drawbacks: the detection performance of single-modal sensors degrades significantly when the train's operating environment undergoes significant changes, such as inclement weather like heavy rain, fog, or strong sunlight, or when entering complex close-range environments like tunnels and construction sites.
[0003] Taking single visual detection as an example, when the light intensity changes suddenly, such as at the entrance and exit of a tunnel or in rainy and foggy weather, the image clarity decreases, the difficulty of extracting target features increases, and missed detection or false detection is very likely to occur; and although a single millimeter-wave radar has strong adaptability to weather, it has difficulty in obtaining semantic information of the target, such as target type and shape characteristics, and cannot accurately distinguish between obstacles and background in complex scenes; the detection range of a single ultrasonic radar is limited and it is almost ineffective in long-distance scenes.
[0004] Furthermore, existing technologies lack a dynamic assessment mechanism for sensor data reliability, making it impossible to adaptively adjust detection strategies based on environmental changes. When a sensor's performance degrades under specific circumstances, the system cannot promptly switch to a more optimal detection mode, resulting in a decrease in overall detection accuracy and a serious threat to rail transit safety. Therefore, to address these issues, the present invention proposes a rail transit vehicle obstacle detection method based on multimodal data fusion. Summary of the Invention
[0005] In view of the shortcomings of the existing technology, the purpose of the present invention is to provide a rail transit vehicle obstacle detection method based on multimodal data fusion.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] The rail transit vehicle obstacle detection method based on multimodal data fusion is characterized by comprising the following steps:
[0008] Environmental perception and classification: Environmental sensors collect real-time data on the train's operating environment and combine it with historical data to classify the current detection environment into categories such as good visual environment, adverse weather environment, and close-range complex environment.
[0009] The trust calculation and mode decision step sets the initial trust of different detection modes, calculates the trust of different detection modes based on the environment type and sensor data collected in combination with the initial trust, and selects a detection mode based on the trust. The detection modes include visual image detection mode, millimeter wave radar detection mode, and ultrasonic detection mode;
[0010] a detection data collection step, determining the amount of data collected by each detection sensor according to different detection modes and performing data collection and data processing, wherein the detection sensors include visual image sensors, millimeter wave radars, and ultrasonic radars;
[0011] The obstacle detection step performs multimodal data fusion on the processed data according to the detection mode to obtain corresponding data to be detected, and performs obstacle detection based on the data to be detected and outputs the detection results.
[0012] As a further improvement of the present invention, the environmental sensor includes a light sensor and a temperature and humidity sensor. The environmental sensor collects ambient light intensity, temperature data and humidity data as environmental collection data. An environmental judgment model is constructed based on the environmental collection data and historical data, and an environmental judgment value is calculated. The environmental type judgment result is obtained based on the comparison result of the environmental judgment value and the preset threshold.
[0013] As a further improvement of the present invention, the trust calculation includes dynamically adjusting the trust based on the environment type, presetting the sensor trust adjustment coefficient for each environment type, the adjustment coefficient of the visual image sensor in a good visual environment is a positive value, the adjustment coefficient of the millimeter-wave radar in a severe weather environment is a positive value, and the adjustment coefficient of the ultrasonic radar in a close-range complex environment is a positive value. The DS evidence theory is used to perform a conflict analysis on the detection results of the same target by the visual image, millimeter-wave radar and ultrasonic radar. If the position deviation of the detection results of the three is within a preset threshold, the trust of the three increases synchronously; if the visual image and millimeter-wave radar detection conflict and the ultrasonic radar supports one of them, the trust of the conflicting party is reduced, and a weighted calculation is performed in combination with the initial trust to obtain the final trust value of each detection mode.
[0014] As a further improvement of the present invention, when the detection mode is visual image detection, the visual image sensor collects full-resolution image data and extracts obstacle features through the improved FasterR-CNN algorithm; the millimeter-wave radar collects distance data and azimuth data of detection targets in the long-distance area within the current detection range; and the ultrasonic radar collects target position data of detection targets in the short-distance area within the current detection range.
[0015] As a further improvement of the present invention, when the detection mode is millimeter wave radar detection, the millimeter wave radar collects the full amount of point cloud data within the current detection range and extracts feature data through the CFAR algorithm; the visual image sensor collects the type recognition data of the moving target within the current detection range; the ultrasonic radar collects the contour data of the detection target in the close range area within the current detection range.
[0016] As a further improvement of the present invention, when the detection mode is ultrasonic detection, the ultrasonic radar collects TDOA positioning data and reconstructs the three-dimensional coordinates of close-range obstacles; the visual image sensor collects edge feature data of the detection target within the current detection range; and the millimeter wave radar collects motion speed vector data of the detection target within the current detection range.
[0017] As a further improvement of the present invention, the obstacle detection method in the visual image detection mode includes: calculating the image blurriness by image quality evaluation of the obstacle features extracted from the full-resolution image data; when the image blurriness is higher than a preset blurriness threshold, combining millimeter-wave radar and ultrasonic radar detection, spatially matching the distance data collected by the millimeter-wave radar with the full-resolution image data collected by the visual image sensor, correcting the depth information of the full-resolution image data, and superimposing the detection target detected by the ultrasonic radar on the full-resolution image data with a highlighted mark according to the target position data; and concatenating the data features of the full-resolution image data, the distance features obtained by combining the distance data and azimuth data collected by the millimeter-wave radar, and the position features of the target object collected by the ultrasonic radar into a multi-layer perceptron for feature-level fusion.
[0018] As a further improvement of the present invention, the obstacle detection method in the millimeter wave radar detection mode includes using the CFAR processing results of the special data of the millimeter wave radar as the main detection data, generating the distance, speed and angle information of the target, associating the target type recognition data provided by the visual image sensor with the motion trajectory obtained by analyzing the point cloud data of the millimeter wave radar detection target, and when the type recognition data successfully matches the spatial position of the millimeter wave target, assigning the target a semantic category identifier.
[0019] As a further improvement of the present invention, the obstacle detection method in the ultrasonic detection mode includes fusing the three-dimensional coordinate data collected by the ultrasonic radar and the edge feature data collected by the visual image sensor to obtain a reconstructed obstacle geometry, and analyzing the motion speed vector data provided by the millimeter-wave radar to obtain motion trend data, predicting the displacement path of close-range obstacles, and fusing the displacement path with the geometric shape.
[0020] As a further improvement of the present invention, the obstacle detection step includes setting environmental weighting values according to different environment types, performing environmental weighting processing on the fused detection data, inputting the processed data into a preset obstacle existence probability model and outputting a comprehensive obstacle confidence score, comparing the comprehensive obstacle confidence score with a preset threshold value, and if the comprehensive obstacle confidence score is greater than the preset threshold value, determining and outputting a result indicating the presence of an obstacle.
[0021] The beneficial effects of the present invention are:
[0022] By integrating three sensors—visual imagery, millimeter-wave radar, and ultrasonic radar—and leveraging their complementary strengths in different environments, a full-scenario detection system has been built. Even when one sensor is affected by environmental interference, the others can still provide effective data support, avoiding detection blind spots caused by failure of a single mode and improving the robustness of obstacle detection.
[0023] Through the trust calculation model that combines environmental perception classification with DS evidence theory, adaptive switching of detection modes is achieved, which improves the environmental adaptability of obstacle detection.
[0024] By using data from different modalities as the main detection data in different weather environments, the impact of weather on obstacle detection results can be reduced. The best and most accurate detection method can be used in each weather condition, thereby improving the detection accuracy of obstacle detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 This is a flow chart of the rail transit vehicle obstacle detection method using multimodal data fusion according to the present invention;
[0026] Figure 2 is a flow chart of the visual image detection mode of the present invention;
[0027] Figure 3 is a flow chart of the millimeter wave radar detection mode of the present invention;
[0028] Figure 4 It is a flow chart of the ultrasonic detection mode of the present invention. DETAILED DESCRIPTION
[0029] The present invention will be described in further detail below with reference to the accompanying drawings and embodiments. Identical components are denoted by the same reference numerals. It should be noted that the terms "front," "rear," "left," "right," "upper," and "lower" used in the following description refer to directions in the accompanying drawings, and the terms "bottom," "top," "inner," and "outer" refer to directions toward or away from the geometric center of a particular component, respectively.
[0030] The embodiment of the present invention proposes a rail transit vehicle obstacle detection method based on multimodal data fusion, such as Figure 1 As shown, including:
[0031] The rail transit vehicle obstacle detection method based on multimodal data fusion is characterized by comprising the following steps:
[0032] Environmental perception and classification: Environmental sensors collect real-time data on the train's operating environment and combine it with historical data to classify the current detection environment into categories such as good visual environment, adverse weather environment, and close-range complex environment.
[0033] A good visual environment refers to an environment in which the visual image sensor can clearly capture target features. The specific quantitative standards are: the light intensity is greater than the typical value of 50klx on a sunny day at noon; the environmental characteristics are free of obvious obstructions, and the target contour, color, texture and other visual features can be effectively extracted. The image blur is less than 5%, calculated based on the Laplace operator.
[0034] Severe weather conditions refer to environments where the performance of visual sensors is significantly affected, but millimeter-wave radars can operate stably. The quantitative standards are: light intensity, 10-50klx on cloudy days, dusk, or light dust; environmental characteristics, the visual image is blurred, noisy, or has reduced contrast. The image blur is 5%-30%, but the millimeter-wave radar echo signal signal-to-noise ratio is ≥15dB.
[0035] A close-range complex environment refers to an environment that is close to the train and has densely distributed obstacles and multiple obstructions. The quantitative standards are: distance range, within the effective detection range of the ultrasonic radar of 0-50m in front of the train; obstacle density, ≥5 / 100㎡, such as facilities, construction equipment, and accumulations in tunnels; environmental characteristics, with multiple target obstructions and sudden changes in lighting, such as tunnel entrances and exits or complex reflective surfaces. Visual and millimeter-wave radars are susceptible to multipath interference, and ultrasonic radars have significant advantages in close-range positioning.
[0036] While the train is running, sensors collect light intensity data ranging from 0 to 100 klx at a frequency of 100 Hz. Historical data is stored in a SQLite database, retaining the last 30 days of data. The environmental judgment model is trained using the XGBoost algorithm. The environmental judgment value is calculated based on the deviation between real-time data and historical data from the same period. A value greater than 80 indicates a good visual environment, 40-80 indicates inclement weather, and less than 40 indicates a complex close-range environment.
[0037] The trust calculation and mode decision step sets the initial trust of different detection modes, calculates the trust of different detection modes based on the environment type and sensor data collected in combination with the initial trust, and selects a detection mode based on the trust. The detection modes include visual image detection mode, millimeter wave radar detection mode, and ultrasonic detection mode;
[0038] Initially, the confidence level for the visual image detection mode is set to 0.7, for the millimeter-wave radar to 0.3, and for the ultrasonic mode to 0.2. Adjustments are applied based on the environment type, such as a +0.2 adjustment for the millimeter-wave radar in severe weather. Sensor conflict data is processed using DS evidence theory. The position deviation threshold is set to 3 meters, and the confidence level is updated every 500ms. The mode with the highest confidence level is selected when the difference exceeds 0.1.
[0039] a detection data collection step, determining the amount of data collected by each detection sensor according to different detection modes and performing data collection and data processing, wherein the detection sensors include visual image sensors, millimeter wave radars, and ultrasonic radars;
[0040] If the current mode is visual image detection, the visual sensor's 12MP camera collects 1920×1080 resolution images, the millimeter wave radar at 77GHz collects distance and azimuth data in the range of 50-200 meters, and the ultrasonic radar at 40kHz collects target position data in the range of 0-5 meters.
[0041] The obstacle detection step performs multimodal data fusion on the processed data according to the detection mode to obtain corresponding data to be detected, and performs obstacle detection based on the data to be detected and outputs the detection results.
[0042] The collected data is fused according to the corresponding mode. For example, in visual mode, a modified Faster R-CNN is used to extract features, combined with millimeter wave data for depth correction, and ultrasonic data for highlighting. The data is then input into the MLP classification system. The fused data is weighted according to the environment, such as a weight of 0.4 for severe weather conditions. The data is then input into a logistic regression model to calculate the confidence level. A confidence level greater than 0.7 is considered an obstacle.
[0043] Specifically, such as Figures 1 to 4 As shown, the environmental sensors include a light sensor and a temperature and humidity sensor. These sensors collect ambient light intensity, temperature, and humidity data as collected environmental data. This collected data, combined with historical data, is used to construct an environmental judgment model. This environmental judgment value is calculated, and the environmental type judgment result is determined by comparing this environmental judgment value with a preset threshold. The environmental sensors utilize industrial-grade, high-precision equipment: the light sensor has a spectral response range of 300-800nm and a measurement accuracy of ±5%; the visibility meter, based on the forward scattering principle, has a measurement range of 1-10,000m and an accuracy of ±2%; and the temperature and humidity sensor has a temperature range of -40-125°C and a humidity range of 0-100%RH with an accuracy of ±1%RH. The environmental judgment model utilizes a transfer learning strategy, using the ResNet50 network as the base network and fine-tuning it on a historical environmental dataset containing 100,000 sets of annotated data.
[0044] During construction, real-time features are collected, including light intensity normalized to 0-100, visibility normalized to 0-1000m, shower intensity normalized to 0-100mm / h, and temperature normalized to -40-125℃; historical features are also obtained, including the mean, variance, and transformation rate of the data for the same period in the past 7 days, to form a training set containing 100,000 sets of labeled data. The model is judged to be valid when the accuracy reaches the preset value through the test set test.
[0045] Data acquisition and preprocessing: The sensor collects raw data at a frequency of 200 Hz. Noise is removed by Kalman filtering. The light intensity is logarithmically transformed to adapt to nonlinear characteristics. The haze data is normalized using the visibility-extinction coefficient conversion formula β = 3.912 / V, where V is visibility.
[0046] Historical data processing: Extract the data of the same period 1 hour, 1 day, 7 days, and 30 days before the current time from the database, calculate the mean, variance, and rate of change, and form a historical feature vector.
[0047] Model Training and Application: The model was trained using the XGBoost algorithm with parameters set to n_estimators=100, learning_rate=0.1, and max_depth=5. A vector of real-time and historical features was input, and the output was an environmental judgment value. Preset thresholds were determined through grid search: the first threshold of 80 corresponds to a clear, fog-free scene, and the second threshold of 40 corresponds to light rain and fog.
[0048] Environment type determination: When the calculated environment judgment value is ≥80, it is judged to be a good visual environment, suitable for visual image detection; when the value is 40≤<80, it is judged to be a bad weather environment, and millimeter-wave radar detection has significant advantages; when the value is <40, it is judged to be a close-range complex environment such as a tunnel or construction area, and ultrasonic radar is more suitable.
[0049] Specifically, such as Figures 1 to 4As shown, the trust calculation includes dynamic trust adjustment based on the environment type. A sensor trust adjustment coefficient is preset for each environment type. The adjustment coefficient for the visual image sensor is positive in good visual environments, the adjustment coefficient for the millimeter-wave radar is positive in poor weather environments, and the adjustment coefficient for the ultrasonic radar is positive in close-range complex environments. Conflict analysis is performed on the same target detection results of the visual image, millimeter-wave radar, and ultrasonic radar using DS evidence theory. If the position deviation of the three detection results is within a preset threshold, the trust of all three is increased simultaneously. If the visual image and millimeter-wave radar detections conflict and the ultrasonic radar supports one of them, the trust of the conflicting party is reduced. The final trust value for each detection mode is obtained through a weighted calculation based on the initial trust. The trust calculation adopts a three-layer architecture: an initial trust setting module, a dynamic adjustment module, and a conflict analysis module. The initial trust is stored in matrix form (vision = [0.7, 0.3, 0.2]); the dynamic adjustment module contains an environmental coefficient table, such as the visual adjustment coefficient for a good visual environment + 0.15; the conflict analysis module constructs a basic probability distribution function (BPA) based on DS evidence theory and uses the Yager combination rule to handle conflicts.
[0050] Initial trust initialization: When the system starts, the initial trust matrix is loaded through the configuration file. The initial value of 0.7 for the visual image detection mode corresponds to its high reliability under good lighting conditions. 0.3 for millimeter-wave radar and 0.2 for ultrasonic are used for initial transition.
[0051] Environmental coefficient adjustment: Based on the environment type determined by the embodiment, the corresponding coefficient is retrieved from the preset counter table. For example, in severe weather, the adjustment coefficient for millimeter-wave radar is +0.2, for visual imagery it is -0.1, and for ultrasonic wave it is -0.05. The current confidence level = initial value + environmental coefficient × 0.8 weighting coefficient.
[0052] DS Evidence Theory Conflict Analysis: A BPA is constructed using the detection results of the same target from visual imagery, millimeter-wave radar, and ultrasonic radar as input. For example, if the target location is (x1, y1) for visual detection, (x2, y2) for millimeter-wave detection, and (x3, y3) for ultrasonic detection, the Euclidean distances are calculated as d1 = √[(x1-x2)² + (y1-y2)²] and d2 = √[(x1-x3)² + (y1-y3)²]. If d1 is less than 3 meters and d2 is less than 3 meters, the confidence level of each is increased by 0.1. If d1 is greater than 5 meters and ultrasonic detection supports vision, the confidence level of millimeter-wave detection is reduced by 0.1, while the confidence levels of visual and ultrasonic detection remain unchanged.
[0053] Sensor status monitoring adjustments: Real-time monitoring of sensor signal strength, such as the PSNR value of visual images, and an update frequency standard of 25fps. If the visual sensor PSNR is less than 20 for 10 seconds or the update frequency is less than 15fps, the confidence level is -0.1.
[0054] Comprehensive trust calculation: Final trust = initial adjusted trust × 0.6 + DS adjustment value × 0.3 + state adjustment value × 0.1. Recalculated every 1 second. When the trust of a mode exceeds the next highest mode by 0.2, a mode switch is triggered.
[0055] Specifically, such as Figures 1 to 4 As shown in the figure, when the detection mode is visual image detection, the visual image sensor collects full-resolution image data and extracts obstacle features through the improved FasterR-CNN algorithm; the millimeter-wave radar collects distance data and azimuth data of the detection target in the long-distance area within the current detection range; the ultrasonic radar collects target position data of the detection target in the short-distance area within the current detection range.
[0056] The visual image sensor uses a Basle RACE2 series industrial camera with a resolution of 2448×2048, a frame rate of 60fps, and a 12mm focal length lens. The millimeter-wave radar uses the Continental ARS408-21, with a detection range of 0-210 meters and an angular resolution of 1°. The ultrasonic radar uses the HC-SR04, with a measurement range of 2cm-400cm and an accuracy of 3mm. Data acquisition is synchronized and triggered via an FPGA, with a clock accuracy of 1μs.
[0057] Visual Image Acquisition: In visual image detection mode, the camera captures images at full resolution (2448×2048) in RGB format, with exposure time automatically adjusted from 1 / 10,000 to 1 / 30 second based on ambient lighting. Each image frame is timestamped to the millisecond and transmitted to the data processing unit via the GigE interface.
[0058] Millimeter-wave radar data acquisition: The radar operates in the 77 GHz frequency band and scans targets within a range of 50-200 meters. Each scan generates point cloud data with target range accuracy of ±0.5 meters, azimuth angle of ±1°, and radial velocity of ±0.1 meters per second. Data is output in frames at a 10 Hz rate, with each frame containing 100-500 points.
[0059] Ultrasonic radar data acquisition: An array of eight ultrasonic sensors is located on either side of the vehicle's front. Each sensor transmits 40kHz pulses at a 50Hz frequency, receives reflected waves, and calculates the target distance. Position data for targets within a range of 0-5 meters is collected and output in polar coordinates, with an angular resolution of 5° and a distance accuracy of 3mm.
[0060] Data synchronization: The FPGA synchronously triggers the three sensor acquisitions, aligning timestamps using the hardware clock. Millimeter-wave and ultrasonic data are converted to coordinates and unified into a visual image coordinate system with the camera's optical center as the origin. Error calibration is performed using a pre-calibrated checkerboard pattern with an accuracy of <0.5 pixels.
[0061] Specifically, such as Figures 1 to 4 As shown, when the detection mode is millimeter-wave radar detection, the millimeter-wave radar collects the full amount of point cloud data within the current detection range and extracts feature data through the CFAR algorithm; the visual image sensor collects the type recognition data of the moving target within the current detection range; the ultrasonic radar collects the contour data of the detection target in the close range within the current detection range.
[0062] The millimeter-wave radar uses the Vayyar WM128, which transmits 76-81 GHz frequency-modulated continuous wave (FMCW) signals. It generates 4D point clouds showing distance, velocity, angle, and reflection intensity, with a detection range of 0-300 meters and a point cloud density of 128×128. The visual image sensor uses the Mobileye EyeQ6, which integrates a neural network accelerator and provides real-time output of target types such as pedestrians, vehicles, and obstacles. The ultrasonic radar uses the Panasonic MA40S4S, with a measurement range of 0-8 meters and outputs point cloud data of target outlines.
[0063] Millimeter-wave radar full point cloud acquisition: The radar transmits FMCW signals at a 10Hz frequency, generating approximately 16,384 points (128×128 pixels) per frame. Each point has a range accuracy of ±0.3 meters, a velocity of ±0.05 m / s, a horizontal angle of ±0.5°, a vertical angle of ±1°, and a reflection intensity of 0-255. Denoising and clustering are performed using a dedicated signal processing chip, such as the NXPS32R294, to extract target point cloud clusters.
[0064] Visual image type recognition data acquisition: The EyeQ6 processor processes 720P resolution images (1280×720) in real time. The built-in YOLOv8 neural network model weight file size is 28MB. It can recognize 15 target types, including pedestrians, bicycles, cars, trucks, etc. The confidence threshold is set to 0.6, and the target category, bounding box coordinates and confidence level are output.
[0065] Ultrasonic radar profile data acquisition: An array of 16 ultrasonic sensors, each scanning at 20Hz, calculates the target profile based on phase difference. This collects a 3D profile point cloud of targets within a range of 0-8 meters, with approximately 50-200 points per target. The point cloud is formatted as (x, y, z) coordinates with an accuracy of ±1cm and transmitted to the main control unit via the I2C bus.
[0066] Data preprocessing: The millimeter-wave point cloud is processed using the CFAR algorithm with a constant false alarm rate (CFAR) algorithm. Noise thresholds for different scenarios, such as sea clutter and ground clutter, are set, such as -60dBm for urban road scenarios, to filter out non-target points. The visual type recognition results are then correlated with the millimeter-wave point cloud using the Hungarian algorithm, matching spatial position and motion trajectory. Association is established when the matching error is less than 2 meters.
[0067] Specifically, such as Figures 1 to 4 As shown, when the detection mode is ultrasonic detection, the ultrasonic radar collects TDOA positioning data and reconstructs the three-dimensional coordinates of close-range obstacles; the visual image sensor collects edge feature data of the detection target within the current detection range; and the millimeter wave radar collects motion speed vector data of the detection target within the current detection range.
[0068] Ultrasonic radar uses a Time-of-Flight (ToF) array sensor, such as the MaxBotix HRXL-MaxSonar-EZ. This array consists of 12 sensors arranged in a 3×4 array, covering a 180-degree angle in front of the vehicle. Its measurement range is 0.1-10 meters and its accuracy is 1mm. Visual image sensors utilize edge computing modules such as the NVIDIA Jetson AGX Orin, equipped with the Canny edge detection algorithm to extract target edge features in real time. The millimeter-wave radar uses the Delphi Evasion sensor, which outputs the x, y, and z components of the target's velocity vector with an accuracy of ±0.01 m / s.
[0069] Ultrasonic TDOA positioning data acquisition: The array sensor sequentially transmits ultrasonic pulses at a frequency of 40kHz. The receiver records the arrival time of each sensor with an accuracy of 1μs. Target position is calculated using the TDOA algorithm. First, a reference sensor is identified. The time difference between the other sensors and the reference sensor is calculated. Substituting this into the hyperbolic positioning equation in three-dimensional space, the target's three-dimensional coordinates (x, y, z) are solved. Each frame takes 20ms to acquire, and up to 10 targets can be located simultaneously.
[0070] 3D coordinate reconstruction: Kalman filtering is performed on at least three sets of multiple measurement data for the same target to eliminate random errors and reconstruct the target's 3D coordinates. The coordinate system is based on the vehicle coordinate system, with the vehicle's front direction being the positive x-axis and the origin at the center of the vehicle's front. The error is less than 5mm.
[0071] Visual edge feature acquisition: Jetson AGX Orin processes 1080P images (1920×1080), first removing noise with a Gaussian filter with a kernel size of 5×5. Then, Canny edge detection is applied with a high-low threshold ratio of 3:1 to extract the pixel coordinates of the target edge. RANSAC line fitting is performed on the edge points to obtain the target edge contour parameters, such as the line equation and arc parameters.
[0072] Millimeter-wave velocity vector acquisition: The millimeter-wave radar outputs a target's velocity vector at a 20Hz frequency. This vector includes velocity components in the x-axis, y-axis, and z-axis directions, and is used to predict target motion trends. The velocity vector is smoothed using a Kalman filter, and its update frequency matches the radar frame rate.
[0073] Data alignment: Ultrasonic 3D coordinates, visual edge features, and millimeter-wave velocity vectors are aligned using timestamps and spatial coordinate transformation. The spatial transformation matrix is pre-calibrated using a high-precision 3D calibration plate. The rotation matrix and translation vector have accuracies of 0.1° and 1cm, respectively, and the time synchronization error is less than 1ms.
[0074] Specifically, such as Figures 1 to 4 As shown, the obstacle detection method in the visual image detection mode includes: calculating the image blurriness by performing image quality evaluation on the obstacle features extracted from the full-resolution image data; when the image blurriness is higher than a preset blurriness threshold, combining millimeter-wave radar and ultrasonic radar detection, spatially matching the distance data collected by the millimeter-wave radar with the full-resolution image data collected by the visual image sensor, correcting the depth information of the full-resolution image data, and superimposing the detection target detected by the ultrasonic radar on the full-resolution image data with a highlighted mark according to the target position data; and concatenating the data features of the full-resolution image data, the distance features obtained by combining the distance data and azimuth data collected by the millimeter-wave radar, and the position features of the target object collected by the ultrasonic radar into a multi-layer perceptron for feature-level fusion.
[0075] Image quality evaluation includes locking potential target areas based on the previous feature extraction results, screening areas of interest in the image that may contain obstacles, excluding meaningless background areas such as the sky, track bed and other areas without obstacle features; and performing noise reduction on the screened areas.
[0076] The edge clarity is determined through a logical edge detection algorithm. The smoother the edge transition, the higher the possibility of blur. The frequency characteristics of the image are analyzed to distinguish between the high-frequency components (corresponding to details) and the low-frequency components (corresponding to the overall outline) of the image. If the proportion of high-frequency components is lower than that of a conventional clear image, blur is indicated.
[0077] Convert the extracted fuzzy features into fuzziness indicators that can be used for judgment:
[0078] Comprehensively quantify edge features and frequency characteristics, and strengthen the weight of features that are sensitive to fuzzy through logical weighted integration;
[0079] Generate an index that reflects the overall blur level. The level of this index is positively correlated with the blur level of the image, that is, the higher the index, the blurrier the image.
[0080] The calculated ambiguity index is compared with the preset threshold. The threshold setting logic is based on the "minimum clarity standard to ensure the recognition of obstacle features":
[0081] If the blur index is lower than the threshold, the image clarity is determined to meet the visual detection requirements. Obstacle detection is performed directly using the current visual features without relying on other sensor data.
[0082] If the blur index is higher than the threshold: the image is judged to be blurry, resulting in unreliable visual features, and the multi-sensor collaboration mechanism is triggered.
[0083] When the ambiguity index is higher than the threshold, other sensor data are integrated according to the following logic to correct the detection result:
[0084] The distance data of the millimeter-wave radar is spatially matched with the visual image. Through the logical coordinate correspondence, the target position detected by the radar is mapped to the image to supplement the missing depth information of the image.
[0085] Using the close-range position data of ultrasonic radar, the target is highlighted in the visual image, and the recognition of obstacles in the blurred image is enhanced by logical position superposition;
[0086] Finally, the visual features, the distance features of the millimeter-wave radar, and the position features of the ultrasonic radar are logically connected in series to form a fusion feature for subsequent obstacle detection.
[0087] Visual feature extraction: The full-resolution 1920×1080 image is fed into a modified Faster R-CNN model. This model adds a CBAM attention mechanism to the original network to improve small object detection. It extracts features such as obstacle shape (HOG) features, color (HSV) histograms, and texture (LBP) features, outputting a feature vector with a dimension of 2048.
[0088] Millimeter-wave data spatial matching: The target range and azimuth detected by the millimeter-wave radar are converted into three-dimensional coordinates x, y, and z. These are then projected onto the visual image plane using a perspective transformation matrix to obtain the predicted target position. The intersection over union (IoU) between the predicted position and the visually detected target is calculated, and a match is considered successful when IoU > 0.5. After a successful match, the millimeter-wave distance data is used to correct the depth information in the visual image: original depth = pixel depth × millimeter-wave distance / visually estimated distance.
[0089] Ultrasonic Data Highlighting: For close-range targets detected by ultrasonic waves, less than 5 meters away, if they are not recognized in the visual image (i.e., if there is no corresponding target in the Faster R-CNN output), their 3D coordinates are projected onto the image plane, generating a rectangular marker box. This box is filled with red and labeled "Suspected Obstacle" to alert the system to focus attention. The marker box size is adjusted based on the target distance measured by ultrasonic measurement, increasing when closer and decreasing when farther away.
[0090] Feature-level fusion: The 2048-dimensional visual features, the 1-dimensional distance and 1-dimensional azimuth of the millimeter-wave distance features, and the 2-dimensional image coordinates of the ultrasonic position features are concatenated into a 2048+2+2=2052-dimensional feature vector, which is then input into the MLP network. The first layer of the MLP uses the ReLU activation function, the second layer uses a dropout probability of 0.5 to prevent overfitting, and the third layer outputs a 2-dimensional vector of obstacle probability and non-obstacle probability. Confidence is calculated using the Softmax function.
[0091] Specifically, such as Figures 1 to 4 As shown, the obstacle detection method in the millimeter-wave radar detection mode includes using the CFAR processing results of the special data of the millimeter-wave radar as the main detection data, generating the distance, speed and angle information of the target, correlating the target type recognition data provided by the visual image sensor with the motion trajectory obtained by analyzing the point cloud data of the millimeter-wave radar detection target, and when the type recognition data successfully matches the spatial position of the millimeter-wave target, assigning the target a semantic category identifier.
[0092] Millimeter-wave CFAR Processing: Millimeter-wave point cloud data is processed using CA-CFAR. First, the data is divided into 20 reference cells (front edge and back edge) and 20 protection cells (front edge and back edge). The mean and variance of the background noise power are calculated. The detection threshold is set to the noise mean multiplied by the CFAR coefficient (6 dB in bad weather and 4 dB in good conditions). Points exceeding the threshold are identified as targets. After clustering, the target's distance r, velocity v, and angle θ are determined.
[0093] Visual type association: The spatial position r,θ of the mmWave target is converted to image coordinates and then compared with the visually detected target bounding box using an Intersection over Union (IoU) threshold of 0.3. If a match is successful, the visually recognized type, such as "pedestrian," is assigned to the mmWave target and the confidence level is recorded. For example, if the confidence level for visual identification of a pedestrian is 0.85, the mmWave target's speed is 1.2 m / s, and its direction is perpendicular to the train's travel direction, the target is classified as a "moving pedestrian."
[0094] Trajectory Analysis: Kalman filtering is performed on at least three points of five consecutive frames of millimeter-wave target data to predict the next moment's position. Trajectory fitting uses a quadratic polynomial, x(t) = at² + bt + c, with coefficients calculated using the least squares method. The prediction error is less than 0.5 meters. If the trajectory intersects the train's route and the distance is less than 50 meters, an alert is triggered.
[0095] Ultrasonic Verification: For close targets detected by millimeter wave (mmWave) less than 10 meters, ultrasonic radar data is used for verification. The mmWave target position is compared with the contour point cloud detected by ultrasonic detection, and a Hausdorff distance threshold of 0.8 meters is calculated. If the distance is greater than 0.8 meters, the results are considered inconsistent, and the mmWave radar's CFAR processing and clustering are restarted.
[0096] Specifically, such as Figures 1 to 4 As shown, the obstacle detection method in the ultrasonic detection mode includes fusing the three-dimensional coordinate data collected by the ultrasonic radar and the edge feature data collected by the visual image sensor to obtain a reconstructed obstacle geometry, and analyzing the motion velocity vector data provided by the millimeter-wave radar to obtain motion trend data, predicting the displacement path of close-range obstacles, and fusing the displacement path with the geometry.
[0097] Ultrasonic 3D coordinate reconstruction: The MLS algorithm is used to smoothly reconstruct the discrete points (approximately 30-50 points per target) obtained by ultrasonic TDOA positioning. A local coordinate system is defined, and the k-nearest neighbors (k=10) of each point are calculated. A quadratic polynomial surface is constructed and fitted to obtain a continuous 3D contour. The reconstruction error is less than 2mm, and the basic shape of the target, such as a cuboid or cylinder, can be restored.
[0098] Visual edge feature fusion: The 2D contour points (x,y) obtained by visual edge detection are converted to 3D coordinates using camera intrinsic and extrinsic parameters and then merged with the 3D point cloud reconstructed by ultrasound. The PointNet++ network is used to extract features from approximately 200-500 points in the merged point cloud, outputting a 1024-dimensional feature vector representing the target's geometric shape, such as angles and curvature.
[0099] Millimeter-wave motion trend prediction: The millimeter-wave velocity vectors vx, vy, and vz and the past 10 frames of historical position data are fed into an LSTM network to predict the position in one, two, and three seconds. The prediction model was trained on 100,000 historical motion trajectories, achieving an average prediction error of less than 0.3 meters.
[0100] Feature-level fusion: The shape features output by PointNet++ are matched with a preset obstacle shape feature library containing 10 categories, including rails, rocks, and pedestrians. The cosine similarity is calculated with a threshold of 0.6.
[0101] Decision-level fusion: When the conflict probability between the millimeter wave predicted path and the train route is greater than 0.5, it is judged to be risky;
[0102] Final decision: If feature matching is successful and the collision probability is greater than 0.5, it is determined to be an obstacle; or if feature matching is successful and the ultrasonic distance is less than 5 meters, it is determined to be an obstacle. The fusion result is updated every 100ms.
[0103] Specifically, such as Figures 1 to 4As shown, the obstacle detection step includes setting environmental weighting values based on different environmental types, performing environmental weighting processing on the fused detection data, inputting the processed data into a preset obstacle existence probability model, and outputting a comprehensive obstacle confidence score. The existence probability model calculation includes invoking a preset environmental weighting strategy based on the current environmental type (good visual environment, bad weather environment, or close-range complex environment) determined in the environmental perception classification step. This strategy defines a weighting rule for each modal feature in the multimodal fusion data based on the reliability characteristics of different sensors in each environmental type. For example, in a good visual environment, the weight of visual image features is higher than that of radar features; in a bad weather environment, the weight of millimeter-wave radar features is dominant; and in a close-range complex environment, the weight of ultrasonic radar features is prioritized. The multimodal data processed and fused in the detection data acquisition step, including visual image features, millimeter-wave radar features, and ultrasonic radar features, are weighted and integrated according to the weighting rule determined in step 1. During the integration process, the contributions of highly reliable modal features are strengthened and the influence of modal features affected by environmental interference is weakened by logically correlating the relevance of each modal feature, such as spatial position matching and feature consistency. The integrated data, after environmental weighting, is then fed into a pre-set obstacle presence probability model. This model, constructed based on the joint distribution of multimodal features, calculates the probability of an obstacle's presence by analyzing the integrated data for common features that reflect obstacle characteristics, such as specific geometric shapes, motion trends, and reflective properties. The model logically determines the degree of match between each feature and prior "obstacle" characteristics, such as feature overlap and trend consistency, gradually accumulating evidence supporting the presence or absence of an obstacle. Based on these feature matching analysis results, the model outputs a comprehensive confidence score that quantifies the probability of the obstacle's presence. This confidence score is generated by summarizing the strength of evidence supporting the presence of an obstacle from each modal feature and combining it with the environmentally weighted feature contributions to form an overall probability indicator. This indicator is positively correlated with the probability of the obstacle's presence, and the output comprehensive confidence score is then compared with a pre-set threshold. The threshold setting logic is based on the need to balance detection accuracy and false alarm rate. For example, both "no missed detection" and "no false alarm" need to be taken into account. In scenarios with high security requirements, the threshold can be appropriately lowered to reduce missed detections, and in normal scenarios, the threshold can be moderately increased to reduce false alarms.
[0104] The obstacle comprehensive confidence is compared with a preset threshold value. If the obstacle comprehensive confidence is greater than the preset threshold value, it is determined and output that there is an obstacle.
[0105] According to the environmental judgment value V obtained in the embodiment, a weighting coefficient W is generated through a fuzzy logic system:
[0106] Good visual environment V ≥ 80: W visual = 0.5, W millimeter wave = 0.3, W ultrasonic = 0.2;
[0107] Severe weather environment 40≤V<80: W visual = 0.3, W millimeter wave = 0.5, W ultrasonic = 0.2;
[0108] Close-range complex environment V<40: W vision = 0.2, W millimeter wave = 0.3, W ultrasonic = 0.5.
[0109] The membership function of fuzzy logic adopts Gaussian type, and the rule base contains 9 rules such as "If the environment is poor and the millimeter wave trust is high, the millimeter wave weight is high."
[0110] Weighted processing of fusion data: The fusion results in each mode, such as the MLP output of the visual mode, the decision result of the millimeter wave mode, and the mixed fusion result of the ultrasonic mode, are converted into confidence values C visual, C millimeter wave, and C ultrasonic between 0 and 1, and the comprehensive confidence C = W visual × C visual + W millimeter wave × C millimeter wave + W ultrasonic × C ultrasonic is calculated.
[0111] Probabilistic model calculation: The comprehensive confidence C is used as evidence input into the Bayesian network. The network structure includes:
[0112] Environment node: status {good, bad, complex};
[0113] Sensor node: state {reliable, general, unreliable}, converted by trust value. Trust value > 0.6 is reliable, 0.3-0.6 is general, and < 0.3 is unreliable;
[0114] Obstacle node: state {exists, does not exist}.
[0115] The posterior probability of the existence of an obstacle is calculated through forward propagation, and the comprehensive confidence of the obstacle Cprob is output.
[0116] Decision output: Compare Cprob with the preset threshold T. The default T=0.6:
[0117] Cprob>T: Determine the presence of an obstacle, output information such as location, type, and confidence level, and trigger an early warning;
[0118] Cprob≤T: It is determined that there is no obstacle and monitoring continues.
[0119] The threshold T can be adjusted from 0.4 to 0.8 through the vehicle's human-machine interface to meet different safety level requirements.
[0120] The above shows and describes the basic features, principles, and advantages of the present invention. It should be noted that the present invention is not limited to the above embodiments, which are only some embodiments. Without departing from the spirit and scope of the present invention, various improvements and supplements made are considered to be within the scope of protection of the present invention.
Claims
1. A rail transit vehicle obstacle detection method based on multimodal data fusion, characterized in that: The steps include: Environmental perception and classification: Environmental sensors collect real-time data on the train's operating environment and combine it with historical data to classify the current detection environment into categories such as good visual environment, adverse weather environment, and close-range complex environment. The trust calculation and mode decision step sets the initial trust of different detection modes, calculates the trust of different detection modes based on the environment type and sensor data collected in combination with the initial trust, and selects a detection mode based on the trust. The detection modes include visual image detection mode, millimeter wave radar detection mode, and ultrasonic detection mode; a detection data collection step, determining the amount of data collected by each detection sensor according to different detection modes and performing data collection and data processing, wherein the detection sensors include visual image sensors, millimeter wave radars, and ultrasonic radars; The obstacle detection step is to perform multimodal data fusion on the processed data according to the detection mode to obtain corresponding data to be detected, and then perform obstacle detection based on the data to be detected and output the detection result; The trust calculation includes dynamically adjusting the trust based on the environment type, presetting the sensor trust adjustment coefficient for each environment type, the adjustment coefficient of the visual image sensor in a good visual environment is a positive value, the adjustment coefficient of the millimeter-wave radar in a bad weather environment is a positive value, and the adjustment coefficient of the ultrasonic radar in a close-range complex environment is a positive value. The DS evidence theory is used to perform a conflict analysis on the detection results of the same target by the visual image, millimeter-wave radar and ultrasonic radar. If the position deviation of the detection results of the three is within a preset threshold, the trust of the three increases simultaneously; if the visual image and millimeter-wave radar detection conflict and the ultrasonic radar supports one of them, the trust of the conflicting party is reduced, and a weighted calculation is performed based on the initial trust to obtain the final trust value of each detection mode.
2. The rail transit vehicle obstacle detection method based on multimodal data fusion according to claim 1, characterized in that: The environmental sensor includes a light sensor and a temperature and humidity sensor. The environmental sensor collects ambient light intensity, temperature data and humidity data as environmental collection data. An environmental judgment model is constructed based on the environmental collection data combined with historical data, and an environmental judgment value is calculated. The environmental type judgment result is obtained based on the comparison result of the environmental judgment value and the preset threshold.
3. The rail transit vehicle obstacle detection method based on multimodal data fusion according to claim 1, characterized in that: When the detection mode is visual image detection, the visual image sensor collects full-resolution image data and extracts obstacle features through the Faster R-CNN algorithm; The millimeter wave radar collects distance data and azimuth data of the detection target in the long-distance area within the current detection range; the ultrasonic radar collects target position data of the detection target in the short-distance area within the current detection range.
4. The rail transit vehicle obstacle detection method based on multimodal data fusion according to claim 1, characterized in that: When the detection mode is millimeter-wave radar detection, the millimeter-wave radar collects the full amount of point cloud data within the current detection range and extracts feature data through the CFAR algorithm; the visual image sensor collects the type recognition data of the moving target within the current detection range; the ultrasonic radar collects the contour data of the detection target in the close range area within the current detection range.
5. The rail transit vehicle obstacle detection method based on multimodal data fusion according to claim 1, characterized in that: When the detection mode is ultrasonic detection, the ultrasonic radar collects TDOA positioning data and reconstructs the three-dimensional coordinates of close-range obstacles; the visual image sensor collects edge feature data of the detection target within the current detection range; and the millimeter wave radar collects motion speed vector data of the detection target within the current detection range.
6. The rail transit vehicle obstacle detection method based on multimodal data fusion according to claim 3, characterized in that: The obstacle detection method in the visual image detection mode includes calculating image blurriness through image quality evaluation of obstacle features extracted from the full-resolution image data. When the image blurriness is higher than a preset blurriness threshold, the distance data collected by the millimeter-wave radar is spatially matched with the full-resolution image data collected by the visual image sensor in combination with millimeter-wave radar and ultrasonic radar detection, and the depth information of the full-resolution image data is corrected. The detection target detected by the ultrasonic radar is superimposed on the full-resolution image data with a highlighted mark based on the target position data. The data features of the full-resolution image data, the distance features obtained by combining the distance data and azimuth data collected by the millimeter-wave radar, and the position features of the target object collected by the ultrasonic radar are concatenated and input into a multi-layer perceptron for feature-level fusion.
7. The rail transit vehicle obstacle detection method based on multimodal data fusion according to claim 4, characterized in that: The obstacle detection method in millimeter-wave radar detection mode includes using the CFAR processing results of the special data of the millimeter-wave radar as the main detection data, generating the target's distance, speed and angle information, correlating the target type recognition data provided by the visual image sensor with the motion trajectory obtained by analyzing the point cloud data of the millimeter-wave radar detection target, and assigning a semantic category identification to the target when the type recognition data successfully matches the spatial position of the millimeter-wave target.
8. The rail transit vehicle obstacle detection method based on multimodal data fusion according to claim 5, characterized in that: The obstacle detection method in ultrasonic detection mode includes fusing the three-dimensional coordinate data collected by the ultrasonic radar and the edge feature data collected by the visual image sensor to obtain a reconstructed obstacle geometry, analyzing the motion velocity vector data provided by the millimeter-wave radar to obtain motion trend data, predicting the displacement path of close-range obstacles, and fusing the displacement path with the geometry.
9. The rail transit vehicle obstacle detection method based on multimodal data fusion according to claim 1, characterized in that: The obstacle detection step includes setting environmental weighting values according to different environment types, performing environmental weighting processing on the fused detection data, inputting the processed data into a preset obstacle existence probability model and outputting a comprehensive obstacle confidence score, comparing the comprehensive obstacle confidence score with a preset threshold, and if the comprehensive obstacle confidence score is greater than the preset threshold, determining and outputting a result that an obstacle exists.
Citation Information
Patent Citations
Evidence distance optimization-based evidence fusion method in D-S (Dempster-Shafer) evidence theory
CN108428008A
Obstacle detection method and obstacle detection system for track
CN120288091A