Radar-image fusion obstacle identification method, device, equipment, medium and product

CN122551316APending Publication Date: 2026-08-11西北铁道电子股份有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-12
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0006]鉴于现有技术存在的上述缺陷与不足,本申请期望提供一种雷达-图像融合障碍物识别方法、装置、设备、介质及产品,可有效解决复杂调车环境下近距离障碍物精准识别与威胁分级问题

Benefits of technology

本申请提供了一种雷达-图像融合障碍物识别方法、装置、设备、介质及产品,通过根据定位信息加载轨道几何参数并对多源数据(雷达点云数据和图像数据)进行自适应预处理,有效利用了轨道交通的场景先验知识,滤除了轨道环境杂波与非关注区域背景,解决了复杂调车环境下虚警率高的问题,实现了计算资源的聚焦与干扰的显著抑制;通过对雷达与视觉信息进行特征层融合与关联,综合了雷达精确的测距测速能力与视觉丰富的语义分类能力,解决了单一传感器对近距离、小尺寸或静态障碍物易漏检的问题,提升了对各类目标的整体识别率与系统鲁棒性;通过结合轨道行驶包络空间约束与多维度威胁量化进行决策验证,解决了传统方案难以有效排除轨旁非威胁目标及威胁评估维度单一的问题,实现了对侵入轨道障碍物的精准威胁感知,从而显著提升了在复杂调车环境下的安全性、可靠性与作业效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
Patent Text Reader

Abstract

This application discloses a radar-image fusion obstacle recognition method, apparatus, device, medium, and product for LFK systems, relating to the field of active safety technology in rail transit. The method includes: acquiring radar point cloud data, image data, and positioning information in front of the locomotive; clustering the filtered point cloud to generate radar candidate targets and their feature vectors; detecting regions of interest in the image to generate image candidate targets and their feature vectors; generating fused targets and their fused feature vectors; verifying whether the target is located within the track travel envelope space; and calculating the threat level of targets within the track travel envelope space based on their distance, speed, size, and category information. This application effectively improves the recognition accuracy of close-range, small-sized obstacles in complex shunting environments by introducing prior information about the track scene for data preprocessing and focusing, combined with feature layer fusion and multi-dimensional threat quantification, reducing false alarms and false negatives, and achieving accurate threat classification and early warning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of active safety technology for rail transit, and in particular to a radar-image fusion obstacle recognition method, device, equipment, medium and product for LFK systems. Background Technology

[0002] The Train Intelligent Observation and Collision Avoidance Monitoring System (hereinafter referred to as the LFK system) is a key piece of equipment for ensuring the safe operation of dedicated railway lines, shunting operations, and railway junctions. Currently, obstacle detection technology in the rail transit field mainly relies on single sensors or multi-sensor fusion solutions.

[0003] Chinese patent CN112406960A discloses "An active collision avoidance system and method for subways using multi-sensor fusion". This solution uses multiple sensors such as lidar, millimeter-wave radar, cameras, and secondary radar, and attempts to fuse them at the data level (i.e., pre-fusion). It aims to solve the collision avoidance problem of subway trains running at long distances (such as more than 950 meters), focusing on ranging and rear-end collision prevention warnings for non-cooperative targets such as moving trains ahead.

[0004] However, when applying solutions designed for long-distance operation on subway mainlines to complex operational scenarios such as shunting, shortcomings remain. Shunting operations are characterized by close proximity (0-100 meters), numerous small obstacles, and variable scenarios (e.g., backlighting, switch areas), and existing solutions lack targeted optimization. For example, they fail to fully utilize high-precision prior information about the track to strongly constrain the sensing area, making it difficult to effectively filter out interference from non-threatening targets near the track (e.g., adjacent vehicles, fixed facilities), resulting in a high false alarm rate. Furthermore, the detection accuracy and reliability for static or slow-moving small obstacles (e.g., abandoned tools, gravel on the track) need improvement. In addition, threat assessment relies heavily on simple distance thresholds, failing to comprehensively consider obstacle distance, size, type, and relative speed for refined classification, thus failing to meet the demand for synergistic improvement in safety and efficiency in shunting operations.

[0005] Therefore, in the existing technology, the LFK system still lacks the ability to accurately identify and finely classify threats to intruding obstacles on the track in complex shunting environments, which poses a risk of missed or false alarms and affects operational safety and efficiency. Summary of the Invention

[0006] In view of the above-mentioned defects and deficiencies in the existing technology, this application aims to provide a radar-image fusion obstacle recognition method, device, equipment, medium and product, which can effectively solve the problem of accurate identification and threat classification of close-range obstacles in complex shunting environments.

[0007] To achieve the above objectives, this application provides the following solution: In a first aspect, this application provides a radar-image fusion obstacle recognition method, including: Acquire radar point cloud data, image data, and current locomotive positioning information from the radar in front of the locomotive; The track geometry parameters of the current road segment are loaded based on the positioning information, and adaptive filtering is performed on the radar point cloud data and the image data is processed based on the track geometry parameters to generate the region of interest in the track area. The filtered radar point cloud data is clustered to generate at least one radar candidate target, and a corresponding radar feature vector is constructed for each radar candidate target; the region of interest in the image is targeted to generate at least one image candidate target, and a corresponding image feature vector is constructed for each image candidate target. The radar candidate target and the image candidate target are feature-mapped and associated, and a fused target and its corresponding fused feature vector are generated based on the radar feature vector and the image feature vector. Based on the position information in the fused feature vector of the fused target and the orbital geometry parameters, it is verified whether the fused target is located within the orbital travel envelope space. For the fused target located within the orbital travel envelope space, the threat level is calculated based on the distance, speed, size and category information in its fused feature vector.

[0008] Secondly, this application provides a radar-image fusion obstacle recognition device, comprising: The multi-source perception module is used to acquire radar point cloud data, image data, and the current positioning information of the locomotive in front of it. The data preprocessing module is used to load the track geometry parameters of the current road segment according to the positioning information, and based on the track geometry parameters, to perform adaptive filtering on the radar point cloud data and process the image data to generate the region of interest in the track area. The feature fusion module is used to cluster the filtered radar point cloud data to generate at least one radar candidate target and construct a corresponding radar feature vector for each radar candidate target; to perform target detection on the region of interest in the image to generate at least one image candidate target and construct a corresponding image feature vector for each image candidate target; and to perform feature mapping and association between the radar candidate target and the image candidate target, and generate the fused target and its corresponding fused feature vector based on the radar feature vector and the image feature vector. The decision verification module is used to verify whether the fused target is located within the orbital travel envelope space based on the position information in the fused feature vector of the fused target and the orbital geometric parameters, and to calculate the threat level of the fused target located within the orbital travel envelope space based on the distance, speed, size and category information in its fused feature vector. The data interaction module is used to output the information in the target's fused feature vector and its threat level to the outside world.

[0009] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the radar-image fusion obstacle recognition method described in any one of the above.

[0010] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the radar-image fusion obstacle recognition method described above.

[0011] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the radar-image fusion obstacle recognition method described above.

[0012] According to the specific embodiments provided in this application, the following technical effects are disclosed: This application provides a radar-image fusion obstacle recognition method, device, equipment, medium, and product. By loading track geometric parameters based on positioning information and adaptively preprocessing multi-source data (radar point cloud data and image data), it effectively utilizes prior knowledge of rail transit scenarios, filters out track clutter and background noise in non-interested areas, solves the problem of high false alarm rates in complex shunting environments, and achieves focused computing resources and significant suppression of interference. By fusing and associating radar and visual information at the feature layer, it combines the precise ranging and velocity measurement capabilities of radar with the rich semantic classification capabilities of vision, solving the problem of single sensors easily missing close-range, small-sized, or static obstacles, and improving the overall recognition rate and system robustness for various targets. By combining track travel envelope space constraints and multi-dimensional threat quantification for decision verification, it solves the problem that traditional solutions are difficult to effectively exclude non-threatening targets along the track and the problem of single threat assessment dimensions, achieving accurate threat perception of obstacles intruding into the track, thereby significantly improving safety, reliability, and operational efficiency in complex shunting environments. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 This is a flowchart illustrating a radar-image fusion obstacle recognition method according to an embodiment of this application. Detailed Implementation

[0015] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0016] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0017] In one exemplary embodiment, a radar-image fusion obstacle recognition method is provided. This method is executed by a computer device, specifically, it can be executed by a computer device such as a terminal or a server alone, or it can be executed by both a terminal and a server. In this embodiment of the application, such as... Figure 1 As shown, the method is applied to a server as an example for illustration, including the following steps 1] to 5], wherein: Step 1】Acquire radar point cloud data, image data, and current locomotive positioning information in front of the locomotive; Step 2】Load the track geometry parameters of the current road segment according to the positioning information, and based on the track geometry parameters, perform adaptive filtering on the radar point cloud data and process the image data to generate the region of interest in the track area. Step 3: Cluster the filtered radar point cloud data to generate at least one radar candidate target, and construct a corresponding radar feature vector for each radar candidate target; perform target detection on the region of interest in the image to generate at least one image candidate target, and construct a corresponding image feature vector for each image candidate target; Step 4: Perform feature mapping and association between the radar candidate target and the image candidate target, and generate the fused target and its corresponding fused feature vector based on the radar feature vector and the image feature vector; Step 5: Based on the position information in the fused feature vector of the fused target and the orbital geometry parameters, verify whether the fused target is located within the orbital travel envelope space, and calculate the threat level of the fused target located within the orbital travel envelope space based on the distance, speed, size and category information in its fused feature vector.

[0018] By executing steps 1 through 5 above, adaptive filtering of radar point cloud data is performed using track geometry parameters, and a region of interest (ROI) is generated from the image data. This effectively filters out track clutter and background interference, focusing the computational area. By mapping and fusing radar features and image features at the feature layer, the system comprehensively utilizes the radar's precise ranging and velocity measurement capabilities with the rich semantic classification capabilities of vision. This significantly improves the overall recognition rate and system robustness for static, small-sized targets that are easily missed by traditional single sensors. By combining position verification in the track travel envelope space with multi-dimensional threat quantification of distance, speed, size, and category information from the fusion features, non-threat targets outside the track clearance can be intelligently excluded, and real threat targets can be finely risk-classified. This enables high-precision, low-false-alarm obstacle perception and early warning in complex shunting environments, ensuring operational safety and efficiency.

[0019] In another exemplary embodiment of this application, the steps of obtaining radar point cloud data, image data, and current locomotive positioning information in front of the locomotive are specifically implemented through the following steps 101 to 104: Step 101: Using the vehicle-mounted millimeter-wave radar deployed at the front of the locomotive, collect raw radar point cloud data within a preset detection range directly in front of the locomotive; each data point in the radar point cloud data contains physical information in four dimensions: distance (the straight-line distance of the target point cloud relative to the radar), relative velocity (the radial velocity of the target point cloud measured based on the Doppler effect), azimuth angle (the angle of the target point cloud on the horizontal plane relative to the longitudinal axis of the locomotive), and radar cross section (a physical quantity reflecting the reflection intensity of the target point cloud). Step 102: Using an industrial-grade visible light camera that is calibrated and aligned with the optical axis of the millimeter-wave radar, synchronously acquire continuous frame images of the same forward field of view to form video stream image data. Step 103: Through the vehicle bus of the LFK system, read in real time the high-precision positioning information (including latitude and longitude, speed, and timestamp) output by the Beidou positioning unit integrated on the locomotive, and retrieve the route data identifier corresponding to the current geographical location index from the vehicle-mounted route electronic map database. Step 104: Using the main control unit clock of the LFK system as a unified time reference, for each frame of image data and the radar point cloud data of the radar scanning cycle near the same time, a hard trigger signal or a high-precision timestamp interpolation method is used to time-align the radar point cloud data and the image data on the time axis to ensure that the time deviation between the two is not greater than a preset threshold, so as to achieve spatiotemporal consistency in subsequent fusion processing.

[0020] Steps 101 and 102 acquire radar point cloud data containing precise distance and speed information, and image data rich in texture and semantic information, respectively, providing complete and complementary raw perceptual input for subsequent fusion recognition. Step 103 loads the track geometry parameters of the current road segment in real time using positioning information, introducing prior knowledge of track constraints in rail transit, laying the foundation for subsequent scene focusing and intelligent filtering. Step 104 ensures strict spatiotemporal consistency between radar data and image data through hardware triggering or high-precision timestamp alignment, solving the problem of association errors and accuracy degradation caused by temporal misalignment in multi-sensor fusion. Steps 101 to 104 together constitute a reliable data acquisition and preprocessing front end, effectively reducing noise, asynchronicity, and scene-irrelevant interference in the raw data, providing accurate and consistent input for subsequent feature layer fusion, target association, and threat decision-making, which is the primary guarantee for the high accuracy and high reliability of the entire recognition method.

[0021] For example, A1 uses an onboard millimeter-wave radar to collect raw radar point cloud data within a preset detection range (e.g., 0-100 meters) in front of the locomotive. Radar point cloud data includes information on range, relative velocity, azimuth, and radar cross section. A2. Utilize industrial-grade visible light cameras deployed in the same field of view to synchronously acquire video stream image data. ; A3. Through the LFK system bus, read the current locomotive's BeiDou positioning information or Global Positioning System information and the corresponding line database index information in real time; A4. Using the master clock of the LFK system as a reference, timestamp interpolation is employed to analyze radar point cloud data. With image data Perform time alignment and ensure that the two times are synchronized, and that their error does not exceed a preset threshold (e.g., ).

[0022] Specifically, in the shunting operation environment of the station, when collecting radar point cloud data using radar, the onboard millimeter-wave radar deployed behind the locomotive obstacle clearer is activated. Its emitted frequency-modulated continuous wave scans a range of 0 to 100 meters directly in front of the locomotive. The radar receives echoes reflected from a discarded tool (wrench) on the track ahead and from a crouching construction worker approximately 50 meters away, generating raw radar point cloud data containing hundreds of points. An example of information included in the tool reflection point is: distance = 12.5 meters, relative speed = 0 m / s (static), azimuth = +0.5 degrees (slightly to the right of the track), radar cross-section = -20 dBsm (small target).

[0023] When acquiring image data using the camera, an industrial-grade visible light camera, mounted on the same beam as the radar and optically calibrated to ensure field-of-view overlap, is simultaneously activated. The camera acquires video streams at a rate of 25 frames per second. In the original image of the same scene captured by the camera at dusk with backlighting, the distant sky is overexposed, while the track area is poorly lit, and the outlines of construction workers and tools are blurred.

[0024] The positioning information output by the locomotive's Beidou positioning unit is read in real time using the vehicle bus of the LFK system: longitude 118.3E, latitude 35.1N, speed 5 km / h, and UTC timestamp. This latitude and longitude coordinate is matched with the onboard electronic map database to retrieve the locomotive's current location on the 5th trainset track within the station. The track geometry parameters of that track are then immediately retrieved, including: it is currently a straight track with a nominal track width of 1435 mm.

[0025] Synchronous data acquisition is achieved by the LFK system's main control unit sending hard trigger signals to the radar and camera at full seconds. Upon receiving the trigger signal, the radar begins a scanning cycle; the camera initiates exposure for the current frame on the rising edge of the same trigger signal. Subsequent verification shows that the timestamp difference between the radar point cloud data and the image data is 15 milliseconds, less than the preset threshold of 50 milliseconds, meeting the spatiotemporal consistency requirement. The data from both sensors are marked as environmental snapshots at the same time and sent to subsequent processing steps. Step 1 of this application, in a complex station operation environment, achieves accurate and synchronous acquisition of multi-source heterogeneous data, providing high-quality input data for subsequent feature-level fusion and threat identification.

[0026] In another exemplary embodiment of this application, in order to filter out environmental noise, narrow the calculation range, and improve the real-time performance of the system, the steps of loading the track geometry parameters of the current road segment according to the positioning information, and adaptively filtering the radar point cloud data to remove noise points based on the track geometry parameters, and processing the image data to generate the region of interest of the track area, are specifically implemented through the following steps 201 to 203: Step 201: Load prior information of the orbital scene Based on the locomotive positioning information obtained from the vehicle bus in step 101, the preset line database is queried to retrieve the track geometry parameters corresponding to the current section (for example, information such as turnout type, curve radius, and track width may be included). Step 202: Perform adaptive filtering on the radar point cloud data. Based on the track geometry parameters loaded in step 201, determine the current road segment type (e.g., straight road, curve, or turnout area), and set the dynamic filtering radius according to different road segment types (e.g., use a smaller radius for straight road areas and a larger radius for curve areas). For each data point in the radar point cloud data, calculate the number of other points (i.e., neighborhood point density, which can be dynamically adjusted according to the orbit scene and sensor parameters) within a spherical space centered on that point and with the aforementioned dynamic filtering radius as the radius. If the density of neighboring points of a certain point is lower than the preset density threshold, the point is determined to be an outlier noise point caused by multipath reflection or other reasons and is removed; otherwise, it is retained as a valid echo point. Step 203: Construct the region of interest for image processing. Enhancement processing of raw image data (e.g., using the Retinex algorithm) can improve image quality under backlighting or low-light conditions; Based on edge detection algorithms (such as the Canny operator) and combined with the projection of the orbital geometric parameters obtained in step 201 onto the image plane, an accurate orbital mask is generated. The region enclosed by extending the boundary of this track mask outward by a preset pixel width (corresponding to a safety margin in the physical world, such as 1 meter) is defined as the region of interest in the image, and subsequent image candidate object detection will only be performed within this region.

[0027] Step 202 dynamically adjusts the filtering strategy based on track geometry parameters, employing differentiated filtering radii and density thresholds for different scenarios such as straight sections and curves. This effectively identifies and filters out specific outlier noise points generated by rail multipath reflections, significantly improving the signal-to-noise ratio and data quality of radar point cloud data in complex track environments. Step 203 enhances the original image (using algorithms such as Retinex), effectively improving image quality under harsh lighting conditions such as backlighting and low illumination. This ensures the effectiveness of subsequent visual detection algorithm inputs, enhances adaptability under all-weather conditions, and strengthens robustness in complex lighting environments. Step 203 generates precise track masks and defines regions of interest in the image, strictly focusing computational resources on the track and its surrounding limited safe area. This completely eliminates interference from a large number of irrelevant background areas such as the sky and distant buildings. While ensuring coverage of the perception range, it significantly reduces the computational load of image processing, meeting the stringent real-time requirements of the LFK system and greatly improving processing efficiency. Through the collaborative work of steps 201 to 203, radar data and image data are targeted for purification and focusing based on prior knowledge, producing high-quality, low-noise, and region-focused preprocessing results. This provides crucial high-quality data input for subsequent steps such as reliable feature extraction, accurate cross-modal association, and fusion, and is the key preprocessing guarantee for the entire method to achieve high accuracy and high reliability.

[0028] For example, B1 retrieves the track geometry parameters of the current section from the preset track database based on the current locomotive's positioning information. (Including turnout type, curve radius, and track width); B2. Define radar point cloud data any point in the middle Coordinates in a spatial coordinate system , with point For the center of the ball, The dynamic filtering radius is used to obtain the spherical space; where , ; If the current orbit parameters If the indication is a straight path, then a small dynamic filter radius is set based on the geometric characteristics of the straight path. equal to the width of the straight road ,Right now =0.3m; If the current orbit parameters If the indication is a curve / turnout, then the dynamic filtering radius is set according to the geometric characteristics of the curve / turnout. Equal to the curve radius of the curve / turnout ,Right now 5m; B3. Count the number of neighborhood points within the spherical space. For radar point cloud data Determine the value at each point: like ( For the preset density threshold, Then the decision point Points identified as outliers are removed; otherwise, the decision point is... Points within the spherical space are preserved; B4. Using image enhancement algorithms to process image data. Enhancement processing is performed to obtain enhanced images to cope with backlighting or low-light tunnel environments; B5. Edge detection algorithms are used to extract the edges of the original image, combined with Hough transform and current orbital parameters. Projection to generate orbital mask ; B6. Set the track mask Expanding the image to both sides by a preset pixel width (corresponding to 1m of physical space), the resulting area is defined as the region of interest in the image, and subsequent processing only retains the image data within this region.

[0029] Specifically, during nighttime shunting operations at the station, the onboard track database will be queried based on the locomotive's location information (longitude 118.3E, latitude 35.1N). The database will confirm that the locomotive is located on the 5th train line within the station, and the track geometry parameters for that track will be retrieved: the current section is a straight track, there are no turnouts, the curve radius is infinite (for straight tracks), and the track width is 1435 mm.

[0030] Based on the loaded track geometry (straight track), the dynamic filtering radius is set to 0.3 meters (a smaller radius is suitable for straight tracks). The raw radar point cloud data acquired by the radar contains mixed echoes from the track, abandoned tools, construction workers, and multipath reflections (such as from adjacent rails).

[0031] Density statistics are performed on each radar point cloud data point. For example, a real tool reflection point (coordinates x=12.5, y=0.3, z=0.1) has 25 neighboring points within a 0.3-meter spherical space, which is higher than the preset density threshold of 15, and is therefore retained. However, a false point generated by multiple reflections from a rail (coordinates x=10.1, y=2.5, z=0.2) has only 3 neighboring points within a 0.3-meter spherical space, which is lower than the threshold of 15, and is identified as an outlier and removed.

[0032] In low-light conditions at night, the original image captured by the camera is first enhanced using the Retinex algorithm to improve the contrast and detail visibility of the track area. Then, edge detection is performed on the enhanced image using the Canny operator, and a binary track mask precisely corresponding to the physical track position is generated by projecting the track geometry parameters (straight track, 1435 mm wide) obtained in step 201 onto the image plane. Finally, the left and right boundaries of this track mask are extended outwards by a certain pixel width (e.g., 100 pixels; according to camera calibration parameters, this pixel width corresponds to a preset safety margin in the physical world, such as 1 meter), to cover the potential threat area around the track, forming a rectangular region of interest. Background pixels such as the sky and distant buildings outside this region are completely masked, and subsequent object detection only processes the image within this region, significantly reducing computational load.

[0033] In another exemplary embodiment of this application, in order to leverage the complementary advantages of radar ranging and visual texture, the filtered radar point cloud data is clustered to generate at least one radar candidate target, and a corresponding radar feature vector is constructed for each radar candidate target; target detection is performed on the region of interest in the image to generate at least one image candidate target, and a corresponding image feature vector is constructed for each image candidate target, which is specifically implemented through the following steps 301 to 304: Step 301: Density Clustering Processing A density-based clustering algorithm is used to perform cluster analysis on the preprocessed radar point cloud data, dividing the spatially dense radar point cloud data into different clusters; each cluster is identified as an independent radar candidate target. Step 302: Feature Vector Construction For each radar candidate target, its corresponding radar feature vector is calculated and generated. This radar feature vector is a multi-dimensional mathematical representation, typically including dimensions such as longitudinal range, relative velocity, azimuth angle, and equivalent surface area; where: The longitudinal distance is defined as the longitudinal distance (in meters) between the centroid or the nearest point of the target cluster in the vehicle coordinate system; the relative velocity is the average relative velocity or radial velocity (in meters per second) calculated based on the Doppler velocity information of the point cloud within the target cluster; the azimuth angle is the azimuth angle (in degrees) of the centroid of the target cluster on the horizontal plane; and the equivalent surface area is the estimated two-dimensional or three-dimensional spatial dimension (in square meters) of the target cluster based on its spatial distribution, used to characterize the physical size of the target cluster.

[0034] Step 303, Target Detection Inference The region of interest in the image is fed into a lightweight object detection neural network, which is trained to recognize common object categories in railway scenes. The neural network performs forward propagation inference and outputs the detection results. Step 304: Detection Result Analysis and Feature Vector Construction The detection results output by the neural network are analyzed to generate a series of image candidate targets and their corresponding image feature vectors. Each image feature vector contains the following information: Detection box coordinates: Detection box coordinates in pixels, typically including the x-coordinate of the top-left corner, the y-coordinate of the top-left corner, the x-coordinate of the bottom-right corner, and the y-coordinate of the bottom-right corner; Target category label: The semantic category of the target determined by the target detection neural network, such as "person", "vehicle", "general obstacle"; Confidence score: The probability that the detection result given by the neural network belongs to its predicted category, with a value ranging from 0 to 1; Steps 301 and 302 above, by performing density clustering on the filtered radar point cloud data and constructing radar feature vectors, transform discrete radar echo points into radar candidate targets with clear physical meanings (longitudinal distance, relative velocity, azimuth angle, equivalent surface area), providing precise spatial location, motion state, and size information of obstacles, and achieving precise physical feature extraction from the radar. Steps 303 and 304, by running a lightweight target detection neural network within the region of interest of the focused image and constructing image feature vectors, provide precise pixel-level localization of obstacles, specific semantic categories (such as people, vehicles), and their recognition confidence, compensating for the shortcomings of radar in target classification capabilities. Through steps 301 to 304, radar feature vectors and image feature vectors with complementary representational information are generated from two independent perception dimensions: radar and vision. This separate but structured feature extraction method provides direct and standardized input for subsequent steps of accurate feature mapping, association, and fusion, and is a core prerequisite for achieving highly reliable cross-modal information complementarity. Step 303 explicitly adopts a lightweight target detection neural network to process the reduced image region of interest. Combined with the efficient radar point cloud data clustering algorithm in step 301, it ensures the computational efficiency of the feature extraction stage and meets the requirements of real-time system response in complex shunting operation environments.

[0035] For example, C1 uses the DBSCAN algorithm to perform cluster analysis on the preprocessed radar point cloud data to generate... Radar candidate targets ( ); C2. For each radar candidate target, construct its corresponding radar feature vector; here, the radar feature vector is a set of multiple physical quantities used to characterize the attributes of the radar candidate target. In one embodiment, this radar feature vector... : ; in, Indicates the first The longitudinal distance of each radar candidate target relative to the locomotive; Indicates the first The radial velocity of each radar candidate target relative to the locomotive; Indicates the first The azimuth angle of each radar candidate target on the horizontal plane; Indicates according to the first Equivalent surface area estimated from point cloud clusters of radar candidate targets; C3. Input the region of interest in the image into a lightweight object detection convolutional neural network (e.g., YOLOv8-Nano), and the network outputs... Image candidate targets (m) ); C4. For each image candidate target, construct its corresponding image feature vector; this image feature vector integrates the visual and semantic information of the image candidate target. In one embodiment, this image feature vector... : ; in, These represent the pixel coordinates of the top-left corner of the detection box of the candidate target in the image; These represent the pixel coordinates of the lower right corner of the detection box of the candidate target in the image; This represents the target category label determined by the target detection neural network (e.g., "person", "vehicle", "foreign object"). This represents the confidence level of the detection result given by the object detection neural network.

[0036] Specifically, during nighttime shunting operations, the DBSCAN algorithm was used to perform cluster analysis on the filtered and retained effective radar point cloud data. Based on the spatial density distribution of the radar point cloud data, this algorithm successfully divided it into two independent clusters: Cluster A: Contains approximately 30 spatially concentrated points, corresponding to the abandoned tool at a distance of 12.5 meters on the track.

[0037] Cluster B: Contains approximately 150 spatially concentrated points, corresponding to the semi-squatting construction workers 50 meters away on the track.

[0038] Cluster A is identified as a radar candidate target (tool), and cluster B is identified as another radar candidate target (personnel).

[0039] For the two radar candidate targets mentioned above, their radar feature vectors are constructed respectively: Radar feature vector of the target (cluster A): calculated ; Radar feature vector of personnel targets (cluster B): calculated ; The region of interest (the enhanced track area image after nighttime illumination) was input into the deployed YOLOv8-Nano neural network, which performed forward propagation inference and output the original detection tensor. In this inference, the neural network mainly identified the relatively obvious construction worker targets in the image, but due to insufficient light at night and the small size of the tools, it failed to effectively detect the tool targets.

[0040] The output of YOLOv8-Nano is parsed to generate an image candidate target and its corresponding image feature vector: Image feature vector of the person target: ; Detection box coordinates: top left corner ( (bottom right corner) ); For tools that are not detected by vision, this step does not generate corresponding image candidate targets.

[0041] This example demonstrates the complementarity between radar and visual perception: radar excels at detecting small, static targets and providing precise motion information, while vision plays a crucial role in target semantic classification and accurate localization. The difference in the information composition of the two systems is the direct motivation for subsequent feature-level fusion to improve overall recognition performance.

[0042] In another exemplary embodiment of this application, the steps described above—namely, performing feature mapping and association between the radar candidate target and the image candidate target, and generating the fused target and its corresponding fused feature vector based on the radar feature vector and the image feature vector—are specifically implemented through the following steps 401 to 404: Step 401: Coordinate Space Projection and Matching Calculation Using the pre-calibrated camera intrinsic parameter matrix and the extrinsic parameter transformation matrix between the radar and the camera, the three-dimensional spatial coordinates of each radar candidate target are projected onto the image pixel coordinate system to obtain its two-dimensional projection point on the image. Step 402: Calculation of intersection-union ratio For each radar candidate target's projection point, it can be expanded into a preset pixel region, and the cross-union ratio (CUI) between this region and the detection boxes of all image candidate targets can be calculated; where the CUI is used to quantify the degree of overlap between the radar projection region and the image detection region. Step 403: Matching Determination and Target Association A preset threshold for the intersection-union ratio is used to determine the following for any radar candidate target: If there are one or more image candidate targets such that the crossover ratio (CRO) between the radar candidate target and the corresponding image candidate target detection box is greater than the matching threshold, then the image candidate target with the largest CRO is selected and associated with it, and it is determined that the same physical entity has been observed, that is, a fused target is generated. If the intersection-union ratio of a radar candidate target with all image candidate targets is lower than the matching threshold, it is temporarily designated as a radar candidate target perceived only by radar; similarly, an image candidate target that is not matched by any radar candidate target is temporarily designated as an image candidate target perceived only by vision. Step 404: Feature Fusion and Vector Generation For each successfully associated fused target, its corresponding radar feature vector and image feature vector are concatenated or weighted to generate a higher-dimensional, more information-rich fused feature vector. This fused feature vector contains both the target's precise spatial motion attributes (from radar) and detailed semantic category and appearance attributes (from vision). For tentative single-sensor candidate targets (radar candidate targets perceived only by radar or image candidate targets perceived only by vision), their original radar feature vectors or image feature vectors are retained as their fused feature vectors to maintain the integrity of the information for further verification by subsequent decision-making modules.

[0043] Steps 401 and 402 above, through coordinate space projection and intersection-union ratio calculation, accurately geometrically correlate the spatial location of radar candidate targets with the pixel regions of image candidate targets. This provides an objective and quantitative basis for determining whether radar and vision perceive the same physical entity, solving the core problem of inaccurate correlation of multi-source sensor data and achieving accurate correlation of cross-modal targets. Step 404 generates a fused feature vector by splicing radar feature vectors and image feature vectors for successfully correlated targets. This vector integrates the target's precise motion attributes (distance, velocity) and rich semantic attributes (category, texture), forming a more comprehensive and accurate description of the obstacle, providing a direct data foundation for subsequent refined threat assessment. Steps 403 and 404 are designed with a correlation strategy that includes a fault-tolerant mechanism. For single-sensor targets that do not reach the matching threshold (such as targets visible only to radar or only to vision), the system temporarily sets their original feature vector as a fused feature vector and retains it for subsequent processes. This mechanism ensures that potential threat target information is not lost due to fusion failure when some sensors fail, are obstructed, or have perception errors, significantly improving reliability in complex real-world scenarios. The fused feature vector (or provisional feature vector) output by this step, containing precise location information, is the direct input for subsequent steps to verify the orbital travel envelope space. Only the structured target information generated in this step can be combined with orbital geometric parameters to perform effective spatial location discrimination, thereby achieving intelligent filtering based on scene priors.

[0044] For example, D1 utilizes the camera intrinsic parameter matrix obtained during the system calibration phase. And the extrinsic transformation matrix from the radar coordinate system to the camera coordinate system The three-dimensional spatial coordinates of each radar candidate target Projected onto image pixel coordinate system , thus obtaining the corresponding two-dimensional projection points; D2. For each radar candidate target's projection point, calculate the cross-union ratio (CUI) between it and the detection bounding boxes of all image candidate targets. D3. Set an intersection-union ratio matching threshold (e.g., =0.5); If the cross-union ratio of a radar candidate target and an image candidate target is greater than or equal to the threshold, then the two are determined to correspond to the same physical entity.

[0045] For each successfully associated radar candidate target and image candidate target pair, a fused target is generated; D4. For each generated fused target, concatenate its corresponding radar feature vector with the image feature vector to generate the fused feature vector corresponding to that fused target. : ; Fusion feature vectors It integrates precise distance and velocity information from radar with semantic category and texture information from images.

[0046] Specifically, the camera intrinsic parameter matrix obtained through pre-offline calibration is used. Radar-camera extrinsic transformation matrix The three-dimensional spatial coordinates of the two radar candidate targets are projected onto the image pixel coordinate system: Tool objective: The three-dimensional spatial coordinates (12.5, 0.3, 0.1) are projected to obtain the two-dimensional projection point coordinates (350, 200) on the image. Personnel target: The three-dimensional spatial coordinates (50.2, -0.1, 1.0) are projected to obtain the two-dimensional projection point coordinates (150, 100). Given an image candidate target (person) with a bounding box coordinate of (320, 180, 380, 250), calculate the intersection-union ratio (IoU) between each radar projection point and the bounding box: Intersection over Union (IoU) of the projected target points: Since the projected point (350, 200) is exactly inside the personnel detection box (320, 180, 380, 250), theoretically its IoU is greater than 0. However, in actual matching logic, it is usually necessary to define a small projection area (such as a pixel block) for the radar candidate target and then calculate it with the detection box. For example, the calculated IoU is 0.15. Intersection over Union (IoU) of Person Target Projection Points: When the projection point (150, 100) is far from the person detection box, the calculated IoU is 0.0 (or approximately 0). If the preset matching threshold is 0.5, then a single target is considered associated: Tool target determination: Its intersection-union ratio with the unique image candidate target (personnel) is 0.15 < 0.5, therefore the tool target is tentatively defined as a radar candidate target that is only perceived by radar; Personnel target determination: The intersection-union ratio (IUU) between the personnel target and the image candidate target is 0.0 < 0.5, which does not reach the matching threshold. A critical situation arises here: the personnel point cloud clusters detected by radar and the personnel bounding boxes detected by vision do not directly match in the image (possibly due to minor calibration errors, partial target occlusion, or cluster center deviation). Therefore, the personnel target is tentatively classified as a radar candidate target perceived only by radar, while the image candidate target detected by vision is tentatively classified as an image candidate target perceived only by vision.

[0047] Therefore, no fused target was generated in this frame of data.

[0048] For two tentatively identified radar candidate targets (tools and personnel) detected solely by radar, their respective radar feature vectors are preserved. and As the fused feature vector of its current frame.

[0049] For tentatively selected image candidate targets (people) perceived solely by visual perception, the system retains their image feature vectors. As the fused feature vector of its current frame.

[0050] This example illustrates the challenges of feature layer fusion in complex real-world scenarios (nighttime, potential calibration errors). Although both radar and vision detected the same person, due to insufficient spatial matching, they were processed as two independent single-sensor targets in this frame (radar candidate target and image candidate target). However, this precisely highlights the value of our proposed solution: even without successful association, the original feature information of each sensor (radar and camera) is preserved (as a provisional fusion feature vector), ensuring no potential threat target information is lost (such as tools detected by radar but not by vision). These provisional targets and their feature vectors are fed into subsequent steps, where the orbital travel envelope space can be used for verification to further determine whether the two person targets are the same entity. This corrects single-frame matching errors at a higher level, improving the overall robustness of the system.

[0051] In another exemplary embodiment of this application, the steps described above—verifying whether the fused target is located within the orbital travel envelope space based on the position information in the fused feature vector of the fused target and the orbital geometric parameters, and calculating the threat level of the fused target located within the orbital travel envelope space based on the distance, speed, size, and category information in its fused feature vector—are specifically implemented through the following steps 501 to 504: Step 501: Construct the dynamic orbital travel envelope space Based on the track geometry parameters of the current track segment loaded from the track database, especially the track width, a dynamic track travel envelope space is constructed with the track centerline as the reference. The lateral boundary of this track travel envelope space is determined by the track width plus a preset safety margin. This track travel envelope space defines the physical areas that may pose a direct threat to locomotive operation; Step 502: Target Location Verification and Initial Screening Extract the position information from the fused feature vector of the target to be verified (including the fused target and the tentative single-sensor target), and determine whether the fused target is located within the orbital travel envelope space constructed in step 501: If so, retain the merged target location and treat it as a potential threat target to be assessed, proceeding to the next step.

[0052] Otherwise, the merged target location is determined to be a non-threat target (such as adjacent vehicles or trackside equipment), and is filtered out and not included in the subsequent threat calculation process; Step 503: Calculate the threat index of the target. For each verified threat target, a threat index assessment model (i.e., a predetermined threat index calculation formula) is established based on the distance, speed, size, and category information extracted from its fused feature vector: ; in, Indicates the longitudinal distance (in meters) of the threat target relative to the locomotive; Indicates the maximum distance (e.g., 100 meters). This indicates that a normalization process has been implemented where the closer the distance, the greater the threat. This represents the radial velocity of the threatening target relative to the locomotive (unit: meters per second). The absolute value indicates the approach speed as a threat. This represents the size factor (within a range of 0 to 1) obtained by normalizing the equivalent surface area in the fused feature vector of the threat target after a preset maximum size. This indicates a weighting coefficient preset based on the target category (e.g., people = 1.0, vehicles = 0.8, foreign objects = 0.5). Let represent the weighting coefficients of each component, and satisfy . (For example, values ​​of 0.4, 0.3, 0.2, and 0.1) are used to adjust the importance of different factors in the overall threat assessment.

[0053] As an optional implementation, step 504 is included after step 503, which involves risk classification and decision-making based on the threat index: Based on the calculated threat index T, threat targets are classified into different risk levels, and corresponding response strategies are triggered.

[0054] Steps 501 and 502, by constructing a dynamic track travel envelope space and verifying the target position, fundamentally eliminate non-threatening targets located outside the track clearance (such as adjacent lines or trackside areas). This directly solves the problem of high false alarm rates caused by the lack of track constraints in traditional solutions, significantly improving the reliability of alarms and achieving intelligent false alarm filtering based on scenario priors. Step 503, through the established threat index assessment model, comprehensively considers multi-dimensional features such as target distance, speed, size, and category information, transforming qualitative threat target perception into a quantitative threat index. This achieves a refined and objective measurement of the degree of obstacle danger, overcoming the limitations of traditional solutions that rely on a single distance threshold. Step 504 triggers graded responses (such as emergency braking, audible and visual warnings, and status recording) based on the quantified threat index, enabling the most appropriate measures to be taken for different levels of risk. This mechanism, while ensuring train operation safety, avoids the impact of excessive braking on shunting operation efficiency, achieving synergistic optimization of safety and efficiency. Steps 501 to 504 together constitute a complete decision-making chain from environmental perception and threat assessment to control response. By outputting structured target information with precise threat levels, it provides the LFK system host and train operation monitoring equipment with directly executable decision-making basis, realizing a closed loop of proactive safety protection.

[0055] For example, E1, based on the three-dimensional spatial coordinates (i.e., physical location information) of the target in the fused feature vector, combined with the track geometry parameters (mainly track width) of the current road segment loaded from the route database, constructs a dynamic track travel envelope space. The lateral boundary of this space is half the track width plus a safety margin on each side of the track centerline. E2. Initial screening of targets is performed based on the orbital travel envelope space, with the following logic: If the center point of the target is located outside the orbital travel envelope, the target is determined to be a non-direct threat target (such as adjacent vehicles or trackside facilities) and is filtered out. Otherwise, if the target center point is located within the orbital travel envelope space, the target is retained as a potential threat target to be evaluated, and its threat index is calculated. E3. Based on the calculated threat index Execute a tiered response: If T ≥ Level 1 alarm threshold (e.g., 0.7), it is determined to be a Level 1 alarm (high risk), and an emergency braking command is immediately triggered; If the level 2 alarm threshold is ≤ T < level 1 alarm threshold (e.g., 0.4 ≤ T < 0.7), it is determined to be a [level 2 alarm] (medium danger), triggering an audible and visual warning and requiring driver intervention or automatic deceleration by the system; If T < Level 2 alarm threshold (e.g., T < 0.4), it is judged as "low risk", and only the status is recorded and displayed, without triggering active alarms or braking.

[0056] Specifically, in a nighttime station environment, for all the provisional targets generated above (regardless of whether they are merged), a dynamic track travel envelope space is constructed based on the loaded current track geometry parameters (straight track, track width 1435 mm) and a preset safety margin of 0.5 meters on each side. The lateral boundary of this space is: extending (0.7175 meters + 0.5 meters) = 1.2175 meters to the left and right from the track centerline. Any target whose center point is located within this width range will be considered a potential threat target that may intrude into the travel path.

[0057] Extract the position information (3D spatial coordinates) from the fused feature vector of each provisional target, and determine whether it is located within the aforementioned orbital travel envelope space: Provisional radar candidate target (tool): Position coordinates (12.5, 0.3, 0.1), with a lateral offset of 0.3 meters and an absolute value less than 1.2175 meters, is determined to be located within the orbital travel envelope space. This target is reserved as a potential threat target to be evaluated. Provisional radar candidate target (personnel): Position coordinates (50.2, -0.1, 1.0), with a lateral offset of -0.1 meters and an absolute value less than 1.2175 meters, is determined to be located within the orbital travel envelope space. This target is reserved as a potential threat target to be evaluated. Provisional candidate targets (personnel) in the image: their image feature vectors It does not contain direct 3D spatial coordinates. The center of its detection box needs to be projected back to the world coordinate system (using camera parameters and an assumed ground plane), or determined directly based on its pixel position. Assuming its calculated lateral offset is 2.0 meters (potentially corresponding to another person on an adjacent track or a false detection), and its absolute value is greater than 1.2175 meters, it is determined to be outside the orbital travel envelope. This target is filtered out and not included in subsequent threat target calculations.

[0058] For two verified potential threat targets, their fused feature vectors (i.e., their respective radar feature vectors) are used to determine the target's identity. and Calculate the threat index T: Tool Target Threat Index Calculation: ; Personnel Target Threat Index Calculation: ; Risk classification and decision-making are performed based on the threat index T and preset classification thresholds: Tool Objective: 0.404. Due to The target was identified as a Level 2 Alert. An audible and visual warning was triggered, and a "small foreign object" was highlighted approximately 12 meters away on the driver's human-machine interface, advising the driver to be aware and prepare to slow down.

[0059] Personnel goals: Similarly satisfied The target was also classified as a Level 2 Alert, with an overlay warning message indicating that there were "personnel" approximately 50 meters away, advising the driver to immediately sound the horn and slow down.

[0060] This example clearly demonstrates the core advantages of the proposed solution: it successfully filters out visual false alarm targets located outside the track travel envelope (adjacent track), significantly reducing false alarms; even when both are level two alarms, tools have lower category weights due to their proximity, while personnel have higher category weights due to their distance, resulting in similar threat indices but different constituent factors, demonstrating the value of multidimensional assessment; it issues clear level two alarms for two real threat targets, guiding the driver to take targeted measures rather than directly triggering emergency braking, thus ensuring both safety and operational efficiency.

[0061] In another exemplary embodiment of this application, after step 504 above, the method further includes: taking the fused target located within the track travel envelope space as a potential threat target, and encapsulating the information in the fused feature vector of the potential threat target and its threat level into an alarm data message, and sending it to the LFK system host and / or train operation monitoring equipment via the vehicle bus, which is specifically implemented through the following steps 601 to 604: Step 601: Alarm Data Encapsulation According to the locomotive onboard network communication protocol, the key information and threat level in the fusion feature vector of each potential threat target are encapsulated into a standard alarm data message.

[0062] The alarm data message here includes at least: a unique identifier for the target (used for continuous frame tracking), threat level ([Level 1 Alert], [Level 2 Alert], or [Low Risk]), the target's absolute geographical location (latitude and longitude calculated by fusing the locomotive's Beidou / GPS positioning information with the target's relative position), relative longitudinal distance, lateral offset, relative speed, semantic category, and the absolute timestamp of the LFK system's main control unit when the data was generated; Step 602, Protocol Conversion and Frame Formatting The alarm data message generated in step 601 is converted into a data frame format that conforms to the specific on-board bus physical layer and data link layer specifications by using a bus protocol stack (such as Controller Area Network (CAN) protocol stack or Train Real-Time Data Protocol (TRDP) protocol stack) that matches the target receiving device (LFK system host or train operation monitoring device). Step 603: Reliable Sending and Transmission Confirmation The formatted data frames are sent to the vehicle bus network via its hardware interface (such as the CAN controller and Ethernet PHY chip). Based on the reliability mechanisms of the adopted bus protocol stack (such as CAN message ACK response and TCP acknowledgment mechanism), it is confirmed whether the alarm data message has been successfully delivered to the target receiving device. Step 604: Hierarchical Linkage and Control Execution After receiving an alarm data message, the LFK system host parses the threat level and target information, and executes the predefined corresponding control logic: for a Level 1 alarm, it generates and outputs an emergency braking command; for a Level 2 alarm, it triggers an audible and visual alarm and displays a graphical warning on the driver's human-machine interface; for low-risk targets, it records and stores their status. If the alarm data message is also sent to the train operation monitoring equipment, the train operation monitoring equipment can use this information as an important external environment perception input and incorporate it into its existing safety monitoring and protection logic to achieve synergy and enhancement of protection functions between different safety systems.

[0063] Step 601 above integrates the multi-dimensional target information and threat level obtained from the fusion identification and encapsulates them into a standardized alarm data message. This message has a clear structure and complete information, enabling downstream steps to accurately and unambiguously parse complex perception results, providing a data foundation for reliable interaction. Steps 602 and 603, by calling the standard onboard bus protocol stack and performing reliable transmission confirmation, ensure that the output of this method strictly adheres to the existing communication specifications of the locomotive. This guarantees that alarm information can be seamlessly and reliably integrated into the existing network of the LFK system and train operation monitoring equipment without modifying the underlying onboard network, significantly reducing deployment costs and complexity. Step 604 enables the LFK system host or train operation monitoring equipment to execute graded responses, from emergency braking and audible and visual warnings to status recording, based on the threat level in the message. This directly transforms the intelligent perception results from the front end into specific safety control actions, realizing a complete automated safety closed loop from environmental perception to decision execution. By simultaneously sending alarm data messages to external systems such as train operation monitoring equipment, the high-precision environmental perception information output by this method can serve as a key input, empowering and enhancing the locomotive's existing multi-layered safety monitoring and protection logic, realizing information sharing and functional synergy between different safety systems, thereby improving the overall active safety protection level of the entire locomotive system.

[0064] As an alternative implementation, an infrared thermal imaging camera can be used instead of the industrial-grade visible light camera. Infrared thermal imaging cameras, based on the temperature difference between the target and the environment, can effectively enhance the system's perception capabilities under low-visibility conditions such as nighttime, fog, and haze. In this alternative, the image enhancement algorithm in image preprocessing (i.e., step two) needs to be adjusted accordingly. For example, histogram equalization or contrast stretching algorithms adapted to thermal imaging images can be used to optimize image quality.

[0065] As an alternative implementation, LiDAR can be used instead of the vehicle-mounted millimeter-wave radar. Although LiDAR is more expensive, it can provide higher resolution and density 3D radar point cloud data, making it suitable for scenarios with extremely high requirements for obstacle contour and attitude recognition accuracy. Under this alternative, the parameters of the density clustering algorithm in the radar point cloud data clustering (i.e., step three) need to be adjusted accordingly (such as the neighborhood search radius and minimum point threshold) to adapt to the data characteristics of the LiDAR point cloud data.

[0066] As an optional implementation, in the step of generating the fused target and its corresponding fused feature vector, a cross-attention mechanism based on the Transformer architecture can be used to replace the direct concatenation operation of the radar feature vector and the image feature vector. This mechanism calculates the attention weights between radar features and image features, enabling it to adaptively learn and fuse the most discriminative information from the two modalities. This allows for a more thorough exploration of the complementarity and correlation between multi-source features, potentially further improving the representational power of the fused features and the robustness of subsequent recognition. This implementation typically requires stronger computing power and is suitable for hardware platforms equipped with embedded GPUs or dedicated AI acceleration units.

[0067] As an alternative implementation, the threat index assessment model is not limited to a linear weighted formula. For example, a machine learning-based classifier (such as a random forest classifier or a support vector machine classifier) ​​can be used as an alternative model. This model is trained using a large amount of labeled data from historical shunting operation scenarios, and can automatically learn the complex nonlinear relationship between multi-dimensional features such as obstacle distance, speed, size, and category information and threat level, thereby obtaining a more discriminative decision boundary to improve the accuracy and adaptability of threat classification. Deploying such a model is suitable for systems with sufficient computing resources and a wealth of scenario data.

[0068] Based on the same inventive concept, this application also provides a radar-image fusion obstacle recognition device for implementing the radar-image fusion obstacle recognition method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more radar-image fusion obstacle recognition device embodiments provided below can be found in the limitations of the radar-image fusion obstacle recognition method described above, and will not be repeated here.

[0069] In one exemplary embodiment, a radar-image fusion obstacle recognition device is provided, comprising: The multi-source perception module is used to acquire radar point cloud data, image data, and the current positioning information of the locomotive (such as latitude and longitude, timestamp, etc.) in front of the locomotive. The data preprocessing module is used to load the track geometry parameters of the current road segment according to the positioning information, and based on the track geometry parameters, to perform adaptive filtering on the radar point cloud data and process the image data to generate the region of interest in the track area. The feature fusion module is used to cluster the filtered radar point cloud data to generate at least one radar candidate target and construct a corresponding radar feature vector for each radar candidate target; to perform target detection on the region of interest in the image to generate at least one image candidate target and construct a corresponding image feature vector for each image candidate target; and to perform feature mapping and association between the radar candidate target and the image candidate target, and generate the fused target and its corresponding fused feature vector based on the radar feature vector and the image feature vector. The decision verification module is used to verify whether the fused target is located within the orbital travel envelope space based on the position information in the fused feature vector of the fused target and the orbital geometric parameters, and to calculate the threat level of the fused target located within the orbital travel envelope space based on the distance (e.g., longitudinal distance), velocity (e.g., relative radial velocity), size and category information in its fused feature vector. The data interaction module is used to output the information in the target's fused feature vector and its threat level to the outside world.

[0070] As an optional implementation, a cloud-edge collaborative computing architecture can be adopted. In this architecture, the radar-image fusion obstacle recognition device acts as an on-board edge computing unit, primarily responsible for real-time sensing, fusion, and decision-making reasoning tasks; while some non-real-time or computationally intensive tasks (e.g., incremental updates of the electronic map of the route, retraining and optimization of the target detection neural network model) are deployed on a cloud server. After completing the computation, the cloud server distributes the updated map or model parameters to the on-board unit. This solution effectively utilizes the abundant computing and storage resources of the cloud, reduces the long-term maintenance and computing burden on the on-board embedded platform, and supports continuous online optimization of the algorithm model. The implementation of this architecture relies on a stable, low-latency wireless communication network between the vehicle and the ground (such as in smart depots or locomotive depots with good 5G network coverage).

[0071] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal. The computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is connected to the system bus via the I / O interfaces. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The I / O interfaces of the computer device are used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a radar-image fusion obstacle recognition method.

[0072] In one exemplary embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0073] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0074] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0075] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0076] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0077] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0078] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0079] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A radar-image fusion obstacle recognition method, characterized in that, include: Acquire radar point cloud data, image data, and current locomotive positioning information from the radar in front of the locomotive; The track geometry parameters of the current road segment are loaded based on the positioning information, and adaptive filtering is performed on the radar point cloud data and the image data is processed based on the track geometry parameters to generate the region of interest in the track area. The filtered radar point cloud data is clustered to generate at least one radar candidate target, and a corresponding radar feature vector is constructed for each radar candidate target; the region of interest in the image is targeted to generate at least one image candidate target, and a corresponding image feature vector is constructed for each image candidate target. The radar candidate target and the image candidate target are feature-mapped and associated, and a fused target and its corresponding fused feature vector are generated based on the radar feature vector and the image feature vector. Based on the position information in the fused feature vector of the fused target and the orbital geometry parameters, it is verified whether the fused target is located within the orbital travel envelope space. For the fused target located within the orbital travel envelope space, the threat level is calculated based on the distance, speed, size and category information in its fused feature vector.

2. The radar-image fusion obstacle recognition method according to claim 1, characterized in that, The adaptive filtering of the radar point cloud data based on the orbital geometric parameters specifically includes: Based on the track geometry parameters, determine whether the current section is a straight section or a curve / turnout area, and set the dynamic filtering radius accordingly; For each data point in the radar point cloud data, calculate the density of neighboring points in a spherical space centered on that data point and with the dynamic filtering radius as the radius; If the density of the neighboring points is lower than a preset density threshold, the data point is determined to be a noise point and is removed.

3. The radar-image fusion obstacle recognition method of claim 1, wherein, The process of clustering the filtered radar point cloud data to generate radar candidate targets and their corresponding radar feature vectors specifically includes: A density-based clustering algorithm is used to perform cluster analysis on the filtered radar point cloud data to generate at least one radar candidate target; For each radar candidate target, a corresponding radar feature vector is constructed, which includes longitudinal distance, relative velocity, azimuth angle and equivalent surface area. The step of performing target detection on the region of interest in the image to generate candidate targets and their corresponding image feature vectors specifically includes: The region of interest in the image is input into a target detection neural network, which outputs at least one candidate target in the image. For each candidate target in the image, a corresponding image feature vector is constructed. The image feature vector includes the pixel coordinates of the detection box, the target category label, and the confidence score.

4. The radar-image fusion obstacle recognition method of claim 1, wherein, The step of performing feature mapping and association between the radar candidate target and the image candidate target, and generating the fused target and its corresponding fused feature vector based on the radar feature vector and the image feature vector, specifically includes: The spatial coordinates of each radar candidate target are projected onto the image pixel coordinate system to obtain the corresponding two-dimensional projection point or projection area. Calculate the intersection-union ratio (IUU) between the projection point or projection region and the detection box of each candidate target in the image; If the intersection-union ratio is greater than or equal to the preset matching threshold, the corresponding radar candidate target and the image candidate target are determined to be the same physical entity, a fused target is generated, and the corresponding radar feature vector and image feature vector are concatenated to generate the fused feature vector corresponding to the fused target.

5. The radar-image fusion obstacle recognition method of claim 1, wherein, The verification of whether the fused target is located within the orbital travel envelope space, and the calculation of the threat level of the fused target located within the orbital travel envelope space, specifically includes: Based on the track width and preset safety margin in the track geometry parameters, a track travel envelope space is constructed with the track centerline as the reference. Extract the position information from the fused feature vector of the fused target, and determine whether the position is located within the orbital travel envelope space; If the location is within the orbital travel envelope, the threat level is calculated based on the distance, speed, size, and category information in the fused feature vector using a predetermined threat index calculation formula.

6. The radar-image fusion obstacle recognition method according to any one of claims 1 to 5, characterized in that, After calculating the threat level based on the distance, velocity, size, and category information in its fused feature vector, the method further includes: The fused target located within the orbital travel envelope is taken as a potential threat target, and the information in the fused feature vector of the potential threat target and its threat level are encapsulated into an alarm data message, which is sent to the LFK system host and / or train operation monitoring equipment via the vehicle bus.

7. A radar-image fusion obstacle recognition apparatus characterized by comprising: include: The multi-source perception module is used to acquire radar point cloud data, image data, and the current positioning information of the locomotive in front of it. The data preprocessing module is used to load the track geometry parameters of the current road segment according to the positioning information, and based on the track geometry parameters, to perform adaptive filtering on the radar point cloud data and process the image data to generate the region of interest in the track area. The feature fusion module is used to cluster the filtered radar point cloud data to generate at least one radar candidate target and construct a corresponding radar feature vector for each radar candidate target; to perform target detection on the region of interest in the image to generate at least one image candidate target and construct a corresponding image feature vector for each image candidate target; and to perform feature mapping and association between the radar candidate target and the image candidate target, and generate the fused target and its corresponding fused feature vector based on the radar feature vector and the image feature vector. The decision verification module is used to verify whether the fused target is located within the orbital travel envelope space based on the position information in the fused feature vector of the fused target and the orbital geometric parameters, and to calculate the threat level of the fused target located within the orbital travel envelope space based on the distance, speed, size and category information in its fused feature vector. The data interaction module is used to output the information in the target's fused feature vector and its threat level to the outside world.

8. A computer device comprising: A memory, a processor, and a computer program stored on the memory and loadable on the processor, characterized in that the processor executes the computer program to implement the steps of the radar-image fusion obstacle recognition method of any one of claims 1-6.

9. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the radar-image fusion obstacle recognition method of any one of claims 1-6.

10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the radar-image fusion obstacle recognition method of any one of claims 1-6.

Citation Information

Patent Citations

  • Subway multi-sensor fusion active anti-collision system and method

    CN112406960A