Meighting and thunder integrated train active sensing system and method
The train active perception system, which integrates radar and vision, achieves unified temporal management and spatial alignment of multi-sensor data, solving the problem of insufficient obstacle recognition and positioning accuracy in existing technologies, and improving the accuracy and stability of train environmental perception.
Patent Information
- Application Number
- CN202511842786.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-09
- Publication Date
- 2026-01-06
AI Technical Summary
In existing train environmental perception systems, multiple sensors are configured in a single or loosely stacked manner, lacking unified time-series management and coordinate alignment of multi-source data. This results in insufficient accuracy in obstacle recognition and spatial location calculation, making it difficult to achieve high-precision obstacle recognition and positioning, especially in complex environments.
The train active perception system using radar-visual fusion acquires image data and 3D point cloud data at different focal lengths through the acquisition module, aligns the data with the time synchronization module, establishes the spatial transformation relationship between the image coordinate system and the point cloud coordinate system using the coordinate transformation module, performs obstacle detection by the image perception module, filters the spatial range by the point cloud perception module, and generates structured obstacle information by the radar-visual fusion module.
It improves the utilization rate of multi-source sensing information, enhances the accuracy of track obstacle recognition and three-dimensional positioning, strengthens the system's perception stability in complex scenarios, and ensures accurate perception and real-time early warning of the environment in front of the train.
Smart Images

Figure CN121276531A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of train perception system technology, and in particular to a radar-visual fusion train active perception system and method. Background Technology
[0002] As a high-capacity rail transit tool, trains are widely used in passenger and freight transport, and their operational safety is directly related to the safety of people's lives and property. Trains often operate at high speeds in open track environments, where various unpredictable obstacles such as falling rocks, pedestrians, animals, and foreign objects may appear ahead of the track. If these obstacles are not detected and identified in time, they can easily lead to rear-end collisions, collisions, and derailments. Therefore, continuous and reliable active perception of the environment ahead of the train is urgently needed.
[0003] Early train environmental perception technologies largely relied on passive monitoring methods such as trackside detection devices and onboard vibration sensors, which could only acquire limited state data and were insufficient to directly perceive the track environment ahead. With the development of sensors and computing platforms, some solutions have begun to introduce active sensors such as onboard cameras and lidar to achieve the detection and early warning of targets ahead, while other technologies have proposed a multi-sensor joint architecture.
[0004] However, current sensor configurations tend to be either singular or loosely stacked. Most systems process only single-modal data, and even when using multi-sensor combinations, there is a lack of unified temporal management and coordinate alignment mechanisms for image data and 3D point cloud data. This makes it difficult to ensure that multi-source data are processed under the same time reference and spatial reference. Furthermore, data fusion in existing multi-sensor systems often remains at the rule-level or result-level simple weighting, lacking deep fusion of image features and point cloud features from a bird's-eye view. As a result, in scenarios with complex weather, various obstacle shapes, and coexisting interference targets around the track, the accuracy of obstacle identification and the ability to calculate spatial positions remain insufficient. Summary of the Invention
[0005] In view of this, embodiments of this application provide a train active perception system and method based on radar-visual fusion to solve the problems of insufficient utilization of multi-sensor perception information, lack of unified spatiotemporal alignment of multi-source data, and low accuracy of track obstacle recognition and three-dimensional positioning in the prior art.
[0006] A first aspect of this application provides a radar-visual fusion train active perception system, comprising: a data acquisition module for acquiring multiple image data containing images with different focal lengths and corresponding 3D point cloud data in front of the train, and adding time markers to each data stream; a time synchronization module for time-aligning the multiple image data and 3D point cloud data according to the time markers of each data stream, generating synchronized image data and synchronized 3D point cloud data; and a coordinate transformation module for determining the external parameters between the lidar coordinate system and the imaging coordinate systems of each industrial camera based on the image data containing the calibration target and the 3D point cloud data, and establishing the 3D point cloud coordinate system. The system includes: a spatial transformation relationship between the target system and the image coordinate system; an image perception module, which inputs synchronized image data into the target detection network to obtain obstacle coordinates, and inputs synchronized image data and 3D point cloud data into the target segmentation network to obtain track coordinates; a point cloud perception module, which filters synchronized 3D point cloud data within a preset spatial range to obtain 3D point cloud data in front of the train; and a radar-visual fusion module, which uses spatial transformation relationships to locate the point cloud set corresponding to each obstacle target candidate in the 3D point cloud data, assigns an accurate center representative coordinate to each obstacle, generates structured obstacle information, and outputs it to the onboard train control system.
[0007] A second aspect of this application provides a train active perception method based on the radar-visual fusion of the system of the first aspect, comprising: acquiring multi-channel image data containing images with different focal lengths and corresponding 3D point cloud data of the train ahead, and attaching time markers to each channel of image data and 3D point cloud data; performing time alignment on the multi-channel image data and 3D point cloud data according to the time markers to generate synchronized image data and synchronized 3D point cloud data; based on the image data and 3D point cloud data containing calibration targets, determining the external parameters between the lidar coordinate system and the imaging coordinate systems of each industrial camera, and establishing a spatial transformation relationship between the 3D point cloud coordinate system and the image coordinate system; and performing time alignment on the multi-channel image data and 3D point cloud data. The system performs feature extraction and target recognition on the image data to obtain image feature representations of the scene in front of the train and a set of image target candidates containing obstacle categories. Image segmentation results for the track area are obtained through image semantic segmentation. The synchronous 3D point cloud data is filtered according to a preset spatial range. Based on the image segmentation results, the set of image target candidates is filtered to retain the set of image target candidates located within the track area. The point cloud set corresponding to each image target candidate is located in the synchronous 3D point cloud data using spatial transformation relationships. The 3D position parameters and category information of each obstacle in the train coordinate system are calculated to generate structured obstacle information and output it to the on-board train control system.
[0008] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects: The acquisition module collects multiple image data streams containing images at different focal lengths, along with corresponding 3D point cloud data in front of the train, and adds time stamps to each data stream. The time synchronization module aligns the multiple image data streams and 3D point cloud data according to the time stamps, generating synchronized image data and synchronized 3D point cloud data. The coordinate transformation module calculates the external parameters between the lidar coordinate system and the imaging coordinate systems of each industrial camera based on the image data and 3D point cloud data containing the calibration target, establishing a spatial transformation relationship between the 3D point cloud coordinate system and the image coordinate system. The image perception module inputs the synchronized image data into the target detection network to obtain obstacle coordinates, and inputs the synchronized image data and 3D point cloud data into the target segmentation network to obtain track coordinates. The point cloud perception module filters the synchronized 3D point cloud data within a preset spatial range to obtain 3D point cloud data in front of the train. The lidar-visual fusion module uses the spatial transformation relationship to locate the point cloud set corresponding to each obstacle target candidate in the 3D point cloud data, assigns an accurate center representative coordinate to each obstacle, generates structured obstacle information, and outputs it to the onboard train control system. This application can improve the utilization rate of multi-source sensing information, enhance the accuracy of track obstacle recognition and 3D positioning, and improve the system's sensing stability in complex scenarios. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 This is a schematic diagram of the structure of the radar-visual fusion train active perception system involved in a real-world scenario according to the embodiments of this application; Figure 2 This is a schematic diagram of the structural composition of the radar-visual fusion train active perception system provided in the embodiments of this application; Figure 3 This is a schematic diagram of the coordinates of the lidar used in the system provided in the embodiments of this application; Figure 4 This is a schematic diagram of a segmentation algorithm based on BEV feature fusion provided in an embodiment of this application; Figure 5 This is a flowchart illustrating the train active perception method based on radar-visual fusion provided in this application embodiment. Detailed Implementation
[0011] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0012] Trains have profoundly impacted people's lives and work. In daily life, they have improved travel efficiency, shortened distances between cities, facilitated family visits, tourism, and the flow of goods, while also changing people's perceptions of time and social interactions. In terms of work, trains have improved commuting efficiency, expanded employment options, promoted regional economic integration and industrial upgrading, and created more business and cooperation opportunities.
[0013] However, due to the numerous uncertainties inherent in the open environment of train operation, obstacles such as falling rocks, pedestrians, and animals may appear on the track, all of which can potentially lead to accidents and disasters. Therefore, train environmental perception systems are crucial. Advanced train perception systems can monitor obstacles ahead of the track in real time and issue timely warnings, enabling the driver or automatic driving system to quickly take deceleration or stopping measures, thereby effectively preventing accidents.
[0014] Train sensing systems have made significant progress in recent years. In their early stages, these systems primarily relied on passive sensing technologies, such as sensors deployed along the track and onboard vibration sensors, to perform basic monitoring of train operation and the surrounding environment. With upgrades in basic hardware performance and breakthroughs in software algorithms, a series of new systems with active sensing capabilities have gradually been applied to real-world scenarios. However, existing systems still generally exhibit a tendency towards single-sensor configurations. Even with the introduction of multi-sensor architectures, deep information fusion mechanisms are often lacking, making it difficult to fully extract and utilize multi-source heterogeneous data. This limitation means that the comprehensiveness and accuracy of sensing remain significant challenges when dealing with complex and changing operating environments.
[0015] Furthermore, existing multi-sensor systems have limitations in data fusion and processing capabilities, making it difficult to achieve real-time and efficient decision support. Therefore, future train active sensing systems need to integrate efficient multimodal sensors, such as combining vision and radar sensors, and use advanced algorithms to fuse multi-source information to improve sensing accuracy and robustness, thereby better coping with various challenges in complex operating environments.
[0016] In view of the problems existing in the prior art, this application provides a radar-visual fusion train active perception system and method. First, the overall architecture of the radar-visual fusion train active perception system provided in this application will be introduced with reference to the accompanying drawings. Figure 1 This is a schematic diagram of the structure of the radar-visual fusion train active perception system involved in a real-world scenario according to the embodiments of this application, as shown below. Figure 1 As shown, this radar-visual fusion-based train active perception system mainly consists of a multi-sensor module, a data fusion processing module, and an information output module. The sensor module includes long-range and short-range industrial cameras and LiDAR, capable of simultaneously processing 2D image information and 3D point cloud data. This ensures that each processed data is multi-source information collected at the same time, effectively solving the data alignment problem caused by inconsistent timing. Based on this, the system performs deep fusion processing on the image data and point cloud data, combining target detection and semantic segmentation algorithms from deep learning to achieve accurate identification and localization of obstacles ahead. The final output obstacle information includes category and location key attributes and is provided to other onboard systems in structured data form, providing reliable environmental perception support for the train's intelligent decision-making and active safety protection.
[0017] The specific modules and functions of the radar-visual fusion train active perception system provided in this application will be described in detail below with reference to the accompanying drawings and specific embodiments. Figure 2 This is a schematic diagram of the structural composition of the radar-visual fusion train active perception system provided in the embodiments of this application, as shown below. Figure 2 As shown, the radar-visual fusion train active perception system may specifically include the following modules: The acquisition module 201 is used to acquire multi-channel image data containing images with different focal lengths and corresponding three-dimensional point cloud data in front of the train, and to add time markers to each channel of data. The time synchronization module 202 is used to time-align multiple image data and 3D point cloud data according to the time identifier of each data channel to generate synchronized image data and synchronized 3D point cloud data. The coordinate transformation module 203 is used to obtain the external parameters between the lidar coordinate system and the imaging coordinate system of each industrial camera based on the image data and three-dimensional point cloud data containing the calibration target, and to establish the spatial transformation relationship between the three-dimensional point cloud coordinate system and the image coordinate system. The image perception module 204 is used to input synchronous image data into the target detection network to obtain obstacle coordinates, and input synchronous image data and 3D point cloud data into the target segmentation network to obtain track coordinates; The point cloud sensing module 205 is used to filter the synchronous three-dimensional point cloud data within a preset spatial range to obtain the three-dimensional point cloud data in front of the train. The radar-visual fusion module 206 uses spatial transformation relationships to locate the point cloud set corresponding to each obstacle target candidate in the three-dimensional point cloud data, assigns an accurate center representative coordinate to each obstacle, generates structured obstacle information, and outputs it to the on-board train control system.
[0018] In some embodiments, the acquisition module is specifically used for: By setting up a first industrial camera and a second industrial camera with different focal lengths, multiple image data covering different viewing distances in front of the train are collected. The vehicle-mounted lidar collects three-dimensional point cloud data covering a preset distance range in front of the train. Based on the unified time reference on the train, time markers are added to multiple image data and 3D point cloud data respectively to form image data and 3D point cloud data with time markers.
[0019] Specifically, the acquisition module of the train active sensing system is installed in the equipment compartment of the train head and at the front of the train head shell. It is used to deploy industrial cameras and on-board LiDAR with different line-of-sight ranges facing the direction of train travel. In conjunction with the unified time reference provided by the train on-board control unit, it timestamps various types of raw sensing data, thereby forming the basic data source required for subsequent sensing and fusion processing.
[0020] In this embodiment, the acquisition module includes a first industrial camera, a second industrial camera, and an automotive-grade LiDAR. Both the first and second industrial cameras are 5-megapixel color industrial cameras, sharing a fixed mounting bracket. The bracket is fixed to the train's front frame structure via a vibration-damping mounting base, ensuring that the camera's optical axis is substantially aligned with the train's forward direction. The first industrial camera is equipped with a 25mm focal length lens for acquiring near-to-mid-range track environment images with a wider field of view; the second industrial camera is equipped with a 75mm focal length lens for acquiring higher magnification far-range track environment images within a narrower field of view. This combination of different focal lengths creates an imaging area directly in front of the train, superimposing a wide near-range field of view and a narrow far-range field of view, thus covering different viewing distances from near to far in front of the train and reducing near-range blind spots or insufficient far-range resolution caused by a single focal length.
[0021] Furthermore, the vehicle-mounted LiDAR is installed below or adjacent to the camera. The scanning plane of the LiDAR is basically parallel to the track surface, and its main field of view covers a preset distance range in front of the train, such as the forward space area from 0m to 500m. This LiDAR is a multi-beam automotive-grade product, capable of periodically outputting point cloud data frames containing three-dimensional spatial coordinates and reflection intensity information. Each frame of point cloud data covers a certain width and height range near the center of the track, which is used for subsequent measurement and analysis of the spatial position of obstacles in front of the track.
[0022] Furthermore, during data acquisition, the acquisition module obtains a unified time reference through a time synchronization unit connected to the train's onboard network. This time synchronization unit can be provided by the train's integrated monitoring device or vehicle control unit, and internally maintains a unified clock using a high-precision crystal oscillator, GNSS clock, or PTP protocol. In this embodiment, the system time with millisecond-level time resolution is used as the unified time reference. The first industrial camera, the second industrial camera, and the onboard LiDAR are connected to the control board within the acquisition module via dedicated communication interfaces. The control board periodically sends acquisition triggers or operating rhythm configurations to each sensor, and upon receiving image frame data or point cloud frame data, it reads the current unified time reference and appends it as a time identifier to the corresponding image data and 3D point cloud data.
[0023] Furthermore, for image data, the acquisition module packages each frame of the original image captured by the first industrial camera and each frame of the original image captured by the second industrial camera into image data units with time stamps. The time stamp can be in an absolute time format of "year-month-day hour:minute:second.millisecond" or a cumulative time count format from system power-on. Additional information such as camera number, resolution, and exposure parameters are recorded in the frame header of the data unit. For 3D point cloud data, the acquisition module considers the set of point clouds output by the LiDAR within one scan cycle as a single frame of point cloud data. It also reads a unified time reference and writes the time stamp into the frame header area of that frame of point cloud data to distinguish point cloud data from different acquisition cycles.
[0024] For example, in an application example of a passenger train with a maximum operating speed of 160 km / h, the acquisition module can configure the sampling frequency of the first industrial camera to 20 Hz for high frame rate imaging in the near-track area, configure the sampling frequency of the second industrial camera to 10 Hz for stable imaging in medium- to long-range scenes, and configure the scanning frequency of the LiDAR to 10 Hz, thereby acquiring 10 frames per second of 3D point cloud data covering a distance range of 0 m to 500 m. Whenever any sensor completes the acquisition of a frame of data, its driver immediately transmits the data to the acquisition module control board via a high-speed bus. The control board writes a unified time identifier for that frame of data according to the current system time, and sends the time-identified image data and 3D point cloud data to subsequent buffer queues, forming a raw sensing data stream with clear time information and a clear source.
[0025] Through the above implementation method, the acquisition module simultaneously acquires multiple image data and three-dimensional point cloud data covering a preset distance range within different line-of-sight ranges in front of the train, and adds time markers to each data based on a unified time reference, providing a multi-source perception data foundation with clear temporal relationships for subsequent time synchronization, coordinate transformation and radar-visual fusion processing.
[0026] In some embodiments, the time synchronization module is specifically used for: The multi-channel image data and 3D point cloud data with time stamps are cached in their respective data queues; Using the time identifier of one of the data streams as a reference, within a preset sliding time window, data frames whose time identifier difference does not exceed the first threshold are retrieved from other data queues. The retrieved data frames are then combined to form a candidate synchronization data combination. Statistical analysis is performed on the time difference of candidate synchronized data combinations to eliminate abnormal matching combinations whose time difference exceeds the second threshold, and the corresponding synchronized image data and synchronized 3D point cloud data are output.
[0027] Specifically, the time synchronization module is deployed on the onboard computing unit at the front of the train, sharing the same communication bus environment as the acquisition module. The time-stamped image data and 3D point cloud data output by the acquisition module are input to the data receiving unit of the time synchronization module via Gigabit Ethernet or a high-speed serial bus. Internally, the time synchronization module sets up independent data queues for different types of sensors. For example, a first image queue, a second image queue, and a point cloud queue are set up for the first industrial camera, the second industrial camera, and the onboard LiDAR, respectively. Each queue adopts a first-in-first-out (FIFO) buffer structure, recording the data frame content, corresponding time stamps, and sensor identifiers sequentially according to the data arrival order.
[0028] Specifically, in this embodiment, the time synchronization module first parses each frame of image data and point cloud data sent by the acquisition module into an internally unified data structure, extracts the time identifier field, sensor type field, and data payload pointer, and stores the data at the tail of the corresponding data queue according to the sensor type. To prevent abnormal data from occupying cache space, the time synchronization module can set a maximum number of cached frames for each queue, and automatically discard the oldest expired data frame when the queue length exceeds the preset limit. In this way, the time synchronization module can maintain the data buffer of each sensor for the most recent period in memory for subsequent time matching calculations.
[0029] During time alignment, the time synchronization module selects one data stream as the time reference stream; in this embodiment, it uses LiDAR point cloud data. Whenever a new point cloud data frame appears in the point cloud queue, the time synchronization module reads the time identifier t_ref of that frame and retrieves data frames in the first and second image queues within a preset sliding time window whose time difference with t_ref does not exceed a first threshold. The size of the sliding time window can be configured according to the train speed and sensor frame rate, for example, set to ±50ms, and the first threshold can be set to 30ms.
[0030] The time synchronization module finds the image frames with the smallest absolute time difference from t_ref that is less than the first threshold in each of the two image queues, and combines these three frames into a candidate synchronization data set. If no image frame meeting the conditions is found in a certain image queue within the time window, it is considered that the current point cloud frame cannot be synchronized with the image data of the corresponding viewpoint. The time synchronization module can skip this synchronization attempt or only form a partial combination with the found matching image frames.
[0031] To further suppress network transmission jitter and anomalous matching caused by asynchronous acquisition, the time synchronization module performs statistical analysis on the time differences of candidate synchronization data combinations formed over a period of time. Specifically, for each candidate synchronization data combination, the time synchronization module calculates the time differences Δt1 and Δt2 between the reference point cloud frame and the first image frame and the second image frame, respectively. It then incorporates Δt1 and Δt2 of several consecutive candidate synchronization combinations into a sliding statistical window to calculate the moving average and dispersion of the time differences.
[0032] Based on the statistical results, the time synchronization module can adaptively update the second threshold, for example, by adding several times the standard deviation to the average time difference to form the anomaly detection threshold under the current operating conditions. For a single candidate synchronization data combination, if the absolute value of any time difference is greater than the second threshold, or deviates too much from the statistical average, the combination is determined to be an abnormal matching combination and discarded, and is not considered a valid synchronization output.
[0033] Furthermore, after the candidate synchronized data combination passes the anomaly check, the time synchronization module marks the image data frame in the combination as synchronized image data and the point cloud data frame as synchronized 3D point cloud data, and outputs them to the subsequent image perception module and point cloud perception module in chronological order.
[0034] To facilitate subsequent modules in tracking the relationships between different synchronization combinations, the time synchronization module can also assign a unified synchronization sequence number to each set of synchronized image data and synchronized 3D point cloud data, serving as an index identifier in subsequent fusion calculations. In this way, subsequent modules can ensure that the input data corresponds to the train's forward scene at the same moment when performing feature extraction, semantic segmentation, and radar-visual fusion processing.
[0035] Through the time synchronization module design in this embodiment, when multiple industrial cameras and vehicle-mounted LiDARs with different frame rates and communication link characteristics are working simultaneously in front of the train, it can effectively align multi-channel image data and 3D point cloud data with time stamps by using a unified time reference, sliding time window matching, and an anomaly elimination mechanism based on statistical analysis. This outputs synchronized image data and synchronized 3D point cloud data with a unified time reference, thereby reducing the timing deviation between multi-source data and providing a stable and reliable timing foundation for subsequent multimodal feature fusion and obstacle 3D position calculation.
[0036] In some embodiments, the coordinate transformation module is specifically used for: Multiple sets of image data and 3D point cloud data containing the calibrated target are simultaneously acquired by various industrial cameras and lidar at different positions and attitudes. Based on the camera calibration algorithm, the coordinates of two-dimensional feature points of the calibration target are extracted from the image data, and the internal parameters and distortion parameters of each industrial camera are calculated to obtain the corresponding image coordinate system. Based on point cloud clustering and geometric constraints, the coordinates of three-dimensional feature points of the calibration target are extracted from three-dimensional point cloud data, and the correspondence between two-dimensional feature points and three-dimensional feature points is constructed. Based on the correspondence between two-dimensional and three-dimensional feature points, the rotation and translation parameters between the lidar coordinate system and the imaging coordinate system of each industrial camera are estimated to obtain the external parameter matrix, and a spatial transformation relationship is established based on the external parameter matrix.
[0037] Specifically, the coordinate transformation module is integrated into the onboard computing unit at the front of the vehicle and is connected to the acquisition module and time synchronization module via an internal bus. This module is mainly responsible for completing the joint calibration between each industrial camera and the onboard LiDAR during the system installation and maintenance phases, obtaining the spatial external parameters between the LiDAR coordinate system and the imaging coordinate systems of each industrial camera, and providing coordinate transformation capabilities between the point cloud coordinate system and the image coordinate system during the operation phase, which can be used by the image perception module, the point cloud perception module, and the radar-visual fusion module.
[0038] Specifically, in the calibration phase, this embodiment uses a calibration board with a regular checkerboard pattern or dot array as the calibration target. The calibration board is sequentially placed at different distances, heights, and deflection angles in front of the train, ensuring that the calibration board is simultaneously within the field of view of two industrial cameras and the vehicle-mounted LiDAR. The coordinate transformation module controls the acquisition module to simultaneously trigger the first industrial camera, the second industrial camera, and the LiDAR to acquire data in each calibration posture, obtaining multiple sets of image data containing the calibration target and corresponding 3D point cloud data. Through alignment processing by the time synchronization module, each set of calibration board images and corresponding point cloud data are stored as a group.
[0039] During camera calibration, the coordinate transformation module first performs intrinsic parameter calibration on both industrial cameras. For each industrial camera, the module inputs the multi-frame calibration board images it has acquired into the camera calibration algorithm, using corner detection or center detection methods to extract the two-dimensional pixel coordinates of regular grid points on the calibration board, forming a multi-frame set of two-dimensional feature points. Combining the known geometric dimensions of the calibration board and the three-dimensional coordinates of the grid points in the calibration board coordinate system, the coordinate transformation module calls the camera calibration function to solve for the intrinsic parameters of the industrial camera, including the pixel scale of the focal length in the horizontal and vertical directions, the principal point position coordinates, and the radial and tangential distortion coefficients, and establishes the corresponding image coordinate system. Through the above processing, the intrinsic parameter matrices and distortion parameters of the first and second industrial cameras can be obtained, providing the imaging model basis for subsequently projecting the three-dimensional point cloud onto the image plane.
[0040] In the point cloud-based calibration process, the coordinate transformation module extracts the coordinates of the 3D feature points of the calibration target for each set of 3D point cloud data corresponding to the calibration board image, using point cloud clustering and geometric constraint methods. Specifically, the module first sets an approximate spatial search region in the point cloud containing the calibration board based on the approximate position and size of the calibration board in the scene. Then, it uses a region growing algorithm to perform planar clustering of the point cloud within the search region, selecting candidate point cloud clusters that match the plane of the calibration board.
[0041] Furthermore, after obtaining the candidate point cloud clusters, the module combines prior geometric information such as the calibration board normal vector and aspect ratio to further fit and filter the candidate point cloud clusters, eliminating noise points and interfering planes that do not meet the geometric constraints, and finally determining the point cloud set corresponding to the calibration board surface. On this basis, the module identifies the point cloud positions corresponding to each grid point or circle point in the coordinate system of the calibration board according to the relative positions of each grid point or circle point in the board surface, and calculates their three-dimensional coordinates in the lidar coordinate system, thereby forming a three-dimensional feature point set that corresponds one-to-one with the two-dimensional feature points in the camera image.
[0042] Furthermore, after obtaining the correspondence between the two-dimensional and three-dimensional feature points, the coordinate transformation module performs extrinsic parameter estimation for the first and second industrial cameras respectively. Taking the first industrial camera as an example, the module inputs the coordinates of the two-dimensional feature points corresponding to the camera, the coordinates of the three-dimensional feature points in the lidar coordinate system, and the intrinsic parameter matrix of the camera into the extrinsic parameter solving algorithm. It uses a robust solving strategy to estimate the rotation matrix and translation vector between the lidar coordinate system and the imaging coordinate system of the first industrial camera, thus obtaining the first extrinsic parameter matrix.
[0043] Similarly, for the second industrial camera, the module utilizes its two-dimensional feature point set and the same three-dimensional feature point set, combined with the intrinsic parameter matrix of the second industrial camera, to solve for the rotation and translation parameters between the lidar coordinate system and the imaging coordinate system of the second industrial camera, thus obtaining the second extrinsic parameter matrix. To improve the accuracy and stability of the extrinsic parameters, the coordinate transformation module can perform nonlinear optimization on the initial extrinsic parameter results under multiple sets of calibration attitudes, adjusting the rotation and translation parameters by minimizing the reprojection error, ultimately obtaining an extrinsic parameter matrix that meets the accuracy requirements.
[0044] Furthermore, after calibration, the coordinate transformation module saves the intrinsic parameter matrix, distortion parameters, and corresponding extrinsic parameter matrix in the system configuration file, and explicitly constructs the spatial transformation relationship between the 3D point cloud coordinate system and the image coordinate systems of each industrial camera based on the extrinsic parameter matrix. During system operation, when the image perception module and the laser-visual fusion module need to project synchronous 3D point cloud data onto the image plane, they can call the interface provided by the coordinate transformation module to convert the 3D point cloud coordinates in the laser radar coordinate system into pixel coordinates in the corresponding image coordinate system through the extrinsic and intrinsic parameter matrices, or, as needed, inversely transform the pixel positions in the image coordinate system to the laser radar coordinate system, realizing bidirectional coordinate transformation between the radar coordinate system and the camera image coordinate system.
[0045] Through the coordinate transformation module design in this embodiment, the system completes the precise joint calibration between each industrial camera and the vehicle-mounted LiDAR during the installation phase. During the operation phase, it can establish a stable and reliable spatial transformation relationship based on the external parameter matrix and the internal parameter matrix. This enables the subsequent processing unit to associate and transform image data and 3D point cloud data under a unified spatial reference, achieving a consistent mapping between the 2D detection results and the 3D point cloud structure of the same obstacle target. This provides an accurate geometric basis for the precise 3D position calculation of obstacles and the fusion processing of radar and vision.
[0046] In some embodiments, inputting synchronized image data into a target detection network to obtain obstacle coordinates includes: Synchronous image data with different focal lengths are normalized and resized according to a preset format to obtain an image sequence; The image sequence is input into the image feature extraction network, which sequentially extracts intermediate feature maps at multiple scales, and generates an image feature representation that characterizes the semantic information of the scene in front of the train through a feature fusion structure. The image feature representation is input into the target recognition network to determine the two-dimensional position information of multiple candidate target regions on the image plane, and to assign a corresponding obstacle category label to each candidate target region, forming an image target candidate set containing the category information and two-dimensional position parameters of each candidate target.
[0047] Specifically, the target recognition submodule within the image perception module is deployed within the vehicle's onboard computing unit at the front end, interacting with the time synchronization module and coordinate transformation module via an internal bus. The synchronized image data output by the time synchronization module includes near-field images captured by the first industrial camera and far-field images captured by the second industrial camera; each frame carries a unified time stamp and camera identifier. Upon receiving the synchronized image data, the image perception module first performs format unification and preprocessing on the raw images output by cameras with different focal lengths. Then, it inputs the data into the image feature extraction network and the target recognition network to generate image feature representations for subsequent radar-visual fusion and a candidate set of image targets containing obstacle category information and two-dimensional position parameters.
[0048] Specifically, in the preprocessing stage, in this embodiment, the original image resolution output by the first and second industrial cameras is 1920×1080. The image perception module checks and converts the color channel format of each frame of synchronized image data, unifying the image into an RGB three-channel format. Then, according to the input requirements of the target recognition network, the image is normalized and resized. For example, by using aspect ratio-preserving scaling and edge padding, the image is uniformly adjusted to an input size of 640×384, and the pixel values are mapped to the range of [0,1] or [-1,1] according to preset rules, forming an image sequence adapted to the network input. At the same time, the module organizes the images corresponding to the first and second industrial cameras into two sub-sequences according to the camera identifier in the frame header, which facilitates the subsequent differentiation of detection results for near and far-range fields of view.
[0049] During the feature extraction stage, the image perception module inputs the aforementioned image sequence into the image feature extraction network. In this embodiment, the image feature extraction network can employ a backbone network with multi-layer convolution and downsampling structures, such as a residual network or a convolutional network containing an attention mechanism. Through layer-by-layer convolution, activation, and pooling operations, intermediate feature maps at different scales are extracted sequentially.
[0050] For example, the network can output several layers of feature maps with progressively decreasing spatial resolution and progressively increasing channel count. For instance, the first layer might have a resolution of 1 / 4 of the input size, the second 1 / 8, and the third 1 / 16. Subsequently, the image perception module uses a feature fusion structure to upsample and laterally connect these intermediate feature maps at different scales, fusing high-level semantic information with low-level detail information to obtain a set of image feature maps with multi-scale representation capabilities. These feature maps are then integrated through channel mapping or convolution to form an image feature representation that characterizes the semantic information of the scene ahead of the train. This image feature representation is uniformly organized into a fixed-size feature tensor within the network and stored in association with the corresponding camera and time identifiers.
[0051] In the target recognition stage, the image perception module inputs the aforementioned image feature representations into the target recognition network. In this embodiment, the target recognition network can adopt a single-stage detection structure based on YOLOv5, mapping multi-scale image features to detection branches of different scales, and predicting the bounding box offset, class probability, and confidence score corresponding to the candidate anchor point on each branch. The initial detection results output by the target recognition network include bounding box parameters of multiple candidate target regions, such as the bounding box center pixel coordinates (x, y), width w, and height h, as well as the corresponding obstacle class score and confidence score. The image perception module performs threshold filtering and non-maximum suppression processing on the initial detection results, eliminating candidate boxes with confidence scores below a preset threshold and highly overlapping redundant candidate boxes, retaining a set of representative target detection results.
[0052] Regarding category labeling, in this embodiment, the target recognition network performs multi-category classification for common obstacle types in train operation scenarios, including at least pedestrians, trains, falling rocks, animals, and other categories. For each candidate target region that passes the screening, the image perception module reads its category label c and its two-dimensional position parameters (x, y, w, h) in the image plane from the detection results, and combines this information with the corresponding camera identifier and time identifier to assemble this information into image target candidate entries. Finally, the image perception module organizes all target candidate entries within the same camera's field of view at the same time into an image target candidate set, the data structure of which can be represented as several records, each containing a target category label and its corresponding two-dimensional position parameters.
[0053] For example, in a train operating at a maximum speed of 160 km / h, the image perception module can process the synchronized image sequences output by the first and second industrial cameras at an update frequency of 20 Hz. At each moment, it generates a set of near-field and a set of far-field image target candidates, which, along with the corresponding image feature representations, are provided to the radar-visual fusion module. Throughout the processing, the image feature extraction network and the target recognition network can be trained offline and their parameters fixed based on pre-collected trackside and simulation data. During operation, they operate in inference mode, thus enabling continuous target detection of synchronized image data within the limits of the train's onboard computing resources.
[0054] Through the design of the image perception module in this embodiment, the synchronous image data with different focal lengths is uniformly normalized and resized. The multi-scale image feature extraction network is used to extract image features representing the scene in front of the train. The target recognition network outputs a candidate set of image targets containing obstacle categories and two-dimensional position parameters. This enables the system to obtain structured image perception results in both near and far fields of view, providing a reliable image detection foundation for subsequent target screening based on track region segmentation and radar-visual fusion and three-dimensional position calculation combined with three-dimensional point clouds.
[0055] In some embodiments, inputting synchronized image data and 3D point cloud data into a target segmentation network to obtain orbital coordinates includes: Synchronous image data and 3D point cloud data are input into a semantic segmentation network to extract intermediate feature maps of images at least two different spatial scales. The intermediate feature map is fused and spatially transformed to generate an image bird's-eye view feature representation of the scene in front of the train in the bird's-eye view coordinate system. The image bird's-eye view feature representation and the point cloud bird's-eye view feature representation obtained based on synchronous 3D point cloud data are fused at the feature level in the bird's-eye view coordinate system to obtain a fused feature map. The fused feature map is input into the segmentation output layer to obtain pixel-level classification results corresponding to the track region. The classification results are then processed according to preset judgment rules to generate image segmentation results for the track region.
[0056] Specifically, the semantic segmentation submodule within the image perception module is also deployed within the vehicle-mounted computing unit at the front of the train, interacting with the time synchronization module, point cloud perception module, and coordinate transformation module. The synchronized image data output by the time synchronization module includes close-range field-of-view images acquired by the first industrial camera and long-range field-of-view images acquired by the second industrial camera. The semantic segmentation submodule performs track region segmentation processing on each frame of synchronized image data and performs feature-level fusion with the point cloud bird's-eye view feature representation obtained based on the synchronized 3D point cloud data, thereby obtaining a relatively accurate track region image segmentation result in the bird's-eye view coordinate system.
[0057] First, in the feature extraction stage, this embodiment normalizes and resizes the synchronous image data using the same preprocessing method as the target recognition process, for example, uniformly adjusting it to an input resolution of H×W, and then inputs it into an image semantic segmentation network with a multi-scale structure. The encoding part of this network can use a Swin Transformer or other backbone networks with window attention and hierarchical feature extraction capabilities to process the input image hierarchically, sequentially outputting at least two levels, preferably four levels, of intermediate feature maps at different spatial scales. For example, the resolutions are approximately H / 4×W / 4, H / 8×W / 8, H / 16×W / 16, and H / 32×W / 32, with the number of channels increasing progressively. Each level of intermediate feature map complements the other in terms of spatial resolution and semantic abstraction, and is used for subsequent multi-scale fusion and bird's-eye view feature construction.
[0058] Furthermore, in the multi-scale feature fusion and spatial transformation processing, the semantic segmentation submodule first inputs multi-level intermediate feature maps into the feature fusion module. Through operations such as upsampling, skip connections, and channel alignment, high-level semantic features are progressively upsampled to a uniform spatial scale and fused with low-level detail features to generate a fused feature map with a uniform resolution. For example, feature maps of H / 8×W / 8, H / 16×W / 16, and H / 32×W / 32 are upsampled to H / 4×W / 4, and then concatenated or added point-by-point with the H / 4×W / 4 feature map to form an image fusion feature representation containing multi-scale information.
[0059] Subsequently, the semantic segmentation submodule combines camera intrinsic parameters and LiDAR extrinsic parameters to map the image fusion features from the camera imaging plane to a grid plane in the bird's-eye view coordinate system. Following preset ground projection rules, features at different pixel locations are aggregated or interpolated into the corresponding ground grid cells, resulting in an image bird's-eye view feature representation with the same resolution as the point cloud bird's-eye view features. This image bird's-eye view feature representation is organized in memory as a feature tensor of size Hb×Wb×C_I, where Hb and Wb are the number of rows and columns of the bird's-eye view plane, respectively, and C_I is the number of channels.
[0060] Furthermore, on the point cloud side, the point cloud perception module, based on the synchronous 3D point cloud data, after spatial range filtering, ground projection rasterization, and point cloud feature extraction network processing, has obtained a point cloud bird's-eye view feature representation aligned with the aforementioned image bird's-eye view feature representation on the raster scale, denoted as Hb×Wb×C_P. The semantic segmentation submodule, in the bird's-eye view coordinate system, concatenates the image bird's-eye view feature representation and the point cloud bird's-eye view feature representation in the channel dimension to form a joint feature map of size Hb×Wb×(C_I+C_P), and inputs this joint feature map into the fusion weight generation unit.
[0061] The fusion weight generation unit can consist of several pointwise convolutional layers or small fully connected layers. It performs channel compression and mapping on the joint features at each grid location, outputting two sets of weight coefficients corresponding to image features and point cloud features. Subsequently, the semantic segmentation submodule uses these weight coefficients to perform pointwise weighted combination of the image bird's-eye view features and point cloud bird's-eye view features at each grid location, and further performs local spatial context modeling through two-dimensional convolution, finally obtaining the fusion feature map for semantic segmentation, denoted as Hb×Wb×C_F.
[0062] Furthermore, in the segmentation output stage, the semantic segmentation submodule inputs the fused feature map into the segmentation output layer. The segmentation output layer can employ several convolutional layers or 1×1 convolutions combined with Sigmoid or Softmax activation functions to map the fused features at each grid location to pixel-level or grid-level classification probabilities corresponding to orbital and non-orbital categories.
[0063] In some examples, the semantic segmentation submodule obtains classification results of size Hb×Wb×1 or Hb×Wb×2 from the segmentation output layer, where each grid position corresponds to a probability value indicating the presence or absence of an orbital region. Subsequently, the module processes this probability map according to preset decision rules, such as using a fixed threshold or an adaptive threshold to adjust the probability. Figure 2 The binary mask is first quantized to obtain a preliminary track region mask. Then, combined with track geometric priors (such as the approximate direction of the track, the range of track gauge, and the left and right offset constraints relative to the train center), the binary mask is subjected to connected component filtering, morphological closure, and longitudinal continuity checks. Scattered small regions and regions that do not conform to the track morphology constraints are eliminated, and only the main track region that is consistent with the track direction in front of the train is retained, finally generating the image segmentation result of the track region.
[0064] For example, in a double-track railway scenario, the semantic segmentation submodule of this embodiment can simultaneously process the near-field image of the first industrial camera and the far-field image of the second industrial camera, outputting a track area mask corresponding to the current line of the train at each moment, and representing the mask as a track area grid set in the bird's-eye view coordinate system for the radar-visual fusion module to call.
[0065] Through this design, the semantic segmentation submodule utilizes the joint representation of multi-scale image features and point cloud bird's-eye view features to output track region image segmentation results that conform to track geometry constraints in the bird's-eye view coordinate system. This not only improves the stability of track region segmentation under complex lighting and background interference conditions, but also provides a reliable spatial region reference for subsequent target selection and obstacle 3D position calculation based on track region constraints.
[0066] In some embodiments, the point cloud sensing module is specifically used for: The three-dimensional coordinates of each point in the synchronous three-dimensional point cloud data are analyzed according to the preset coordinate system. The point cloud is then filtered according to the horizontal and vertical ranges of the track, and point cloud points that exceed the preset spatial range are removed to obtain the point cloud data within the target range.
[0067] Specifically, the point cloud perception module is deployed in the vehicle-mounted computing unit at the front of the vehicle and is connected to the time synchronization module and the coordinate transformation module. It is used to perform spatial range filtering, ground projection and raster encoding, and feature extraction on the synchronized 3D point cloud data output by the time synchronization module, and to convert the original dense point cloud into a point cloud feature representation with a fixed raster resolution in the bird's-eye view coordinate system for use by the subsequent radar-view fusion module.
[0068] In some examples, the synchronous 3D point cloud data output by the vehicle-mounted LiDAR is described using a right-handed coordinate system. Each point consists of four elements: (x, y, z, r). The x-axis is parallel to the ground and points forward of the train, the y-axis is parallel to the ground and points to the left of the train, the z-axis is perpendicular to the ground and points upward, and r represents the reflection intensity of the point. The point cloud perception module first analyzes each frame of synchronous 3D point cloud data according to the preset coordinate system, reads the 3D coordinates and reflection intensity of each point, and combines this with the geometric parameters of the line and the position of the track center to determine the spatial range that needs to be focused on.
[0069] For example, a three-dimensional spatial box is constructed, encompassing the track area in front of the train and its surrounding environment. The longitudinal range is defined as 0m to 200m in front of the train; the lateral range is defined as 7m to the left and right of the track centerline; and the vertical range is defined as 0.5m below the track surface to 6m above the track surface. The point cloud perception module traverses all points in the current frame's point cloud, retaining only those points whose x, y, and z coordinates all fall within the aforementioned range. Points outside this range are removed, resulting in target point cloud data confined to the target spatial region. This spatial range filtering effectively removes background point clouds far from the track area, reducing the computational burden on subsequent operations.
[0070] Furthermore, after spatial filtering of the target point cloud data, the point cloud perception module projects the target point cloud data onto the ground projection plane using a bird's-eye view coordinate system as a reference. The bird's-eye view coordinate system takes a reference point near the center of the orbit as its origin, and its plane coordinate axes are aligned with the x-axis and y-axis in the radar coordinate system. To facilitate network processing, this embodiment divides the ground projection plane into several regular grid units, for example, dividing it into 2000 equally spaced intervals (each grid length 0.1m) along the forward direction and 280 equally spaced intervals (each grid width 0.05m) along the lateral direction, thus imaged as a grid of approximately 2000×280 pixels.
[0071] The point cloud perception module calculates the corresponding raster index (i,j) on the bird's-eye view plane based on the x and y coordinates of each point and assigns the point to the corresponding raster cell. For each non-empty raster cell, the module statistically analyzes the spatial distribution and reflection intensity characteristics of the point cloud within that cell, such as calculating the maximum height, minimum height, average height, point count, maximum reflection intensity, and average reflection intensity. It can also record whether a point cloud occupancy flag exists within the cell. Through these statistical operations, the point cloud perception module generates a point cloud encoded feature vector containing several dimensions of feature values for each raster.
[0072] Subsequently, the point cloud perception module reorganizes the point cloud encoded features of all grid cells according to their row and column positions in the bird's-eye view coordinate system into a two-dimensional feature map, namely a point cloud encoded feature map of size 2000×280×C_P, where C_P is the number of channels of a single grid feature vector. In this embodiment, the point cloud feature extraction network can adopt a two-dimensional convolutional network similar to the PointPillars structure, treating the aggregated point cloud encoded features within each grid as "pillar" features. By performing convolution, batch normalization, and nonlinear activation on the bird's-eye view plane, the spatial relationship between adjacent grid cells is modeled. The network can be composed of several layers of two-dimensional convolution and downsampling modules stacked together. By progressively extracting high-level semantic features, the original point cloud encoded feature map is gradually transformed into a point cloud feature map with higher semantic abstraction, suitable for fusion with image bird's-eye view features.
[0073] In practical configurations, the point cloud feature extraction network processes the input 2000×280×C_P feature map through multiple convolutions and downsampling to obtain a point cloud feature representation with reduced spatial size and increased channel count. For example, the output point cloud feature map has a size of H_b×W_b×C_F, where H_b and W_b are the downsampled bird's-eye view resolution, and C_F is the number of extracted channels. This point cloud feature map is aligned with the image bird's-eye view features generated by the image semantic segmentation module on a raster scale, providing a unified spatial reference for subsequent feature-level weighted fusion in the bird's-eye view coordinate system. The point cloud perception module outputs this point cloud feature map along with its corresponding time stamp to the Rave-Vision fusion module, where it participates in obstacle recognition and 3D position calculation together with the image feature representation.
[0074] Through the point cloud perception module design in this embodiment, the system can perform spatial range filtering on synchronous 3D point cloud data within a preset track lateral and height range, eliminating point cloud interference that is irrelevant to train operation, and performing raster encoding and feature extraction on the target point cloud data in the bird's-eye view coordinate system, converting the irregular and dense 3D point set into a regular 2D feature map representation. This not only preserves the spatial structure and reflection intensity information near the track area, but also provides a unified, compact and easy-to-compute point cloud feature foundation for subsequent 2D convolutional network processing and multimodal feature fusion.
[0075] In some embodiments, the radar-visual fusion module is also used for: Based on the image segmentation results and the two-dimensional positional relationship of each image target candidate in the image coordinate system, the overlap between each image target candidate and the track region is determined, image target candidates located outside the track region are eliminated, and the target candidate set within the track region is retained.
[0076] Specifically, the radar-view fusion module is deployed within the vehicle's onboard computing unit and interacts with the image perception module, point cloud perception module, and coordinate transformation module via an internal bus. The image bird's-eye view feature representation output by the image perception module in the bird's-eye view coordinate system and the point cloud feature representation output by the point cloud perception module in the bird's-eye view coordinate system are regarded by the radar-view fusion module as two feature maps aligned on the same ground grid coordinate system. Each feature map can be represented as a three-dimensional tensor of size H_b×W_b×C, where H_b and W_b are the number of grid cells in the bird's-eye view plane, and C is the number of channels.
[0077] In the feature-level weighted fusion stage, the Rave-Vision fusion module first concatenates the image feature representation and the point cloud feature representation in the channel dimension under the bird's-eye view coordinate system. For example, for each grid position (i,j), the image feature vector F_img(i,j) at that position is concatenated with the point cloud feature vector F_pts(i,j) to form a joint feature vector F_cat(i,j) of length 2C, and a joint feature map of size H_b×W_b×2C is obtained on the entire bird's-eye view plane.
[0078] Subsequently, the Rayvision fusion module inputs the joint feature map into the fusion weight generation unit. The fusion weight generation unit can consist of several 1×1 convolutional layers or a small multilayer perceptron structure, used to perform channel compression and nonlinear mapping on the joint features at each grid location, outputting two weight coefficients w_img(i,j) and w_pts(i,j) corresponding to the image features and point cloud features, respectively. To ensure that the weight values are within a controllable range, the fusion weight generation unit uses a Sigmoid or normalization operator at the output, ensuring that the two weight coefficients at each grid location fall within the (0,1) interval, and can be further normalized proportionally.
[0079] Furthermore, after obtaining the weight coefficient matrix, the Rave-Vision fusion module performs a weighted combination of the image feature representations and point cloud feature representations for each grid location based on the weight coefficient matrix. Specifically, for grid location (i,j), the module calculates the fusion feature vector according to preset rules. The operation is then performed point-by-point across the entire H_b×W_b grid area to obtain a fused feature map of size H_b×W_b×C. To further explore the spatial context relationships between adjacent grids, the Rave-Vision fusion module can also concatenate several layers of two-dimensional convolution and nonlinear activation operations on the weighted fused feature map to smooth local features and enhance feature consistency within continuous track areas, thereby forming the final fused feature map representing the scene ahead of the train.
[0080] When filtering the image target candidate set based on the track region, the Rave-Vision fusion module obtains the image target candidate set and the image segmentation result of the track region at the corresponding time from the image perception module. For each synchronized image, the image segmentation result gives the track region range in the form of a pixel-level mask, and the image target candidate set consists of several target candidate boxes in the image after being filtered by the target recognition network. Each candidate box is composed of two-dimensional position parameters (x, y, w, h) and obstacle category labels.
[0081] In the image coordinate system, the Rave-Vision fusion module calculates the area of the overlapping region between each target candidate box and the track region mask, as well as the area of the candidate box itself, and obtains the overlap ratio accordingly. If the overlap ratio is lower than a preset threshold (e.g., 10% or 20%), the candidate target is determined to be mainly located outside the track region, marked as an off-track target, and removed from the image target candidate set; otherwise, the candidate target is retained in the target candidate set within the track region.
[0082] For a multi-path image target candidate set from both the first and second industrial cameras, the Rayvision fusion module can perform the above-mentioned filtering steps in their respective image coordinate systems to form target candidate sets within their respective orbits in the near-field and far-field fields of view. These sets are then associated with and stored with the corresponding fusion feature maps for subsequent location of the corresponding point cloud set and calculation of 3D position parameters in the 3D point cloud.
[0083] For example, in a single-track railway scenario, when there are people on the platform, roadside buildings, and foreign objects on the track in the vicinity of the track area ahead during train operation, the radar-visual fusion module of this embodiment can use feature-level weighted fusion to highlight the joint feature expression of point cloud and image in the track area under the bird's-eye view coordinate system. At the same time, it performs geometric constraint screening on the image target candidate set based on the track area image segmentation results, retaining only candidate targets that have sufficient overlap with the track area.
[0084] The design of the above embodiments enables the subsequent 3D point cloud localization and obstacle information generation stages to focus on processing effective targets within the track, reducing the impact of interference from non-track areas in the image and point cloud, improving the relevance and robustness of image and point cloud feature fusion in complex forward scenes, and providing a more reliable input basis for accurate identification of obstacles in front of the train and subsequent 3D position calculation.
[0085] In some embodiments, the radar-visual fusion module is specifically used for: Based on the spatial transformation relationship, the synchronous three-dimensional point cloud data is transformed from the lidar coordinate system to the image coordinate system corresponding to each industrial camera. Point cloud points whose projection positions fall within the two-dimensional position range of each image target candidate are selected in the image plane to form a point cloud set corresponding to each image target candidate. Based on the coordinate distribution of point cloud sets in the three-dimensional point cloud coordinate system, clustering and geometric parameter calculation are performed on each point cloud set to determine the three-dimensional position parameters of each obstacle in the train coordinate system. The three-dimensional position parameters are associated with the obstacle category labels of each image target candidate. Structured obstacle information containing obstacle category identifiers and three-dimensional position parameters is generated according to the preset data structure and output to the on-board train control system in a preset data interface format.
[0086] Specifically, the obstacle calculation and information output process is jointly completed by the obstacle 3D calculation submodule and the information output submodule in the radar-visual fusion module. This process takes the synchronized 3D point cloud data output by the time synchronization module, the image target candidate set output by the image perception module, and the spatial transformation relationship provided by the coordinate transformation module as input. By locating the point cloud set corresponding to each image target candidate in the 3D point cloud, the 3D position parameters of each obstacle in the train coordinate system are calculated, and these parameters are assembled with category labels to form structured obstacle information. Finally, the information is output to the on-board train control system in a preset data interface format.
[0087] During the point cloud localization phase, the obstacle 3D solution submodule first reads the extrinsic parameter matrix between the LiDAR coordinate system and the imaging coordinate system of each industrial camera, as well as the camera's intrinsic parameter matrix, from the coordinate transformation module. At the current synchronization moment, the module acquires the corresponding synchronized 3D point cloud data and the image target candidate set filtered by the track region. For a synchronized 3D point cloud under the field of view of a certain industrial camera, the module traverses each point in the point cloud data, uses the extrinsic parameter matrix to transform the point from the LiDAR coordinate system to the camera coordinate system of the industrial camera, and then completes perspective projection through the intrinsic parameter matrix to obtain the pixel coordinates (u,v) of the point in the image coordinate system. If the depth value of the point is positive and (u,v) falls within the effective imaging area of the image, then the point is considered to be within the imaging range of the camera.
[0088] Subsequently, the module compares the projected position of the point with the two-dimensional position parameters of each candidate box in the current camera's image target candidate set to determine whether the projected pixel of the point falls within the boundary range of a candidate box. All points that meet the conditions are recorded, and an initial point cloud set is built for each image target candidate based on this. If the number of projected point clouds in a target candidate box is less than a preset lower limit, the candidate can be marked as having insufficient point cloud support and will not participate in the 3D solution for the time being.
[0089] Furthermore, in the 3D geometric parameter calculation stage, the obstacle 3D solution submodule performs clustering and geometric parameter calculations for each target candidate with sufficient point cloud support, based on the coordinate distribution of its corresponding point cloud set in the 3D point cloud coordinate system. Specifically, the module can first perform clustering processing on the point cloud set based on a distance threshold, dividing spatially distinct subclusters into multiple subsets to handle situations where the candidate box contains multiple neighboring targets simultaneously. Within each subset, the module calculates the 3D geometric center of the point cloud (e.g., using the average or median of all point coordinates) and the corresponding 3D bounding box parameters, including the minimum and maximum x-coordinates along the train's forward direction, the minimum and maximum y-coordinates along the lateral direction, and the minimum and maximum z-coordinates along the height direction.
[0090] If a rigid transformation relationship between the lidar coordinate system and the train coordinate system is pre-established in the system, the module further utilizes this transformation relationship to convert the geometric center coordinates and bounding box coordinates from the lidar coordinate system to the train coordinate system, so that each obstacle has a uniform longitudinal distance, lateral offset, and height representation in the train coordinate system. In this embodiment, the train coordinate system takes a reference point at the front of the train as its origin, with the x-axis along the forward direction of the train, the y-axis pointing to the left side of the train, and the z-axis pointing upwards.
[0091] Furthermore, in the structured information generation stage, the obstacle 3D solution submodule associates the aforementioned 3D position parameters with the category labels of the image target candidates. For each subset of obstacles obtained through point cloud clustering and geometric calculation, the module reads the obstacle category label (e.g., pedestrian, train, falling rock, animal, or others) from its corresponding image target candidates, and extracts at least the (x_c, y_c, z_c) coordinates of the obstacle's center point in the train coordinate system, the obstacle's size parameters (e.g., length, width, height) in the forward, lateral, and height directions, and optional point cloud confidence indices from the 3D geometric parameters, based on the actual application requirements.
[0092] Subsequently, the module organizes the aforementioned category identifiers and 3D position parameters into structured obstacle information entries according to a preset data structure format. Each entry may include fields such as obstacle number, category code, center coordinates, bounding box size, coordinate system identifier, and time identifier. For the same physical obstacle from different industrial camera fields of view, the module can associate and merge results based on the 3D position proximity and category consistency in the train coordinate system to avoid duplicate output.
[0093] Furthermore, during the information output phase, the information output submodule packages the structured information of all obstacles within the same timeframe into obstacle information messages according to a preset period, and sends them to the train control system via onboard Ethernet, MVB, or other onboard communication links. The messages contain a unified time identifier and a data version identifier, enabling onboard control units such as the automatic protection system and automatic train operation system to parse the obstacle category and its three-dimensional position parameters in the train coordinate system according to a preset interface protocol, and to call upon these parameters in the upper-level control logic.
[0094] Through the obstacle 3D solution and information output design of this embodiment, the system can use the pre-established spatial transformation relationship to accurately correspond the image target candidate and the 3D point cloud data at the same obstacle target level. Based on the 3D distribution of the target corresponding point cloud set, the system obtains its 3D position parameters in the train coordinate system, and generates structured obstacle information by combining the category labels obtained from image recognition. Thus, the system obtains a 3D description of obstacles with a unified coordinate system and a unified data format on the vehicle control system side, which improves the accuracy and consistency of obstacle 3D position solution and provides directly callable basic data for the subsequent train control unit to perform distance assessment, braking strategy calculation and related decision processing.
[0095] The above embodiments provide a detailed description of the framework and functions of the train active perception system based on radar-visual fusion of this application. The following section uses a practical application scenario as an example to illustrate the specific implementation process and principles of the train active perception system of this application.
[0096] Both cameras are 5MP industrial cameras, equipped with 25mm and 75mm lenses respectively, to perceive the track environment at different distances and avoid blind spots in image information. The 25mm lens is responsible for close-range perception, while the 75mm lens is responsible for long-range perception. The lidar is an automotive-grade radar with a maximum detection range of 500 meters, used to detect the specific coordinates of obstacles in front of the train.
[0097] To eliminate the heterogeneous and asynchronous issues in data output rate and volume among the three sensors, the system employs a high-precision time alignment algorithm at the data entry point. This algorithm maps all data to a unified time base, laying a solid foundation for subsequent deep fusion processing. Specifically, data from each sensor is temporarily stored in an independent queue. The time alignment algorithm approximates different data sets each time to achieve synchronization, dynamically matching the data frames with the closest timestamps through a sliding window to minimize synchronization errors. To enhance robustness, the algorithm integrates a moving average and outlier removal mechanism, effectively suppressing data jitter and ensuring a continuous output of high-confidence, low-jitter synchronized data stream even in extreme scenarios such as network congestion or momentary loss of synchronization. This provides a reliable timing foundation for subsequent fusion computation.
[0098] During the data processing phase, the point cloud data generated by the LiDAR needs to be preprocessed first. The collected raw point cloud data is... ,in This indicates the number of point clouds, and 4 represents the three-dimensional spatial coordinates and reflectivity of the point clouds. The coordinate system used in the system is as follows: Figure 3 As shown, Figure 3 This is a schematic diagram of the coordinates of the lidar used in the system provided in this application embodiment. The x-axis is parallel to the ground and points forward, the y-axis is parallel to the ground and points to the left, and the z-axis is perpendicular to the ground and points upward. Since the lidar has a wide field of view, point cloud data within the useless range needs to be removed. Therefore, pass-through filtering is performed on the y-axis and z-axis. The filtering range varies depending on the application scenario; here, the filtered point cloud data is set to... , This indicates the number of points in the filtered point cloud.
[0099] Furthermore, for collaborative sensing using different sensors, the coordinate systems of long-focus and short-focus cameras and radar need to be aligned, meaning their coordinates can be transformed into each other. The joint calibration process for camera and radar coordinates involves: first, simultaneously using two cameras and a LiDAR to acquire calibration board data at different positions and angles; then, obtaining the coordinates of the calibration board in the image and point cloud; and finally, using a random sample consensus algorithm to calculate the coordinate transformation matrix between the two, i.e., the extrinsic parameter matrix. This extrinsic parameter matrix allows for the transformation between their coordinates. The intrinsic parameter matrix for the image is as follows: in, and It is the focal length (in pixels), which represents... x and y Scaling factor of direction u 0、 v0 represents the actual location of the principal point, i.e., the intersection of the optical axis and the image plane; the distortion coefficients are divided into radial distortion and tangential distortion, with a total of 5 parameters. The intrinsic parameter matrix and distortion matrix reflect the camera's own properties, which vary from camera to camera. These parameters are calculated using relevant functions in MATLAB or OpenCV.
[0100] Finally, a corner detection algorithm is used to obtain the image coordinates of the calibration board, i.e., the 2D coordinates of the calibration board. For point cloud processing, a region growing method + scale features + prior knowledge is used to determine the point cloud of the calibration board. Then, a DIoU-based algorithm is used to determine the 3D coordinates of the corners of the calibration board. Finally, based on the 2D-3D coordinate pairs of the corners, a random sampling consensus algorithm is used to calculate the extrinsic parameter matrix. in, It is a 3×3 rotation matrix. It is a 3×1 translation matrix, therefore the overall extrinsic parameter matrix is 3×4. The intrinsic parameter matrices for the telephoto camera and LiDAR calibration are... and The intrinsic parameter matrices for the short-focus camera and LiDAR calibration are as follows: and .
[0101] The image data processing involves object detection and semantic segmentation algorithm reasoning. The object detection algorithm uses the YOLOv5 algorithm to identify the category and location of obstacles in the image. Categories include pedestrians, trains, falling rocks, animals, and others. Location includes the obstacle's center coordinates and its length and width. The semantic segmentation algorithm employs a segmentation algorithm based on BEV feature fusion, the framework of which is as follows: Figure 4 As shown, Figure 4 This is a schematic diagram of a segmentation algorithm based on BEV feature fusion provided in an embodiment of this application. The segmentation algorithm includes the following operations: Input image data is H and W represent the length and width of the image, respectively, and 3 represents the number of RGB channels. The image data is first processed by the Swin Transformer, a multi-scale image feature extraction network capable of outputting image features at four different scales. ,in C1, C2, C3, C4 These represent the number of channels for each scale of the feature [192, 384, 768, 1536]. Then, to ensure the features simultaneously possess information from multiple scales, these four scales of features are input into a multi-scale feature fusion module, which outputs features containing multiple scales. Finally, the multi-scale features... The input is fed into the BEV feature extraction module to extract the BEV features from the image. , The number of channels for the BEV feature is 384.
[0102] Furthermore, point cloud data is used to extract features through Point Pillars, a network that is a point cloud feature extraction network capable of directly outputting the BEV features of the point cloud. ,in equal After obtaining the BEV features of the image and point cloud, the two types of features are fused using the BEV feature fusion module. Specifically, the feature concatenation operation is first performed on the image and point cloud. and Perform channel feature splicing Then, based on the stitched features, the fusion weights of the image BEV features and the point cloud BEV features are calculated separately, and an embedding vector is further introduced. Multiply by the concatenated features, and finally use Sigmoid The activation function and channel normalization are used to obtain the weights. The calculation process is as follows: in, This is a matrix of attention coefficients where all terms are greater than 0. To standardize channel features, E It is a learnable parameter vector.
[0103] for and The adaptive fusion formula is: in, for( i , j The channel features after location fusion; This represents the final characteristic representation of the merged BEV space; and They are respectively and The middle position is ( i , j The corresponding eigenvalues; and For BEV features ( i , j )Location and The corresponding fusion weight value.
[0104] Finally, using two fully connected layers and one Sigmoid The activation function yields the final segmentation probability, as shown in the following formula: in, For the segmentation probability; This is the first fully connected layer, and the dimension of the number of feature channels varies as follows: ; This is the second fully connected layer, and the dimension of the number of feature channels changes as follows: Z represents the final feature representation of the merged BEV space.
[0105] The telephoto camera, based on object detection and semantic segmentation algorithms, can output a list of obstacles. and orbital coordinates The short-focus camera, based on object detection and semantic segmentation algorithms, can output a list of obstacles. and orbital coordinates . These represent the center pixel coordinates, width, height, and semantic category label of the target detection box, respectively.
[0106] The system first constructs a position discrimination model based on the spatial coordinate mapping relationship between the target detection box and the preset track area. It determines whether the target is located within the track area using geometric constraints and filters out interference from targets outside the track. Subsequently, the system imports 3D point cloud data that is strictly time-synchronized with the image frames and loads precisely calibrated sensor intrinsic and extrinsic parameter matrices. Based on the above multi-source data, the system projects the 3D point cloud of each detection point onto the 2D image coordinate system through a coordinate system transformation process. Based on the 2D pixel coordinates of the target obstacle identified in the image, the system locates and extracts the corresponding point cloud cluster in the 3D point cloud data through the strict spatial mapping relationship between the image coordinate system and the point cloud coordinate system. Then, based on the 3D coordinate information of the point cloud, the system accurately calculates the Euclidean distance between the target obstacle and the sensor, thereby achieving a consistent conversion from 2D perception to 3D measurement.
[0107] Compared to traditional sensing methods that rely on a single sensor or passively receive information, the radar-camera fusion-based train active perception system of this application achieves high-precision real-time perception of the train's operating environment by deeply fusing radar and camera data. This system can effectively overcome the perception limitations of a single sensor in complex scenarios such as extreme weather and insufficient light, and can proactively identify potential risks and provide early warnings, thereby significantly improving the safety and intelligence level of train operation.
[0108] The implementation process of the radar-visual fusion train active perception method of this application will be described in detail below with reference to specific embodiments. Figure 5 This is a flowchart illustrating the radar-visual fusion-based active train perception method provided in this application embodiment, as shown below. Figure 5 As shown, the train active perception method based on radar-visual fusion may specifically include the following steps: S501 collects multi-channel image data containing images with different focal lengths and corresponding 3D point cloud data in front of the train, and adds time markers to each channel of image data and 3D point cloud data. S502, performs time alignment of multi-channel image data and 3D point cloud data according to time stamp, and generates synchronized image data and synchronized 3D point cloud data; S503, based on image data and 3D point cloud data containing the calibration target, calculates the external parameters between the lidar coordinate system and the imaging coordinate system of each industrial camera, and establishes the spatial transformation relationship between the 3D point cloud coordinate system and the image coordinate system. S504 performs feature extraction and target recognition on synchronous image data to obtain image feature representations of the scene in front of the train and a candidate set of image targets containing obstacle categories, and obtains image segmentation results of the track area through image semantic segmentation. S505, preset spatial range filtering for synchronous 3D point cloud data; S506, Based on the image segmentation results, the image target candidate set is filtered, and the image target candidate set located in the orbital region is retained; S507 uses spatial transformation relationships to locate the point cloud set corresponding to each image target candidate in synchronous three-dimensional point cloud data, calculates the three-dimensional position parameters and category information of each obstacle in the train coordinate system, generates structured obstacle information, and outputs it to the on-board train control system.
[0109] Specifically, the radar-visual fusion-based active train perception method of this application relies on the aforementioned radar-visual fusion-based active train perception system. For example, the system can be installed on the front of a high-speed train with a maximum operating speed of 160 km / h on a mainline railway. Its hardware includes two industrial cameras with different focal lengths, an automotive-grade LiDAR, and an onboard computing and time synchronization device. The following is a combination of... Figure 5 The flowchart shown provides a detailed explanation of the implementation process of each step.
[0110] In step S501, the system first performs multi-source data acquisition and time stamping. The acquisition module controls a first industrial camera and a second industrial camera mounted on the front of the train to periodically acquire image data of the area in front of the train. The first industrial camera is equipped with a 25mm focal length lens to acquire near-to-medium distance images of the track environment with a wide field of view; the second industrial camera is equipped with a 75mm focal length lens to acquire mid-to-long distance images of the track environment with a narrower field of view but clearer details. Both cameras have a resolution of 5 megapixels and output RGB format images.
[0111] The vehicle-mounted LiDAR is installed near the camera, with a main field of view covering a distance of 0m to 500m in front of the train. It periodically outputs 3D point cloud frame data containing (x, y, z, r). The acquisition module reads the current system time when each frame of image data and each frame of point cloud data is formed, using a unified time reference shared with the train integrated monitoring unit or vehicle control unit. This time is written into the frame header as a time identifier, while retaining additional information such as sensor number, exposure parameters, and scanning cycle, forming multi-channel image data and 3D point cloud data with time identifiers.
[0112] In step S502, the time synchronization module performs time alignment processing on the multi-channel image data and 3D point cloud data. The time synchronization module maintains independent data queues for different sensors, writing image frames and point cloud frames with time stamps into the corresponding queues in arrival order. Using the point cloud frames output by the LiDAR as the time reference stream, whenever a new frame arrives in the point cloud queue, its time stamp t_ref is read. Within a preset sliding time window (e.g., ±50ms), image frames whose absolute difference between the time stamp and t_ref is less than a first threshold (e.g., 30ms) are searched in the first and second image queues. The image frame with the smallest time difference is selected and combined with the point cloud frame to form a candidate synchronization data combination.
[0113] The time synchronization module performs statistical analysis on the time difference sequence of candidate combinations formed over a period of time. It adaptively sets a second threshold based on the moving average and dispersion. Combinations with time differences greater than the second threshold or significantly deviating from the statistical mean are considered abnormal matches and are removed. The image frames and point cloud frames in the retained combinations are marked as synchronized image data and synchronized 3D point cloud data, respectively, and a unified synchronization sequence number is attached for subsequent processing.
[0114] In step S503, the coordinate transformation module constructs a spatial transformation relationship using calibration data collected during the system installation and debugging phase. Specifically, with the train stationary, a calibration board with a checkerboard pattern or dot array is positioned at different distances, heights, and deflection attitudes in front of the train, ensuring the calibration board is simultaneously within the field of view of two industrial cameras and a LiDAR. The acquisition module simultaneously acquires multiple sets of calibration board images and corresponding 3D point cloud data at various calibration attitudes. Based on the camera calibration algorithm, the coordinate transformation module extracts the 2D pixel coordinates of the checkerboard corner points or center points from each frame of the calibration board image, and, combined with the calibration board's geometric dimensions, solves for the intrinsic parameter matrices and distortion parameters of the two industrial cameras, establishing their respective image coordinate systems.
[0115] For point cloud data, the module employs a combination of region growing and geometric constraints to segment point cloud clusters corresponding to the calibration plate plane. It further identifies the coordinates of 3D feature points corresponding to each grid point or dot, forming a set of 3D feature points that correspond one-to-one with the 2D feature points in the image. Based on the correspondence between 2D and 3D feature points, extrinsic parameter solutions are performed on the first and second industrial cameras respectively, obtaining the rotation matrix and translation vector from the LiDAR coordinate system to the imaging coordinate system of each camera, thus forming the extrinsic parameter matrix. By performing nonlinear optimization on the extrinsic parameter results under multiple calibration attitudes, the final extrinsic parameters that meet the accuracy requirements are obtained. Based on this, a spatial transformation relationship between the 3D point cloud coordinate system and each image coordinate system is established, providing a unified geometric model for point cloud projection and back projection during the runtime phase.
[0116] In step S504, the image perception module performs feature extraction, target recognition, and track region semantic segmentation on the synchronized image data. First, the synchronized images from two industrial cameras are normalized and resized, for example, uniformly scaled to a preset resolution H×W, mapping pixel values to a specified range to form an image sequence. Subsequently, the image sequence is input into a multi-scale image feature extraction network, which sequentially extracts intermediate feature maps at different spatial scales. A feature fusion structure (such as a top-down feature pyramid) is then used to fuse high-level semantic features with low-level edge details to obtain an image feature representation that characterizes the semantic information of the scene ahead of the train.
[0117] Image feature representations are fed into a target recognition network (e.g., a single-stage detection network based on YOLOv5). The network predicts bounding boxes and class scores of candidate target regions on multi-scale feature maps. After confidence thresholding and non-maximum suppression, an image target candidate set is formed, containing the center pixel coordinates, width, height, and obstacle class labels of each candidate box. Simultaneously, a semantic segmentation sub-network synchronously inputs image data into a multi-scale encoding network (e.g., based on SwinTransformer and feature fusion structure), outputting pixel-level or raster-level track region classification results on the image plane or a raster plane after bird's-eye view transformation. After binarization and track geometric constraint processing, image segmentation results of the track regions are generated.
[0118] In step S505, the point cloud perception module performs spatial range filtering, ground projection and rasterization encoding, and two-dimensional feature extraction on the synchronous 3D point cloud data. For each frame of synchronous 3D point cloud data, the module first analyzes the (x, y, z) coordinates of each point according to the preset radar coordinate system, and sets spatial filtering intervals according to the track lateral range, longitudinal distance range, and height range, retaining only the point cloud data falling on the track in front of the train and its vicinity, thus obtaining the target point cloud data.
[0119] Subsequently, using the bird's-eye view coordinate system as a reference, the target point cloud data is projected onto the ground plane, and regular grids are divided along the forward and lateral directions. Each point is assigned to its corresponding grid cell according to its projected coordinates, and point cloud coding features are generated based on statistical information such as the height distribution, quantity, and reflection intensity of the point cloud within the grid. The point cloud coding features of all grids are recombined into a two-dimensional feature map according to their spatial location and input into a point cloud feature extraction network (such as a PointPillars-type network). Through convolution and downsampling operations on the bird's-eye view plane, the spatial relationship between grids is modeled, and finally, a point cloud feature representation of size H_b×W_b×C_P is output in the bird's-eye view coordinate system, providing a point cloud-side feature foundation for subsequent radar-visual fusion.
[0120] In step S506, the radar-view fusion module performs feature-level weighted fusion of image feature representations and point cloud feature representations in the bird's-eye view coordinate system, and filters the candidate set of image targets based on the track region segmentation results. Specifically, the module aligns the image bird's-eye view features generated by the image semantic segmentation module with the point cloud features output by the point cloud perception module at the raster level, concatenates them in the channel dimension, and inputs them into the fusion weight generation unit. For each raster position, it calculates the weight coefficients corresponding to the image features and point cloud features. After Sigmoid or normalization processing, it uses these weights to perform weighted summation of the two features and combines it with convolution operations to form a fusion feature map that integrates image texture semantic information and point cloud geometric shape information.
[0121] Meanwhile, the Rave-Vision fusion module uses the image segmentation results of the track region and the two-dimensional positional relationship of each image target candidate in the image coordinate system to calculate the overlap ratio between the candidate box and the track region mask. Target candidates with an overlap ratio lower than a preset threshold are eliminated, and only the set of image target candidates located in the track region or significantly overlapping with the track region is retained. The set of target candidates in the track region, along with the fusion feature map and the synchronization sequence number, are then transmitted to the obstacle three-dimensional solution stage.
[0122] In step S507, the system utilizes the spatial transformation relationship established in step S503 to accurately locate the point cloud set corresponding to each image target candidate in the synchronous 3D point cloud data, calculates the 3D position parameters of each obstacle in the train coordinate system, and generates structured obstacle information. The obstacle 3D solution submodule first projects the points in the synchronous 3D point cloud data from the lidar coordinate system to the image coordinate system of each camera based on the extrinsic and intrinsic parameter matrices, determines whether the projected position of the point falls within the 2D bounding box of the corresponding image target candidate, and adds the points that meet the conditions to the point cloud set of that target. For each image target candidate's point cloud set, the module performs clustering and geometric parameter calculation in the 3D point cloud coordinate system to obtain the 3D geometric center coordinates and spatial boundary range of the obstacle, and maps these parameters to the train coordinate system through a preset rigid transformation to form longitudinal distance, lateral offset, and height parameters with the train head as a reference.
[0123] Then, the module associates the 3D position parameters with the obstacle category labels of the image target candidates, forming a structured obstacle information record according to a preset data structure. This record includes at least the obstacle category identifier, center coordinates in the train coordinate system, bounding box size, time stamp, and data validity flag. The information output submodule packages multiple obstacle information entries from the same moment into a message according to the onboard communication protocol and sends it to the onboard train control system via onboard Ethernet or other bus interfaces for subsequent logical processing by the train automatic protection system or train automatic operation system.
[0124] The active train perception method based on radar-visual fusion in this embodiment processes multiple images with different focal lengths and high-precision 3D point clouds at a unified time reference and a unified spatial coordinate system. It successively completes perception data acquisition and time alignment, joint calibration of camera and lidar, image detection and track area segmentation, point cloud bird's-eye view feature extraction, and feature-level radar-visual fusion from the bird's-eye view. Under the constraints of the track area, it filters target candidates and calculates their 3D positions. Finally, it outputs the structured obstacle information to the on-board train control system, enabling the train control unit to make subsequent decisions and controls based on obstacle categories and 3D position data in a unified coordinate system.
[0125] It should be understood that the sequence number of each step in the above method embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0126] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although the technical solutions of this application have been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A radar-vision fusion-based active train perception system, characterized in that, The application relates to a train front obstacle detection method and device. The application comprises: a collection module for collecting multi-channel image data containing images of different focal lengths and three-dimensional point cloud data corresponding to the front of a train, and adding time identifiers to each channel of data; a time synchronization module for time aligning the multi-channel image data and the three-dimensional point cloud data according to the time identifiers of each channel of data, and generating synchronized image data and synchronized three-dimensional point cloud data; a coordinate transformation module for calculating external parameters between a laser radar coordinate system and imaging coordinate systems of each industrial camera based on image data and three-dimensional point cloud data containing a calibration target, and establishing a spatial transformation relationship between a three-dimensional point cloud coordinate system and an image coordinate system; an image perception module for inputting the synchronized image data into a target detection network to obtain obstacle coordinates, and inputting the synchronized image data and three-dimensional point cloud data into a target segmentation network to obtain track coordinates; a point cloud perception module for performing preset spatial range screening on the synchronized three-dimensional point cloud data to obtain three-dimensional point cloud data in front of the train; 2. The system of claim 1, wherein, a radar and vision fusion module for locating a point cloud set corresponding to each obstacle target candidate in the three-dimensional point cloud data by using the spatial transformation relationship, assigning an accurate central representative coordinate to each obstacle, and generating structured obstacle information and outputting the information to a vehicle-mounted train control system. The collection module is specifically used for: collecting multi-channel image data covering different visual distance ranges in front of the train by setting first and second industrial cameras with different focal lengths; collecting three-dimensional point cloud data covering a preset distance range in front of the train by a vehicle-mounted laser radar; 3. The system of claim 1, wherein, based on a unified time reference of the train, adding time identifiers to the multi-channel image data and the three-dimensional point cloud data respectively to form image data and three-dimensional point cloud data with time identifiers. The time synchronization module is specifically used for: buffering the multi-channel image data and the three-dimensional point cloud data with time identifiers in corresponding data queues respectively; taking the time identifier of one channel of data as a reference, searching for data frames with a time identifier difference not exceeding a first threshold value in other data queues within a preset sliding time window, combining the searched data frames to form a candidate synchronous data combination; 4. The system of claim 1, wherein, statistically analyzing the time difference of the candidate synchronous data combination, eliminating abnormal matching combinations with a time difference exceeding a second threshold value, and outputting corresponding synchronized image data and synchronized three-dimensional point cloud data. The coordinate transformation module is specifically used for: synchronously collecting multiple sets of image data and three-dimensional point cloud data containing a calibration target by each industrial camera and the laser radar in different positions and postures; extracting two-dimensional feature point coordinates of the calibration target from the image data based on a camera calibration algorithm, and calculating internal parameters and distortion parameters of each industrial camera to obtain corresponding image coordinate systems; extracting three-dimensional feature point coordinates of the calibration target from the three-dimensional point cloud data based on point cloud clustering and geometric constraints, and constructing a corresponding relationship between the two-dimensional feature points and the three-dimensional feature points. Based on the correspondence between the two-dimensional feature points and the three-dimensional feature points, rotation and translation parameters between a laser radar coordinate system and imaging coordinate systems of the industrial cameras are estimated to obtain an external parameter matrix, and the spatial transformation relationship is established according to the external parameter matrix.
5. The system of claim 1, wherein, The synchronous image data is input into a target detection network to obtain obstacle coordinates, including: The synchronous image data of different focal lengths is normalized and size-adjusted according to a preset format to obtain an image sequence; The image sequence is input into an image feature extraction network to sequentially extract intermediate feature maps of multiple scales, and an image feature representation representing semantic information of a scene in front of the train is generated through a feature fusion structure; The image feature representation is input into a target recognition network to determine two-dimensional position information of multiple candidate target regions on an image plane, and assign corresponding obstacle class labels to each candidate target region to form an image target candidate set containing candidate target class information and two-dimensional position parameters.
6. The system of claim 5, wherein, The synchronous image data and the three-dimensional point cloud data are input into a target segmentation network to obtain track coordinates, including: The synchronous image data and the three-dimensional point cloud data are input into a semantic segmentation network to extract image intermediate feature maps of at least two different spatial scales; The intermediate feature maps are subjected to feature fusion and spatial transformation processing to generate an image bird's eye view feature representation representing a scene in front of the train in a bird's eye view coordinate system; The image bird's eye view feature representation and a point cloud bird's eye view feature representation obtained based on the synchronous three-dimensional point cloud data are subjected to feature-level fusion in the bird's eye view coordinate system to obtain a fusion feature map; The fusion feature map is input into a segmentation output layer to obtain a pixel-level classification result corresponding to a track region, and the classification result is processed according to a preset determination rule to generate an image segmentation result of the track region.
7. The system of claim 1, wherein, The point cloud perception module is specifically configured to: The three-dimensional coordinates of each point in the synchronous three-dimensional point cloud data are analyzed according to a preset coordinate system, and the point cloud is subjected to spatial range screening according to a track lateral range and a height range, and point cloud points exceeding the preset spatial range are removed to obtain point cloud data within a target range.
8. The system of claim 6, wherein, The radar and vision fusion module is further configured to: According to the image segmentation result and the two-dimensional position relationship of each image target candidate in the image coordinate system, the overlap of each image target candidate and the track region is determined, and image target candidates located outside the track region are removed to retain a target candidate set within the track region.
9. The system of claim 1, wherein, The radar and vision fusion module is specifically configured to: Based on the spatial transformation relationship, the synchronous three-dimensional point cloud data is converted from a laser radar coordinate system to image coordinate systems corresponding to the industrial cameras, and point cloud points whose projection positions fall within the two-dimensional position range of each image target candidate are selected in the image plane to form a point cloud set corresponding to each image target candidate; Based on the coordinate distribution of the point cloud set in the three-dimensional point cloud coordinate system, each point cloud set is clustered and geometric parameter calculation is performed to determine three-dimensional position parameters of each obstacle in the train coordinate system; The three-dimensional position parameter is associated with the obstacle category label of each image target candidate, a structured obstacle information containing the obstacle category identification and the three-dimensional position parameter is generated according to a preset data structure, and is output to the vehicle-mounted train control system in a preset data interface format.
10. A train active perception method based on the radar-vision fusion of the system according to any one of claims 1 to 9, characterized in that, Comprise: Collecting multi-channel image data containing images of different focal lengths and corresponding three-dimensional point cloud data in front of the train, and adding time identifiers to each channel of image data and three-dimensional point cloud data; According to the time identifier, the multi-channel image data and the three-dimensional point cloud data are time-aligned to generate synchronized image data and synchronized three-dimensional point cloud data; Based on the image data and three-dimensional point cloud data containing the calibration target, the external parameters between the laser radar coordinate system and each industrial camera imaging coordinate system are calculated, and the spatial transformation relationship between the three-dimensional point cloud coordinate system and the image coordinate system is established; Feature extraction and target recognition are performed on the synchronized image data to obtain image feature representation representing the scene in front of the train and image target candidate set containing obstacle categories, and image segmentation results of the track area are obtained through image semantic segmentation; The synchronized three-dimensional point cloud data is screened according to a preset spatial range; Based on the image segmentation results, the image target candidate set is screened to retain the image target candidate set located in the track area; Using the spatial transformation relationship, the point cloud set corresponding to each image target candidate is located in the synchronized three-dimensional point cloud data, the three-dimensional position parameter and the category information of each obstacle in the train coordinate system are calculated, the structured obstacle information is generated and output to the vehicle-mounted train control system.
Citation Information
Patent Citations
Urban rail barrier detecting device and method
CN110329316A
Perception result acquisition method and device, computer equipment and storage medium
CN115240168A
Multi-data fusion train active obstacle detection method and system
CN118038412A
Joint calibration method and device, computer storage medium and terminal
CN118823134A
Railway track obstacle detection method and system
CN119291717A
Cited By
Semantic annotation method and device, model training method and device, electronic equipment and storage medium
CN122223721A