Method and system for detecting train obstacles based on four-eye sensor
By constructing a fused three-dimensional perception space using a four-eye sensor system and combining it with orbital coordinate system constraints, the problem of inaccurate obstacle recognition under extreme lighting and inclement weather conditions in existing technologies has been solved, achieving efficient, low-cost, and full-coverage obstacle detection in complex environments.
Patent Information
- Application Number
- CN202511983826.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2045-12-26
AI Technical Summary
Existing obstacle detection solutions for trains lack accuracy in extreme lighting and adverse weather conditions. Traditional vision and lidar fusion solutions fail to fully exploit the complementarity of heterogeneous data and have high hardware costs, making it impossible to provide accurate 3D obstacle information in complex environments.
A quad-sensor system is adopted, which combines binocular visible light sensors and binocular far-infrared sensors to simultaneously acquire stereo image data. Through cross-validation of heterogeneous visual information and spatial correlation processing, a fused three-dimensional perception space is constructed. Dynamic perception boundaries are generated by combining orbital coordinate system constraints, the perception space is cropped, and obstacles are identified and labeled using a multi-attribute decision network.
Maintaining stable perception under extreme lighting and inclement weather improves the robustness and accuracy of obstacle recognition, reduces hardware costs, avoids the shortcomings of traditional solutions, and ensures full coverage and safety of the track area.
Smart Images

Figure CN121392800B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of train operation safety monitoring technology, specifically to a method and system for detecting train obstacles based on a four-eye sensor. Background Technology
[0002] Existing obstacle detection solutions for trains largely rely on single-type sensors, such as visible light cameras, lidar, or single infrared thermal imagers. Visible light vision is highly susceptible to failure under complex lighting conditions such as nighttime, fog, and strong light, making it difficult to reliably acquire three-dimensional geometric information of the scene. While infrared thermal imaging can penetrate some smoke and dust and detect temperature differences, its imaging lacks texture details, has low spatial resolution, and struggles to identify objects at normal temperatures or with small temperature differences from the background. Some current technologies attempt to fuse visible light and infrared images at a two-dimensional level to improve recognition, but this fusion fails to create a unified and accurate three-dimensional spatial description, unable to provide precise, usable three-dimensional position and contour information for obstacle detection. Furthermore, a common technique combines visual sensors with lidar to generate three-dimensional point clouds for obstacle detection. While lidar can provide direct distance information, its detection range is significantly reduced in adverse weather conditions such as rain, snow, and dense fog. Furthermore, its hardware cost is high, which is not conducive to large-scale engineering applications. At the same time, the fusion of vision and lidar in this type of solution is mostly a simple information superposition, which fails to fully explore the texture details of visual data and the complementarity of lidar distance data. For obstacles with low reflectivity or small size, it is still difficult to achieve accurate identification, and there is a certain risk of missed detection.
[0003] Another common approach is to define a fixed detection area based on the sensor's own coordinate system, such as a rectangular area with a fixed forward distance and width. This method does not consider the actual spatial constraints of the train's running track. On curves, slopes, and other similar tracks, the fixed area may contain a large amount of irrelevant trackside scenery, introducing interference and increasing the computational burden. Furthermore, it may not completely cover the entire curved track area, leading to blind spots or omissions in threat perception within the track zone. Current technologies lack an intelligent perception mechanism that can closely integrate with the track's geometric features and dynamically focus on the potential threat space. Summary of the Invention
[0004] The purpose of this invention is to provide a method and system for detecting train obstacles based on a quad-sensor, so as to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides a method for detecting train obstacles based on a four-eye sensor, the method comprising:
[0006] The system acquires stereo natural light image data synchronously collected by binocular visible light sensors and stereo thermal radiation image data synchronously collected by binocular far-infrared sensors. The stereo natural light image data records the color and texture details of the detection scene, and the stereo thermal radiation image data records the temperature distribution information of each object in the detection scene.
[0007] Using the stereo natural light image data and the stereo thermal radiation image data, cross-validation and spatial correlation processing of heterogeneous visual information are performed to construct a fused three-dimensional perception space corresponding to the forward track area. This fused three-dimensional perception space contains a multi-dimensional description of geometric structure, surface properties and thermodynamic state.
[0008] In the fused three-dimensional perception space, dynamic perception boundary generation processing under the constraints of the orbital coordinate system is performed to delineate the dynamic perception boundary of interest to the detection system;
[0009] The fused 3D perception space is cropped according to the dynamic perception boundary. The cropped perception space slices are input into a multi-attribute decision network for analysis. All potential obstacles inside the perception space slices are identified and extracted. Each potential obstacle is assigned a type label, a 3D outline, and real-time spatial coordinates.
[0010] For the type label, 3D outline, and real-time spatial coordinates, perform obstacle threat level calculation based on the runtime environment context to generate an obstacle list report containing obstacle identity, precise location, and threat level.
[0011] Preferably, the step of using the stereo natural light image data and the stereo thermal radiation image data to perform cross-validation and spatial association processing of heterogeneous visual information to construct a fused three-dimensional perception space corresponding to the forward track area includes:
[0012] Dense stereo matching operations are performed on the stereo natural light image data to generate a high-resolution natural light depth field;
[0013] The stereo thermal radiation image data is subjected to a consistency matching operation based on the thermal radiation intensity gradient to generate an infrared thermal radiation depth field.
[0014] The natural light depth field and the infrared thermal radiation depth field are projected onto the same spatial coordinate system to form a natural light three-dimensional point cloud and a thermal radiation three-dimensional point cloud.
[0015] In the spatial coordinate system, a point-by-point physical attribute attachment process is performed on the natural light three-dimensional point cloud and the thermal radiation three-dimensional point cloud. Specifically, for each three-dimensional point in the spatial coordinate system, the nearest corresponding point in the natural light three-dimensional point cloud is found, and the color and brightness attributes of the point are obtained. At the same time, the nearest corresponding point in the thermal radiation three-dimensional point cloud is found, and the thermal radiation intensity attribute of the point is obtained. The two attributes are merged and bound to the three-dimensional point.
[0016] By integrating all the 3D points that have completed the physical property attachment process, the fused 3D perception space is generated.
[0017] Preferably, the step of performing point-by-point physical attribute attachment processing on the natural light 3D point cloud and the thermal radiation 3D point cloud in the spatial coordinate system includes:
[0018] A spatial three-dimensional voxel mesh is established, and the points in the natural light three-dimensional point cloud and the thermal radiation three-dimensional point cloud are assigned to the corresponding three-dimensional voxels.
[0019] For each three-dimensional voxel that is assigned at least one natural light three-dimensional point and at least one thermal radiation three-dimensional point, calculate the average color attribute and average brightness attribute of the natural light three-dimensional point within the three-dimensional voxel, and at the same time calculate the average thermal radiation intensity attribute of the thermal radiation three-dimensional point within the three-dimensional voxel.
[0020] The calculated average color attribute, average brightness attribute, and average thermal radiation intensity attribute are collectively assigned to the geometric center point of the three-dimensional voxel.
[0021] The set of all geometric center points, which are assigned average color, average brightness, and average thermal radiation intensity attributes, is used as the basic data for the simplified fused three-dimensional perception space.
[0022] Preferably, the dynamic sensing boundary generation process under the constraints of the orbital coordinate system includes:
[0023] The precise track geometry parameters of the current track are obtained from the train control system. These track geometry parameters include the track centerline equation, track gauge, and superelevation information.
[0024] With the center of the train head as the origin, a forward fan-shaped scanning area is defined along the direction of the track centerline. The lateral boundary of the fan-shaped scanning area is obtained by extending a dynamic safety margin from the track centerline to both sides. The dynamic safety margin is dynamically adjusted according to the current speed of the train and the curvature of the track.
[0025] The three-dimensional spatial model of the sector scanning area is mapped to the coordinate system of the fused three-dimensional perception space to generate a three-dimensional sector volume space;
[0026] Calculate the inclusion relationship of all three-dimensional points in the fused three-dimensional perception space relative to the sector volume space of the three-dimensional space, retain all three-dimensional points located inside and on the surface of the sector volume space of the three-dimensional space and their multidimensional descriptions, and form the perception space after boundary clipping.
[0027] The smallest circumscribed cuboid of the perception space after the boundary trimming is used as the dynamic perception boundary.
[0028] Preferably, the step of cropping the fused 3D perception space based on the dynamic perception boundary and inputting the cropped perception space slices into a multi-attribute decision network for analysis includes:
[0029] According to the preset longitudinal depth interval, the dynamic sensing boundary is divided into several continuous sensing space slices in the track extension direction.
[0030] For each of the aforementioned perceptual space slices, a multidimensional description of the geometric structure, surface properties, and thermodynamic state of all three-dimensional points within it is extracted and combined to form the feature tensor of the perceptual space slice.
[0031] The feature tensor corresponding to each perceptual space slice is sequentially input into the spatial segmentation module of the multi-attribute decision network. The spatial segmentation module outputs several three-dimensional subspace regions in the perceptual space slice that may contain obstacles.
[0032] For each identified three-dimensional subspace region, the attribute classification module of the multi-attribute decision network is invoked. Based on the multi-dimensional description of all three-dimensional points in the three-dimensional subspace region, the attribute classification module analyzes its spatial morphology, material and thermal characteristics, and outputs the label of the most likely obstacle type in the region, a three-dimensional outline describing its shape and size, and its real-time spatial coordinates relative to the train.
[0033] Preferably, the step of invoking the attribute classification module of the multi-attribute decision network for each identified three-dimensional subspace region includes:
[0034] The three-dimensional points within the three-dimensional subspace region are divided into high-temperature point clusters and low-temperature point clusters according to their thermodynamic state, i.e., thermal radiation intensity properties.
[0035] Spatial distribution and geometric morphology analysis were performed on the high-temperature point clusters and low-temperature point clusters respectively to determine whether the high-temperature point clusters exhibited typical biological heating characteristics and whether the low-temperature point clusters exhibited typical inorganic geometric structures.
[0036] Combining the surface properties of the three-dimensional points within the three-dimensional subspace region, namely color and texture details, the materials of the high-temperature point clusters and low-temperature point clusters are analyzed to determine whether they belong to metal, non-metal, biological epidermis or other materials.
[0037] Based on the combined spatial morphology, thermal characteristic distribution, and material analysis results, the most suitable type label is matched from a predefined obstacle type library;
[0038] Based on the standard size model corresponding to the successfully matched type label, and the actual fitting of the point cloud of the three-dimensional subspace region, an actual three-dimensional contour box that fits the three-dimensional point cloud in the region is generated, and the real-time spatial coordinates of the center of the three-dimensional contour box in the orbital coordinate system are calculated.
[0039] Preferably, the step of performing obstacle threat level calculation based on the runtime environment context to generate an obstacle list report containing obstacle identity, precise location, and threat level includes:
[0040] Based on the real-time spatial coordinates, calculate the straight-line distance between each potential obstacle and the current position of the train, as well as the lateral offset of the center of the potential obstacle relative to the centerline of the track.
[0041] Based on the train's current operating speed and braking capacity parameters, estimate the braking distance required for the train to decelerate to a stop using the maximum service braking.
[0042] Compare the straight-line distance with the braking distance, and combine the lateral offset to determine whether the potential obstacle is on the train's expected running path;
[0043] For potential obstacles on the expected running path, determine whether they have autonomous movement capability based on their type label. If they do, calculate their motion vector and trajectory by combining their real-time spatial coordinates from several historical frames.
[0044] Based on straight-line distance, braking distance, lateral offset, potential obstacle type, and calculated motion vector and trajectory, a quantitative threat level value is calculated.
[0045] The obstacle list report is generated by integrating the type labels, real-time spatial coordinates, and threat level values of all potential obstacles.
[0046] Preferably, the step of comprehensively calculating a quantified threat level value based on straight-line distance, braking distance, lateral offset, potential obstacle type, and calculated motion vector and trajectory includes:
[0047] A threat level calculation model is established, which takes straight distance, braking distance, lateral offset, potential obstacle type, movement speed and movement direction as input variables.
[0048] A base threat coefficient is assigned to each input variable, with the highest base threat coefficient when the straight-line distance is less than the braking distance and the lowest base threat coefficient when the straight-line distance is much greater than the braking distance; the smaller the lateral offset, the higher the base threat coefficient.
[0049] Type correction coefficients are set for different types of potential obstacles, with different type correction coefficients assigned to stationary large obstacles, moving vehicles, and living organisms.
[0050] For potential obstacles with autonomous movement capabilities, a motion correction coefficient is set based on the angle between their motion vector and the train's direction of travel. The correction coefficient for motion directions toward the track is higher than the correction coefficient for motion directions away from the track.
[0051] The final quantitative threat level value is obtained by weighting and fusing the basic threat coefficient of straight-line distance, basic threat coefficient of lateral offset, type correction coefficient and motion correction coefficient.
[0052] Preferably, the method further includes:
[0053] Continuously monitor potential obstacles in the obstacle list report whose threat level exceeds a predetermined threshold;
[0054] Extract the real-time spatial coordinates of the potential obstacle in multiple consecutive detection cycles to form a discrete motion trajectory point sequence of the potential obstacle;
[0055] By performing smoothing filtering and curve fitting on the discrete motion trajectory point sequence, the possible location area of the potential obstacle in the next few detection cycles is predicted;
[0056] Spatial interference checks are performed between the predicted possible location area and the train's expected running path at the same future point in time;
[0057] If the spatial interference inspection results indicate a collision risk, the threat level of the potential obstacle in the obstacle inventory report will be recalculated and increased.
[0058] Preferably, the present invention also includes a train obstacle detection system based on a four-eye sensor, the system including a processor and a memory, the memory and the processor being connected, the memory being used to store programs, instructions or code, and the processor being used to execute the programs, instructions or code in the memory to implement the train obstacle detection method based on a four-eye sensor as described above.
[0059] Compared with the prior art, the beneficial effects of the present invention are:
[0060] Stereo image pairs acquired simultaneously by binocular visible light and binocular far-infrared sensors are used to independently generate 3D point cloud data with depth information. Cross-validation and spatial correlation processing of heterogeneous visual information are performed, essentially registering and fusing the precise geometric coordinates and texture information of the visible light point cloud with the temperature attributes of the far-infrared point cloud at the same location in 3D space at the pixel level. This constructs a unified 3D spatial model, where each effective spatial unit simultaneously contains geometric location, surface features, and thermal radiation value. This model allows obstacle recognition to rely not only on shape and texture but also on its thermodynamic state. Even when visible light information degrades due to extreme lighting or severe weather, the 3D spatial structure driven by thermal radiation information can still maintain stable perception; in the detection of room-temperature objects that are not sensitive to temperature, high-resolution visible light 3D contours and textures provide the dominant judgment basis. The complementarity and verification of two heterogeneous information sources at the three-dimensional level not only improves the robustness and accuracy of perceiving the three-dimensional shape and attributes of various obstacles in complex and ever-changing environments, but also avoids the hardware cost pressure of LiDAR in traditional vision-LiDAR combined solutions, as well as its limitation of significant range attenuation in adverse weather conditions such as rain, snow, and dense fog. Compared to the simple information superposition of vision and LiDAR in traditional solutions, the pixel-level fusion of this technology fully explores the complementary value of different sensor data. For low-reflectivity obstacles, accurate identification can be achieved through the collaborative analysis of their thermal radiation characteristics and texture details. For small-sized obstacles, the high-resolution depth field generated by dense stereo matching and the detail preservation mechanism after voxelization effectively make up for the recognition shortcomings caused by insufficient information fusion in traditional solutions.
[0061] In the fused 3D perception space, based on the precise parameters of the train's current pose and the track line, a dynamically changing perception boundary is calculated in real time in the track coordinate system. This boundary is a 3D volume region that strictly conforms to the spatial orientation of the track area ahead, and its shape and orientation can dynamically adjust with curves and slopes. Based on this boundary, the global 3D perception space is cropped, eliminating a large amount of non-threatening scene data outside the track area, retaining only the "perception space slice" within the track clearance for subsequent analysis. This dynamic spatial focusing mechanism of track constraints precisely allocates limited computing resources to the space range where collision threats are most likely to occur. It eliminates interference from fixed or moving backgrounds such as trackside buildings, vegetation, and moving vehicles, reducing the complexity and false alarm rate of subsequent recognition algorithms, while ensuring full coverage of hazards in the track area by the detection system, fundamentally strengthening the targeted nature of perception and its direct correlation with train safety. Attached Figure Description
[0062] Figure 1 This is a schematic diagram illustrating the working principle of the train obstacle detection method based on a quad-eye sensor described in this invention.
[0063] Figure 2 A flowchart for constructing a fused 3D perceived space;
[0064] Figure 3 A flowchart for generating dynamically perceived boundaries;
[0065] Figure 4 A thermal radiation characteristic analysis diagram of a sensory space slice in train obstacle detection;
[0066] Figure 5 A scatter plot of temperature distribution in the track area ahead of the train, integrated with a three-dimensional perception space. Detailed Implementation
[0067] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0068] Please see Figure 1This invention provides a method for detecting train obstacles based on a quad-camera sensor. The method includes: simultaneously acquiring stereo natural light image data and stereo thermal radiation image data of the track area ahead using a binocular visible light sensor consisting of two visible light cameras and a binocular far-infrared sensor consisting of two far-infrared cameras. The stereo natural light image data records rich color and texture details of the detected scene, while the stereo thermal radiation image data records the temperature distribution information of various objects in the scene. Using these two sets of heterogeneous stereo image data, cross-validation and spatial correlation processing are performed to spatially align and fuse information from different physical domains, constructing a unified, multi-dimensional description of the track area ahead, including geometric structure, surface properties, and thermodynamic state. Based on this, combined with the current train operating state and track geometry parameters, dynamic perception boundary generation processing under track coordinate system constraints is performed to delineate a perception area boundary that dynamically adjusts with train speed and track conditions, and which the system needs to focus on. The vast fused three-dimensional perception space is then cropped based on this dynamic perception boundary, and the resulting slices are input into a pre-trained multi-attribute decision network for analysis. This network can identify and extract all potential obstacles within a perceived space slice, assigning each obstacle a type label, a 3D bounding box describing its occupied space, and its real-time spatial coordinates in the track coordinate system. Finally, based on this identified obstacle information, it performs obstacle threat level calculation processing based on the operating environment context, comprehensively considering factors such as obstacle type, location, motion state, and the train's own braking performance, generating an obstacle list report containing obstacle identity, precise location, and quantified threat level, providing information input for train safety assurance decisions.
[0069] Example 1: See Figure 2Dense stereo matching operations are performed on stereo natural light image data synchronously acquired by binocular visible light sensors to calculate the disparity of each pixel between the two natural light images, thereby generating a high-resolution natural light depth field. Consistency matching operations based on thermal radiation intensity gradients are performed on stereo thermal radiation image data synchronously acquired by binocular far-infrared sensors. Corresponding points are searched using the unique thermal radiation distribution characteristics of infrared images to generate an infrared thermal radiation depth field. Using known binocular sensor calibration parameters, the natural light depth field and the infrared thermal radiation depth field are back-projected and transformed into the same spatial coordinate system centered on the sensor group, forming a natural light 3D point cloud and a thermal radiation 3D point cloud. In the spatial coordinate system, point-by-point physical attribute attachment processing is performed on the natural light 3D point cloud and the thermal radiation 3D point cloud. A 3D voxel mesh covering the foreground space is established, and all spatial points in the natural light 3D point cloud and the thermal radiation 3D point cloud are assigned to their corresponding 3D voxels according to their coordinates. For each 3D voxel containing at least one natural light 3D point and at least one thermal radiation 3D point, calculate the average color attribute and average brightness attribute of all natural light 3D points within the voxel, and simultaneously calculate the average thermal radiation intensity attribute of all thermal radiation 3D points within the voxel. Assign the calculated average color attribute, average brightness attribute, and average thermal radiation intensity attribute to the geometric center point of the 3D voxel. Iterate through all 3D voxels that meet the conditions, and integrate the set of all geometric center points assigned average color attribute, average brightness attribute, and average thermal radiation intensity attribute as the basic data for the simplified fused 3D perception space. Each point in this data set carries information on geometric position, visible light attribute, and thermal radiation attribute.
[0070] In specific implementation, the method involves a process of constructing a fused three-dimensional perception space. Stereo natural light image data comes from binocular visible light sensors positioned on both sides of the train's front, while stereo thermal radiation image data comes from binocular far-infrared sensors positioned along the same baseline. The two sets of sensors achieve microsecond-level synchronous acquisition via hardware triggering. Taking the track area in front of the train as the detection scene, a dense stereo matching operation based on semi-global matching is performed on the two synchronously acquired natural light images. The disparity value is calculated for each pixel in the image, and the disparity is converted into depth based on the binocular camera's geometric model, generating a natural light depth field with the same resolution as the input image. Each pixel value in the natural light depth field represents the longitudinal distance of a scene point in the camera coordinate system. A consistency matching operation based on thermal radiation intensity gradient is performed on the two synchronously acquired far-infrared images. Pixels with similar thermal radiation intensity gradient characteristics are searched in the infrared image pair to establish a correspondence and calculate the disparity, thereby generating an infrared thermal radiation depth field.
[0071] In some embodiments, dense stereo matching is implemented using a convolutional neural network model trained on large-scale scene data, which directly outputs pixel-level disparity maps. Specifically, the convolutional neural network model used in dense stereo matching is trained using a large number of acquired stereo images of orbital scenes. This training data covers various lighting conditions, weather conditions, and orbital environments to ensure the model's generalization ability. The model architecture typically includes a feature extraction module, a cost volume construction module, and a disparity regression module. The feature extraction module uses convolutional layers to extract multi-level features from the left and right natural light images. The cost volume construction module forms a three-dimensional cost volume by calculating the matching cost between features. The disparity regression module decodes the pixel-level continuous disparity map from the cost volume using a softmax operation or a regression network. During training, a supervised learning approach is used, employing real disparity maps as labels. The network weights are adjusted by optimizing loss functions such as smoothed L1 loss or cross-entropy loss, enabling the model to directly output accurate disparity estimates end-to-end from the input images. In some embodiments, the consistency matching operation based on thermal radiation intensity gradient completes the matching cost calculation by comparing the distribution characteristics of thermal radiation intensity values within a local window centered on the candidate point.
[0072] In practical implementation, using the intrinsic parameter matrices, distortion coefficients, and mutual rotation and translation matrices of the binocular visible light sensor and binocular far-infrared sensor obtained through joint calibration beforehand, the coordinates and depth values of each pixel in the natural light depth field are projected onto a unified spatial coordinate system with the sensor group installation position as the origin, forming a set of a large number of three-dimensional spatial points, namely the natural light three-dimensional point cloud. Each point in the natural light three-dimensional point cloud carries the color and brightness attributes from the corresponding pixel position of the original natural light image. Similarly, the coordinates and depth values of each pixel in the infrared thermal radiation depth field are projected onto the same unified spatial coordinate system to form a thermal radiation three-dimensional point cloud. Each point in the thermal radiation three-dimensional point cloud carries the thermal radiation intensity attribute from the corresponding pixel position of the original far-infrared image.
[0073] In practice, point-by-point physical attribute attachment is performed on the natural light 3D point cloud and thermal radiation 3D point cloud after coordinate space alignment. This process is achieved by constructing a spatial 3D voxel mesh. In a unified spatial coordinate system, a cubic mesh covering the volume of interest in the forward orbital region is defined, with the side lengths of the mesh set to fixed values along the X, Y, and Z axes, for example, 0.1 meters. All spatial points in the natural light and thermal radiation 3D point clouds are then assigned to corresponding 3D voxels according to their coordinates.
[0074] In practice, for each 3D voxel assigned at least one natural light 3D point and at least one thermal radiation 3D point, attribute averaging is performed. The average color attribute and average brightness attribute of the natural light 3D points are calculated. The average color attribute is obtained by calculating the arithmetic mean of the red, green, and blue components of all natural light 3D points within the voxel. The average brightness attribute is obtained by calculating the arithmetic mean of the brightness values of all natural light 3D points. The average thermal radiation intensity attribute of the thermal radiation 3D points is calculated by calculating the arithmetic mean of the thermal radiation intensity values of all thermal radiation 3D points within the voxel. The formula for calculating the average thermal radiation intensity attribute is expressed as:
[0075] in: Representing the three-dimensional voxels currently being processed, Representing three-dimensional voxels The total number of three-dimensional thermal radiation points included. Representing the Thermal radiation intensity properties of a three-dimensional point. This represents the calculated average thermal radiation intensity attribute.
[0076] In practice, after the attribute averaging calculation is completed, the calculated average color attribute, average brightness attribute, and average thermal radiation intensity attribute are jointly assigned to the geometric center point of the current 3D voxel. The spatial coordinates of the geometric center point are determined by calculating the average of the coordinates of the eight vertices of the 3D voxel. By traversing all 3D voxels in space and integrating all the attribute-assigned geometric center points, this set of geometric center points carrying geometric position, average color attribute, average brightness attribute, and average thermal radiation intensity attribute constitutes the basic data of the simplified fused 3D perception space.
[0077] It is understandable that the side length parameter of a 3D voxel mesh is a configurable parameter. Smaller side lengths retain more detail but require more computation, while larger side lengths simplify the data but result in some loss of spatial resolution. It is also understandable that for 3D voxels containing only natural light 3D points or only thermal radiation 3D points, when constructing the base data for a fused 3D perceptual space, one can choose not to include them in the final set or use interpolation methods to supplement missing attributes.
[0078] Example 2: See Figure 3The system acquires precise track geometry parameters of the current operating track in real time from the train control system. These parameters include the track centerline equation, standard gauge, and superelevation information of curve segments, described mathematically. A forward-facing fan-shaped scanning area is defined with the train's head center as the origin, extending along the track centerline. The lateral boundary of this fan-shaped scanning area is obtained by extending a dynamic safety margin to both sides of the track centerline. This dynamic safety margin is not a fixed value but is dynamically adjusted according to the train's current speed and the radius of curvature of the track segment; the higher the speed or the smaller the curve radius, the larger the safety margin. The three-dimensional spatial model of the fan-shaped scanning area is mapped to the unified coordinate system of the fused three-dimensional perception space through coordinate transformation, generating a three-dimensional sector volume space. The inclusion relationship of all three-dimensional points in the fused three-dimensional perception space relative to this three-dimensional sector volume space is calculated. All three-dimensional points located inside and on the surface of this sector volume space, along with their attached multi-dimensional descriptive information, are retained, forming the perception space after boundary trimming. Based on the spatial distribution of all points in the perception space after clipping to this boundary, calculate its minimum circumscribed cuboid. The boundary of this cuboid is the final determined dynamic perception boundary, which is used to constrain the range of subsequent processing.
[0079] In practical implementation, the method involves dynamic sensing boundary generation processing under track coordinate system constraints. Precise track geometry parameters are obtained from the train control system. These parameters are stored digitally and updated in real time, including the track centerline equation defined as a three-dimensional spatial curve, the standard gauge value, and the superelevation information of the curve segment. The track centerline equation is typically represented parametrically, such as a spline curve or piecewise linear equation based on track design data. In practice, the track centerline equation is expressed as a discrete sequence of three-dimensional coordinate points within a certain distance ahead of the train. Each point contains longitudinal distance, lateral offset, and elevation information relative to the geodetic coordinate system or the track's starting reference point.
[0080] In practical implementation, a temporary local coordinate system is established with the projection point of the train's head center onto the track centerline as the origin. The positive X-axis of the temporary local coordinate system points along the tangent of the track centerline towards the train's forward direction, the Y-axis points to the left normal direction of the track, and the Z-axis is perpendicular to the track plane and points upward. Taking the forward track extension direction as the main body, a fan-shaped scanning area extending forward is defined in the horizontal plane. The central axis of symmetry of the fan-shaped scanning area coincides with the X-axis of the temporary local coordinate system, and the angle of the fan-shaped scanning area is determined based on the sensor's maximum effective field of view. The lateral boundary of the fan-shaped scanning area is obtained by extending a dynamic safety margin to the left and right sides from the track centerline. The dynamic safety margin is a width value calculated in real time based on the train's current operating speed and the track curvature. The formula for calculating the dynamic safety margin is defined as follows:
[0081]
[0082] in: Represents dynamic safety margin (unit: ), Represents the basic safety width (unit: ), Represents the current train speed obtained from the train control system (unit: ), Represents the current orbital radius of curvature obtained from the track geometry parameters (unit: ), It is a speed-related adjustment coefficient (unit: ), It is an adjustment coefficient related to curvature (unit: Basic safety width It is usually set to a fixed value slightly greater than half the standard gauge to ensure coverage of the track structure itself and the adjacent area. Adjustment factor and A pre-set positive constant provides a dynamic safety margin. With speed It increases linearly with the radius of curvature. It decreases while increasing linearly.
[0083] In some embodiments, the sector scanning area also has a range constraint in the vertical direction. Its lower boundary is set at a certain distance below the track plane to include track surface obstacles, and its upper boundary is dynamically set according to track environment information such as tunnel clearance or bridge clearance. In some embodiments, the calculation of dynamic safety margin also considers factors such as train type and weather conditions, which is achieved by introducing additional correction terms. In specific implementation, the three-dimensional spatial model of the sector scanning area defined above is mapped to the unified coordinate system of the fused three-dimensional perception space to generate a three-dimensional sector volume space. The mapping process involves coordinate transformation, converting the key feature points on the boundary of the sector scanning area from the temporary local coordinate system to the unified coordinate system. The three-dimensional sector volume space can be mathematically described as a three-dimensional volume formed by stretching a sector base and a certain longitudinal depth, and its boundary surface is defined by equations or a system of inequalities.
[0084] In practical implementation, the inclusion relationship of all 3D points in the fused 3D perception space relative to the 3D sector volume space is calculated, and it is determined whether the coordinates of each 3D point satisfy the inequality constraints of the boundary surface of the 3D sector volume space. All 3D points located inside and on the surface of the 3D sector volume space are retained, along with multi-dimensional descriptive information about the geometry, surface properties, and thermodynamic state to which these 3D points are attached. This retained set of points collectively constitutes the boundary-trimmed perception space. It can be understood that the calculation of inclusion relationships can be accelerated using spatial indexing structures, such as storing the 3D points of the fused 3D perception space in an octree structure and then performing rapid collision detection with the 3D sector volume space. It can also be understood that the boundary description method of the 3D sector volume space directly affects the calculation efficiency of inclusion relationships; using axially aligned bounding boxes for initial screening followed by precise geometric judgment is an acceptable approach.
[0085] In practice, the coordinate distribution of all 3D points in the perception space after boundary clipping is used as input to calculate its minimum bounding cuboid. The goal of calculating the minimum bounding cuboid is to find a cuboid whose faces are parallel to the axes of the temporary local coordinate system and can enclose all points. The range of the three axes of this cuboid is determined by the minimum and maximum values of the point set in the X, Y, and Z directions of the temporary local coordinate system, respectively.
[0086] Example 3: According to a preset longitudinal depth interval, the dynamic sensing boundary is divided into several continuous sensing space slices with a certain thickness along the track extension direction. For each sensing space slice, the geometric coordinates, surface properties, and thermodynamic state of all three-dimensional points within it are extracted as a multidimensional description, and this information is combined into a feature tensor corresponding to that sensing space slice. The feature tensor corresponding to each sensing space slice is sequentially input into the spatial segmentation module of the multi-attribute decision network. This module analyzes the spatial and attribute relationships in the feature tensor and outputs several three-dimensional sub-space regions containing obstacles in the sensing space slice and their spatial ranges. For each identified three-dimensional sub-space region, the attribute classification module of the multi-attribute decision network is called for analysis. This module first performs cluster analysis on the three-dimensional points within the three-dimensional sub-space region according to their thermodynamic state, i.e., thermal radiation intensity attribute, dividing them into high-temperature point clusters and low-temperature point clusters. The spatial distribution and geometric morphology of the high-temperature point clusters and low-temperature point clusters are analyzed respectively to determine whether the high-temperature point clusters exhibit typical biological heating characteristics and whether the low-temperature point clusters exhibit typical inorganic geometric structures. By combining the surface attributes (color and texture details) of 3D points within a 3D subspace region, the materials of high-temperature and low-temperature point clusters are analyzed to determine whether they belong to metal, non-metal, biological skin, or other materials. Integrating the spatial morphology analysis results, thermal characteristic distribution patterns, and material analysis results, the most suitable type label is matched from a predefined obstacle type library. Based on the standard size model corresponding to the successfully matched type label and the actual spatial distribution of the point cloud in the 3D subspace region, an actual 3D contour box that fits the 3D point cloud within that region is generated, and the real-time spatial coordinates of the center point of this 3D contour box in the orbital coordinate system are calculated.
[0087] In practice, the dynamic sensing boundary is a three-dimensional cuboid space defined in the orbital coordinate system. The dynamic sensing boundary is divided into several continuous sensing space slices along the orbital extension direction according to a preset longitudinal depth interval. Each sensing space slice has a fixed longitudinal length but covers the entire lateral and vertical range of the dynamic sensing boundary. The longitudinal depth interval is set to a fixed value, such as 2 meters, to ensure that each slice contains sufficient spatial information for analysis. For each sensing space slice, a multidimensional description of the geometric structure, surface properties, and thermodynamic state of all three-dimensional points within it is extracted. The geometric structure description includes the set of spatial coordinates of the three-dimensional points; the surface property description includes the color and brightness attributes of the three-dimensional points; and the thermodynamic state description includes the thermal radiation intensity attributes of the three-dimensional points. These multidimensional descriptions are combined into a high-dimensional feature tensor. The dimension of the feature tensor corresponds to the number of three-dimensional points within the sensing space slice and the attribute dimension of each point. The feature tensor, as a standardized data representation of the sensing space slice, is input into a multi-attribute decision network.
[0088] In specific implementations, the spatial segmentation module of the multi-attribute decision network receives feature tensors from perceptual spatial slices. Based on a convolutional neural network structure, the module encodes and decodes these feature tensors, outputting several three-dimensional sub-spatial regions containing obstacles within the perceptual spatial slices. Each three-dimensional sub-spatial region is defined by a set of spatial coordinate ranges and corresponds to a continuous point cloud cluster in the fused three-dimensional perceptual space. The spatial segmentation module identifies point cloud clusters significantly different from the background environment by analyzing the spatial distribution and attribute associations of points in the feature tensors. In some embodiments, the spatial segmentation module employs an attention-based graph neural network to model the topological relationships between three-dimensional points, thereby more accurately segmenting potential obstacle regions.
[0089] In practical implementation, the attribute classification module of the multi-attribute decision network is invoked for each identified 3D subspace region. This module analyzes the spatial morphology, material, and thermal characteristics of all 3D points within the subspace region based on their multi-dimensional descriptions. First, the module categorizes the 3D points within the subspace region into high-temperature clusters and low-temperature clusters according to their thermodynamic state, i.e., thermal radiation intensity. This clustering process is achieved by setting an adaptive threshold, calculated based on the thermal radiation intensity distribution of the entire sensing space slice. High-temperature clusters include all points with thermal radiation intensity higher than the adaptive threshold, while low-temperature clusters include all points with thermal radiation intensity lower than or equal to the adaptive threshold. Spatial distribution and geometric morphology analysis are then performed on the high-temperature and low-temperature clusters to determine whether they exhibit typical biological heating characteristics and whether they exhibit typical inorganic geometric structures. The biological heating characteristics of high-temperature clusters are evaluated by calculating the uniformity and spatial compactness of the thermal radiation intensity, while the typical inorganic geometric structure of low-temperature clusters is evaluated by calculating the principal component analysis characteristics, such as comparing the extent of extension of the clusters along different principal axes to determine whether they exhibit regular geometric shapes.
[0090] In practice, the surface attributes of 3D points within a 3D subspace region—namely, color and texture details—are combined to analyze the materials of high-temperature and low-temperature point clusters. Color attributes are used to distinguish between metallic and non-metallic materials; metallic materials typically exhibit high color saturation and reflectivity. Texture details are obtained by analyzing the normal changes or gray-level co-occurrence matrix features of local surface areas in the point cloud, used to determine whether the material belongs to biological skin, rock, asphalt, or other categories. The results of spatial morphology analysis, thermal feature distribution patterns, and material analysis are combined to match the most suitable type label from a predefined obstacle type library. This library includes various types such as pedestrians, animals, vehicles, falling rocks, and trees. Each type is associated with a set of typical feature descriptions of spatial morphology, thermal features, and materials. The matching process is achieved by calculating the similarity score between the features of the 3D subspace region and the feature descriptions of each category in the type library, with the type with the highest score being used as the output type label.
[0091] In practical implementation, based on the standard size model corresponding to the successfully matched type label and the actual fitting of the point cloud in the 3D subspace region, an actual 3D contour box that fits the 3D point cloud in that region is generated. The standard size model provides the common size range of this type of obstacle. The actual fitting adopts the minimum bounding box algorithm or a 3D reconstruction algorithm based on the convex hull of the point cloud surface. The generated 3D contour box is a cuboid with its axis aligned with the orbital coordinate system. Its center position and size are determined by the point cloud distribution. The real-time spatial coordinates of the center of the 3D contour box in the orbital coordinate system are calculated by transforming the center of the contour box from the fused 3D perception spatial coordinate system to the orbital coordinate system. Optionally, the clustering adaptive threshold calculation formula for high temperature point clusters and low temperature point clusters is defined as:
[0092]
[0093] in: Represents an adaptive threshold. This represents the average value of the thermal radiation intensity attribute of all three-dimensional points within a slice of the perception space. The standard deviation represents the thermal radiation intensity attribute. This is an adjustable coefficient used to control the degree of deviation of the threshold relative to the average value. Optionally, the attribute classification module introduces a metallicity index based on color attributes when analyzing materials. The metallicity index is quantified by calculating the component values of points in a specific color space. It is understandable that training a multi-attribute decision network requires a large amount of labeled fused 3D perception spatial data, including obstacle type labels, 3D bounding boxes, and spatial coordinates.
[0094] See Figure 4 This is a thermal radiation feature analysis map of a sensor space slice in train obstacle detection. The proportion of high-temperature point clusters (38%) and the adaptive threshold (38.2℃) both reach their peak values, indicating the presence of significant high-temperature objects in this area. In other slices, the proportion of high-temperature point clusters is below 20%, with low-temperature point clusters dominating, corresponding to inanimate or low-temperature obstacles. This map is a visualization result of the "thermal feature clustering analysis" stage in the "four-eye sensor train obstacle detection method." By correlating the proportion of thermal radiation point clusters with the adaptive threshold, high-risk high-temperature obstacle areas can be quickly located, providing a core basis for subsequent obstacle type identification and threat assessment.
[0095] Example 4: Based on the real-time spatial coordinates of each identified potential obstacle, the straight-line distance between the obstacle and the train's current location is calculated, along with the lateral offset of the obstacle's center point relative to the track centerline. Combining the current operating speed and braking capacity parameters obtained from the train control system, the braking distance required for the train to decelerate to a complete stop using maximum service braking is estimated through a dynamic model. The straight-line distance between the obstacle and the train is compared with the train's braking distance, and the obstacle's lateral offset is used to determine whether it is on the train's expected operating path. For potential obstacles on the expected operating path, their assigned type label is used to determine whether they possess autonomous movement capability. If the type label indicates autonomous movement capability, their motion vector and trajectory are calculated based on their real-time spatial coordinate sequence over several consecutive historical detection frames.
[0096] A threat level calculation model is established, taking straight-line distance, braking distance, lateral offset, potential obstacle type, speed, and direction of movement as input variables. A base threat coefficient is assigned to each input variable, with the highest coefficient when the straight-line distance is less than the braking distance and the lowest when the straight-line distance is much greater than the braking distance; the smaller the lateral offset, the higher the base threat coefficient. Type correction coefficients are set for different types of potential obstacles, with different coefficients assigned to stationary large obstacles, mobile vehicles, and living organisms. For potential obstacles with autonomous movement capabilities, a motion correction coefficient is set based on the angle between their motion vector and the train's direction of travel, with a higher correction coefficient for those moving towards the track than those moving away from the track. The base threat coefficient for straight-line distance, the base threat coefficient for lateral offset, the type correction coefficient, and the motion correction coefficient are weighted and fused to obtain the final quantitative threat level value. A structured obstacle inventory report is generated by integrating the type labels of all potential obstacles, their real-time spatial coordinates, and the calculated threat level values.
[0097] In practice, based on the real-time spatial coordinates output by the multi-attribute decision network for each potential obstacle, the straight-line distance between each potential obstacle and the train's current position is calculated. The train's current position is obtained from the fusion positioning results of the Global Navigation Satellite System and the Inertial Measurement Unit. The straight-line distance is obtained by calculating the Euclidean distance between the real-time spatial coordinates of the potential obstacle and the coordinates of the train's current position. Simultaneously, the lateral offset of the center point of the potential obstacle relative to the track centerline is calculated. The lateral offset is calculated by first determining the point on the track centerline closest to the center point of the potential obstacle, and then calculating the absolute value of the coordinate difference between the center point of the potential obstacle and this closest point in the Y-direction of the track coordinate system.
[0098] In practice, by combining real-time data from the train control system, including the train's current speed, weight, and braking system efficiency, a train braking dynamics model is used to estimate the braking distance required for the train to decelerate to a stop using maximum service braking. This model considers factors such as basic resistance, gradient resistance, and curve resistance. The calculated straight-line distance is compared with the estimated braking distance, and lateral offset is used to determine whether a potential obstacle is on the train's expected path. The logic is that if the straight-line distance is less than or equal to the braking distance and the lateral offset is less than the track intrusion threshold, the potential obstacle is considered to be on the expected path. For potential obstacles determined to be on the expected path, a multi-attribute decision network assigns them a type label to determine whether they possess autonomous movement capabilities. Type labels such as "pedestrian," "animal," and "vehicle" are categorized as having autonomous movement capabilities, while type labels such as "falling rocks," "cargo," and "trees" are categorized as not having autonomous movement capabilities.
[0099] In some embodiments, the braking distance is estimated using a lookup table method. The corresponding braking distance values are pre-calculated and stored based on different speed levels and track gradient levels. During runtime, interpolation is performed based on the current speed and gradient information. In some embodiments, the track intrusion determination threshold is determined comprehensively based on the current track gauge and dynamic safety margin.
[0100] In practical implementation, for potential obstacles classified as having autonomous movement capabilities, their motion vectors and trajectories are calculated by combining their real-time spatial coordinate sequences over several consecutive historical detection frames. The coordinate sequences are stored chronologically. The motion vector is obtained by calculating the difference between the coordinates of the two most recent frames and dividing by the frame time interval, containing velocity magnitude and direction information. The motion trajectory is obtained by polynomial fitting of the coordinate sequences. A threat level calculation model is established, with straight-line distance, braking distance, lateral offset, potential obstacle type, motion speed, and motion direction as input variables. A basic threat coefficient is assigned to each input variable. The allocation rule for the basic threat coefficient of straight-line distance is that the basic threat coefficient is highest when the straight-line distance is less than the braking distance, and lowest when the straight-line distance is much greater than the braking distance. Specifically, a piecewise linear function maps the ratio of straight-line distance to braking distance to a threat coefficient value between 0 and 1. The allocation rule for the basic threat coefficient of lateral offset is that the smaller the lateral offset, the higher the basic threat coefficient, and the mapping is achieved through a negatively correlated linear or nonlinear function.
[0101] In practice, type correction coefficients are set for different types of potential obstacles. These coefficients are multiplicative factors based on the obstacle's static risk and dynamic behavioral characteristics. See Table 1 for an example of a type correction coefficient configuration.
[0102] Table 1: Correction Coefficients for Potential Obstacle Types
[0103]
[0104] For potential obstacles with autonomous movement capabilities, a motion correction coefficient is set based on the angle between their motion vector and the train's direction of travel. When the angle between the motion direction and the train's direction of travel is less than 90 degrees (i.e., towards the track), the motion correction coefficient is higher than when the angle is greater than 90 degrees (i.e., away from the track). The motion correction coefficient also takes into account the magnitude of the motion speed; the higher the speed, the higher the correction coefficient.
[0105] In practice, the final quantitative threat level value is obtained by weighted fusion calculation of the basic threat coefficient of straight-line distance, basic threat coefficient of lateral offset, type correction coefficient, and motion correction coefficient. The weighted fusion calculation adopts a linear weighted summation formula. The formula for calculating the quantitative threat level value is defined as follows:
[0106] in: This represents the calculated quantitative threat level value. Represents the basic threat coefficient based on straight-line distance. Braking distance The dimensionless coefficient obtained by the ratio mapping has a value range of 0-1; when hour, =1.0 indicates that the obstacle is within the braking range and poses the highest threat; when hour, Linearly decreasing; when hour, This indicates that the obstacle is outside the safe zone and poses no direct threat. Represents the lateral offset base threat coefficient, based on straight-line distance. Braking distance The dimensionless coefficient obtained by the ratio mapping has a value range of 0-1; when hour, , indicates that the obstacle is within the braking range and poses the highest threat; when hour, Linearly decreasing; when hour, This indicates that the obstacle is outside the safe zone and poses no direct threat. Representative type correction factor, Represents the motion correction factor. and These are the weighting factors for the basic threat coefficient of straight-line distance and the basic threat coefficient of lateral offset, respectively. These weighting factors are preset positive constants used to adjust the importance of distance threat and offset threat. Both are dimensionless coefficients, ranging from 0 to 1, and satisfy the following conditions: Value selection rules: Default settings , The impact of distance threats on overall risk is weighted higher than that of lateral offset threats, which aligns with the safety logic of prioritizing distance in train braking decisions. A structured obstacle list report is generated by integrating the type labels, real-time spatial coordinates, and calculated quantified threat level values of all potential obstacles, according to a predefined report format. The report is presented in list form, with each record containing a unique obstacle identifier, type label, 3D outline coordinates, real-time spatial coordinates of the center point, and a quantified threat level value.
[0107] Example 5: Continuously monitor potential obstacles in the obstacle inventory report whose threat level exceeds a predetermined threshold. Extract the real-time spatial coordinates of the potential obstacle over multiple consecutive detection cycles to form a discrete motion trajectory point sequence. Perform smoothing filtering and curve fitting on this discrete motion trajectory point sequence to predict the location area of the potential obstacle in the next few detection cycles. Perform a spatial interference check between this predicted location area and the train's expected running path at the same future time point, calculated based on the current speed and track. If the spatial interference check indicates a collision risk, recalculate and increase the threat level value of the potential obstacle in the obstacle inventory report.
[0108] In practice, the obstacle list report continuously monitors potential obstacles whose quantified threat level exceeds a predetermined threshold. This threshold is a constant pre-set based on the train operation safety level, used to filter out objects requiring focused tracking and risk assessment. The real-time spatial coordinates of potential obstacles meeting the criteria are extracted over multiple consecutive detection cycles, forming a discrete motion trajectory point sequence for each obstacle. Each point in the discrete motion trajectory point sequence contains a timestamp and three-dimensional coordinate information, with the timestamp corresponding to the acquisition time of the detection cycle.
[0109] In practice, the discrete motion trajectory point sequence undergoes smoothing filtering and curve fitting. Smoothing filtering employs a moving average filter or a Kalman filter to reduce the impact of measurement noise on the trajectory point coordinates. Curve fitting uses the least squares method to fit the time-space coordinate relationship into a polynomial function or spline curve. Based on the fitted curve model, the location region of potential obstacles is predicted over several future detection cycles. The prediction time span is set according to the train braking distance and current speed. The predicted location region is represented by a three-dimensional spatial range, which takes into account the prediction error of the fitted curve. This range is determined by calculating the covariance matrix of the predicted coordinates and taking a certain confidence interval.
[0110] In practice, a spatial interference check is performed between the predicted location region and the train's expected running path at the same future time point. The expected running path is calculated based on the train's current speed, track geometry parameters, and train length. The expected running path is represented as a three-dimensional tubular spatial region with a certain width and height, centered on the track centerline. The spatial interference check is performed by calculating the geometric intersection between the three-dimensional bounding box of the predicted location region and the three-dimensional tubular structure of the expected running path. If the two three-dimensional volumes intersect, a collision risk is identified. In some embodiments, the smoothing filtering process employs a Kalman filter based on a uniform motion model or a uniformly accelerated motion model, with the filter state variables including position and velocity.
[0111] In practice, if the spatial interferometry inspection results indicate a collision risk, the quantitative threat level of the potential obstacle in the obstacle inventory report is recalculated and increased. The recalculation process calls the threat level calculation model and updates the input variables, using the predicted future position coordinates and motion state as new inputs. Based on the updated input variables, the basic threat coefficients for straight-line distance and lateral offset are recalculated, and the motion correction coefficients are adjusted according to the predicted motion trend, ultimately resulting in a higher quantitative threat level. The increase in the quantitative threat level is achieved by adding a risk increment based on the predicted collision time or distance; the risk increment is inversely proportional to the predicted collision time. The logical formula for judging collision risk in spatial interferometry is defined as follows:
[0112]
[0113] in: This represents a collision risk indicator, with a value of 1 indicating a collision risk and a value of 0 indicating no collision risk. The three-dimensional volume representing the predicted location region. The three-dimensional volume representing the train's expected route. This represents the geometric intersection operation between two volumes.
[0114] See Figure 5 This is a scatter plot of temperature distribution in the track area ahead of the train, fused with a 3D perception space. Low-temperature points (blue / cyan) predominate, corresponding to the track area background; high-temperature points (yellow / red) are concentrated in areas with a longitudinal distance of 50-200m, a lateral distance of -5 to 5m, and a height of 0-4m, consistent with the spatial distribution characteristics of high-temperature obstacles such as pedestrians / animals. This plot is a visualization result of the "fusion of 3D perception space construction" stage in the "quad-eye sensor train obstacle detection method." Through a multi-dimensional display of 3D space and temperature, it intuitively presents the thermal distribution characteristics of the track area, providing basic perception data for subsequent obstacle identification and threat assessment.
[0115] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for detecting train obstacles based on a quad-sensor, characterized in that, The method includes: The system acquires stereo natural light image data synchronously collected by binocular visible light sensors and stereo thermal radiation image data synchronously collected by binocular far-infrared sensors. The stereo natural light image data records the color and texture details of the detection scene, and the stereo thermal radiation image data records the temperature distribution information of each object in the detection scene. Using the stereo natural light image data and the stereo thermal radiation image data, cross-validation and spatial correlation processing of heterogeneous visual information are performed to construct a fused three-dimensional perception space corresponding to the forward track area. This fused three-dimensional perception space contains a multi-dimensional description of geometric structure, surface properties and thermodynamic state. In the fused three-dimensional perception space, dynamic perception boundary generation processing under the constraints of the orbital coordinate system is performed to delineate the dynamic perception boundary of interest to the detection system; The fused 3D perception space is cropped according to the dynamic perception boundary. The cropped perception space slices are input into a multi-attribute decision network for analysis. All potential obstacles inside the perception space slices are identified and extracted. Each potential obstacle is assigned a type label, a 3D outline, and real-time spatial coordinates. For the type label, 3D outline, and real-time spatial coordinates, perform obstacle threat level calculation based on the runtime environment context to generate an obstacle list report containing obstacle identity, precise location, and threat level.
2. The method for detecting train obstacles based on a quad-eye sensor according to claim 1, characterized in that, The process of utilizing the stereo natural light image data and the stereo thermal radiation image data to perform cross-validation and spatial association processing of heterogeneous visual information, and constructing a fused three-dimensional perception space corresponding to the forward track area, includes: Dense stereo matching operations are performed on the stereo natural light image data to generate a high-resolution natural light depth field; The stereo thermal radiation image data is subjected to a consistency matching operation based on the thermal radiation intensity gradient to generate an infrared thermal radiation depth field. The natural light depth field and the infrared thermal radiation depth field are projected onto the same spatial coordinate system to form a natural light three-dimensional point cloud and a thermal radiation three-dimensional point cloud. In the spatial coordinate system, a point-by-point physical attribute attachment process is performed on the natural light three-dimensional point cloud and the thermal radiation three-dimensional point cloud. Specifically, for each three-dimensional point in the spatial coordinate system, the nearest corresponding point in the natural light three-dimensional point cloud is found, and the color and brightness attributes of the point are obtained. At the same time, the nearest corresponding point in the thermal radiation three-dimensional point cloud is found, and the thermal radiation intensity attribute of the point is obtained. The two attributes are merged and bound to the three-dimensional point. By integrating all the 3D points that have completed the physical property attachment process, the fused 3D perception space is generated.
3. The method for detecting train obstacles based on a quad-eye sensor according to claim 2, characterized in that, The step of performing point-by-point physical attribute attachment processing on the natural light 3D point cloud and the thermal radiation 3D point cloud in the spatial coordinate system includes: A spatial three-dimensional voxel mesh is established, and the points in the natural light three-dimensional point cloud and the thermal radiation three-dimensional point cloud are assigned to the corresponding three-dimensional voxels. For each three-dimensional voxel that is assigned at least one natural light three-dimensional point and at least one thermal radiation three-dimensional point, calculate the average color attribute and average brightness attribute of the natural light three-dimensional point within the three-dimensional voxel, and at the same time calculate the average thermal radiation intensity attribute of the thermal radiation three-dimensional point within the three-dimensional voxel. The calculated average color attribute, average brightness attribute, and average thermal radiation intensity attribute are collectively assigned to the geometric center point of the three-dimensional voxel. The set of all geometric center points, which are assigned average color, average brightness, and average thermal radiation intensity attributes, is used as the basic data for the simplified fused three-dimensional perception space.
4. The method for detecting train obstacles based on a quad-eye sensor according to claim 1, characterized in that, The dynamic sensing boundary generation process under the constraints of the orbital coordinate system includes: The precise track geometry parameters of the current track are obtained from the train control system. These track geometry parameters include the track centerline equation, track gauge, and superelevation information. With the center of the train head as the origin, a forward fan-shaped scanning area is defined along the direction of the track centerline. The lateral boundary of the fan-shaped scanning area is obtained by extending a dynamic safety margin from the track centerline to both sides. The dynamic safety margin is dynamically adjusted according to the current speed of the train and the curvature of the track. The three-dimensional spatial model of the sector scanning area is mapped to the coordinate system of the fused three-dimensional perception space to generate a three-dimensional sector volume space; Calculate the inclusion relationship of all three-dimensional points in the fused three-dimensional perception space relative to the sector volume space of the three-dimensional space, retain all three-dimensional points located inside and on the surface of the sector volume space of the three-dimensional space and their multidimensional descriptions, and form the perception space after boundary clipping. The smallest circumscribed cuboid of the perception space after the boundary trimming is used as the dynamic perception boundary.
5. The method for detecting train obstacles based on a quad-eye sensor according to claim 4, characterized in that, The process of cropping the fused 3D perception space based on the dynamic perception boundary, and then inputting the cropped perception space slices into a multi-attribute decision network for analysis, includes: According to the preset longitudinal depth interval, the dynamic sensing boundary is divided into several continuous sensing space slices in the track extension direction. For each of the aforementioned perceptual space slices, a multidimensional description of the geometric structure, surface properties, and thermodynamic state of all three-dimensional points within it is extracted and combined to form the feature tensor of the perceptual space slice. The feature tensor corresponding to each perceptual space slice is sequentially input into the spatial segmentation module of the multi-attribute decision network. The spatial segmentation module outputs several three-dimensional subspace regions in the perceptual space slice that may contain obstacles. For each identified three-dimensional subspace region, the attribute classification module of the multi-attribute decision network is invoked. Based on the multi-dimensional description of all three-dimensional points in the three-dimensional subspace region, the attribute classification module analyzes its spatial morphology, material and thermal characteristics, and outputs the label of the most likely obstacle type in the region, a three-dimensional outline describing its shape and size, and its real-time spatial coordinates relative to the train.
6. The method for detecting train obstacles based on a quad-eye sensor according to claim 5, characterized in that, For each identified 3D subspace region, the attribute classification module of the multi-attribute decision network is invoked, including: The three-dimensional points within the three-dimensional subspace region are divided into high-temperature point clusters and low-temperature point clusters according to their thermodynamic state, i.e., thermal radiation intensity properties. Spatial distribution and geometric morphology analysis were performed on the high-temperature point clusters and low-temperature point clusters respectively to determine whether the high-temperature point clusters exhibited typical biological heating characteristics and whether the low-temperature point clusters exhibited typical inorganic geometric structures. Combining the surface properties of the three-dimensional points within the three-dimensional subspace region, namely color and texture details, the materials of the high-temperature point clusters and low-temperature point clusters are analyzed to determine whether they belong to metal, non-metal, biological epidermis or other materials. Based on the combined spatial morphology, thermal characteristic distribution, and material analysis results, the most suitable type label is matched from a predefined obstacle type library; Based on the standard size model corresponding to the successfully matched type label, and the actual fitting of the point cloud of the three-dimensional subspace region, an actual three-dimensional contour box that fits the three-dimensional point cloud in the region is generated, and the real-time spatial coordinates of the center of the three-dimensional contour box in the orbital coordinate system are calculated.
7. The method for detecting train obstacles based on a quad-eye sensor according to claim 1, characterized in that, The process of performing obstacle threat level calculation based on the runtime environment context generates an obstacle list report containing obstacle identity, precise location, and threat level, including: Based on the real-time spatial coordinates, calculate the straight-line distance between each potential obstacle and the current position of the train, as well as the lateral offset of the center of the potential obstacle relative to the centerline of the track. Based on the train's current operating speed and braking capacity parameters, estimate the braking distance required for the train to decelerate to a stop using the maximum service braking. Compare the straight-line distance with the braking distance, and combine the lateral offset to determine whether the potential obstacle is on the train's expected running path; For potential obstacles on the expected running path, determine whether they have autonomous movement capability based on their type label. If they do, calculate their motion vector and trajectory by combining their real-time spatial coordinates from several historical frames. Based on straight-line distance, braking distance, lateral offset, potential obstacle type, and calculated motion vector and trajectory, a quantitative threat level value is calculated. The obstacle list report is generated by integrating the type labels, real-time spatial coordinates, and threat level values of all potential obstacles.
8. The method for detecting train obstacles based on a quad-eye sensor according to claim 7, characterized in that, The threat level is calculated by comprehensively considering factors such as straight-line distance, braking distance, lateral offset, potential obstacle type, and the calculated motion vector and trajectory. This includes: A threat level calculation model is established, which takes straight distance, braking distance, lateral offset, potential obstacle type, movement speed and movement direction as input variables. A base threat coefficient is assigned to each input variable, with the highest base threat coefficient when the straight-line distance is less than the braking distance and the lowest base threat coefficient when the straight-line distance is much greater than the braking distance; the smaller the lateral offset, the higher the base threat coefficient. Type correction coefficients are set for different types of potential obstacles, with different type correction coefficients assigned to stationary large obstacles, moving vehicles, and living organisms. For potential obstacles with autonomous movement capabilities, a motion correction coefficient is set based on the angle between its motion vector and the train's direction of travel. The correction coefficient for motion directions toward the track is higher than the correction coefficient for motion directions away from the track. The final quantitative threat level value is obtained by weighting and fusing the basic threat coefficient of straight-line distance, basic threat coefficient of lateral offset, type correction coefficient and motion correction coefficient.
9. The method for detecting train obstacles based on a quad-eye sensor according to claim 7, characterized in that, The method further includes: Continuously monitor potential obstacles in the obstacle list report whose threat level exceeds a predetermined threshold; Extract the real-time spatial coordinates of the potential obstacle in multiple consecutive detection cycles to form a discrete motion trajectory point sequence of the potential obstacle; By performing smoothing filtering and curve fitting on the discrete motion trajectory point sequence, the possible location area of the potential obstacle in the next few detection cycles is predicted; Spatial interference checks are performed between the predicted possible location area and the train's expected running path at the same future point in time; If the spatial interference inspection results indicate a collision risk, the threat level of the potential obstacle in the obstacle inventory report will be recalculated and increased.
10. A train obstacle detection system based on a quad-sensor, characterized in that, The device includes a processor and a memory, the memory being connected to the processor. The memory is used to store programs, instructions, or code, and the processor is used to execute the programs, instructions, or code in the memory to implement the train obstacle detection method based on a quad-eye sensor as described in any one of claims 1-9.
Citation Information
Patent Citations
Obstacle detection method and obstacle detection system for track
CN120288091A
Depth data set construction method based on visual scene of unmanned aerial vehicle
CN120808062A