Vehicle-mounted height-limiting anti-collision identification alarm system
By using multimodal sensor data fusion technology, the system can accurately identify obstacles and dynamically assess their height, thus solving the problems of insufficient perception and poor environmental adaptability of existing vehicle height restriction and collision avoidance systems, and providing accurate and reliable collision warnings.
Patent Information
- Application Number
- CN202511891994.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-16
- Publication Date
- 2026-03-06
AI Technical Summary
Existing vehicle height restriction and collision avoidance systems rely on a single sensor, which has limited perception capabilities, is susceptible to environmental interference, and cannot accurately identify obstacle types and dynamic changes in vehicle height, resulting in insufficient accuracy and reliability of warnings.
Employing multimodal sensor data fusion technology, including lidar, millimeter-wave radar, image sensors, and vehicle height sensors, and through synchronous data frame acquisition, obstacle perception module, vehicle environment determination module, and collision risk analysis module, it achieves accurate obstacle identification and dynamic height assessment, generating reliable graded alarm signals.
It improves the accuracy and all-weather adaptability of obstacle perception, provides accurate and reliable collision warnings, overcomes the warning failure problem caused by incomplete information and fixed parameters, and provides safety assurance for drivers.
Smart Images

Figure CN121617280A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicle collision avoidance alarms, and more specifically, to a vehicle-mounted height restriction collision avoidance recognition alarm system. Background Technology
[0002] With the rapid development of modern logistics and transportation, road networks are becoming increasingly complex, and the demand for large passenger, freight, and special vehicles to travel on urban roads, bridges, and tunnels is growing. However, because these vehicles are typically tall, collisions with height restriction facilities (such as height restriction bars and bridge underpasses) occur frequently due to drivers misjudging height restriction information or failing to see clearly in low visibility conditions such as inclement weather or at night. Such accidents not only cause serious vehicle damage and property loss but can also lead to traffic congestion and even threaten public safety. Therefore, developing a system that can proactively identify and warn of height restriction risks is of paramount importance.
[0003] Currently, preventing such height-restriction collisions mainly relies on the driver's visual observation and experience-based judgment. However, this method is highly susceptible to driver fatigue, distraction, and adverse weather conditions such as rain, snow, and fog, resulting in low reliability. Existing vehicle assistance systems, such as those using a single ultrasonic or laser ranging sensor, can detect obstacles to some extent, but their technical solutions have significant flaws. First, a single sensor has limited perception capabilities in complex environments and is easily affected by environmental interference, leading to inaccurate ranging, missed detections, or false alarms. Second, these systems typically only perform simple distance detection and cannot effectively identify the type of obstacle. For example, they cannot distinguish between genuine height-restriction obstacles and non-dangerous hanging objects such as tree branches, nor can they intelligently identify the specific numerical information on height restriction signs. Furthermore, most existing solutions ignore the crucial factor that vehicle height dynamically changes due to load, tire pressure, or road surface disturbances, using fixed vehicle height parameters for calculations, which further reduces the accuracy and reliability of warnings.
[0004] Therefore, existing technologies are significantly insufficient in terms of the accuracy of perception, environmental adaptability, and comprehensiveness of information, making it difficult to meet the collision avoidance and warning needs of large vehicles in complex and ever-changing road conditions. Summary of the Invention
[0005] To address the aforementioned technical problems, this application is proposed. A vehicle-mounted height restriction and collision avoidance alarm system according to this application includes: The sensor data frame acquisition module is used to acquire synchronous sensor data frames, which include raw lidar point cloud data, raw millimeter-wave radar target list, synchronous image pairs, and raw vehicle height values. The obstacle perception module is used to extract and identify features of obstacles ahead based on the original lidar point cloud data, synchronous image pairs, original millimeter-wave radar target list and original vehicle height value in the synchronous sensor data frame, so as to obtain the list of perceived obstacles and the real-time accurate height of the vehicle. The vehicle environment determination module is used to comprehensively determine the vehicle and environmental states from the perceived obstacle list to obtain an obstacle list that incorporates the context. The collision risk analysis module is used to assess the collision risk level of each obstacle in the obstacle list based on the vehicle's real-time accurate height and in conjunction with the context, so as to obtain a collision risk profile. The alarm drive module is used to input the collision risk profile and the obstacle list combined with the context into the hierarchical alarm engine to obtain the alarm drive signal.
[0006] Compared with existing technologies, this application provides a vehicle-mounted height restriction collision avoidance and alarm system. To address the limitations of single-sensor perception capabilities and susceptibility to environmental interference in the prior art, this system simultaneously acquires LiDAR, millimeter-wave radar, and image data through a sensor data frame acquisition module, achieving multimodal information complementarity and significantly improving the accuracy and all-weather adaptability of obstacle perception. Addressing the core shortcomings of existing technologies—the inability to identify obstacle types and the neglect of dynamic changes in vehicle height—the system performs deep processing of multi-source data through an obstacle perception module to identify obstacle features. Simultaneously, a vehicle environment determination module incorporates the original vehicle height value to obtain the vehicle's real-time accurate height. Finally, a collision risk analysis module performs a comprehensive assessment based on the vehicle's dynamic height and accurate information about obstacles ahead, generating reliable graded alarm signals. This effectively overcomes the warning failure problems caused by incomplete information and fixed parameters, providing drivers with accurate and reliable safety assurance. Attached Figure Description
[0007] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0008] Figure 1 This is a block diagram of a vehicle-mounted height restriction and collision avoidance alarm system according to an embodiment of this application.
[0009] Figure 2 This is a schematic diagram of the data flow of a vehicle-mounted height restriction and collision avoidance alarm system according to an embodiment of this application.
[0010] Figure 3This is a block diagram of the obstacle sensing module in a vehicle-mounted height restriction and collision avoidance alarm system according to an embodiment of this application.
[0011] Figure 4 This is a block diagram of the collision risk analysis module in a vehicle-mounted height restriction and collision avoidance alarm system according to an embodiment of this application. Detailed Implementation
[0012] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0013] This application is made in response to the problems existing in the aforementioned prior art. Figure 1 This is a block diagram of a vehicle-mounted height restriction and collision avoidance alarm system according to an embodiment of this application. Figure 2 This is a schematic diagram of the data flow of a vehicle-mounted height restriction and collision avoidance alarm system according to an embodiment of this application. Specifically, as shown... Figure 1 and Figure 2 As shown, the vehicle-mounted height restriction collision avoidance alarm system 100 according to an embodiment of this application includes: a sensor data frame acquisition module 110, used to acquire synchronous sensor data frames, the synchronous sensor data frames including original LiDAR point cloud data, original millimeter-wave radar target list, synchronous image pairs, and original vehicle height values; an obstacle perception module 120, used to extract and identify features of obstacles ahead based on the original LiDAR point cloud data, synchronous image pairs, original millimeter-wave radar target list, and original vehicle height values in the synchronous sensor data frames to obtain a perceived obstacle list and the vehicle's real-time accurate height; a vehicle environment determination module 130, used to comprehensively determine the vehicle and environmental status of the perceived obstacle list to obtain an obstacle list combined with the context; a collision risk analysis module 140, used to assess the collision risk level of each obstacle in the obstacle list combined with the context based on the vehicle's real-time accurate height to obtain a collision risk profile; and an alarm driving module 150, used to input the collision risk profile and the obstacle list combined with the context into a hierarchical alarm engine to obtain an alarm driving signal.
[0014] Specifically, the sensor data frame acquisition module 110 is used to acquire synchronous sensor data frames, which include raw LiDAR point cloud data, raw millimeter-wave radar target list, synchronous image pairs, and raw vehicle height values. It is understandable that in complex road traffic environments, vehicle assistance systems relying on a single sensor suffer from insufficient perception dimensions, leading to significant reductions in accuracy and reliability in situations such as adverse weather or poor lighting conditions. This limitation results in misjudgments of the distance, type, and height of obstacles ahead, while also ignoring the dynamic height differences caused by changes in vehicle load, thus failing to provide accurate and reliable collision warnings. Therefore, to construct a system capable of comprehensively and accurately perceiving the driving environment and conducting reliable risk assessments, it is necessary to acquire a comprehensive dataset containing multiple dimensions and strictly aligned in time, providing a high-quality, unambiguous data foundation for subsequent fusion analysis, feature extraction, and decision-making.
[0015] In one specific implementation, the synchronized image pair is acquired by a binocular camera group, which includes an infrared camera and an RGB camera, and the original vehicle height value is acquired by a vehicle height sensor.
[0016] In the above embodiment, the sensor data frame acquisition module 110 is implemented as follows: First, a hardware platform integrating multiple sensors needs to be constructed, and a high-precision time synchronization mechanism needs to be established. This platform has a high-precision lidar installed at the front of the vehicle roof for acquiring three-dimensional spatial point clouds; a millimeter-wave radar deployed inside the front grille for detecting long-range moving targets; a binocular camera group consisting of an infrared thermal imaging camera and a high-resolution RGB camera installed side-by-side next to the lidar for acquiring image information; and vehicle height sensors installed at the four suspension positions of the vehicle. The time synchronization mechanism is the core of the entire data acquisition process, and a pulses per second (PPS) signal based on the Global Positioning System (GPS) can be used as the master clock source. All sensors and data acquisition units are connected to this clock source, ensuring that each sensor simultaneously acquires data the instant it receives the synchronization trigger signal, thereby guaranteeing strict alignment of the data in the time dimension.
[0017] At each synchronization clock cycle, for example every 100 milliseconds, the sensor data frame acquisition module performs a complete data acquisition process. The LiDAR begins a 360-degree rotational scan, with its internal laser emitter emitting a laser beam, and the return signal being captured by the receiver. By calculating the time of laser flight, the three-dimensional coordinates (x, y, z) and reflection intensity value of each reflection point can be accurately calculated. A single scan generates hundreds of thousands of such data points, and the collection of these points constitutes the original LiDAR point cloud data, which precisely depicts the three-dimensional contours of the vehicle's surrounding environment in digital form. For example, a height restriction pole will appear in the point cloud data as a cluster of dense points distributed laterally at a specific height.
[0018] Simultaneously, the millimeter-wave radar transmits millimeter-wave electromagnetic waves forward and receives the reflected echoes from targets. Its internal signal processor analyzes the frequency, phase, and time delay of the echoes, directly outputting a structured raw millimeter-wave radar target list. This list is not the raw echo signal, but rather target-level information after preliminary processing. Each entry in the list represents a detected target, containing information such as the target's unique identifier, longitudinal and lateral distances, relative speed, azimuth, and radar cross-section (RCS). For example, a car traveling in the same direction 50 meters ahead might be recorded in the target list as {ID:001, distance: 50.2 meters, speed: -1.5 meters / second, azimuth: 0.5 degrees}.
[0019] At the same moment, the hardware triggering circuit of the binocular camera array is activated, and the infrared and RGB cameras simultaneously complete an exposure, generating one infrared image and one color image respectively. These two images together constitute a synchronized image pair. The RGB image records rich visual information such as the color and texture of the scene. For example, it can clearly capture the red circular border and the white background with black lettering indicating a height limit of 4.5 meters on a height restriction sign. The infrared image records the temperature distribution of the scene, and even at night or in foggy weather, it can effectively detect heat-generating components such as engines and tires of vehicles ahead through differences in thermal radiation, providing important supplementary information for target detection.
[0020] Finally, four vehicle height sensors mounted on the vehicle suspension measure the vertical distance between the chassis and the axles in real time. These analog signals are processed by an analog-to-digital converter to generate four independent digital values, i.e., the raw vehicle height values. For example, at any given moment, the readings from the four sensors might be 30.1 cm for the left front, 30.3 cm for the right front, 32.5 cm for the left rear, and 32.4 cm for the right rear. These values directly reflect the real-time attitude and ground clearance changes caused by uneven load distribution or road surface undulations. Ultimately, the module packages the raw LiDAR point cloud data, the raw millimeter-wave radar target list, the synchronized image pairs, and the four raw vehicle height values acquired at that moment into a data structure with a unified timestamp, i.e., a synchronized sensor data frame.
[0021] Specifically, the obstacle perception module 120 is used to extract and identify features of obstacles ahead based on the original LiDAR point cloud data, synchronized image pairs, original millimeter-wave radar target list, and original vehicle height value in the synchronized sensor data frame, in order to obtain a list of perceived obstacles and the vehicle's real-time accurate height. Correspondingly, after acquiring the synchronized sensor data frame, what is obtained is a raw data stream from different physical media and in various formats. This data itself has not been refined and cannot be directly used for decision-making. For example, the LiDAR point cloud consists of massive amounts of three-dimensional coordinate points, while image data is a pixel matrix. In their most primitive state, they lack clear semantic information and cannot indicate what the specific obstacles ahead are or where they are. The reason why existing solutions pointed out in the background art have defects is fundamentally due to insufficient processing capabilities for this raw information, making effective feature extraction and target recognition impossible. Therefore, this application extracts and identifies features of obstacles ahead based on the original LiDAR point cloud data, synchronized image pairs, original millimeter-wave radar target list, and original vehicle height value in the synchronized sensor data frame. This low-level, disordered raw data is processed through a series of specialized algorithms and transformed into a structured obstacle list containing clear semantic information. This process is a crucial step in moving from basic perception to environmental understanding. It directly solves the technical challenges of existing technologies being unable to distinguish obstacle types or identify key information, laying a solid data foundation for subsequent accurate risk assessment.
[0022] As can be understood, raw LiDAR point cloud data is a collection of unprocessed 3D spatial points directly acquired by the sensor, depicting the environment around the vehicle in an extremely primitive and complex form. This dataset contains information on roads, buildings, vegetation, and all potential obstacles, but this information is mixed together, lacks structure and semantics, and cannot be directly used for judgment. As mentioned in the background, existing solutions are flawed in their inability to effectively distinguish between genuine height-restricted obstacles and non-dangerous objects such as roadside branches. Directly processing the entire point cloud is not only computationally intensive but also fails to separate individual physical entities from the environment. Therefore, to transform this disordered and massive raw point cloud data into a structured, easily processed list of independent physical objects through systematic filtering, aggregation, and abstraction, this application calculates a LiDAR target set.
[0023] In one specific implementation, Figure 3 This is a block diagram of the obstacle sensing module in a vehicle-mounted height restriction and collision avoidance alarm system according to an embodiment of this application. Figure 3 As shown, the obstacle perception module 120 includes: a point cloud data segmentation unit 120-1, used to segment the original lidar point cloud data into ground points to obtain non-ground point clouds; a clustering analysis unit 120-2, used to perform clustering analysis on the non-ground point clouds to obtain multiple point cloud clusters; and a lidar target set generation unit 120-3, used to calculate the minimum bounding rectangle of each point cloud cluster in the multiple point cloud clusters to obtain a lidar target set.
[0024] Specifically, raw LiDAR point cloud data is a massive data stream composed of hundreds of thousands or even millions of data points. Each point contains its three-dimensional coordinates (x, y, z) in the LiDAR coordinate system, as well as its reflection intensity value. For example, when a vehicle is driving on a city road and there is a height-restricted gantry spanning the road 100 meters ahead, this frame of raw point cloud data will simultaneously contain numerous points forming the flat road surface, points of the columns on both sides of the gantry and the top beam, points of the exterior walls of distant buildings, and even points of tree branches and leaves along the roadside. These points are mixed together, forming a complete three-dimensional snapshot of the scene.
[0025] First, the point cloud data segmentation unit 120-1 receives the raw LiDAR point cloud data and performs ground point segmentation. Since the ground itself is not an obstacle to be avoided, removing ground points significantly reduces the amount of data for subsequent processing and eliminates interference from the ground itself in obstacle detection. An implementable ground segmentation algorithm is based on a random sampling consistency plane fitting method. This unit randomly selects a very small number of points, such as three, from the entire input point cloud to define an initial planar model. Then, it iterates through the entire point cloud, calculating the vertical distance from each point to this assumed plane. All points with a distance less than a preset threshold are considered interior points of the planar model. This threshold needs to be determined based on the road surface smoothness and the LiDAR measurement accuracy; for example, it can be set to 10 centimeters. This random sampling and verification process is repeated for a preset number of iterations, such as 100 times. In all iterations, the planar model with the most interior points is ultimately determined as the best model representing the ground. Once the ground planar model is determined, all interior points belonging to this model are marked as ground points and removed from the dataset. All the remaining points, those that do not conform to the ground plane model, constitute the non-ground point cloud. In the example above, after processing, tens of thousands of points that make up the road surface were successfully filtered out, and the remaining non-ground point cloud mainly consists of the point set that makes up the gantry and distant buildings and trees.
[0026] Next, the clustering analysis unit 120-2 receives the non-ground point cloud as input and performs clustering analysis on these discrete points to group points that are spatially adjacent and logically belong to the same physical entity together. For example, all the points constituting the Longmengjia should be grouped into one cluster, while a distant tree should be grouped into another independent cluster. A density-based clustering algorithm, such as DBSCAN (Density-Based Noisy Spatial Clustering), can be used here. This algorithm requires two key parameters: neighborhood radius (Eps) and minimum number of points (MinPts). The neighborhood radius defines the search range for a point, while the minimum number of points defines the minimum number of points required to form a dense region. The settings of these two parameters are crucial to the clustering effect and need to be calibrated according to the density characteristics of the LiDAR point cloud. For example, for a typical vehicle-mounted LiDAR, at a distance of 50 meters, the neighborhood radius can be set to 0.5 meters and the minimum number of points to 20. The clustering process starts with any unvisited point in the non-ground point cloud and checks the number of points within its neighborhood radius. If the number of points is greater than or equal to the minimum number of points, the point is marked as a core point, and a new cluster is created. Then, starting from the core point, all points in its neighborhood that meet the density requirements are recursively added to this cluster. This process continues until all core points and their density-reachable points are assigned to the corresponding clusters. Sparse points that cannot be covered by the neighborhood of any core point are considered noise points and are discarded. After processing by the clustering analysis unit 120-2, the input non-ground point cloud is successfully divided into multiple point cloud clusters. In the example of this application, all the points constituting the gantry are aggregated into a large point cloud cluster due to spatial continuity; distant buildings and trees also form their own independent point cloud clusters. The final output is a list, where each item is a point cloud cluster, i.e., a set of points.
[0027] Finally, the lidar target set generation unit 120-3 receives these multiple point cloud clusters and calculates a compact geometric representation for each cluster, thus obtaining the final lidar target set. A point cloud cluster itself is still a collection of hundreds or thousands of points, which is inconvenient for kinematic analysis and data fusion. Abstracting it into a bounding box with a defined location, size, and orientation is an efficient method for data dimensionality reduction and structuring. To calculate the minimum bounding box of each point cloud cluster, principal component analysis (PCA) can be used. For each input point cloud cluster, the unit first calculates the covariance matrix of the three-dimensional coordinates of all points within the cluster. Then, by performing eigenvalue decomposition on the covariance matrix, three mutually orthogonal eigenvectors and their corresponding eigenvalues are obtained. These three eigenvectors precisely define the principal directions of the point cloud cluster's spatial distribution, i.e., the three coordinate axes of the bounding box. The eigenvector corresponding to the largest eigenvalue points in the direction of the longest point cloud distribution. After determining the orientation of the bounding box, all points in the point cloud cluster are projected onto these three new coordinate axes. By finding the maximum and minimum values of the projected coordinates on each axis, the length, width, and height of the bounding box can be accurately calculated. The center point of the bounding box can be taken as the geometric center of all points in the point cloud cluster. After the calculation is completed, each point cloud cluster is replaced by structured data containing its center point coordinates, dimensions (length, width, height), and rotation matrix (or quaternion, representing orientation). The processing results of all point cloud clusters are summarized to form the LiDAR target set. For example, for the point cloud cluster representing the gantry, the unit will output a LiDAR target, which may contain the following information: {Target ID: Lidar_001, Center position: {x:100.1, y:0.5, z:5.0}, Dimensions: {Length:12.0, Width:0.4, Height:5.5}, Orientation: {...}}. This target set describes the key physical properties of all non-ground objects sensed by the LiDAR in a very concise and standardized way.
[0028] Correspondingly, sensors such as LiDAR and millimeter-wave radar can accurately perceive the geometry, position, and motion of objects in the environment, but they have an inherent deficiency in understanding the semantic attributes of objects. For example, they cannot distinguish between a suspended height restriction pole and a harmless roadside tree branch, let alone read the key numerical information marked on a height restriction sign. The background art clearly points out that existing solutions lack this deep understanding capability, leading to insufficient reliability in complex scenarios. Therefore, to bridge this semantic gap, this application uses deep learning and computer vision techniques with a computer vision object set to decode and translate the raw, meaningless pixel matrix into a high-dimensional, structured information list containing object categories, specific identifiers, and even textual information, thereby endowing the entire perception system with cognitive capabilities.
[0029] Based on this, in one specific embodiment, the obstacle perception module 120 further includes: a depth map generation unit 120-4, used to perform stereo matching on the synchronized image pair to obtain a depth map; a target detection unit 120-5, used to input the left camera image or the right camera image in the synchronized image pair into a pre-trained convolutional neural network target detection model to obtain a detection result list, wherein each detection result in the detection result list includes the two-dimensional bounding box of the target in the image and its corresponding object category; a height restriction sign judgment unit 120-6, used to traverse each detection result in the detection result list, such as if the object category of a certain target is a height restriction sign, inputting the image region corresponding to the two-dimensional bounding box of the target in the image into the OCR engine to obtain the height restriction sign value; a target position and distance calculation unit 120-7, used to calculate the three-dimensional position and distance of each target in the visual coordinate system based on the two-dimensional bounding box and depth map of the target in the image; and a data encapsulation unit 120-8, used to encapsulate the three-dimensional position and distance of each target in the visual coordinate system, the object category, and the height restriction sign value to obtain a visual target set.
[0030] Specifically, the synchronized image pair comprises two images captured simultaneously by an infrared camera and an RGB camera, both physically mounted side-by-side. These two images are strictly aligned at the pixel level, laying the foundation for subsequent 3D information extraction. For example, in a specific scenario, when there is a gantry with a 4.5-meter height restriction sign 60 meters in front of a vehicle, the RGB image clearly records the structure and color of the gantry, as well as the text on the sign, while the infrared image presents a heat map of the scene based on temperature differences.
[0031] First, the depth map generation unit 120-4 receives a synchronized image pair as input. This image pair consists of images captured simultaneously by an RGB camera and an infrared camera. Using binocular stereo matching technology, it calculates the depth information for each pixel in the scene, thereby generating a depth map. This application employs an advanced existing stereo matching algorithm, such as the semi-global block matching (SGBM) algorithm. This algorithm first corrects the RGB image (left image) and the infrared image (right image) so that the search for corresponding points can be constrained to the same horizontal line. Specifically, in this binocular stereo matching process, the left and right images are defined based on the physical left-right layout of the two cameras. One image serves as the matching reference (left image), and the other is used to search for corresponding points (right image). Because it contains richer texture details, the RGB image is used as the left image, while the infrared image becomes the right image accordingly. The task of this unit is... Next, it takes a small pixel block, such as a 5x5 window, centered on each pixel in the left image, and then slides a window of the same size along the corresponding horizontal line in the right image. By calculating the structural similarity between the two pixel blocks, such as using normalized cross-correlation, it finds the best match. This method effectively handles the significant differences in brightness and color between RGB and infrared images. This process calculates a disparity value for each pixel in the left image, which is the difference in its horizontal pixel coordinates between the two images. The disparity values of all pixels are combined to form a disparity map. The SGBM algorithm not only considers the matching cost of individual pixels but also introduces a smoothness constraint through dynamic programming in multiple directions, meaning that the disparity values of adjacent pixels should tend to be consistent. This makes the generated disparity map more robust and accurate in weakly textured regions and occlusion boundaries. Finally, based on the pre-calibrated intrinsic parameters of the binocular camera system—namely, focal length and extrinsic parameters (the physical distance between the optical centers of the two cameras, i.e., the baseline length)—the disparity map is converted pixel-by-pixel into a depth map using the formula: Depth = (Focal Length * Baseline Length) / Disparity. This depth map is a single-channel image, where the grayscale value of each pixel directly corresponds to the physical distance of that point from the camera in the real world. In the example above, the pixel region in the depth map corresponding to the gantry sign will have values concentrated around 60 meters.
[0032] Next, the object detection unit 120-5 receives RGB camera images from the synchronized image pair as input, locates and identifies all objects of interest in the two-dimensional image, and outputs a list of detection results. This is achieved using a pre-trained convolutional neural network object detection model, such as YOLOv5. This model consists of three main parts: a backbone network, a neck network, and a head network. The backbone network uses a CSPDarknet53 structure, extracting feature maps from the input image layer by layer, from low-level features like edges and corners to high-level features like textures and parts, through a series of convolutional layers, batch normalization layers, and activation function layers. The neck network uses a path aggregation network structure, which effectively fuses feature maps output from different depth layers of the backbone network, combining rich semantic information from deep feature maps with precise positional information from shallow feature maps. The head network then performs the final prediction on the fused feature map. It divides the feature map into a grid and predicts the bounding box of the object it may contain, the object confidence (i.e., the probability that an object exists within the bounding box), and the conditional probability that the object belongs to each predefined category (e.g., vehicle, pedestrian, height restriction sign) for each grid cell. The model's massive parameters, such as weights and biases, were obtained through supervised learning on a large-scale dataset containing millions of labeled traffic scene images. During training, the parameters were continuously adjusted using backpropagation and gradient descent optimizers to minimize a composite loss function consisting of localization loss, confidence loss, and classification loss. When the RGB image of the scene is input into this trained model, the model performs a forward propagation calculation and outputs a list of detection results. Each item in the list represents a detected object, including the object's two-dimensional bounding box coordinates in the image (e.g., the pixel coordinates of the top-left and bottom-right corners [x1, y1, x2, y2]), its corresponding object category (e.g., a height restriction sign), and a confidence score indicating the accuracy of the classification. In this embodiment, the list may contain an entry whose bounding box precisely encloses a height restriction sign in the image, classifying it as a height restriction sign, with a confidence score of 0.95.
[0033] Subsequently, the height limit sign judgment unit 120-6 processes the detection result list, traverses each detection result in the list, and extracts depth information for specific categories of objects. It checks the object category field of each detection result. If the object category of a certain target is identified as a height limit sign, this unit will perform a special operation: First, based on the two-dimensional bounding box coordinates of this target, it precisely crops out the image area of the height limit sign from the original RGB image. Then, for this image slice that only contains the height limit sign, this OCR engine itself is also an end-to-end trained deep learning model, adopting an advanced CRNN architecture. This architecture cleverly combines the advantages of convolutional networks, recurrent networks, and temporal classification to achieve high-precision recognition without character segmentation. The first part of this architecture is a deep convolutional neural network (CNN) whose function is feature extraction. In this CNN part, for example, a lightweight VGG or ResNet variant, through the stacking of multiple 3x3 convolutional kernels and max pooling layers, transforms the input two-dimensional image slice into a sequence of high-level feature maps. During this process, the height of the image is greatly compressed, while the width information is retained, thus effectively encoding the spatial features into a sequence of temporal feature vectors from left to right. Subsequently, this feature sequence is fed into the second part of the architecture, a Bi-LSTM. At each time step of the sequence, the Bi-LSTM outputs a probability distribution that includes all possible characters (including a special 'blank' character). Finally, the third part of the architecture connects to the connectionist temporal classification (CTC) layer, which receives the sequence of probability distributions from the Bi-LSTM. The subtlety of the CTC layer lies in its ability to automatically handle alignment problems. It uses a dynamic programming algorithm to find all paths that can map to the final correct text, and sums the probabilities of these paths to calculate the loss. In the inference stage, it decodes the redundant and repetitive sequences output by the RNN (for example, "4-4-..-5-meter-meter") into the final concise and correct text string through a greedy decoding or beam search algorithm. The entire OCR engine has also undergone end-to-end pre-training on a large image dataset containing various traffic sign texts. In this example, when the OCR engine processes the input height limit sign image slice, the internal CRNN model therein executes the above complete process, and finally outputs the recognized string, such as the height limit sign value: 4.5 meters. This value will be precisely associated with the corresponding detection result.
[0034] Next, the target position distance calculation unit 120-7 begins operation. This unit's task is to calculate the true position and distance in three-dimensional space for each detected two-dimensional object in the list. For each target in the list, the unit first calculates the pixel coordinates (u, v) of the center point of its two-dimensional bounding box. For example, for the previously detected height restriction sign, its center point coordinates in the bounding box are calculated as (976, 490). Then, it uses these center point coordinates as an index to look up the corresponding depth value D in the depth map. In this scene, the depth map value at this pixel location is 60.0, indicating that the straight-line distance from the object's center point to the camera is 60.0 meters. After obtaining the two-dimensional image coordinates (u, v) and depth D, the unit uses a pre-calibrated camera intrinsic parameter matrix, which includes the camera's focal length (e.g., fx=1200, fy=1200) and principal point coordinates (e.g., cx=960, cy=540), to backproject the points in the image coordinate system to the three-dimensional space in the camera coordinate system through the inverse operation of perspective projection. The specific calculation process is as follows: the X coordinate is calculated using the formula (u-cx)*D / fx, i.e., (976-960)*60.0 / 1200, yielding 0.8 meters; the Y coordinate is calculated using the formula (v-cy)*D / fy, i.e., (490-540)*60.0 / 1200, yielding -2.5 meters; and the Z coordinate is directly equal to the depth value D, i.e., 60.0 meters. Thus, the three-dimensional position coordinates of the object in the camera coordinate system are obtained as {X: 0.8 meters, Y: -2.5 meters, Z: 60.0 meters}, thereby completing the precise conversion from two-dimensional image detection to three-dimensional spatial positioning.
[0035] Finally, the data encapsulation unit 120-8 performs the final step. It receives a list of detection results containing all calculated information and encapsulates it into a standardized visual target set. This unit iterates through each fully processed target in the list, extracting its key information, including the calculated 3D position and distance, the object category and confidence level given by the object detection model, and the height restriction sign value obtained by the height restriction sign judgment unit and the OCR engine. It then organizes this information into a structured data format. For example, for a height restriction sign in the scene, the final encapsulated visual target might look like this: {Target ID:Vision_001, 3D Position:{X:0.8m,Y:-2.5m,Z:60.0m}, Distance:60.0m, Object Category: Height Restriction Sign, Confidence Level:0.95, Height Restriction Sign Value:4.5m}. For other objects detected in the scene, such as a car, the data entry does not include the height restriction sign value field. All these structured data entries are combined to form the final visual target set. This object set comprehensively describes the external environment as understood by the visual sensor in a machine-readable and information-rich manner.
[0036] Understandably, while independent processing of multi-sensor information yields preliminary perception results, these results are fragmented, heterogeneous, and uncertain. LiDAR provides precise geometric contours, vision imparts semantic meaning to objects, and millimeter-wave radar excels at speed measurement, but their descriptions of the same physical entity are not yet unified. More importantly, all these measurements are performed in a relative coordinate system with the vehicle itself as the reference. Background technology profoundly points out that ignoring the dynamic height changes of the vehicle due to load and road conditions is a key flaw leading to the failure of existing warning solutions. Therefore, this application further calculates the absolute height of obstacles based on LiDAR target sets, visual target sets, the original millimeter-wave radar target list, and the vehicle's real-time precise height. This transforms multi-source, heterogeneous relative perception information through deep fusion and coordinate transformation into a unified, accurate, and highly reliable absolute height judgment based on the ground, thus providing the most critical and reliable decision-making basis for the final collision risk assessment.
[0037] In one specific embodiment, the obstacle perception module further includes: a multimodal target fusion unit 120-9, used to perform multimodal target association and credibility enhancement on the lidar target set, visual target set and original millimeter-wave radar target list to obtain a fused candidate obstacle list; and an absolute height calculation unit 120-10, used to calculate the absolute height of obstacles in the fused candidate obstacle list based on the vehicle's real-time accurate height to obtain the perceived obstacle list.
[0038] Specifically, before fusion begins, coordinate system registration is performed first. All sensors are precisely calibrated during installation, resulting in their respective coordinate systems relative to the vehicle coordinate system—a fixed coordinate system with the rear axle center as the origin, known as the extrinsic parameter matrix. The multimodal target fusion unit 120-9 first applies these extrinsic parameter matrices to uniformly transform the spatial information of all input targets, such as position, size, and velocity, to the vehicle coordinate system, ensuring all data is compared and calculated within the same spatial reference. After coordinate transformation, the data association process officially begins. This process aims to identify different sensor detection results pointing to the same physical entity. A data association algorithm based on gating strategy and optimal allocation is employed. First, the set of LiDAR targets with the highest measurement accuracy and most stable information is used as the reference. For each target in this set, a three-dimensional ellipsoidal association gate is set around its current position. The size of this association gate is dynamically calculated based on the target's motion uncertainty, determined by its velocity and the covariance of the previous moment, and the measurement errors of each sensor. Subsequently, the unit iterates through all targets in the visual target set and the millimeter-wave radar target list, calculating their Mahalanobis distances to the current LiDAR target. Mahalanobis distance is an effective statistical distance that considers the covariance of the data distribution and can effectively measure the consistency between different measurement results. An association threshold is set, which is determined based on the confidence interval of the chi-square distribution, for example, the value corresponding to the 95% confidence interval. All candidate targets with Mahalanobis distances less than this threshold are initially considered to be associated with the current LiDAR target. To resolve the ambiguity of multiple target matches, the Hungarian algorithm or a more advanced joint probabilistic data association algorithm is used to find an optimal matching scheme that maximizes the global association probability, ultimately finding a unique cross-sensor matching relationship for each target, forming an association cluster. Once the target association is completed, the information fusion and credibility enhancement process is initiated. The unit maintains an Extended Kalman Filter (EKF) for each continuously tracked obstacle, i.e., an association cluster. The core of this filter is a state vector used to describe the comprehensive properties of the obstacle, such as 3D position, velocity, size, orientation angle, and object category. In each time period, the filter first performs a prediction step based on a preset motion model (such as a uniform velocity or uniform acceleration model) to obtain a prior state estimate. Then, in the update step, measurements from different sensors associated with the current moment are used to correct this prior estimate. The weight of the measurements from different sensors in the update process is determined by their measurement noise covariance matrix, which is obtained through extensive experimental statistics during the sensor calibration phase. For example, lidar is extremely accurate in measuring position and size, and its corresponding measurement noise value is very small; therefore, its measurements dominate when updating the position and size state. Millimeter-wave radar is accurate in measuring radial velocity, and its measurements have the highest weight when updating the velocity state.Visual sensors provide crucial semantic information. When a visual target is associated, its identified object category, such as a height restriction sign, and confidence level, are updated probabilistically to enhance the reliability of the category attribute in the state vector. For example, if an obstacle is simultaneously identified as a slender rod-like structure by LiDAR and a height restriction sign by visual sensors, the probability of it ultimately being identified as a height restriction sign will be much higher than the result of a single sensor. Consider a specific scenario: there is a height restriction gantry ahead. The LiDAR target set outputs a precise 3D bounding box; the visual target set outputs a 2D box, categorized as a height restriction sign, with a confidence level of 0.95, and includes an OCR-generated value of 4.5 meters; the millimeter-wave radar target list outputs a target with zero velocity and a large radar cross-section. After receiving these three sets of data, the multimodal target fusion unit first unifies them into the vehicle coordinate system. Through data association, it discovers that these three targets highly overlap in space, successfully associating them with the same physical entity. Subsequently, the Kalman filter for the gantry maintenance was updated: its position and size status primarily relied on precise data from the lidar; its velocity status relied on zero-velocity data from millimeter-wave radar and was confirmed as stationary; its category status relied on visual information and was updated to a height restriction sign, with the confidence level further increased to 0.99. Ultimately, the fused candidate obstacle list output by this unit is a structured dataset, where each entry represents a highly reliable physical obstacle confirmed and enhanced by multi-sensor information, with overall reliability far exceeding that of any single sensor's independent observation. Taking a height restriction gantry as an example, it will exist in the fused list as a unified target ID such as F_001, with its data entry format as follows: {ID:F_001, 3D position:{x:100.1,y:0.5,z:5.0}, dimensions:{length:12.0,width:0.4,height:5.5}, motion state: stationary, object category: height restriction sign, confidence level:0.99, height restriction sign value:4.5 meters}. Its precise position and dimension information are mainly derived from the high-precision 3D geometric measurement results provided by LiDAR, as it can most reliably describe the physical contours of the obstacle; its category is determined as a height restriction sign, and the crucial 4.5-meter height value is entirely derived from the semantic recognition capabilities of the vision system; while its stationary motion state is determined by comprehensively integrating precise velocity measurement information from millimeter-wave radar. More importantly, the confidence level was able to increase from 0.95 to 0.99 given by the vision system alone because the physical structure detected by the lidar, which is highly consistent in space, provides strong evidence for visual recognition, thus significantly enhancing the overall credibility of the judgment.
[0039] In one specific embodiment, the absolute height calculation unit 120-10 includes: a first candidate obstacle extraction subunit, used to extract a first candidate obstacle from the fused candidate obstacle list; a vertical distance data extraction subunit, used to extract the vertical distance data of the bottom of the first candidate obstacle relative to the installation height of the lidar; and an obstacle absolute height calculation subunit, used to calculate the sum of the real-time accurate height of the vehicle and the vertical distance data of the bottom of the first candidate obstacle relative to the installation height of the lidar to obtain the absolute height of the first candidate obstacle.
[0040] Specifically, firstly, the first candidate obstacle extraction subunit extracts obstacles one by one from the fused candidate obstacle list for processing. Taking the height restriction gantry as an example, this subunit extracts the target with ID F_001, whose data entries contain highly reliable information after fusion, such as its precise location, size, and object category in the vehicle coordinate system.
[0041] Next, the vertical distance data extraction subunit operates on the extracted F_001 target. This unit's task is to extract the vertical distance of the obstacle's bottom relative to the lidar's installation height from its rich fused data. Since the obstacle's 3D bounding box information primarily comes from the highest-precision lidar and has already been uniformly represented in the vehicle coordinate system (i.e., with the lidar's own position as the origin or at a known fixed position), this calculation is very straightforward. This subunit reads the minimum value of the F_001 target's 3D bounding box along the Z-axis (vertical direction) of the vehicle coordinate system. For example, if the Z-coordinate of F_001's center position is 5.0 meters and its height is 5.5 meters, then the Z-axis range of its bounding box is from (5.0 - 5.5 / 2) = 2.25 meters to (5.0 + 5.5 / 2) = 7.75 meters. Therefore, the vertical distance data of the obstacle's bottom relative to the lidar's installation height extracted by this subunit is +2.25 meters. This value indicates that the bottom crossbeam of the height-restricted Longmeng frame is 2.25 meters above the roof-mounted lidar (the origin of the coordinate system).
[0042] Then, the obstacle absolute height calculation subunit performs the final calculation. But before summing, it requires a crucial input: the vehicle's real-time accurate height. This value is obtained through a parallel calculation process. This process receives raw vehicle height values from four sensors mounted on the vehicle's suspension. These raw values reflect relative distances in local suspension states; for example, at a given moment, the readings from the four sensors might be 30.1 cm to the left front, 30.3 cm to the right front, 32.5 cm to the left rear, and 32.4 cm to the right rear. Because these raw values can fluctuate rapidly due to road bumps and sensor noise, to obtain a smooth, stable value that accurately reflects the vehicle's true ground clearance, in one embodiment, the calculation process for the vehicle's real-time accurate height includes: applying a low-pass filter or Kalman filter to the raw vehicle height values to obtain the vehicle's real-time accurate height. Low-pass filtering is a relatively simple and effective method that effectively filters out rapid, high-frequency fluctuations caused by minor road bumps or sensor noise by weighted averaging of continuous measurements, resulting in a smoother height curve. However, simple low-pass filtering has inherent latency and limited ability to distinguish between slow, real-time height changes and continuous low-frequency noise. Therefore, a more advanced and adaptive approach is to employ Kalman filtering, which tracks the vehicle's true height as an internal state. At each time step, the filter first predicts the current height based on the vehicle's kinematic model (a stochastic process model assuming relatively smooth height changes). Then, it refines this prediction with an observed, noisy instantaneous vehicle height, initially calculated from raw readings from four sensors and combined with inherent vehicle geometry parameters (such as wheel radius, fixed distance from suspension mounts to the roof, etc.). The Kalman filter dynamically calculates an optimal gain based on pre-defined process noise covariance (reflecting the drasticness of vehicle height changes) and measurement noise covariance (reflecting the reliability of sensor readings), determining whether to place more weight on the predicted or measured values. For example, when the preliminary calculated instantaneous height estimate, which has not yet been filtered, may fluctuate slightly between [3.98, 3.99, 4.01, 4.02] meters, after processing by a Kalman filter, a stable and accurate real-time vehicle height, such as 4.00 meters, can be output.
[0043] After obtaining the vehicle's real-time accurate height of 4.00 meters, the obstacle absolute height calculation subunit performs its core summation operation. It algebraically sums the vehicle's real-time accurate height of 4.00 meters with the previously extracted vertical distance data of the obstacle's bottom relative to the lidar installation height, i.e., +2.25 meters. The calculation formula is: Obstacle Absolute Height = Vehicle Real-time Accurate Height + Vertical Distance of Obstacle's Bottom Relative to LiDAR Installation Height. The result is: 4.00 meters + 2.25 meters = 6.25 meters. This 6.25 meters is the actual physical height of the bottom crossbeam of the height-restricted gantry from the ground, i.e., the absolute height of the first candidate obstacle. This calculated absolute height value is added back to the data entry for target F_001 as a new data field. Subsequently, this complete process is repeated for the next obstacle in the fused candidate obstacle list until all obstacles are assigned accurate absolute height values. Ultimately, the obstacle perception list output by this unit is an enhanced list that adds an absolute obstacle height field to each target, building upon the original rich information. For example, the final entry for F_001 will become: {ID:F_001, Position:{x:100.1,y:0.5,z:5.0}, Dimensions:{Length:12.0,Width:0.4,Height:5.5}, Motion Status: Stationary, Category: Height Restriction Sign, Confidence:0.99, Height Restriction Sign Value:4.5 meters, Obstacle Absolute Height:6.25 meters}.
[0044] Specifically, the vehicle environment determination module 130 is used to comprehensively determine the vehicle and environmental states of the perceived obstacle list to obtain a context-sensitive obstacle list. It should be understood that this perceived obstacle list may include drones flying high above the road, billboards in adjacent lanes, and height restriction barriers directly ahead. For a system about to make a decision, this information is redundant and even distracting. The core pain point pointed out in the background technology—the inability to effectively distinguish between real threats and non-dangerous objects—is exposed precisely at this stage. Therefore, it is necessary to comprehensively determine the vehicle and environmental states of the perceived obstacle list to introduce the crucial dimension of context. Through precise understanding of the vehicle's own state and a deep understanding of the relationships between obstacles and between obstacles and vehicles, information is filtered, correlated, and reconstructed, thereby refining a complex perceived list into a tactical situation map highly relevant to the current driving task and highlighting key points.
[0045] In one specific implementation, the vehicle environment determination module 130 is implemented as follows: the perceived obstacle list is transformed into an obstacle list combined with context. This process is mainly implemented through a context analysis engine based on rules and logical reasoning. The engine traverses each obstacle in the perceived obstacle list and performs a series of association, correction, and filtering operations based on its attributes and its spatiotemporal relationship with other obstacles.
[0046] Taking a specific scenario as an example, the input list of perceived obstacles may contain two related entries: the first is a height restriction sign with ID F_001, whose data is {ID:F_001, location:{x:100.1,...}, object category: height restriction sign, confidence level: 0.99, height restriction sign value: 4.5 meters, obstacle absolute height: 6.25 meters}; the second is a bridge with ID F_002, whose data is {ID:F_002, location:{x:115.0,...}, object category: bridge, obstacle absolute height: 5.0 meters}.
[0047] The first step of the context analysis engine is to perform semantic and spatial association. It searches the list for pairs of obstacles with specific semantic relationships. For example, the engine's rule base predefines height restriction signs as strongly associated pairs with bridges, tunnels, and gantry structures. This rule base is predefined based on general understanding of traffic regulations and road facilities. When the engine detects a height restriction sign (F_001) and a bridge (F_002), it then examines their spatial relationship. The rule base sets an association distance threshold, an empirical value derived from analyzing relevant regulations and combining statistical data from a large amount of real-world road scenarios to ensure the effectiveness of the association. For example, if a height restriction sign appears within 50 meters in front of a bridge, they are considered related. In this example, the bridge is approximately 15 meters behind the height restriction sign, perfectly meeting the association criteria.
[0048] After successful association, the engine proceeds to the second step: information correction and enhancement. This is crucial for realizing the value of context. The engine compares information from different sources and makes decisions based on preset safety principles. For bridge F_002, its absolute height, measured by lidar, is 5.0 meters. However, the associated height restriction sign F_001 clearly indicates a legal height restriction of 4.5 meters. In this case, the safety principle dictates that the stricter restriction applies. Therefore, the engine creates a new attribute field for obstacle object F_002, such as the effective height restriction, and sets its value to 4.5 meters. Simultaneously, the information from the height restriction sign F_001 itself has been absorbed by F_002. In the new list, it can be marked as processed or removed directly to avoid information redundancy.
[0049] The engine's third step is to filter irrelevant or low-threat targets. For example, the perception list might also include a tree branch with an absolute height of 4.2 meters detected by LiDAR, i.e., object category: vegetation. The engine can determine that it is a non-rigid, traversable obstacle based on its category attribute and combined with the LiDAR's reflection intensity value (non-rigid objects such as leaves have lower reflection intensity), thereby removing it from the high-priority threat list or significantly reducing its risk level.
[0050] After the above processing, the module's final output of an obstacle list, combined with context, is a highly refined and more focused list. It is no longer a simple list of objects, but a collection of threats containing logical relationships. For example, the output list will primarily contain a modified bridge object, which might take the following form: {ID:F_002, Location:{x:115.0,...}, Original Category: Bridge, Threat Category: Height-Restricted Obstacle, Physical Absolute Height: 5.0 meters, Effective Height Restriction: 4.5 meters}. This list clearly indicates the existence of a height-restricted threat ahead with an effective height of 4.5 meters.
[0051] Specifically, the collision risk analysis module 140 is used to assess the collision risk level of each obstacle in the context-based obstacle list based on the vehicle's real-time accurate height to obtain a collision risk profile. It is understandable that after the vehicle environment determination module performs contextual association and refinement of the perceived information, although a logically clearer and more focused obstacle list is obtained, this list itself is still descriptive and does not directly answer the driver's most pressing question: "Am I safe? Which obstacle ahead poses the greatest threat to me?" As mentioned in the background technology, the fundamental flaw of existing solutions lies in their inability to provide accurate and reliable warnings, due to the lack of a quantitative and dynamic risk assessment mechanism. Therefore, performing a collision risk level assessment of each obstacle in the context-based obstacle list based on the vehicle's real-time accurate height transforms a static environmental description into a dynamic, quantifiable, and prioritized risk situation map. By accurately calculating every millimeter of clearance and combining it with the distance of obstacles, it provides the driver with a clear, intuitive, and absolutely focused decision-making aid on the highest threat, thereby achieving true proactive safety warnings.
[0052] In one specific implementation, Figure 4 This is a block diagram of the collision risk analysis module in a vehicle-mounted height restriction and collision avoidance alarm system according to an embodiment of this application. Figure 4As shown, the collision risk analysis module 140 includes: an absolute height and distance extraction unit 140-1, used to extract the absolute height and distance of a first obstacle from an obstacle list combined with context; a height clearance calculation unit 140-2, used to calculate the height clearance of the first obstacle based on the absolute height of the first obstacle and the real-time accurate height of the vehicle; an effective clearance calculation unit 140-3, used to calculate the effective clearance of the first obstacle based on the height clearance of the first obstacle and a safety margin value; and an individual risk level generation unit 140-4, used to compare the effective clearance and distance of the first obstacle with a warning distance threshold to generate an individual risk level for the first obstacle.
[0053] In the above embodiment, the collision risk analysis module 140 is implemented as follows: Taking a specific scenario as an example, the input obstacle list combined with the context includes a modified bridge object with the following data: {ID:F_002, Threat Category: Height-restricted obstacle, Physical absolute height: 5.0 meters, Effective height restriction: 4.5 meters, Location: {x:115.0,...}}. Simultaneously, the input vehicle's real-time accurate height is 4.00 meters.
[0054] First, the absolute height and distance extraction unit 140-1 begins processing. This unit extracts the obstacle with ID F_002 from the list as the first obstacle. Its task is to extract the two most critical values for risk calculation from the obstacle's data entry: the obstacle's effective height and distance. This unit has an internal rule: if an effective height restriction field exists in the obstacle entry, that value is used as the obstacle's absolute height; otherwise, the physical absolute height field is used. In this example, since there is an effective height restriction of 4.5 meters, this unit extracts the obstacle's absolute height as 4.5 meters. Simultaneously, it extracts the obstacle's longitudinal distance from the location field, which is 115.0 meters.
[0055] Next, the height clearance calculation unit 140-2 receives the height value from these two values, as well as the input real-time accurate vehicle height. The task of this unit is to calculate the original vertical clearance between the vehicle and the obstacle. In one specific embodiment, the height clearance calculation unit 140-2 is used to calculate the height clearance of the first obstacle using the following formula: in, The absolute height of the first obstacle. For the vehicle's real-time accurate altitude, The height clearance of the first obstacle. Substitute the extracted values into the formula: =4.5 meters - 4.00 meters = 0.5 meters. This 0.5 meters is the height clearance of the first obstacle, meaning that, without considering any additional factors, the physical vertical distance between the top of the vehicle and the bottom of the obstacle is 0.5 meters.
[0056] Then, the effective clearance calculation unit 140-3 further processes this height clearance value. In one specific embodiment, the effective clearance calculation unit 140-3 is used to calculate the effective clearance of the first obstacle using the following formula: in, Clearance at the height of the first obstacle, As a safety margin value, This unit provides effective clearance for the first obstacle. A crucial preset parameter is introduced: the safety margin value. This value is a safety buffer set to absorb various uncertainties, such as instantaneous bouncing caused by suspension fluctuations during vehicle operation, minor measurement errors from sensors, and slight changes in road slope. The setting of this value depends on the vehicle type and safety standards. For example, for large trucks, it can be set to a fixed value, such as 0.3 meters, based on experience and regulatory requirements. The task of this unit is to subtract this safety buffer from the theoretical clearance to obtain the truly usable safe space. Substitute the values: =0.5m - 0.3m = 0.2m. This 0.2m is the calculated effective clearance of the first obstacle, which represents the vehicle's true safety margin after considering all dynamic uncertainties.
[0057] Finally, the individual risk level generation unit 140-4 generates the final risk level of the obstacle based on the calculated effective clearance and the extracted distance. This unit contains a rule set for a multi-level warning strategy. This rule set integrates both effective clearance and distance dimensions, classifying risks into different levels. These thresholds are scientifically set based on the reaction time of a typical driver and the braking performance of a vehicle. This rule set can be preset as follows: Level 0 (Safe): When the effective clearance is greater than 0.2 meters, regardless of distance, it is considered safe. Level 1 (Alert): When the effective clearance is between 0 and 0.2 meters, and the distance is greater than 100 meters, it is defined as an alert level, reminding the driver of a potential risk ahead and requiring continued attention. Level 2 (Warning): When the effective clearance is less than or equal to 0 meters, but the distance is still considerable, such as between 50 and 100 meters, it is defined as a warning level, indicating that a collision will occur if no action is taken. Level 3 (Danger): When the effective clearance is less than or equal to 0 meters, and the distance is very close, such as less than 50 meters, it is defined as a danger level, indicating that a collision is imminent and requiring the highest priority emergency alarm. Based on the current calculations, the effective clearance is 0.2 meters and the distance is 115.0 meters, which fully meets the triggering conditions for Level 1 (hint). Therefore, the individual risk level generated for obstacle F_002 in this unit is Level 1.
[0058] This complete four-step process is executed for every obstacle in the list. Once all obstacles have undergone risk level assessment, the module aggregates the information of these obstacles with risk level labels to form the final output collision risk profile. This collision risk profile is a structured dataset. It is essentially an enhanced version of the input obstacle list combined with context, where each obstacle object is additionally labeled with its height clearance, effective clearance, and final individual risk level. For example, the final entry for F_002 would be: {ID:F_002,...,Effective Height Restriction:4.5 meters,...,Height Clearance:0.5 meters,Effective Clearance:0.2 meters,Individual Risk Level:1}. This profile depicts all the height restriction risks currently faced by the vehicle and their severity in a clear and quantitative manner.
[0059] Specifically, the alarm driving module 150 is used to input the collision risk profile and the obstacle list combined with the context into the hierarchical alarm engine to obtain an alarm driving signal. In other words, after the collision risk analysis module completes the quantitative assessment of all potential threats, a clear and accurate risk profile has been formed internally. However, this digital risk profile is invisible to the driver; it is merely a string of data stored in the processor and has not yet been transformed into any effective information that the driver can perceive and understand. Therefore, inputting the collision risk profile and the obstacle list combined with the context into the hierarchical alarm engine completes the transformation from internal calculation to external manifestation. It is the final output of the entire perception, analysis, and decision-making chain, aiming to translate the abstract risk level into specific, clear, and human-friendly visual, auditory, and even tactile warnings, ensuring that the analysis results reach the driver in the most effective way, thereby triggering the correct evasive action.
[0060] In one specific implementation, the alarm driving module 150 is implemented as follows: The engine first filters out the most urgent threat from the input collision risk profile. If the profile contains multiple obstacles with different risk levels, the engine follows a simple priority principle: selecting the obstacle with the highest individual risk level as the subject of the current alarm. For example, if there is both a level 1 alert risk and a level 2 warning risk, the engine will prioritize the level 2 risk.
[0061] Once the highest risk level is determined, the engine generates a corresponding set of alarm drive signals according to a preset alarm strategy. This alarm strategy is carefully designed based on human factors engineering to ensure that the alarm information is clearly received without causing unnecessary interference to the driver in low-risk situations. A typical tiered alarm strategy might look like this: Risk level 0 (safe): The engine does not produce any output signals.
[0062] For Risk Level 1 (Warning): The engine will generate a set of low-intensity warning signals. Taking the previously calculated bridge obstacle F_002 (Risk Level 1) as an example, the engine will: 1. Generate a visual drive signal, which is a data packet sent to the controller of the instrument panel or central control display via the vehicle's Controller Area Network (CAN) bus. This data packet contains instructions to display a yellow, non-flashing height restriction sign icon in a specific area of the screen. Simultaneously, it extracts detailed information about F_002 from the obstacle list in context, such as distance 115.0 meters, effective height restriction 4.5 meters, dynamically combining this information into a text string, such as "Height restriction obstacle 115 meters ahead, height restriction 4.5 meters, please note," and instructs the screen to display this text next to the icon. 2. Generate an auditory drive signal, which is another data packet sent to the vehicle's audio system, instructing it to play a soft, crisp warning sound, such as a "ding," to attract the driver's attention, but not enough to cause them alarm.
[0063] Corresponding to Risk Level 2 (Warning): When the risk escalates, the engine will generate a moderately strong warning signal. For example, if the vehicle continues to move forward, the effective clearance decreases, and the risk level of F_002 becomes 2, the engine will: 1. Generate a visual drive signal, instructing the screen to change the original yellow icon to orange and begin flashing at a moderate frequency. The displayed text content remains unchanged, but the font may be bolded or enlarged. 2. Generate an auditory drive signal, instructing the audio system to play a repetitive, more warning-oriented voice prompt, such as "Attention, height restriction risk! Attention, height restriction risk!"
[0064] Corresponding to Risk Level 3 (Danger): In the most urgent situations, the engine will generate the strongest hazard warning signal. For example, when the vehicle is very close to an obstacle and the effective clearance is negative, the risk level rises to 3, and the engine will: 1. Generate a visual drive signal, instructing the screen to display a red, high-frequency flashing stop or collision warning icon in full screen or in the largest area, with the text simplified to the core message: "Danger! Height Restriction Collision!" 2. Generate an auditory drive signal, instructing the audio system to play a continuous, piercing alarm at maximum volume, accompanied by the most direct voice command, such as "Danger! Brake Immediately!" 3. Generate a tactile drive signal; if the vehicle is equipped with relevant hardware, the engine will also send a drive signal to the steering wheel vibration motor or seat vibrator to reinforce the urgency of the warning through tactile feedback.
[0065] Ultimately, the alarm drive signal output by the alarm drive module is a collection of precisely coded data packets used to control different human-machine interface (HMI) devices. These signals propagate on the vehicle's internal network, are received and executed by the corresponding electronic control units (ECUs), and thus present the intangible risk calculation results to the driver in real time and in a hierarchical manner, forming the last line of defense for ensuring driving safety.
[0066] In summary, the vehicle-mounted height restriction collision avoidance alarm system 100 based on the embodiments of this application is explained. To address the limitations of single-sensor perception capabilities and susceptibility to environmental interference in the prior art, this system simultaneously acquires LiDAR, millimeter-wave radar, and image data through a sensor data frame acquisition module, achieving multi-modal information complementarity and significantly improving the accuracy of obstacle perception and all-weather adaptability. Addressing the core shortcomings of existing technologies that cannot identify obstacle types and ignore dynamic changes in vehicle height, the system performs deep processing of multi-source data through an obstacle perception module to identify obstacle features. Simultaneously, the vehicle environment determination module incorporates the original vehicle height value to obtain the vehicle's real-time accurate height. Finally, the collision risk analysis module performs a comprehensive evaluation based on the vehicle's dynamic height and accurate information about obstacles ahead, thereby generating a reliable graded alarm signal. This effectively overcomes the warning failure problem caused by incomplete information and fixed parameters, providing drivers with accurate and reliable safety assurance.
[0067] Various implementations of this disclosure have been described above. The foregoing description is exemplary and not exhaustive. Furthermore, it is not limited to the disclosed implementations, and many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations.
Claims
1. A vehicle-mounted anti-collision identification and alarm system, characterized in that, Comprise: A sensor data frame acquisition module for acquiring a synchronized sensor data frame, the synchronized sensor data frame comprising raw lidar point cloud data, raw millimeter wave radar target list, synchronized image pair and raw vehicle body height value; An obstacle perception module for performing front obstacle feature extraction and identification based on the raw lidar point cloud data, synchronized image pair, raw millimeter wave radar target list and raw vehicle body height value in the synchronized sensor data frame to obtain a perceived obstacle list and a real-time accurate vehicle height; A vehicle environment determination module for comprehensively determining the vehicle and environment state of the perceived obstacle list to obtain a context-integrated obstacle list; A collision risk analysis module for performing collision risk level evaluation on each obstacle in the context-integrated obstacle list based on the real-time accurate vehicle height to obtain a collision risk portrait; An alarm driving module for inputting the collision risk portrait and the context-integrated obstacle list into a hierarchical alarm engine to obtain an alarm driving signal.
2. The vehicle-mounted height-limiting anti-collision identification and alarm system according to claim 1, characterized in that, The synchronized image pair is collected by a binocular camera group comprising an infrared camera and an RGB camera, and the raw vehicle body height value is collected by a vehicle body height sensor.
3. The vehicle mounted height limiting anti-collision identification alarm system according to claim 1, characterized in that, The obstacle perception module comprises: A point cloud data segmentation unit for performing ground point segmentation on the raw lidar point cloud data to obtain non-ground point cloud; A clustering analysis unit for performing clustering analysis on the non-ground point cloud to obtain a plurality of point cloud clusters; A lidar target set generation unit for calculating the minimum circumscribed rectangle bounding box of each point cloud cluster in the plurality of point cloud clusters to obtain a lidar target set.
4. The vehicle mounted height limiting anti-collision identification alarm system according to claim 1, characterized in that, The obstacle perception module further comprises: A depth map generation unit for performing binocular stereo matching on the synchronized image pair to obtain a depth map; A target detection unit for inputting the left camera image or right camera image in the synchronized image pair into a pre-trained convolutional neural network target detection model to obtain a detection result list, each detection result in the detection result list comprising a two-dimensional bounding box of the target in the image and its corresponding object class; A height limit plate judgment unit for traversing each detection result in the detection result list, such as a target whose object class is a height limit plate, inputting the image region corresponding to the two-dimensional bounding box of the target in the image into an OCR engine to obtain a height limit plate value; A target position distance calculation unit for calculating the three-dimensional position and distance of each target in the visual coordinate system based on the two-dimensional bounding box of the target in the image and the depth map; A data packaging unit for packaging the three-dimensional position and distance of each target in the visual coordinate system, object class and height limit plate value to obtain a visual target set.
5. The vehicle mounted height limiting anti-collision identification and alarm system according to claim 4, characterized in that, The obstacle perception module further comprises: A multi-modal target fusion unit for performing multi-modal target association and credibility enhancement on the lidar target set, visual target set and raw millimeter wave radar target list to obtain a fused candidate obstacle list; An absolute height solving unit for performing obstacle absolute height solving on the fused candidate obstacle list based on the real-time accurate vehicle height to obtain the perceived obstacle list.
6. The vehicle mounted height limiting anti-collision identification alarm system according to claim 5, characterized in that, The absolute height solving unit comprises: a first candidate obstacle extraction sub-unit configured to extract a first candidate obstacle from the fused candidate obstacle list; a vertical distance data extraction sub-unit configured to extract vertical distance data of a bottom of the first candidate obstacle relative to a laser radar installation height; an obstacle absolute height calculation sub-unit configured to calculate a sum value between a vehicle real-time accurate height and the vertical distance data of the bottom of the first candidate obstacle relative to the laser radar installation height to obtain an obstacle absolute height of the first candidate obstacle.
7. The vehicle mounted height limiting anti-collision identification alarm system according to claim 6, characterized in that, The calculation process of the vehicle real-time accurate height comprises: performing low-pass filtering or Kalman filtering on an original vehicle body height value to obtain the vehicle real-time accurate height.
8. The vehicle mounted height restriction collision prevention identification alarm system as claimed in claim 1, wherein, The collision risk analysis module comprises: an absolute height and distance extraction unit configured to extract an obstacle absolute height and distance of the first obstacle from the context-integrated obstacle list; a height clearance calculation unit configured to calculate a height clearance of the first obstacle based on the obstacle absolute height of the first obstacle and the vehicle real-time accurate height; an effective clearance calculation unit configured to calculate an effective clearance of the first obstacle based on the height clearance of the first obstacle and a safety margin value; an individual risk level generation unit configured to compare the effective clearance and distance of the first obstacle with a warning distance threshold value to generate an individual risk level of the first obstacle.
9. The vehicle mounted height limiting anti-collision identification alarm system according to claim 8, characterized in that, The height clearance calculation unit is configured to calculate the height clearance of the first obstacle according to the following formula: , wherein, the absolute height of the obstacle being the first obstacle, the real-time accurate height of the vehicle, the height clearance of the first obstacle.
10. The vehicle mounted height limiting anti-collision identification alarm system according to claim 8, wherein, The effective clearance calculation unit is configured to calculate the effective clearance of the first obstacle according to the following formula: , wherein is a height clearance of the first obstacle, is a safety margin value, is an effective clearance of the first obstacle.
Citation Information
Patent Citations
Special vehicle height limit measuring method
CN112630791A
Height limit detection vehicle control system based on binocular stereo vision and 4D millimeter wave radar and control method thereof
CN113561894A
Vehicle-mounted obstacle accurate sensing method and system and storage medium
CN116385997A
Alarm system and method for car roof placed article passing through height-limited area
CN116901834A