Intelligent rear traffic passing early warning method, device and equipment and storage medium

By using multimodal data fusion and a spatiotemporal graph neural network model, the future trajectories and behavioral intentions of traffic participants are predicted, solving the problem of insufficient perception reliability in existing technologies and achieving more efficient rear traffic crossing warnings.

CN121982931APending Publication Date: 2026-05-05CHINA FAW CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA FAW CO LTD
Filing Date
2026-02-02
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing rear traffic crossing warning technologies lack sufficient reliability in complex environments, are prone to false alarms or missed alarms, and are difficult to meet the needs of safety warnings.

Method used

Employing multimodal perception data fusion technology, this system acquires vehicle-side data using solid-state LiDAR, wide-angle cameras, and short-range millimeter-wave radar. Combining roadside blind spot perception and cloud-based traffic information, it predicts the future trajectories and behavioral intentions of traffic participants through a spatiotemporal neural network model, calculates the comprehensive collision risk level, and executes proactive safety responses.

Benefits of technology

It improves the reliability, accuracy, and timeliness of rear traffic crossing warnings, ensuring safety and driving experience in complex traffic scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121982931A_ABST
    Figure CN121982931A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent rear traffic passing early warning method, device and equipment and a storage medium, and the method comprises the steps: obtaining vehicle end multi-mode sensing data, roadside blind area sensing data and cloud traffic information; carrying out data-level pre-fusion on the vehicle-end multi-modal sensing data to generate a three-dimensional target with a semantic tag and a motion track of the three-dimensional target; based on the motion trail and the scene context information, utilizing a space-time diagram neural network model to predict a plurality of future probabilistic trails of the three-dimensional target and corresponding behavior intentions of the probabilistic trails; calculating an initial risk matrix based on a preset driving path of the vehicle and the plurality of probabilistic trajectories; based on the initial risk matrix, combined with the correction of the behavior intention on the collision probability, the roadside blind area sensing data and the cloud traffic information, determining a comprehensive collision risk level; and executing a corresponding active safety response operation based on the comprehensive collision risk level. By adopting the method, the reliability, the accuracy and the timeliness of the rear traffic passing early warning can be comprehensively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of traffic management, and more specifically, to an intelligent rear traffic crossing early warning method, device, equipment, and storage medium. Background Technology

[0002] Rear Cross Traffic Alert (RCTA) is one of the core safety functions in the field of intelligent driving. It is mainly used in scenarios such as reversing out of a parking space, changing lanes, and turning at intersections. By sensing traffic participants behind and to the sides of the vehicle, it predicts the risk of collision and alerts the driver to reduce the incidence of collision accidents.

[0003] The current mainstream RCTA technology adopts a combination of millimeter-wave radar and camera. The millimeter-wave radar is responsible for accurate ranging and speed measurement, while the camera provides target texture and semantic information. The two work together through "post-fusion" or "feature-level fusion" mode, that is, each sensor first independently identifies the target, and then the identification results are compared and summarized.

[0004] The existing technology has obvious defects: the fusion level is shallow, which cannot give full play to the complementary advantages of millimeter-wave radar and camera data. When the performance of a certain sensor deteriorates or fails due to environmental factors such as light and weather, the overall perception reliability of the system will deteriorate sharply, and false alarms or missed alarms are likely to occur, making it difficult to meet the safety early warning needs of complex traffic scenarios. Summary of the Invention

[0005] In view of this, the purpose of this application is to provide an intelligent rear traffic crossing early warning method, device, equipment and storage medium, which can comprehensively improve the reliability, accuracy and timeliness of rear traffic crossing early warning.

[0006] In a first aspect, embodiments of this application provide an intelligent rear traffic crossing early warning method, the method comprising: Acquire multimodal perception data from the vehicle, roadside blind spot perception data, and cloud-based traffic information; The vehicle-side multimodal perception data is fused at the data level to generate a 3D target with semantic labels and its motion trajectory. Based on the motion trajectory and scene context information, a spatiotemporal graph neural network model is used to predict multiple probabilistic trajectories of the three-dimensional target in the future and their corresponding behavioral intentions. Based on the vehicle's preset driving path and the multiple probabilistic trajectories, an initial risk matrix is ​​calculated, wherein the initial risk matrix represents the collision risk corresponding to each probabilistic trajectory. Based on the initial risk matrix, combined with the correction of the collision probability by the behavioral intent, the roadside blind spot perception data, and the cloud-based traffic information, the comprehensive collision risk level is determined. Based on the comprehensive collision risk level, the corresponding proactive safety response operation is executed.

[0007] Optionally, acquiring vehicle-side multimodal perception data, roadside blind spot perception data, and cloud-based traffic information includes: The vehicle-side multimodal perception data is acquired by solid-state LiDAR, wide-angle cameras, and short-range millimeter-wave radar deployed at the rear and sides of the vehicle. The blind spot sensing data is received from the roadside unit via V2I communication; The cloud-based traffic information, including real-time traffic flow, historical accident data, and high-precision maps, is obtained from the cloud-based traffic platform via an in-vehicle T-Box.

[0008] Optionally, the step of performing data-level pre-fusion of the vehicle-side multimodal perception data to generate a 3D target with semantic labels and its motion trajectory includes: The 3D point cloud data collected by solid-state LiDAR is projected into the image coordinate system of a wide-angle camera to achieve spatiotemporal alignment between point cloud data and image pixels. A deep learning network based on the PointPainting architecture is used to process the spatiotemporally aligned fused data to achieve joint target recognition and semantic segmentation. The identified target is tracked across multiple frames to generate the 3D target with semantic labels and its motion trajectory.

[0009] Optionally, the step of predicting multiple probabilistic trajectories of a 3D target and their corresponding behavioral intentions using a spatiotemporal graph neural network model based on motion trajectory and scene context information includes: A spatiotemporal graph model is constructed using the motion trajectory and scene context information, where traffic participants are nodes and the interaction relationships between participants are edges. The historical trajectory data, motion state, and scene context information in the motion trajectory are used as input features; The input features are processed by a spatiotemporal graph neural network model, and the multiple probabilistic trajectories of each target and their corresponding behavioral intention probability distributions within a preset future time period are output.

[0010] Optionally, the calculation of the initial risk matrix based on the vehicle's preset driving path and multiple probabilistic trajectories includes: For each of the multiple probabilistic trajectories, calculate its collision time and collision probability with the vehicle's preset driving path; Based on the correspondence between collision time and collision probability, an initial risk matrix containing risk scores for each trajectory is constructed.

[0011] Optionally, the comprehensive collision risk level is determined based on the initial risk matrix, combined with corrections to the collision probability based on behavioral intent, roadside blind spot perception data, and cloud-based traffic information, including: The collision probabilities in the initial risk matrix are weighted and corrected according to the probability distribution of the stated behavioral intent; The global decision-making suggestions provided by the roadside blind spot perception data and the regional risk coefficients provided by the cloud-based traffic information are weighted and fused with the modified risk matrix; Based on the weighted fusion results, the comprehensive collision risk level is determined by a preset risk level mapping rule.

[0012] Optionally, the step of performing corresponding proactive safety response operations based on the comprehensive collision risk level includes: Based on the different ranges of the overall collision risk level, corresponding proactive safety response actions will be performed: When the overall collision risk level is in the first range, target visualization prompts are provided via AR-HUD; When the overall collision risk level is in the second range, activate the sound, light, and tactile level three alarm. When the overall collision risk level is in the third range, autonomous emergency braking is triggered and the roadside warning devices are activated.

[0013] Secondly, embodiments of this application provide an intelligent rear traffic crossing warning device, the device comprising: The data acquisition module is used to acquire vehicle-side multimodal perception data, roadside blind spot perception data, and cloud-based traffic information; The data-level pre-fusion module is used to perform data-level pre-fusion on the vehicle-side multimodal perception data to generate a three-dimensional target with semantic labels and its motion trajectory. The trajectory intent prediction module is used to predict multiple probabilistic trajectories of the three-dimensional target and their corresponding behavioral intents based on the motion trajectory and scene context information using a spatiotemporal graph neural network model. The risk matrix calculation module is used to calculate an initial risk matrix based on the vehicle's preset driving path and the multiple probabilistic trajectories, wherein the initial risk matrix represents the collision risk corresponding to each probabilistic trajectory. The risk level determination module is used to determine the comprehensive collision risk level based on the initial risk matrix, combined with the correction of the collision probability by the behavioral intention, the roadside blind spot perception data, and the cloud traffic information. The safety response operation execution module performs corresponding proactive safety response operations based on the comprehensive collision risk level.

[0014] Optionally, acquiring vehicle-side multimodal perception data, roadside blind spot perception data, and cloud-based traffic information includes: The vehicle-side multimodal perception data is acquired by solid-state LiDAR, wide-angle cameras, and short-range millimeter-wave radar deployed at the rear and sides of the vehicle. The blind spot sensing data is received from the roadside unit via V2I communication; The cloud-based traffic information, including real-time traffic flow, historical accident data, and high-precision maps, is obtained from the cloud-based traffic platform via an in-vehicle T-Box.

[0015] Optionally, the step of performing data-level pre-fusion of the vehicle-side multimodal perception data to generate a 3D target with semantic labels and its motion trajectory includes: The 3D point cloud data collected by solid-state LiDAR is projected into the image coordinate system of a wide-angle camera to achieve spatiotemporal alignment between point cloud data and image pixels. A deep learning network based on the PointPainting architecture is used to process the spatiotemporally aligned fused data to achieve joint target recognition and semantic segmentation. The identified target is tracked across multiple frames to generate the 3D target with semantic labels and its motion trajectory.

[0016] Optionally, the step of predicting multiple probabilistic trajectories of a 3D target and their corresponding behavioral intentions using a spatiotemporal graph neural network model based on motion trajectory and scene context information includes: A spatiotemporal graph model is constructed using the motion trajectory and scene context information, where traffic participants are nodes and the interaction relationships between participants are edges. The historical trajectory data, motion state, and scene context information in the motion trajectory are used as input features; The input features are processed by a spatiotemporal graph neural network model, and the multiple probabilistic trajectories of each target and their corresponding behavioral intention probability distributions within a preset future time period are output.

[0017] Optionally, the calculation of the initial risk matrix based on the vehicle's preset driving path and multiple probabilistic trajectories includes: For each of the multiple probabilistic trajectories, calculate its collision time and collision probability with the vehicle's preset driving path; Based on the correspondence between collision time and collision probability, an initial risk matrix containing risk scores for each trajectory is constructed.

[0018] Optionally, the comprehensive collision risk level is determined based on the initial risk matrix, combined with corrections to the collision probability based on behavioral intent, roadside blind spot perception data, and cloud-based traffic information, including: The collision probabilities in the initial risk matrix are weighted and corrected according to the probability distribution of the stated behavioral intent; The global decision-making suggestions provided by the roadside blind spot perception data and the regional risk coefficients provided by the cloud-based traffic information are weighted and fused with the modified risk matrix; Based on the weighted fusion results, the comprehensive collision risk level is determined by a preset risk level mapping rule.

[0019] Optionally, the step of performing corresponding proactive safety response operations based on the comprehensive collision risk level includes: Based on the different ranges of the overall collision risk level, corresponding proactive safety response actions will be performed: When the overall collision risk level is in the first range, target visualization prompts are provided via AR-HUD; When the overall collision risk level is in the second range, activate the sound, light, and tactile level three alarm. When the overall collision risk level is in the third range, autonomous emergency braking is triggered and the roadside warning devices are activated.

[0020] Thirdly, embodiments of this application provide a computer device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, the steps of the intelligent rear traffic crossing warning method described in any of the optional embodiments of the first aspect are performed.

[0021] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the intelligent rear traffic crossing warning method described in any of the optional embodiments of the first aspect.

[0022] The technical solution provided in this application includes, but is not limited to, the following beneficial effects: The steps of acquiring vehicle-side multimodal perception data, roadside blind spot perception data, and cloud-based traffic information can break through the information limitations of traditional single-vehicle perception. By integrating multi-source heterogeneous data, it can achieve comprehensive coverage of the traffic environment behind and to the sides of the vehicle, ensuring the integrity and richness of perception information and laying a data foundation for subsequent accurate analysis.

[0023] The step of performing data-level pre-fusion of multimodal perception data from the vehicle to generate 3D targets with semantic labels and their motion trajectories can give full play to the complementary advantages of different modal data, significantly improve the accuracy of target recognition and the stability of trajectory tracking, enable the system to clearly distinguish different types of traffic participants and their motion states, and reduce recognition bias caused by independent data processing.

[0024] Based on motion trajectory and scene context information, the spatiotemporal graph neural network model is used to predict multiple probabilistic trajectories of a three-dimensional target in the future and the steps of its corresponding behavioral intentions. This enables forward-looking judgment of the behavior of traffic participants, no longer limited to the perception of the current state, but also able to predict potential risks in advance, allowing drivers sufficient reaction time and transforming passive warning into proactive prevention.

[0025] The step of calculating the initial risk matrix based on the vehicle's preset driving path and multiple probabilistic trajectories can quantify the collision risk, clarify the risk level corresponding to different probabilistic trajectories, and provide a clear and quantifiable basis for subsequent risk level determination, avoiding the ambiguity and subjectivity of risk assessment.

[0026] Based on the initial risk matrix, and combined with the correction of collision probability by behavioral intent, roadside blind spot perception data and cloud traffic information, the steps to determine the comprehensive collision risk level can integrate multi-dimensional risk influencing factors, achieve a comprehensive and objective assessment of collision risk, make the risk level determination more in line with actual traffic scenarios, and avoid the one-sidedness of decision-making caused by a single information source.

[0027] Based on the comprehensive collision risk level, the corresponding active safety response steps can be executed to achieve precise matching of warning and intervention. Different intensity response measures can be taken according to the risk level to ensure effective protection in high-risk scenarios and avoid excessive intervention in low-risk scenarios, thereby improving the balance between driving experience and safety.

[0028] In summary, the progressive and synergistic effects of each step, from the comprehensiveness of data acquisition, the accuracy of perception fusion, the foresight of behavior prediction, the objectivity of risk assessment, to the adaptability of response operations, comprehensively improve the reliability, accuracy, and timeliness of rear traffic crossing warnings, effectively avoid the limitations of traditional warning methods, and provide stronger safety support for intelligent driving.

[0029] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0030] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0031] Figure 1 The flowchart of an intelligent rear traffic crossing early warning method provided in Embodiment 1 of this application is shown; Figure 2 A flowchart of a data acquisition method provided in Embodiment 1 of this application is shown; Figure 3 A flowchart of a data-level pre-fusion method provided in Embodiment 1 of this application is shown; Figure 4 A flowchart of a probabilistic trajectory and behavioral intent method provided in Embodiment 1 of this application is shown; Figure 5 A flowchart of an initial risk matrix calculation method provided in Embodiment 1 of this application is shown; Figure 6 A flowchart of a comprehensive collision risk level determination method provided in Embodiment 1 of this application is shown; Figure 7 This diagram illustrates the structure of an intelligent rear traffic crossing warning device provided in Embodiment 2 of this application; Figure 8 A schematic diagram of the structure of a computer device provided in Embodiment 3 of this application is shown. Detailed Implementation

[0032] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0033] Example 1 To facilitate understanding of this application, the following is combined with... Figure 1The flowchart of the intelligent rear traffic crossing early warning method provided in Embodiment 1 of this application illustrates Embodiment 1 in detail.

[0034] See Figure 1 As shown, Figure 1 A flowchart of an intelligent rear traffic crossing warning method provided in Embodiment 1 of this application is shown, wherein the method includes steps S101 to S106: S101: Acquire vehicle-side multimodal perception data, roadside blind spot perception data, and cloud-based traffic information.

[0035] Specifically, the multimodal perception data on the vehicle comes from three types of core sensors deployed at the rear and sides of the vehicle: solid-state LiDAR is responsible for collecting high-precision 3D point clouds, which can accurately obtain the geometric position of the target and is not affected by extreme weather such as light, rain, snow and fog; wide-angle high-definition camera provides high-resolution images, capturing semantic information such as target texture and color, making up for the lack of semantic recognition capability of LiDAR; short-range millimeter-wave radar mainly collects target speed data, and at the same time serves as a redundant backup to avoid perception interruption caused by the failure of a single sensor.

[0036] Blind spot perception data is acquired from the roadside unit (RSU) at the intersection via V2I (vehicle-to-infrastructure) communication. The RSU itself is also equipped with LiDAR and cameras, which can cover the blind spots formed by the vehicle's sensors due to physical obstruction (such as large trucks parked next to it), providing perception information from a global perspective of the intersection, and solving the pain point of traditional single-vehicle perception that "cannot see" blind spot targets.

[0037] The cloud-based traffic information is obtained from the cloud-based traffic platform through the vehicle-mounted T-Box (Telematics Processor) and includes three types of key data: real-time traffic flow information reflects the current regional congestion status (such as traffic density during peak hours), historical accident data is used to mark accident-prone areas (such as school gates and intersections), and high-precision maps provide accurate geographic coordinate references for subsequent route planning and risk calculation. Together, these three provide information support for decision-making based on "historical experience + real-time global perspective".

[0038] S102: Perform data-level pre-fusion on the vehicle-side multimodal perception data to generate a three-dimensional target with semantic labels and its motion trajectory.

[0039] Specifically, data-level pre-fusion is a core innovation that distinguishes it from traditional "post-fusion" and "feature-level fusion"—in traditional solutions, each sensor first independently identifies the target and then summarizes the results, which cannot fully leverage the complementarity of the data. However, this invention fuses at the raw data level: first, it aligns the 3D point cloud of the LiDAR with the image pixels of the camera in time and space, and then processes them together, fundamentally solving the problems of low perception accuracy and many false alarms and missed alarms caused by the "each sweeping its own door" approach in traditional fusion.

[0040] The fusion processing relies on a deep learning network based on the PointPainting architecture: this network can assign semantic information of the image (such as category labels for "pedestrian", "electric vehicle", and "roadblock") to the corresponding point cloud data, realizing joint recognition and semantic segmentation of targets. It can not only determine the location of the target, but also accurately distinguish between "pedestrians crossing the lane" and "pedestrians standing on the roadside", "moving delivery vehicles" and "stationary road bollards", which greatly reduces false alarms for non-threatening targets such as ground locks and bushes, as well as false alarms for stationary obstacles.

[0041] When generating motion trajectories, the identified targets are continuously tracked through a multi-frame data association algorithm: noise interference from single-frame data (such as image noise from rainy cameras and sporadic noise from lidar) is eliminated to form a continuous and stable trajectory, ensuring the accuracy of subsequent predictions and risk calculations, and avoiding misjudgments caused by fluctuations in single-frame data.

[0042] S103: Based on the motion trajectory and scene context information, use a spatiotemporal graph neural network model to predict multiple probabilistic trajectories of the three-dimensional target in the future and their corresponding behavioral intentions.

[0043] Specifically, the contextual information includes two key categories: first, static environmental information, such as the distance of the target from the stop line at the intersection, the current traffic light status (red / green), and the lane type (motor vehicle lane / pedestrian lane); second, dynamic interactive information, such as the relative position of the target with other traffic participants (e.g., the distance between an electric vehicle and a pedestrian). This information is an important basis for judging the target's behavioral intentions.

[0044] The core of the Spatiotemporal Graph Neural Network (ST-GNN) is to construct a dynamic traffic scene graph: each traffic participant (such as pedestrians, electric vehicles, and vehicles) is treated as a "node", and the node attributes include the target's historical trajectory, speed, and acceleration; the interaction relationships between participants (such as following, avoiding, and crossing) are treated as "edges", and the weight of the edges is dynamically adjusted according to the interaction intensity, thereby accurately modeling the complex interaction relationships in the traffic scene - this is different from the limitations of traditional physical models (such as constant velocity models) that cannot consider the interaction between targets.

[0045] The prediction output consists of two parts: first, multiple probabilistic trajectories within the next 5 seconds, covering the possible movement paths of the target (such as an electric vehicle that may go straight or cross the lane, or slow down to avoid it), with each trajectory corresponding to a probability of occurrence; second, the probability distribution of behavioral intentions, such as "80% probability of continuing to cross the lane and 20% probability of slowing down and stopping", providing a forward-looking basis for subsequent risk assessment and transforming the traditional "post-event response" into "proactive early warning".

[0046] S104: Based on the vehicle's preset driving path and the multiple probabilistic trajectories, calculate an initial risk matrix, wherein the initial risk matrix represents the collision risk corresponding to each probabilistic trajectory.

[0047] Specifically, the vehicle's preset driving path is determined based on the actual scenario, such as the reversing path when reversing into a parking space and the turning path when turning. This is calculated in real-time by the onboard system using parameters such as steering wheel angle and vehicle speed to ensure it matches the actual driving intention. Figure 1 To.

[0048] The initial risk matrix calculation includes three core indicators: first, the probability of collision (P), which is calculated based on the degree of intersection between the target trajectory and the vehicle's path, and the target's motion state (such as speed). The closer the intersection and the faster the speed, the higher the probability; second, the time to collision (TTC), which is obtained by dividing the relative distance between the two by the relative speed. The shorter the TTC, the higher the urgency; and third, the severity of collision (S), which is determined based on the target type (such as pedestrian collision severity > electric vehicle > motor vehicle, stationary target > moving target). These three indicators together form the basis of risk assessment.

[0049] The initial risk matrix is ​​presented in tabular form, with each row corresponding to a probabilistic trajectory and each column corresponding to the collision probability, collision time, severity, and comprehensive risk score (the three indicators are quantified into a score of 0-1 through a preset algorithm), which intuitively reflects the risk differences of different trajectories and provides an initial risk basis for subsequent collaborative decision-making.

[0050] S105: Based on the initial risk matrix, combined with the correction of the collision probability by the behavioral intent, the roadside blind spot perception data, and the cloud-based traffic information, a comprehensive collision risk level is determined.

[0051] Specifically, the correction of collision probability based on behavioral intent follows the "probability weighting" principle: if the probability of "high-risk behavior" (such as crossing the lane) in the target's behavioral intent is high (such as 90%), the collision probability of the corresponding trajectory is corrected upward (such as from 0.6 to 0.8); if the probability of "low-risk behavior" (such as slowing down to avoid a collision) is high (such as 80%), the collision probability is corrected downward (such as from 0.6 to 0.3), ensuring that the collision probability matches the actual behavioral trend of the target.

[0052] The role of roadside blind spot perception data is to provide "global perspective correction": for example, if the roadside RSU detects a blind spot target that the vehicle has not perceived (such as a pedestrian obscured by a large truck), it will add the risk information of the target to the decision; if the roadside provides a global decision suggestion such as "a vehicle is approaching at high speed on the left, it is recommended to wait", the system will increase the risk weight of the corresponding direction to avoid the one-sidedness of a single vehicle's perception.

[0053] The role of cloud-based traffic information is to provide "regional risk correction": if the cloud marks the current area as a "school area" (where children frequently pass through) or "peak hours" (where traffic is dense), it will output a higher regional risk coefficient (such as 1.2 times). The system will then weight and fuse this coefficient with the corrected risk matrix to improve the sensitivity of the warning. If the area is low-risk (such as an empty parking lot), the coefficient will be reduced (such as 0.8 times) to avoid over-warning.

[0054] Specifically, the comprehensive collision risk level is determined through "weighted fusion + threshold mapping": First, the corrected collision probability, roadside decision weight, and cloud-based risk coefficient are weighted according to a preset ratio (e.g., 4:3:3) to obtain the comprehensive collision risk coefficient (CCR, value 0-1); then, the level is mapped according to the range of CCR, such as CCR<0.3 for low risk, 0.3≤CCR<0.7 for medium risk, 0.7≤CCR<0.9 for high risk, and CCR≥0.9 for extremely high risk.

[0055] S106: Based on the comprehensive collision risk level, execute the corresponding active safety response operation.

[0056] Specifically, the response operation follows the "risk matching" principle to avoid low-risk high response causing driver resentment, or high-risk low response missing the opportunity to avoid danger. It also covers the dual dimensions of "in-vehicle warning + external coordination" to build a comprehensive protection system - which is different from the traditional single protection mode that only relies on in-vehicle alarms.

[0057] Different risk levels correspond to different responses: Low risk (CCR<0.3) only uses AR-HUD (Augmented Reality Head-Up Display) for visual prompts, overlaying the target and trajectory onto the real field of vision without sound or light alarms, reducing driving interference; Medium risk (0.3≤CCR<0.7) activates AR prompts with a yellow box and trajectory line to help the driver quickly locate the risky target; High risk (0.7≤CCR<0.9) triggers a three-level alarm system with sound, light, and touch, namely, a flashing orange prompt (visual), a high-frequency alarm sound (auditory), and seat vibration (tactile), comprehensively waking up the driver's attention; Very high risk (CCR≥0.9) in addition to the three-level alarm, directly triggers autonomous emergency braking (AEB), forcing the vehicle to slow down or stop, and simultaneously links with roadside warning devices (such as smart warning signs flashing the words "Caution: Avoid") through V2I to remind road users outside the vehicle (such as electric vehicle drivers) to avoid the risk.

[0058] In an optional implementation, see Figure 2 As shown, Figure 2 The flowchart of a data acquisition method provided in Embodiment 1 of this application is shown, wherein the acquisition of vehicle-side multimodal perception data, roadside blind spot perception data, and cloud-based traffic information includes steps S201-S203: S201: Acquire the vehicle-side multimodal perception data by deploying solid-state LiDAR, wide-angle cameras, and short-range millimeter-wave radar at the rear and sides of the vehicle.

[0059] Specifically, the sensor deployment locations have been optimized: the rear sensor (1 LiDAR + 1 wide-angle camera + 1 millimeter-wave radar) mainly covers the area directly behind, while the side sensors (1 LiDAR + 1 wide-angle camera each) cover the left and right rear areas respectively. The three work together to achieve 360° rear-view perception without blind spots, solving the blind spot problem of incomplete coverage by traditional single-side sensors.

[0060] The functional complementarity of the three types of sensors is reflected in extreme scenarios: such as rainy nights, when cameras suffer from low light and water fog on the lens, resulting in more image noise and reduced recognition accuracy, LiDAR is not affected by light and weather and can still accurately collect point cloud data; if LiDAR has missing point cloud data due to obstruction, millimeter-wave radar can use velocity data to help confirm the existence of the target, ensuring that the system can still maintain basic perception capabilities when a single sensor fails.

[0061] Sensor parameter selection to meet early warning requirements: Solid-state lidar with a detection range ≥50m and angular resolution ≤0.1° ensures accurate long-range capture of small targets (such as pedestrians); wide-angle camera with a field of view ≥120° and resolution ≥1080P ensures wide coverage and clear details; short-range millimeter-wave radar with a speed measurement range of 0-100km / h and a ranging range of 0.5-30m is suitable for speed monitoring of short-to-medium distance targets in the rear.

[0062] S202: Receive the blind spot perception data collected from the roadside unit via V2I communication.

[0063] Specifically, the deployment scenarios for roadside units (RSUs) focus on high-risk areas, such as intersections, parking lot entrances and exits, and entrances to residential areas where blind spots are likely to occur. Each RSU has a perception range of ≥100m, which can cover traffic participants in multiple lanes and avoid missed detections caused by "view of sight obstruction" (such as buildings or large vehicles) when a single vehicle perceives the traffic.

[0064] The data transmitted for blind spot perception includes: the type of the target (pedestrian / electric vehicle / vehicle), real-time location (latitude and longitude coordinates), speed, acceleration, and the relative distance and relative speed between the target and the vehicle. The data transmission delay is ≤100ms to ensure the real-time nature of the information. If the delay is too high, it may lead to delayed warnings and missed opportunities for hazard avoidance.

[0065] V2I communication uses automotive-grade protocols (such as LTE-V2X or 5G-V2X) and has anti-interference capabilities: even in urban environments with complex signals, it can ensure the stability of data transmission and avoid the loss of blind spot information due to communication interruption. This is the key foundation for realizing "beyond line of sight early warning".

[0066] S203: Obtain the cloud-based traffic information, including real-time traffic flow, historical accident data, and high-precision maps, from the cloud-based traffic platform via the in-vehicle T-Box.

[0067] Specifically, the real-time traffic flow information is updated every 1-5 minutes, including the average vehicle speed, traffic density, and lane occupancy rate in the current area. If the average vehicle speed in a lane is below 20 km / h (congested state), the system will predict that the target may slow down and correct the collision risk. If the traffic density is high, the warning sensitivity will be improved to avoid the risk of being missed due to multiple targets intertwined.

[0068] Historical accident data includes the time, location, and cause of accidents in the past year (such as "reversing into a pedestrian at an intersection" or "electric vehicle crossing the road on a rainy night"). The system will match historical accident characteristics with the current scenario (such as rainy night or reversing at an intersection). If the matching degree is high, the regional risk coefficient will be automatically increased, such as from 1.0 to 1.3, to enhance the targeting of the warning.

[0069] High-precision maps achieve centimeter-level accuracy and include information such as lane line positions, roadside obstacles (e.g., curbs, guardrails), and traffic signs (e.g., no parking, speed limits). When calculating the vehicle's preset driving path, the map accurately matches the lane lines, avoiding path calculation deviations caused by map errors, which in turn affect the accuracy of the risk matrix.

[0070] In an optional implementation, see Figure 3 As shown, Figure 3 The flowchart of a data-level pre-fusion method provided in Embodiment 1 of this application is shown. The step of performing data-level pre-fusion on vehicle-side multimodal perception data to generate a 3D target with semantic labels and its motion trajectory includes steps S301-S303: S301: Projects the 3D point cloud data collected by the solid-state lidar onto the image coordinate system of the wide-angle camera to achieve spatiotemporal alignment between the point cloud data and the image pixels.

[0071] Specifically, spatiotemporal alignment comprises two parts: spatial alignment and temporal alignment. Spatial alignment converts the X / Y / Z coordinates of the 3D point cloud into pixel coordinates (u / v) of the image by calibrating parameters (such as camera intrinsic parameters and LiDAR and camera extrinsic parameters), ensuring the correct pixel position of each point cloud in the corresponding image. Temporal alignment eliminates the time difference caused by the difference in acquisition frequency between the two types of sensors (such as 10Hz for LiDAR and 30Hz for the camera) by synchronizing clock signals, thus avoiding fusion deviations caused by data "asynchronization".

[0072] Aligned point cloud data will include color and texture information of the image: for example, if a LiDAR captures a "point target", after alignment, the color (such as red clothing) and texture (such as human body outline) of the image pixels can be used to determine that the target is a "pedestrian" rather than a "road post", providing a key basis for subsequent semantic segmentation - this solves the pain point of traditional LiDAR that can only acquire geometric information and cannot identify the target type.

[0073] S302: Use a deep learning network based on the PointPainting architecture to process the spatiotemporally aligned fused data to achieve joint target recognition and semantic segmentation.

[0074] Specifically, the PointPainting architecture's processing flow consists of three steps: First, semantic segmentation is performed on the camera image, outputting a category label for each pixel (such as "pedestrian", "vehicle", "background"); second, the pixel labels are "mapped" to the corresponding point cloud data, that is, each point cloud is given a semantic label ("point cloud coloring"); third, the labeled point cloud is input into the point cloud segmentation network to further optimize the target boundary and ensure segmentation accuracy.

[0075] This network can solve two major problems of traditional solutions: First, it can distinguish between "static threat targets" and "non-threat static targets". For example, it can identify "delivery vehicles that have just stopped" (threat targets) and "fixed road bollards" (non-threat targets), avoiding missed detections caused by traditional radar filtering all static point clouds; Second, it can distinguish between "dynamic small targets" and "noise". For example, it can identify "electric vehicles" (dynamic small targets) on rainy nights and "image noise" from cameras, reducing the false alarm rate.

[0076] The network training data covers a variety of extreme scenarios, including rain, snow, fog, nighttime backlight, and strong light, to ensure that the model can still work stably in complex environments. This is the key to improving the robustness of perception. Traditional models are prone to failure in extreme scenarios due to the limited training data.

[0077] S303: Perform multi-frame tracking on the identified target to generate the 3D target with semantic labels and its motion trajectory.

[0078] Specifically, multi-frame tracking uses a "Kalman filter + data association" algorithm: Kalman filter predicts the possible position of the target in the next frame based on the target's current position and velocity; data association matches the target identified in the next frame with the predicted position to determine whether they are the same target, thus avoiding target "jumping" or "losing".

[0079] During the tracking process, the target attributes are continuously updated: in addition to position, speed, and trajectory, the target's size (such as pedestrian height, vehicle length) and motion state (such as acceleration / deceleration / constant speed) are also updated. These attributes provide richer features for subsequent behavior prediction. For example, "small size + constant speed" may indicate a pedestrian, while "large size + acceleration" may indicate a motor vehicle.

[0080] After the trajectory is generated, "outlier removal" will be performed: if the target position in a certain frame deviates too much from the historical trajectory (such as a sudden shift caused by noise), the system will determine it as an outlier and use interpolation to correct it, ensuring the smoothness and continuity of the trajectory. If the trajectory fluctuates too much, it will lead to deviations in subsequent prediction results and affect the accuracy of risk assessment.

[0081] In an optional implementation, see Figure 4 As shown, Figure 4 The flowchart illustrates a probabilistic trajectory and behavioral intent method provided in Embodiment 1 of this application. The method involves predicting multiple probabilistic trajectories of a 3D target and their corresponding behavioral intents using a spatiotemporal graph neural network model based on motion trajectory and scene context information, including steps S401-S403: S401: Construct a spatiotemporal graph model using the motion trajectory and scene context information, where traffic participants are nodes and the interaction relationships between participants are edges.

[0082] Specifically, the nodes have rich feature dimensions: in addition to historical data of motion trajectory (such as the position sequence of the past 3 seconds), they also include the static attributes (type, size) and dynamic attributes (velocity, acceleration, heading angle) of the target, as well as scene context features (such as distance from the stop line, whether it is on the pedestrian crossing). These multi-dimensional features ensure the accuracy of node representation.

[0083] The construction of edges is based on the type of interaction relationship: if two targets have a "follow-the-car" relationship (such as a vehicle following a vehicle in front), the edge type is marked as "follow-the-car", and the weight is set according to the distance (the closer the distance, the higher the weight); if there is a "crossing" relationship (such as an electric vehicle crossing the path of the vehicle itself), the edge type is marked as "crossing", and the weight is set according to the crossing angle (vertical crossing has a higher weight than diagonal crossing); if there is no direct interaction, the edge weight is set to 0, dynamically reflecting the interaction state of the targets in the scene.

[0084] The spatiotemporal graph model has the ability to "update in time": after each frame of data is collected, the node features (such as speed changes) and edge types / weights (such as changes in interaction relationships) are recalculated to ensure that the model can adapt to the dynamic changes of the scene in real time and avoid the limitations of static graph models in reflecting real-time interactions.

[0085] S402: Use the historical trajectory data, motion state, and scene context information in the motion trajectory as input features.

[0086] Specifically, the time window for historical trajectory data is set to 3-5 seconds: if the time window is too short, the features will be incomplete due to insufficient data (such as the inability to determine whether the target is accelerating); if the time window is too long, redundant data will be introduced (such as the target's early uniform speed state being meaningless for the current prediction). 3-5 seconds is the optimal range that balances completeness and efficiency.

[0087] Motion state characteristics include instantaneous state and trend of change: instantaneous state such as current speed and acceleration; trend of change such as rate of change of speed (how fast to accelerate / decelerate) and rate of change of heading angle (how fast to turn). Trend characteristics can reflect the target's behavioral tendencies. For example, "positive rate of change of speed (acceleration) + stable heading angle" may indicate that the target will cross quickly.

[0088] The scene context information is "feature-encoded": such as traffic light status (red light is coded as 1, green light as 0), lane type (pedestrian lane is coded as 1, motor vehicle lane as 0), and distance from the stop line (normalized to a value between 0 and 1). The encoded features are easy for neural networks to process, while ensuring that different types of context information can be compared and fused.

[0089] S403: The input features are processed by a spatiotemporal graph neural network model, and the multiple probabilistic trajectories of each target and their corresponding behavioral intention probability distributions within a preset future time period are output.

[0090] Specifically, the model structure includes "graph convolutional layers" and "temporal network layers": the graph convolutional layers are responsible for extracting the interaction features between targets (such as the interaction effect of "electric vehicle slowing down because it sees a pedestrian"), and updating its own node features by aggregating the features of neighboring nodes; the temporal network layers (such as LSTM or Transformer) are responsible for extracting the temporal features of the trajectory (such as the trend of "electric vehicle continuously accelerating in the past 2 seconds"). The two are combined to achieve joint modeling of "spatial interaction + temporal trend".

[0091] The preset time period is set to 5 seconds: this duration is the optimal value obtained by taking into account the driver's reaction time (about 1.5-2 seconds) and the vehicle braking distance (about 2-3 seconds). If the duration is too short (such as 2 seconds), there will be insufficient reaction time for the driver; if the duration is too long (such as 10 seconds), the prediction accuracy will drop significantly (due to uncontrollable changes in traffic scenarios).

[0092] There is a "correspondence" between probabilistic trajectories and behavioral intentions: each trajectory corresponds to a behavioral intention, such as "Trajectory 1: Straight ahead and cross the lane" corresponding to "Behavioral intention". Figure 1 "Continue forward (probability 80%)", "Track 2: Slow down and stop at the roadside" corresponds to "behavioral intention". Figure 2 "Slow down and avoid (probability 20%)", the probability distributions of the two are consistent, ensuring the consistency of the output results and avoiding contradictions in subsequent decisions.

[0093] In an optional implementation, see Figure 5 As shown, Figure 5 The flowchart of an initial risk matrix calculation method provided in Embodiment 1 of this application is shown. The calculation of the initial risk matrix based on the vehicle's preset driving path and multiple probabilistic trajectories includes steps S501-S502: S501: For each of the multiple probabilistic trajectories, calculate its collision time and collision probability with the vehicle's preset driving path.

[0094] Specifically, the calculation of Time of Collision (TTC) needs to take into account "relative motion": if the vehicle is reversing (speed backward) and the target is moving forward and crossing, the relative speed between the two is "absolute value of the vehicle's speed + target speed"; if the vehicle is moving forward and the target is moving in the same direction, the relative speed is "vehicle speed - target speed" (a positive difference indicates that the vehicle is faster than the target). The direction and magnitude of the relative speed directly affect the calculation result of TTC to ensure that it conforms to the actual motion scenario.

[0095] The collision probability (P) is calculated using a Gaussian probability model: the vehicle's path and the target trajectory are each considered as a "probability distribution interval" (considering positional errors), and the area ratio of the overlapping region is the base probability; then, the probability of the target's behavioral intention (such as "80% probability of moving forward") is weighted to obtain the final collision probability—for example, if the base probability is 0.7 and the behavioral intention probability is 0.8, then the final collision probability is 0.7 × 0.8 = 0.56, ensuring the scientific nature of the probability calculation.

[0096] Specifically, during the calculation process, "impossible collision" trajectories are excluded: if the target trajectory does not intersect with the vehicle's path (such as the target moving away from the vehicle), or the TTC at the time of intersection is greater than 10 seconds (extremely low risk), then its collision probability is set to 0, so as to avoid these trajectories occupying decision resources and improve calculation efficiency.

[0097] S502: Based on the correspondence between the collision time and the collision probability, construct the initial risk matrix containing the risk scores of each trajectory.

[0098] Specifically, the risk score is calculated using a weighted summation formula: Risk Score = (1 - TTC / 10) × 0.4 + Collision Probability × 0.4 + Collision Severity × 0.2, where "1 - TTC / 10" is a value that normalizes TTC (0-10 seconds) to 0-1 (the shorter the TTC, the larger the value); the weights are set according to the degree of risk impact, with collision time and collision probability having a greater impact on urgency (each accounting for 40%), followed by severity (20%).

[0099] The initial risk matrix includes columns for "trajectory number, target type, collision time (TTC), collision probability (P), collision severity (S), and risk score". Each row corresponds to a probabilistic trajectory, and the trajectories are sorted from high to low risk score. This allows for prioritizing high-risk trajectories in subsequent decision-making, such as prioritizing trajectories with a risk score of 0.8 over those with a score of 0.3, thus improving decision-making efficiency.

[0100] The matrix will indicate "Preliminary Risk Level Assessment": Based on the risk score, a preliminary threshold is preset (e.g., a score ≥ 0.7 indicates high risk, 0.4 ≤ score < 0.7 indicates medium risk, and a score < 0.4 indicates low risk). This provides a basic reference for subsequent comprehensive decision-making combining roadside and cloud information, avoiding starting the assessment from scratch and shortening the decision-making time.

[0101] In an optional implementation, see Figure 6 As shown, Figure 6 The flowchart of a comprehensive collision risk level determination method provided in Embodiment 1 of this application is shown. The method involves determining the comprehensive collision risk level based on an initial risk matrix, combined with corrections to the collision probability based on behavioral intent, roadside blind spot perception data, and cloud-based traffic information. The steps include S601-S603: S601: The collision probabilities in the initial risk matrix are weighted and corrected according to the probability distribution of the stated behavioral intent.

[0102] Specifically, the correction logic follows the principle of "weighted increase for high-risk behaviors and weighted decrease for low-risk behaviors": If the sum of probabilities of "high-risk behaviors" (such as crossing or accelerating) in the target behavior intention is P_high, and the sum of probabilities of "low-risk behaviors" (such as decelerating or stopping) is P_low, then the corrected collision probability = initial collision probability × (P_high + 0.5 × P_low) — for example, if the initial probability is 0.6, P_high = 0.8, and P_low = 0.2, then the corrected probability = 0.6 × (0.8 + 0.5 × 0.2) = 0.6 × 0.9 = 0.54, which takes into account both the dominant role of high-risk behaviors and the potential impact of low-risk behaviors.

[0103] The correction process will target "multi-trajectory integration": If a target has 3 probabilistic trajectories (trajectory 1: risk score 0.8, probability 0.7; trajectory 2: score 0.5, probability 0.2; trajectory 3: score 0.2, probability 0.1), then the collision probability of each trajectory will be corrected separately, and then the "weighted risk score" will be calculated as follows: (trajectory 1 score × trajectory 1 probability) + (trajectory 2 score × trajectory 2 probability) + (trajectory 3 score × trajectory 3 probability) = 0.8 × 0.7 + 0.5 × 0.2 + 0.2 × 0.1 = 0.56 + 0.1 + 0.02 = 0.68, avoiding the one-sidedness of single trajectory analysis.

[0104] S602: The global decision-making suggestions provided by the roadside blind spot perception data and the regional risk coefficients provided by the cloud traffic information are weighted and fused with the corrected risk matrix.

[0105] Specifically, the weight of roadside global decision-making suggestions is set higher than that of local information: the roadside provides a global perspective (such as "a pedestrian is approaching in the blind spot"), with a wider information coverage and higher accuracy, so its weight is set to 0.3; the weight of the locally corrected risk score is set to 0.4 to ensure the dominant position of local core perception information; the weight of the cloud-based regional risk coefficient is set to 0.3 to supplement the regional historical and real-time risk characteristics. The total weight of the three is 1 to ensure the rationality of the fusion result.

[0106] The quantification method for global decision-making recommendations is as follows: "Stop immediately" corresponds to a quantification value of 1.0, "Proceed with caution" corresponds to 0.7, "Proceed normally" corresponds to 0.3, and "No risk" corresponds to 0.0. The regional risk coefficient is set based on cloud data: 1.2 for accident-prone areas, school areas, and peak hours, 1.0 for ordinary areas and off-peak hours, and 0.8 for low-risk areas (such as closed parking lots). The quantified information facilitates weighted calculation.

[0107] The weighted fusion formula is: Comprehensive Risk Coefficient (CCR) = Locally Corrected Risk Score × 0.4 + Roadside Decision Recommendation Quantitative Value × 0.3 + Cloud Area Risk Coefficient × 0.3. For example, if the local score is 0.6, the roadside recommendation is 0.7, and the cloud coefficient is 1.2, then CCR = 0.6 × 0.4 + 0.7 × 0.3 + 1.2 × 0.3 = 0.24 + 0.21 + 0.36 = 0.81, which intuitively reflects the comprehensive risk level.

[0108] S603: Based on the weighted fusion result, the comprehensive collision risk level is determined by a preset risk level mapping rule.

[0109] Specifically, the preset mapping rules refer to the risk tolerance of actual application scenarios: CCR<0.3 corresponds to "low risk", at which point there is no direct collision risk between the target and the vehicle, or the risk is extremely low, and no strong warning is required; 0.3≤CCR<0.7 corresponds to "medium risk", there is a potential collision possibility, but there is sufficient time to remind the driver; 0.7≤CCR<0.9 corresponds to "high risk", the collision risk is relatively high, and the driver needs to be reminded to take measures urgently; CCR≥0.9 corresponds to "extremely high risk", the probability of collision is extremely high, and the driver's reaction alone cannot avoid the danger, and the system needs to actively intervene.

[0110] The mapping rules are dynamically adjusted based on the target type: if the target is a pedestrian or non-motorized vehicle (high collision severity), the risk level threshold is adjusted downward (e.g., CCR≥0.8 is considered extremely high risk); if the target is a motorized vehicle (with protective measures, low severity), the threshold is adjusted upward (e.g., CCR≥0.95 is considered extremely high risk), ensuring that the risk level matches the collision consequences and avoiding the problem of "pedestrian risk being underestimated under the same CCR".

[0111] Once the risk level is determined, a "risk cause label" will be generated. For example, the cause label for CCR=0.81 (high risk) is "local perception of electric vehicle crossing (score 0.6) + roadside suggestion to pass with caution (0.7) + cloud-marked accident high-incidence area (1.2)", which facilitates subsequent system traceability and optimization, and provides risk judgment basis to maintenance personnel in diagnostic mode.

[0112] In an optional implementation, the step of performing corresponding proactive safety response operations based on the comprehensive collision risk level includes: Based on the different ranges of the overall collision risk level, corresponding proactive safety response actions will be performed: When the overall collision risk level is in the first range, target visualization prompts are provided via AR-HUD.

[0113] Specifically, the first interval corresponds to CCR<0.3 (low risk). The AR-HUD's visual prompts adopt a "weak marking" design: targets are marked with light green boxes, and trajectories are drawn with thin lines. The color and line intensity are low to avoid obstructing the driver's view. The prompts only include the target type (such as "pedestrian" or "electric vehicle") and relative distance (such as "50m"), without displaying additional alarm information to ensure driver focus.

[0114] This response mode is suitable for open scenarios (such as reversing on a suburban road) or scenarios where the target is far away from the vehicle: for example, when the vehicle is reversing in a parking lot and there is a pedestrian walking slowly 50m behind, it is only necessary to inform the driver that the target exists, without interfering with driving, which conforms to the principle of "minimal intervention".

[0115] When the overall collision risk level is in the second range, a level three alarm with sound, light, and touch will be activated.

[0116] Specifically, the second range includes "medium risk (0.3≤CCR<0.7)" and "high risk (0.7≤CCR<0.9)". The three levels of alarm intensity differ between the two: for medium risk, the visual cue is a yellow box with a medium-thickness trajectory line, the auditory cue is a low-frequency alarm sound (500Hz, 1-second interval), and the tactile cue is slight seat vibration (1 time / second); for high risk, the visual cue is an orange flashing box (2 times / second) with a thick trajectory line, the auditory cue is a high-frequency alarm sound (800Hz, 0.5-second interval), and the tactile cue is strong seat vibration (2 times / second). The intensity increases with the risk level.

[0117] The alarm information will be "associated with the target location": the warning box on the AR-HUD will move with the target to ensure that the driver can quickly locate the source of the risk; the seat vibration will distinguish the direction, such as the left seat vibrating if there is a risk to the left rear, and the right seat vibrating if there is a risk to the right rear, to help the driver quickly judge the location of the risk and shorten the reaction time.

[0118] When the overall collision risk level is in the third range, autonomous emergency braking is triggered and the roadside warning devices are activated.

[0119] Specifically, the third interval corresponds to a CCR ≥ 0.9 (extremely high risk). The intervention intensity of autonomous emergency braking (AEB) is adjusted according to the vehicle speed: if the vehicle speed is ≤ 10 km / h (such as low-speed reversing), AEB adopts "gentle braking" with a braking deceleration of ≤ 3 m / s² to avoid sudden braking that may cause passenger discomfort; if the speed is > 10 km / h, "emergency braking" is adopted with a deceleration of ≥ 5 m / s² to ensure rapid stopping, while simultaneously disengaging the accelerator pedal to prevent driver misoperation.

[0120] The response of the roadside warning equipment includes two types: First, the roadside smart warning sign will flash red lights and display text such as "Caution: Avoid" and "Vehicle braking ahead" to remind targets outside the vehicle (such as electric vehicle drivers and pedestrians) to slow down or avoid it; Second, the roadside RSU will broadcast the message "This vehicle is braking urgently" to other vehicles in the vicinity. If the surrounding vehicles are also equipped with V2X function, they can take deceleration measures in advance to avoid chain collisions.

[0121] If the vehicle is reversing at night in the rain and a large truck on the left creates a blind spot, the RSU detects the approaching electric vehicle (CCR=0.95), the system triggers AEB to stop the vehicle, and at the same time the roadside warning sign flashes. After seeing the warning, the electric vehicle driver slows down and successfully avoids a collision, demonstrating the comprehensive protection value of "in-vehicle active braking + external cooperative warning".

[0122] Example 2 See Figure 7 As shown, Figure 7 The diagram shows a structural schematic of an intelligent rear traffic crossing warning device according to Embodiment 2 of this application, wherein the device includes: The data acquisition module 701 is used to acquire vehicle-side multimodal perception data, roadside blind spot perception data, and cloud-based traffic information; The data-level pre-fusion module 702 is used to perform data-level pre-fusion on the vehicle-side multimodal perception data to generate a three-dimensional target with semantic labels and its motion trajectory. The trajectory intent prediction module 703 is used to predict multiple probabilistic trajectories of the three-dimensional target and their corresponding behavioral intents based on the motion trajectory and scene context information using a spatiotemporal graph neural network model. The risk matrix calculation module 704 is used to calculate an initial risk matrix based on the vehicle's preset driving path and the multiple probabilistic trajectories, wherein the initial risk matrix represents the collision risk corresponding to each probabilistic trajectory. The risk level determination module 705 is used to determine the comprehensive collision risk level based on the initial risk matrix, combined with the correction of the collision probability by the behavioral intention, the roadside blind spot perception data and the cloud traffic information; The safety response operation execution module 706 performs corresponding active safety response operations based on the comprehensive collision risk level.

[0123] In an optional implementation, acquiring vehicle-side multimodal perception data, roadside blind spot perception data, and cloud-based traffic information includes: The vehicle-side multimodal perception data is acquired by solid-state LiDAR, wide-angle cameras, and short-range millimeter-wave radar deployed at the rear and sides of the vehicle. The blind spot sensing data is received from the roadside unit via V2I communication; The cloud-based traffic information, including real-time traffic flow, historical accident data, and high-precision maps, is obtained from the cloud-based traffic platform via an in-vehicle T-Box.

[0124] Optionally, the step of performing data-level pre-fusion of the vehicle-side multimodal perception data to generate a 3D target with semantic labels and its motion trajectory includes: The 3D point cloud data collected by solid-state LiDAR is projected into the image coordinate system of a wide-angle camera to achieve spatiotemporal alignment between point cloud data and image pixels. A deep learning network based on the PointPainting architecture is used to process the spatiotemporally aligned fused data to achieve joint target recognition and semantic segmentation. The identified target is tracked across multiple frames to generate the 3D target with semantic labels and its motion trajectory.

[0125] Optionally, the step of predicting multiple probabilistic trajectories of a 3D target and their corresponding behavioral intentions using a spatiotemporal graph neural network model based on motion trajectory and scene context information includes: A spatiotemporal graph model is constructed using the motion trajectory and scene context information, where traffic participants are nodes and the interaction relationships between participants are edges. The historical trajectory data, motion state, and scene context information in the motion trajectory are used as input features; The input features are processed by a spatiotemporal graph neural network model, and the multiple probabilistic trajectories of each target and their corresponding behavioral intention probability distributions within a preset future time period are output.

[0126] Optionally, the calculation of the initial risk matrix based on the vehicle's preset driving path and multiple probabilistic trajectories includes: For each of the multiple probabilistic trajectories, calculate its collision time and collision probability with the vehicle's preset driving path; Based on the correspondence between collision time and collision probability, an initial risk matrix containing risk scores for each trajectory is constructed.

[0127] Optionally, the comprehensive collision risk level is determined based on the initial risk matrix, combined with corrections to the collision probability based on behavioral intent, roadside blind spot perception data, and cloud-based traffic information, including: The collision probabilities in the initial risk matrix are weighted and corrected according to the probability distribution of the stated behavioral intent; The global decision-making suggestions provided by the roadside blind spot perception data and the regional risk coefficients provided by the cloud-based traffic information are weighted and fused with the modified risk matrix; Based on the weighted fusion results, the comprehensive collision risk level is determined by a preset risk level mapping rule.

[0128] Optionally, the step of performing corresponding proactive safety response operations based on the comprehensive collision risk level includes: Based on the different ranges of the overall collision risk level, corresponding proactive safety response actions will be performed: When the overall collision risk level is in the first range, target visualization prompts are provided via AR-HUD; When the overall collision risk level is in the second range, activate the sound, light, and tactile level three alarm. When the overall collision risk level is in the third range, autonomous emergency braking is triggered and the roadside warning devices are activated.

[0129] Example 3 Based on the same application concept, see [link / reference] Figure 8 As shown, Figure 8 This illustration shows a structural schematic diagram of a computer device provided in Embodiment 3 of this application, wherein, as shown... Figure 8 As shown, the computer device 800 provided in Embodiment 3 of this application includes: The system includes a processor 801, a memory 802, and a bus 803. The memory 802 stores machine-readable instructions that can be executed by the processor 801. When the computer device 800 is running, the processor 801 communicates with the memory 802 through the bus 803. When the machine-readable instructions are executed by the processor 801, the steps of the intelligent rear traffic crossing warning method shown in Embodiment 1 are performed.

[0130] Example 4 Based on the same concept, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the intelligent rear traffic crossing warning method described in any of the above embodiments.

[0131] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system and apparatus described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0132] The computer program product for intelligent rear traffic crossing warning provided in this application includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the preceding method embodiments. For specific implementation details, please refer to the method embodiments, which will not be repeated here.

[0133] The intelligent rear traffic crossing warning device provided in this application embodiment can be specific hardware on a device or software or firmware installed on the device. The implementation principle and technical effects of the device provided in this application embodiment are the same as those in the foregoing method embodiments. For the sake of brevity, any parts not mentioned in the device embodiment can be referred to the corresponding content in the foregoing method embodiments. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can all be referred to the corresponding processes in the above method embodiments, and will not be repeated here.

[0134] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0135] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0136] In addition, the functional units in the embodiments provided in this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0137] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0138] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0139] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be determined by the protection scope of the claims.

Claims

1. A method for intelligent rear traffic crossing early warning, characterized in that, The method includes: Acquire multimodal perception data from the vehicle, roadside blind spot perception data, and cloud-based traffic information; The vehicle-side multimodal perception data is fused at the data level to generate a 3D target with semantic labels and its motion trajectory. Based on the motion trajectory and scene context information, a spatiotemporal graph neural network model is used to predict multiple probabilistic trajectories of the three-dimensional target in the future and their corresponding behavioral intentions. Based on the vehicle's preset driving path and the multiple probabilistic trajectories, an initial risk matrix is ​​calculated, wherein the initial risk matrix represents the collision risk corresponding to each probabilistic trajectory. Based on the initial risk matrix, combined with the correction of the collision probability by the behavioral intent, the roadside blind spot perception data, and the cloud-based traffic information, the comprehensive collision risk level is determined. Based on the comprehensive collision risk level, the corresponding proactive safety response operation is executed.

2. The method according to claim 1, characterized in that, The acquisition of vehicle-side multimodal perception data, roadside blind spot perception data, and cloud-based traffic information includes: The vehicle-side multimodal perception data is acquired by solid-state LiDAR, wide-angle cameras, and short-range millimeter-wave radar deployed at the rear and sides of the vehicle. The blind spot sensing data is received from the roadside unit via V2I communication; The cloud-based traffic information, including real-time traffic flow, historical accident data, and high-precision maps, is obtained from the cloud-based traffic platform via an in-vehicle T-Box.

3. The method according to claim 1, characterized in that, The step of performing data-level pre-fusion of vehicle-side multimodal perception data to generate 3D targets with semantic labels and their motion trajectories includes: The 3D point cloud data collected by solid-state LiDAR is projected into the image coordinate system of a wide-angle camera to achieve spatiotemporal alignment between point cloud data and image pixels. A deep learning network based on the PointPainting architecture is used to process the spatiotemporally aligned fused data to achieve joint target recognition and semantic segmentation. The identified target is tracked across multiple frames to generate the 3D target with semantic labels and its motion trajectory.

4. The method according to claim 1, characterized in that, The method of predicting multiple probabilistic trajectories of a 3D target and their corresponding behavioral intentions based on motion trajectory and scene context information using a spatiotemporal graph neural network model includes: A spatiotemporal graph model is constructed using the motion trajectory and scene context information, where traffic participants are nodes and the interaction relationships between participants are edges. The historical trajectory data, motion state, and scene context information in the motion trajectory are used as input features; The input features are processed by a spatiotemporal graph neural network model, and the multiple probabilistic trajectories of each target and their corresponding behavioral intention probability distributions within a preset future time period are output.

5. The method according to claim 1, characterized in that, The calculation of the initial risk matrix based on the vehicle's preset driving path and multiple probabilistic trajectories includes: For each of the multiple probabilistic trajectories, calculate its collision time and collision probability with the vehicle's preset driving path; Based on the correspondence between collision time and collision probability, an initial risk matrix containing risk scores for each trajectory is constructed.

6. The method according to claim 1, characterized in that, Based on the initial risk matrix, and combined with the correction of collision probability for behavioral intent, roadside blind spot perception data, and cloud-based traffic information, a comprehensive collision risk level is determined, including: The collision probabilities in the initial risk matrix are weighted and corrected according to the probability distribution of the stated behavioral intent; The global decision-making suggestions provided by the roadside blind spot perception data and the regional risk coefficients provided by the cloud-based traffic information are weighted and fused with the modified risk matrix; Based on the weighted fusion results, the comprehensive collision risk level is determined by a preset risk level mapping rule.

7. The method according to claim 1, characterized in that, The process of executing corresponding proactive safety response operations based on the comprehensive collision risk level includes: Based on the different ranges of the overall collision risk level, corresponding proactive safety response actions will be performed: When the overall collision risk level is in the first range, target visualization prompts are provided via AR-HUD; When the overall collision risk level is in the second range, activate the sound, light, and tactile level three alarm. When the overall collision risk level is in the third range, autonomous emergency braking is triggered and the roadside warning devices are activated.

8. An intelligent rear traffic crossing early warning device, characterized in that, The device includes: The data acquisition module is used to acquire vehicle-side multimodal perception data, roadside blind spot perception data, and cloud-based traffic information; The data-level pre-fusion module is used to perform data-level pre-fusion on the vehicle-side multimodal perception data to generate a three-dimensional target with semantic labels and its motion trajectory. The trajectory intent prediction module is used to predict multiple probabilistic trajectories of the three-dimensional target and their corresponding behavioral intents based on the motion trajectory and scene context information using a spatiotemporal graph neural network model. The risk matrix calculation module is used to calculate an initial risk matrix based on the vehicle's preset driving path and the multiple probabilistic trajectories, wherein the initial risk matrix represents the collision risk corresponding to each probabilistic trajectory. The risk level determination module is used to determine the comprehensive collision risk level based on the initial risk matrix, combined with the correction of the collision probability by the behavioral intention, the roadside blind spot perception data, and the cloud traffic information. The safety response operation execution module performs corresponding proactive safety response operations based on the comprehensive collision risk level.

9. A computer device, characterized in that, include: The system includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, they perform the steps of the intelligent rear traffic crossing warning method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the intelligent rear traffic crossing warning method as described in any one of claims 1 to 7.