Traffic scene risk identification method and device considering dynamic and static information fusion, and storage medium
Through deep learning algorithms, the spatial characteristics of static blind spots and dynamic timing behavior characteristics are extracted, and the problem of insufficient fusion of dynamic and static information in the existing technology is solved, and the accurate assessment and identification of traffic scene risks is achieved.
Patent Information
- Application Number
- CN202510977060.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-08-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing traffic scenario risk identification methods are insufficient in handling dynamic and static factors. Inadequate fusion of dynamic information leads to inaccurate risk assessment and insufficient static environmental characteristics, resulting in omissions in risk identification.
Through deep learning algorithms, high-precision maps and dynamic trajectory data are fused, static blind spot spatial features and dynamic timing behavior characteristics are extracted, feature extraction and fusion is used for feature extraction and fusion, and risk factor prediction is combined with fully connected neural networks to achieve nonlinear correlation of dynamic and static information.
It realizes comprehensive perception and accurate assessment of potential hazards in complex traffic scenarios, improves the accuracy and interpretability of risk identification, and captures the nonlinear coupling relationship of dynamic and static risk factors.
Smart Images

Figure CN120496360A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of traffic scene risk identification methods, and specifically to a traffic scene risk identification method, device and storage medium that consider the fusion of dynamic and static information. Background Art
[0002] As the global car population continues to grow, traffic safety issues are becoming increasingly prominent, with frequent traffic accidents causing significant loss of life and property. Accurately predicting risks in complex traffic scenarios is key to preventing accidents. Furthermore, the rapid development of autonomous driving and intelligent assisted driving technologies is placing higher demands on understanding traffic scenarios and assessing risks. Autonomous vehicles and intelligent connected vehicles require more precise perception of their surroundings and early warning of risks.
[0003] Current technical solutions for traffic risk identification are primarily based on single-information approaches. On the one hand, while methods based on static scene information can consider fixed elements such as road infrastructure and traffic signs for risk assessment, they struggle to capture the dynamic changes in traffic flow and the real-time behavior of pedestrians and vehicles, resulting in insufficiently timely and comprehensive risk assessments. On the other hand, while research focused on dynamic scene information can account for the motion and trajectory changes of traffic participants, it is unable to accurately assess the severity and scope of risk in the absence of static environmental information. Furthermore, risk identification and assessment methods often rely on simple rules and statistical models, making it difficult to fully capture the complex relationships between dynamic and static information, resulting in inaccurate and in-depth assessments.
[0004] To address these issues, existing research has considered static elements in the environment. Narksri et al. [Narksri P, Darweesh H, Takeuchi E, et al. Visibility estimation in complex, real-world driving environments using high-definition maps[C] / / 2021IEEE International Intelligent Transportation Systems Conference (ITSC).IEEE, 2021: 2847-2854.] and Easa et al. [Easa S, Ma Y, Elshorbagy A, et al. Visibility-based technologies and methodologies for autonomous driving[M] / / Self-Driving Vehicles and Enabling Technologies. IntechOpen, 2020.] used high-precision maps and 3D point cloud data, combined with a depth buffer algorithm and image projection technology, to estimate visibility and occlusion areas in complex driving environments. They also quantified the visibility of specific locations using a visibility ratio, enabling autonomous vehicles to more effectively identify and respond to risks in complex driving environments.
[0005] However, the above-mentioned prior art methods have the following problems: (1) Existing technical solutions have the problem of insufficient fusion of dynamic and static information in traffic scene risk identification. On the one hand, existing technical solutions focus on the motion state and interaction of dynamic entities. Although they can capture real-time changes in traffic flow and the dynamic evolution of risks, they are limited in assessing the severity and scope of risks due to the lack of static environmental background information. On the other hand, although there are methods that use high-precision maps and three-dimensional point cloud data to estimate visibility and occlusion areas, they are insufficient in processing dynamic factors such as real-time traffic flow changes and cannot fully combine dynamic information to improve the accuracy of risk identification.
[0006] (2) In terms of dynamic risk indicators, some existing technical solutions focus mainly on the basic motion states of the vehicle, such as its trajectory, speed, and acceleration, but do not adequately consider more complex dynamic risk factors such as sudden acceleration, deceleration, and changes in heading angle. At the same time, most existing technologies do not fully consider the impact of the static environment on traffic risks, such as static features such as the geometry of the road, the distribution of obstacles, and the type of lane lines. These factors have a significant impact on the driving safety and risk of vehicles in actual traffic scenarios. The lack of assessment of static risk indicators will lead to omissions in risk identification.
[0007] In summary, the existing traffic scene risk identification methods still have shortcomings in dealing with dynamic and static factors. Summary of the Invention
[0008] The present invention provides a traffic scene risk identification method, device and storage medium that consider the fusion of dynamic and static information to solve the shortcomings of the existing traffic scene risk identification method in processing dynamic and static factors.
[0009] In order to achieve the above object, the technical solution adopted by the present invention is: The traffic scene risk identification method considering the fusion of dynamic and static information is as follows: Step 1: Obtain a high-precision map of the traffic scene and the time-series trajectory of each vehicle in the traffic scene; each track point in the time-series trajectory of each vehicle contains the timestamp, location data, and speed data of the corresponding track point; Convert all track points in each vehicle's time-series trajectory to the high-precision map coordinate system of the traffic scene, and perform lane matching on each track point of each vehicle to obtain a high-precision map of each vehicle's time-series trajectory; Step 2: Predict the position of each vehicle based on the time series trajectory of each vehicle obtained in step 1, obtain the prediction result of each vehicle position and the prediction error of each vehicle position prediction result, and extract the dynamic time series behavior characteristics of each vehicle time series trajectory; Based on the high-precision map of each vehicle's time-series trajectory obtained in step 1, extract the static blind spot spatial features of each vehicle's time-series trajectory; Step 3: The dynamic temporal behavior characteristics and static blind spot spatial characteristics of each vehicle obtained in step 2 are fused to obtain the fusion characteristics of each vehicle; The fusion features of all vehicles are input into the risk factor prediction model, and the risk factor prediction model is used to predict dynamic risk factor prediction results and static risk factor prediction results; Based on the prediction errors between the dynamic risk factor prediction results, the static risk factor prediction results and the corresponding true results, and combined with the prediction error of the vehicle position prediction results obtained in step 2, the comprehensive risk value of the traffic scenario is calculated.
[0010] Furthermore, in step 1, after each trajectory point in each vehicle's time-series trajectory is converted to the high-precision map coordinate system, the candidate lanes for each trajectory point of the corresponding vehicle are determined in the high-precision map, and then the hidden Markov model is used to match each trajectory point of the corresponding vehicle with the corresponding candidate lane. In this way, the trajectory points and lanes are matched in the high-precision map, and a high-precision map of each vehicle's time-series trajectory is obtained.
[0011] Furthermore, in step 2, an LSTM network is used to train the LSTM network separately based on the time series trajectory of each vehicle obtained in step 1. The dynamic time series behavior characteristics of the corresponding vehicle time series trajectory are obtained from the hidden layer of the trained LSTM network, and the trained LSTM network outputs the corresponding vehicle position prediction result and the prediction error of the corresponding vehicle position prediction result.
[0012] Furthermore, in step 2, an obstacle map of each vehicle is obtained based on the high-precision time-series trajectory map of each vehicle obtained in step 1, and a visibility heat map of the corresponding vehicle is obtained based on the obstacle map of each vehicle; The visibility heat map of each vehicle is fused with the high-precision map of the corresponding vehicle's time-series trajectory to form a fused map corresponding to each vehicle. The fused map corresponding to each vehicle is then input separately into a pre-trained CNN network, and the trained CNN network extracts the static blind spot spatial features of the corresponding vehicle's time-series trajectory.
[0013] Furthermore, the process of obtaining the visibility heat map is as follows: First, the semantic layer of each vehicle's time-series trajectory high-precision map is rasterized to obtain a raster map; Obstacles are determined in the grid map, and a plurality of target points are sampled in a traversable area of a corresponding vehicle in the grid map, thereby obtaining an obstacle map of the corresponding vehicle; Next, the time series trajectory of the corresponding vehicle is divided into multiple segments in the obstacle map. The visibility of each target point in each segment is detected by the ray method. Thus, the visibility of each target point is detected in each segment. Finally, the visibility of multiple target points in each segment of the corresponding vehicle time series trajectory is color mapped to obtain the visibility heat map of the corresponding vehicle.
[0014] Furthermore, in step 2, the dynamic temporal behavior features obtained include time features, position and trajectory features, trajectory shape features, motion features, context features, and sequence features, wherein: time features include absolute time features, relative time features, and time interval features; position and trajectory features include absolute position features, relative position features, offset features, and relative distance features; trajectory shape features include curvature features and velocity change rate features; motion features include velocity features, acceleration features, angular velocity features, and heading angle consistency features; context features include neighboring target relationship features, environmental features, and time features; sequence features include sliding window sequence features and multi-target association features; The obtained static blind spot spatial features include blind spot edge shape features, local occlusion relationship features, blind spot area ratio features, and blind spot and lane centerline distance features.
[0015] Furthermore, in step 3, the risk factor prediction model is a trained fully connected neural network.
[0016] Furthermore, in step 3, the dynamic risk factors predicted by the risk factor prediction model include the collision risk index TTC, and the static risk factors include blind spot density, lane proximity, and maximum invisible ratio.
[0017] An electronic device includes a processor and a memory. When program instructions in the memory are read and executed by the processor, the above-mentioned traffic scene risk identification method considering the fusion of dynamic and static information is executed.
[0018] A storage medium stores program instructions, which, when read and executed, execute the above-mentioned traffic scene risk identification method considering the fusion of dynamic and static information.
[0019] In this invention, the static environmental information in the high-precision map is deeply integrated with the temporal behavioral characteristics of dynamic targets to achieve comprehensive perception and accurate assessment of potential dangers under complex traffic conditions. With the extraction of static blind spot spatial features as the core, a rasterized obstacle map is generated based on the semantic layer of the high-precision map. The visibility heat map is generated by combining ray method detection to quantify the blind spot risk. The spatial features of static blind spots are extracted through a convolutional neural network (CNN) to obtain accurate modeling of static occluded areas. The extraction of dynamic temporal behavioral features relies on the long short-term memory network (LSTM) to process the temporal trajectory of the vehicle, capture the long-term dependencies of its motion patterns, and predict future short-term trajectories and interactive risk parameters to effectively identify the collision probability of dynamic entities. The dynamic and static features are spliced and fused and then input into a fully connected neural network to implicitly learn the nonlinear relationship between the two to achieve risk identification and output of traffic scenarios.
[0020] The present invention overcomes the limitations of traditional single data source analysis and achieves a more comprehensive risk assessment by deeply fusing the static blind spot spatial features extracted by CNN and the dynamic temporal behavior features extracted by LSTM. In terms of statics, the visibility heat map generation technology is used to quantify the dynamic changes of blind spots from the driver's perspective, rather than relying solely on the physical position of obstacles; in terms of dynamics, multi-dimensional text features such as trajectory curvature, heading angle consistency, and interaction between neighboring targets are introduced, and combined with high-precision map matching, the association between trajectory prediction and lane semantics is strengthened. Finally, the implicit learning strategy of feature splicing and fully connected neural networks is adopted to adaptively capture the nonlinear coupling relationship between dynamic and static risk factors, thereby improving the accuracy of traffic scene risk identification.
[0021] The present invention simulates the field of view of a real driver by constructing a dynamic perspective model of a local coordinate system, and combines the ray method to detect the obstacle occlusion relationship between the target point and the vehicle trajectory. The invisible frequency of the target point is counted point by point along the vehicle trajectory, and the blind spot is converted from a discrete "visibility" to a continuous thermal distribution map, which intuitively reflects the blind spot risk level. Finally, the thermal map is integrated with the high-precision map and input into the CNN, so that the model can automatically learn the morphological characteristics of the blind spot and improve the interpretability of risk identification.
[0022] For autonomous vehicles or connected vehicles on the road, in addition to moving objects, static obstacles such as buildings and greenery can also affect the vehicle's ability to identify risks in the current traffic scene, posing a threat to road safety. Traditional risk identification solutions often rely solely on the geometric information of high-precision maps or focus solely on the interactive risks of dynamic individual trajectories. This leads to ignoring the interaction between dynamic and static factors, resulting in risk underestimation or even misjudgment.
[0023] The present invention, however, targets complex traffic scenarios and achieves three-dimensional modeling through multi-level fusion of dynamic and static features, effectively resolving the disconnect between static obstacle analysis and dynamic target analysis in traditional risk identification solutions. In terms of static analysis, a ray method is used to map physical obstacles into a driver-perceived blind spot heat map, achieving a comprehensive consideration of map information and vehicle position. When extracting dynamic temporal behavioral features, the trajectory is forced to match the lane, ensuring that the trajectory prediction has inherent lane constraints. A fully connected neural network is used to achieve a fusion evaluation of dynamic temporal behavioral features and static blind spot spatial features, nonlinearly coupling the vehicle driver's blind spot with the dynamic trajectory, realizing the combined impact of dynamic and static information on risk.
[0024] Therefore, the present invention can not only capture the static obstacle information of the traffic scene and effectively evaluate the possible collision risk of dynamic subjects during movement, but also simultaneously consider the impact of dynamic and static factors on risk identification, achieve the unity of static obstacle analysis and dynamic target analysis, and provide new ideas for traffic risk identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a principle diagram of the method of an embodiment of the present invention.
[0026] Figure 2 4 is a data processing flow chart of the LSTM training process in an embodiment of the present invention.
[0027] Figure 3 : This is a perspective diagram of the scanning area using the ray method in an embodiment of the present invention, where: (a) is the vehicle coordinate system with the driver as the coordinate origin; (b) is a top view of the scanning perspective; and (c) is a side view of the scanning perspective.
[0028] Figure 4 Schematic diagram of the ray method in an embodiment of the present invention, wherein: (a) is a schematic diagram of the ray method in a two-dimensional plane; (b) is a schematic diagram of the ray method in the vehicle coordinate system.
[0029] Figure 5 It is a cumulative calculation process diagram of the ray method in an embodiment of the present invention.
[0030] Figure 6 This is an example diagram of a visibility heat map in an embodiment of the present invention.
[0031] Figure 7 This is a risk identification flow chart in an embodiment of the present invention.
[0032] Figure 8 It is a risk assessment logic diagram in an embodiment of the present invention. DETAILED DESCRIPTION
[0033] To help those skilled in the art better understand the present invention, the following detailed description of the embodiments of the present invention is provided in conjunction with the accompanying drawings and examples. This will help those skilled in the art to fully understand and implement the present invention by applying technical means to solve technical problems and achieve corresponding technical effects. The embodiments of the present invention and the various features therein may be combined with each other as long as they do not conflict with each other, and the resulting technical solutions are all within the scope of protection of the present invention.
[0034] Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0035] It should be noted that the terms "include" and "have" in the specification, claims and drawings of the present invention and any variations thereof are intended to cover non-exclusive inclusions.
[0036] like Figure 1As shown, this embodiment discloses a traffic scenario risk identification method that considers the fusion of dynamic and static information. This embodiment uses a deep learning algorithm to collaboratively analyze the interaction between environmental blind spots and dynamic targets, ultimately outputting the risk of individual traffic locations, enabling vehicle-level risk warnings. Static blind spot spatial feature extraction is based on high-precision maps. Obstacle maps are generated through rasterization. Target points are extracted and their visibility is detected using the ray method. This is then accumulated to generate a visibility heat map. Finally, blind spot features are extracted using a CNN. Dynamic temporal behavioral feature extraction focuses on vehicle and pedestrian trajectory data. After outlier cleaning, Kalman filter smoothing, and hidden Markov model map matching, dynamic trajectory features are extracted using an LSTM. Dynamic and static features are concatenated and normalized into a unified vector. Dynamic and static risk factors are introduced and fed into a fully connected neural network to learn nonlinear risk relationships, capturing the temporal evolution of comprehensive risk. Based on the learned risk evolution, the location, speed, and corresponding risk value of traffic trajectories are predicted, forming a corresponding spatiotemporal risk variation curve.
[0037] The traffic scene risk identification method considering the fusion of dynamic and static information in this embodiment includes the following steps: Step 1. In this embodiment, a high-precision map of the traffic scene and the time-series trajectory of each vehicle in the traffic scene are obtained; each trajectory point in the time-series trajectory of each vehicle includes the timestamp of the corresponding trajectory point, the position data of the corresponding trajectory point, and the speed data of the corresponding trajectory point.
[0038] Then, the acquired time series trajectory of each vehicle is preprocessed, which includes outlier processing and Kalman filtering, where: When processing outliers, set an acceleration threshold for the trajectory point (±10m / s² in this embodiment). If the acceleration of a trajectory point calculated based on its velocity is greater than or equal to the acceleration threshold, the trajectory point is considered an outlier and is removed.
[0039] When performing Kalman filtering, the state is updated based on formula (1): (1) In formula (1): Indicates time The state vector of , in this embodiment, is each trajectory point in the time series trajectory of each vehicle; A is the state transfer matrix, which is used to describe the linear relationship between the state and time; is the process noise, Subject to mean 0 and covariance The normal distribution of .
[0040] After preprocessing in this embodiment, all trajectory points in the preprocessed time-series trajectory of each vehicle are converted from their coordinate system to the high-precision map coordinate system of the traffic scene, and each trajectory point of each vehicle is matched with the lane in the high-precision map, thereby obtaining a high-precision map of each vehicle's time-series trajectory. The specific process is as follows: (S1) First, based on the coordinate system of the time-series trajectory source (e.g., a camera), a transformation matrix is established between the time-series trajectory source coordinate system and the high-precision map coordinate system. Based on the transformation matrix, all trajectory points in each vehicle's time-series trajectory are transformed from their coordinate system to the high-precision map coordinate system of the traffic scene, thereby ensuring that the positions of the trajectory points and the high-precision map elements correspond correctly.
[0041] (S2) Next, candidate lane selection is performed. This process is accelerated through spatial indexing, where a grid spatial index is established for the lane centerlines in the HD map. This allows for rapid retrieval of candidate lanes around each trajectory point, with the search radius dynamically adjusted. Lane attribute filtering is then performed to exclude non-drivable lanes (e.g., bus lanes and construction zones), and restricted lanes are filtered based on vehicle type (e.g., trucks vs. cars). This allows for precise determination of the appropriate candidate lane for each vehicle at each trajectory point.
[0042] (S3) Finally, a hidden Markov model is used to match each trajectory point with a candidate lane.
[0043] The state set of the hidden Markov model is the candidate lane set, and the observation sequence is the coordinates, speed, and heading angle of the trajectory point. In the hidden Markov model, the emission probability Indicates that the status (a certain lane), observe the current position and posture of the vehicle The possibility of is obtained from formula (2): (2) In formula (2): Indicates the vertical distance from the trajectory point to the lane centerline; Indicates the angle between the vehicle heading and the lane direction; represents the standard deviation of the heading angle; Indicates the standard deviation of vertical distances.
[0044] In the Hidden Markov Model, the transition probability Indicates that the vehicle is moving from the lane Shift to lane The possibility of is obtained from formula (3): (3) In formula (3): Indicates a time interval; is the lane keeping time constant, The larger it is, the more the vehicle tends to stay in the same lane for a longer period of time; That is, the basic lane change probability, which ranges from 0 to 1, reflecting the influence of lane line type. The solid lane is close to 0, and the dashed lane is close to 1. and No connection, then topological reachability ,on the contrary .
[0045] At the same time, topological consistency constraints are constructed for HD map lanes. Lanes are abstracted as graph nodes, and adjacent lanes that allow passage are connected. Edge weights reflect the difficulty of lane changes, and connectivity checks are performed on matching results to ensure that there are reachable paths between adjacent lanes in the topological graph.
[0046] Based on the above steps (S1)-(S3), accurate matching of each trajectory point and lane is achieved in the high-precision map.
[0047] Step 2: In this embodiment, the dynamic temporal behavior characteristics and static blind spot spatial characteristics of each vehicle temporal trajectory are extracted, as described below.
[0048] (2.1) In this embodiment, an LSTM network is used to extract dynamic temporal behavior features and predict position and speed.
[0049] Specifically, the time series trajectory of each vehicle obtained in step 1 is segmented into a sliding time window. After the basic attributes of the trajectory data (timestamp, location coordinates, speed, acceleration, heading angle) and the context (lane and lane type, target ID and type of surrounding vehicles) are annotated, they are input into the LSTM network for independent training.
[0050] The data processing process during LSTM network training is as follows Figure 2 As shown in the figure, the continuous trajectory data of the corresponding vehicle is segmented into 5-second segments through time windows. Multi-dimensional information is annotated for each trajectory point, including motion features (speed, acceleration, heading angle), spatial features (position coordinates, lane offset), interaction features (type and distance of adjacent objects), and environmental features (lane and lane type). A time step sequence is then constructed. The network then propagates forward, processing the sequence forward and backward through a bidirectional LSTM layer. The outputs are merged and passed through a dropout layer to prevent overfitting. The LSTM outputs are then mapped to prediction targets through a fully connected layer, including the predicted results for the corresponding vehicle position and speed. The mean squared error (MSE) loss function is used to calculate the difference between the predicted value and the true value, i.e., the prediction error. The loss weights for position and speed are set to 1.0 and 0.3 respectively.
[0051] Among them, the prediction error of the corresponding vehicle position prediction result using the mean square error (MSE) loss function is The calculation formula is as follows: (4) In formula (4): Indicates the number of vehicle trajectory points; Respectively represent the horizontal and vertical coordinates of the actual position of the vehicle; are the horizontal and vertical coordinates of the predicted position respectively.
[0052] Subsequently, backpropagation optimization was performed using the Adam optimizer to update the network weights and bias parameters. During training, the data was divided into training, validation, and test sets in a ratio of 70%, 15%, and 15%. Using an early stopping mechanism, training was terminated when the validation set loss did not decrease for 10 consecutive times.
[0053] After training the LSTM network, the hidden state of the last hidden layer of the LSTM network is extracted as the hidden state vector. This hidden state vector can represent the motion pattern of the corresponding vehicle and the interaction between the corresponding vehicle and other vehicles. Specifically, the hidden state contains the historical time information of the sequence, the current historical position information and change pattern, the accumulation of the position sequence change rate, the instantaneous motion state and historical evolution trend, the perception information of the external environment and the historical interaction state information. With this information, the LSTM network can predict the future trajectory of the corresponding vehicle, and the accuracy of the prediction depends largely on the feature information contained in the hidden state. Trajectory prediction using the hidden state can be obtained through formula (5): (5) In formula (5): Represents the LSTM network at time step The hidden state of represents the time from time step 1 to time step Historical data series; 、 Respectively indicate the corresponding vehicles in Two-dimensional coordinates of time ,speed; is the linear transformation matrix that maps the latent state vector to the trajectory prediction output; The base offset used to adjust the forecast values.
[0054] In this embodiment, the LSTM network is ultimately trained using the time series trajectory of each vehicle, and the dynamic time series behavior features of each vehicle's time series trajectory are extracted from the hidden layer of the LSTM network, including time features, position and trajectory features, trajectory shape features, motion features, context features, and sequence features. Among them, time features include absolute time features, relative time features, and time interval features; position and trajectory features include absolute position features, relative position features, offset features, and relative distance features; trajectory shape features include curvature features and velocity change rate features; motion features include velocity features, acceleration features, angular velocity features, and heading angle consistency features; context features include neighboring target relationship features, environmental features, and time features; and sequence features include sliding window sequence features and multi-target association features.
[0055] Among them, time features can be obtained based on the sequence historical time information in the hidden state of the hidden layer; position and trajectory features can be obtained based on the current and historical position information and their change patterns in the hidden state of the hidden layer; trajectory shape features can be obtained based on the cumulative effect of the position sequence change rate in the hidden state of the hidden layer; motion features can be obtained based on the instantaneous motion state and its historical evolution trend in the hidden state of the hidden layer; context features can be obtained based on the perception information of the external environment in the hidden state of the hidden layer and its historical interaction state; sequence features can be obtained based on the temporal pattern and correlation information contained in the compression result of the hidden state of the hidden layer itself as a sequence processing.
[0056] (2.2) In a traffic environment, in addition to the impact of moving dynamic traffic entities on risk, static obstacles also affect risk identification. Therefore, in this embodiment, the risk identification of traffic scenes takes into account the occlusion factor of obstacles and extracts the static blind spot spatial features of each vehicle's time-series trajectory. The process is as follows: 2.2a) Rasterize the semantic layer of each vehicle's time-series trajectory HD map obtained in step 1 (the semantic layer includes lane lines, static obstacle types and heights, and traffic signs) and convert each vehicle's time-series trajectory HD map into a 0.1m resolution raster map.
[0057] The grid annotation attributes in the grid map include lane, obstacle type, and height. Then, a height threshold is set (set to 0.4m in this embodiment). Obstacles in the grid map are classified according to the height threshold. Obstacles below the height threshold are marked as non-visibility-affecting obstacles, and obstacles above or equal to the height threshold are marked as visibility-affecting obstacles.
[0058] Finally, within the passable area of each trajectory point of the corresponding vehicle in the grid map, multiple target points are evenly sampled at a spacing of 0.5 m to ensure that the activity area of potential road users is covered, thereby completing the generation of the obstacle map.
[0059] 2.2b) In each vehicle's obstacle map, the vehicle's time-series trajectory is divided into multiple segments at 10cm intervals. Within each segment of the obstacle map, the visibility of each target point is tested using a raycast method based on obstacles that affect visibility. The visibility of each target point is then measured in each segment.
[0060] Specifically, in the ray method, the center of the corresponding vehicle is established as the origin, The axis points to the right side of the corresponding vehicle, The axis points straight ahead, The coordinate system with the axis vertically upward. The scanning area viewing angle parameters in the ray method are set as follows: the horizontal field of view angle is , covering the front of the corresponding vehicle Range; vertical field of view from (ground) to (sky) to avoid false detection of aerial targets; the angular resolution in the horizontal direction is , vertical direction is , thus achieving a balance between accuracy and computational efficiency; the maximum detection distance is set to 60m to meet the typical sight distance requirements of urban roads. Figure 3 As shown in (a), (b) and (c).
[0061] When scanning the target points based on the ray method for each segment of the obstacle map, an initial visibility label is assigned to each target point. , visibility label The initial value is 0. In each segment, a ray is emitted from the corresponding vehicle position to each target point, and the intersection of the ray and the grid of obstacles affecting visibility is detected. If there is an intersection, the corresponding target point is judged to be invisible at the corresponding vehicle driver position, and the visibility label of the corresponding target point is changed to The value is increased by 1, otherwise it remains unchanged. Figure 4 From (a) and (b), we can see that due to the corresponding vehicle position Figure 4 After the target point emits a ray, if the ray is blocked by an obstacle that affects visibility, the target point is judged to be invisible.
[0062] Execute step by step along each segment of the corresponding vehicle time series trajectory in the current driving lane, and the visibility label of each target point The values are accumulated to give the visibility label of each target point The value is compared with the total number of scans to get the invisible ratio , invisible scale The higher the value, the more the target point is invisible from the driver's position, that is, the larger the blind spot is. Therefore, in this embodiment, the invisible ratio is used as the Characterize the visibility of each target point. Starting from a certain moment, scan four times, the cumulative calculation process of the ray method is as follows Figure 5 shown.
[0063] Finally, the visibility (i.e., invisible ratio) of multiple target points in each segment of the corresponding vehicle time series trajectory is calculated. ) are all color mapped, thereby obtaining the visibility heat map of the corresponding vehicle. In this embodiment, a visibility test is performed every 10cm track interval, multiple frames of data are accumulated, and the invisible ratio is converted into a Jet color spectrum (blue low risk, red high risk). Perform color mapping to generate the visibility heat map of the corresponding vehicle. See the example of visibility heat map for details. Figure 6 .
[0064] 2.2c) The visibility heat map of each vehicle is fused with the high-precision map of the corresponding vehicle's time series trajectory at the pixel level through grid alignment and layer overlay technology to form a fused map corresponding to each vehicle. The specific fusion process is: the invisible ratio in the visibility heat map of each vehicle is converted into the As an independent channel, the grid map corresponding to the high-precision map of the time-series trajectory of the corresponding vehicle is embedded, thereby forming a multi-dimensional input matrix containing obstacle positions, lane structures and blind spot risk values in the grid map, thereby obtaining a fused map of the corresponding vehicle.
[0065] Then, the fusion map corresponding to each vehicle is input separately into the pre-trained CNN network, and the trained CNN network extracts the static blind spot spatial features of the corresponding vehicle time-series trajectory, thereby obtaining the static blind spot spatial features of the time-series trajectory of each vehicle in the traffic scene.
[0066] In this embodiment, the CNN network includes an input layer, a convolutional layer, a pooling layer, and a fully connected layer. The convolutional layer includes a shallow convolutional layer and a deep convolutional layer. The shallow convolutional layer extracts basic spatial features, namely the shape of the blind spot edge and the local occlusion relationship; the deep convolutional layer captures the blind spot area ratio and the distance between the blind spot and the lane centerline. The pooling layer is used to reduce the feature dimension while retaining important feature information. The fully connected layer is used to integrate the extracted features and the invisible ratio of each grid. index.
[0067] The dataset used for CNN network pre-training is a multi-channel fusion of visibility heatmap matrices and rasterized high-precision maps. During CNN network training, an unsupervised learning method is used on the training set samples to drive the network to learn static blind spot features. The input layer receives the multi-channel fused map, and the shallow convolutional layers extract blind spot edges and local occlusion features. Deep convolutional layers are used to capture large-area blind spot structures. The feature maps are then pooled and finally output through the fully connected layers as static spatial features. During training, the map is randomly occluded and fused, and the network is required to reconstruct the heatmap and obstacle distribution based on the unoccluded parts, thereby driving the network to learn static spatial features.
[0068] In this embodiment, the static blind spot spatial features of the corresponding vehicle time-series trajectory extracted from each vehicle fusion map through a pre-trained CNN network include blind spot edge shape features, local occlusion relationship features, blind spot area ratio features, and blind spot and lane centerline distance features.
[0069] Specifically, the fusion map is input into the pre-trained CNN network, and spatial features are extracted through the multi-layer convolution and pooling operations of the CNN network. The shallow convolution kernel of the CNN network captures the edge shape features and local occlusion relationship features of the blind spot from the fusion map, and the deep network of the CNN network models the blind spot area ratio features and the distance features between the blind spot and the lane centerline. Finally, all spatial features and the proportion of each grid are integrated through the fully connected layer of the CNN network. Indicators are used to form the static blind spot spatial characteristics of the corresponding vehicle time-series trajectory. This static blind spot spatial characteristic can comprehensively characterize the spatial distribution, density, and proximity of the blind spot to the lane, thereby providing a static environmental risk factor for subsequent risk assessment.
[0070] Step 3: When autonomous vehicles or intelligent connected vehicles identify traffic scene risks, existing technologies often only focus on dynamic factors and ignore the impact of static scene factors in real scenes on traffic risks. Therefore, this embodiment comprehensively considers the extracted dynamic and static features to achieve risk identification of traffic scenes, such as Figure 7 As shown, the process is as follows: 3.1) Using a two-layer processing mechanism of multimodal feature splicing and standardization, the dynamic temporal behavior characteristics and static blind spot spatial characteristics of each vehicle obtained in step 2 are fused to obtain the fused features of each vehicle.
[0071] Specifically, the static blind spot spatial features extracted by the CNN network in step 2 and the dynamic temporal behavior features extracted by the LSTM network are first spliced in feature dimension to retain the holographic information of the original data and avoid the feature interaction loss that may be caused by early fusion.
[0072] Then, to address the dimensional difference problem of the spliced features, the Z-Score normalization method is used to normalize them, thereby obtaining the fusion features. The normalization is shown in formula (6): (6) In formula (6): Represents the original value of a feature dimension in the concatenated fused feature vector.
[0073] and They represent the mean and standard deviation of a feature dimension in the concatenated fusion feature vector.
[0074] 3.2) Construct a trained fully connected neural network as a risk factor prediction model.
[0075] The fully connected neural network takes fused features as input and maps the input data to a fully connected layer through a linear combination of a weight matrix and a bias vector. Features are gradually extracted using two fully connected layers, with the number of nodes in each layer decreasing. The fully connected layers use the ReLU activation function to perform a nonlinear transformation on the linear combination results, extracting high-order features and compressing the dimensionality. The fully connected neural network design uses branched outputs to predict dynamic and static risk factors separately, avoiding feature coupling while sharing weights from previous layers to capture joint representations.
[0076] Specifically, the fully connected neural network of this embodiment includes an input layer, two fully connected layers, and an output layer (including a dynamic risk factor output layer and a static risk factor output layer).
[0077] The fully connected neural network training process extracts time-aligned feature sequences based on historical data. Current features are input through forward propagation, and predicted values for future dynamic and static features are output. The difference between the predicted values and the actual future features is used as the loss function to train the model. Backward propagation uses the Adam optimizer to update network parameters and minimize the total loss.
[0078] Specifically, the dataset used for training the fully connected neural network consists of a concatenated vector of dynamic and static features and their corresponding true values of dynamic risk factors (including TTC) and static risk factors (including blind spot density, lane proximity, and maximum invisible ratio). During training, the fused features are first Z-score normalized (Formula 6) to eliminate dimensionality differences. Dynamic and static features are then aligned with the true values of the risk factors by time step to form a continuous sequence. The dataset is then divided into training, validation, and test sets with a 70%, 15%, and 15% ratio. The loss function uses the mean squared error (MAE) to calculate the difference between the predicted and true risk factor values. The optimization strategy uses the Adam optimizer to update the network weights. Iterations are terminated if the validation set loss does not decrease for 10 consecutive rounds. Finally, the trained fully connected neural network is used as the risk factor prediction model.
[0079] In this embodiment, the dynamic risk factor includes the collision risk index TTC, and the static risk factor includes blind spot density, lane proximity, and maximum invisible ratio, which are specifically described as follows: (A) To quantify the degree of dynamic collision risk, this embodiment uses the commonly used collision risk index TTC as the dynamic risk factor. The calculation formula of the collision risk index TTC is as shown in formula (7): (7) In formula (7): Indicates the Dynamic risk factors at each time step; is the distance between the two vehicles at this time step; is the relative speed between the two vehicles; 、 Respectively represent the coordinates of the current vehicle's position and speed; 、 Represent the coordinates of the other vehicle's position and speed respectively.
[0080] (B) Blind Spot Density in Static Risk Factors
[0081] The blind spot density reflects the coverage of the blind spots within the path range, which directly affects the perception ability of the vehicle. A circular search area with a certain radius is generated with the path as the center. If the visibility label of a grid in the circular search area is , then the grid is marked as a blind area. Then the total number of grids in the area is counted. and number of blind area grids , and the blind spot density is obtained As shown in formula (8): (8) (C) Lane proximity in static risk factors
[0082] Lane proximity Defined as the distance between the vehicle's current position and the lane boundary, it is used to assess whether the vehicle is following the lane line closely, thereby determining whether the vehicle's driving status is safe. As shown in formula (9): (9) In formula (9): The vehicle position The vertical distance to the left lane line, The vehicle position The vertical distance to the right lane line.
[0083] (D) Maximum unobservable proportion in static risk factors By searching the visibility labels of all rasters in the path range, select the maximum value When the maximum invisible ratio When the risk is high, the risk weight coefficient will be dynamically increased, and the dynamic compensation mechanism is used to ensure the accuracy of risk assessment in high-visibility scenarios. As shown in formula (10): (10) In formula (10): 、 、......、 They represent the invisible proportions of grids 1, 2, ..., N within the path range, which are numbered in the order of the path.
[0084] The actual result of the collision risk indicator TTC as a dynamic risk factor is calculated based on the dynamic time series characteristics of the vehicle trajectory.
[0085] Among the above static risk factors, blind spot density The actual result is calculated based on the position coordinates and the invisible ratio of each grid in the visibility heat map; lane proximity The actual result is calculated based on the vehicle trajectory position coordinates and lane data; the maximum invisible ratio The actual result is calculated based on the invisible ratio of the grid corresponding to the vehicle trajectory coordinates and the visibility heat map.
[0086] 3.3) Figure 8As shown in the figure, the fusion features of all vehicles are input into the risk factor prediction model, and the risk factor prediction model predicts the dynamic risk factor (including the collision risk index TTC) and the static risk factor (including blind spot density, lane proximity, and maximum invisible ratio) prediction results.
[0087] Since the dynamic risk factor collision risk index TTC is only obtained based on the current relative distance and speed, in order to fully characterize the dynamic risk, improve the accuracy of the prediction of the future trajectory, and achieve a full range of risk assessment from "current threat" to "future potential threat", this embodiment also introduces the prediction error of the vehicle position prediction result of the LSTM network trained in step 2 , to supplement the dynamic risk factors.
[0088] In this embodiment, the comprehensive risk value of the traffic scene is calculated based on the prediction error between the dynamic risk factor prediction result and the actual result obtained by the risk factor prediction model, the prediction error between the static risk factor prediction result and the actual result, and the prediction error of the vehicle position prediction result obtained in step 2. R , the calculation formula is shown in formula (11): (11) In formula (11): Represents the prediction errors between the predicted results and the actual results of TTC, blind spot density, lane proximity, and maximum invisible ratio, respectively. The prediction error is expressed as the absolute value of the difference between the true value and the predicted value.
[0089] Represents the prediction error of the vehicle position prediction result of a certain vehicle calculated by the trained LSTM. For example, after the position of a certain vehicle is predicted by the trained LSTM network, the prediction error calculated by the MSE loss function is obtained. ,but The value is the same as .
[0090] is the average prediction error of the risk factor corresponding to the training set. Specifically, It is the average value calculated by summing the prediction errors of the risk factor prediction model for TTC of all samples in the training set; It is the average value calculated by summing the prediction errors of the blind area density of all samples in the training set by the risk factor prediction model; It is the average value calculated by summing the prediction errors of the lane proximity of all samples in the training set by the risk factor prediction model; It is the average value calculated by summing the prediction errors of the risk factor prediction model for the maximum unseen proportion of all samples in the training set; It is the average value of all vehicle position prediction errors obtained by the trained LSTM, which is the same as the value of all vehicles. The average of the sum.
[0091] is the risk weight corresponding to each risk factor.
[0092] In this embodiment, due to the comprehensive risk value R It is an overall risk score obtained by combining different risk factors. In order to unify the values of different ranges into the same range, the comprehensive risk value is R Take the minimum-maximum normalization and convert the comprehensive risk value R Converted into a risk label of 0~1, the conversion is shown in formula (12): (12) In formula (12): 、 They represent the maximum and minimum values of the comprehensive risk value respectively.
[0093] In this embodiment, the risk labels of each time step are organized into a time series through time dimension mapping to form a risk time series. This risk time series reflects the changing trend of risk over time and can be used to dynamically assess the potential risk of a vehicle at different time points.
[0094] At the same time, through spatial dimension mapping, the risk labels are mapped one-to-one with the coordinates of the traffic trajectories to form a risk space distribution map, which intuitively shows the spatial distribution of risks, especially the risk distribution in key areas such as blind spots and lane boundaries.
[0095] On this basis, the combined risk values of time and space are further integrated to generate a spatiotemporal risk curve. This curve uses time as the horizontal axis, spatial location as the vertical axis, and risk value as the third dimension, visually demonstrating the dynamic changes in risk across time and space. Finally, the spatiotemporal risk curve is visualized in the form of a three-dimensional graph, facilitating the assessment of risk distribution in traffic scenarios. Pre-set risk thresholds are then used to issue early warnings for high-risk areas, ensuring the safety of autonomous vehicles.
[0096] This embodiment also discloses an electronic device, including a processor and a memory. When the program instructions in the memory are read and executed by the processor, steps 1 to 3 of the above-mentioned traffic scene risk identification method considering the fusion of dynamic and static information are executed.
[0097] This embodiment further discloses a storage medium storing program instructions. When the program instructions are read and executed, steps 1 to 3 of the above-mentioned traffic scene risk identification method considering the fusion of dynamic and static information are executed.
[0098] The preferred embodiments of the present invention are described in detail above with reference to the accompanying drawings. The embodiments described in the present invention are merely descriptions of the preferred embodiments of the present invention and do not limit the concept and scope of the present invention. The various specific technical features described in the above specific embodiments can be combined in any suitable manner unless there is any contradiction. Such combinations should also be regarded as the contents disclosed in this disclosure as long as they do not violate the concept of the present invention. In order to avoid unnecessary repetition, the present invention will not further describe various possible combinations.
[0099] The present invention is not limited to the specific details of the above-mentioned embodiments. Within the scope of the technical concept of the present invention and without departing from the design concept of the present invention, various modifications and improvements made to the technical solution of the present invention by those skilled in the art should fall within the scope of protection of the present invention. The technical contents for which protection is sought in the present invention have been fully recorded in the claims.
Claims
1. Traffic scene risk identification method considering the fusion of dynamic and static information, characterized by: The process is as follows: Step 1: Obtain a high-precision map of the traffic scene and the time-series trajectory of each vehicle in the traffic scene; each track point in the time-series trajectory of each vehicle contains the timestamp, location data, and speed data of the corresponding track point; Convert all track points in each vehicle's time-series trajectory to the high-precision map coordinate system of the traffic scene, and perform lane matching on each track point of each vehicle to obtain a high-precision map of each vehicle's time-series trajectory; Step 2: Predict the position of each vehicle based on the time series trajectory of each vehicle obtained in step 1, obtain the prediction result of each vehicle position and the prediction error of each vehicle position prediction result, and extract the dynamic time series behavior characteristics of each vehicle time series trajectory; Based on the high-precision map of each vehicle's time-series trajectory obtained in step 1, extract the static blind spot spatial features of each vehicle's time-series trajectory; Step 3: The dynamic temporal behavior characteristics and static blind spot spatial characteristics of each vehicle obtained in step 2 are fused to obtain the fusion characteristics of each vehicle; The fusion features of all vehicles are input into the risk factor prediction model, and the risk factor prediction model is used to predict dynamic risk factor prediction results and static risk factor prediction results; Based on the prediction errors between the dynamic risk factor prediction results, the static risk factor prediction results and the corresponding true results, and combined with the prediction error of the vehicle position prediction results obtained in step 2, the comprehensive risk value of the traffic scenario is calculated.
2. The traffic scene risk identification method considering dynamic and static information fusion according to claim 1 is characterized in that: In step 1, after converting each trajectory point in each vehicle's time-series trajectory to the high-precision map coordinate system, the candidate lanes for each trajectory point of the corresponding vehicle are determined in the high-precision map. Then, the hidden Markov model is used to match each trajectory point of the corresponding vehicle with the corresponding candidate lane. In this way, the trajectory points and lanes are matched in the high-precision map, and a high-precision map of each vehicle's time-series trajectory is obtained.
3. The traffic scene risk identification method considering dynamic and static information fusion according to claim 1 is characterized in that: In step 2, an LSTM network is used to train each vehicle's time series trajectory obtained in step 1. The dynamic time series behavior characteristics of the corresponding vehicle's time series trajectory are obtained from the hidden layer of the trained LSTM network. The trained LSTM network then outputs the corresponding vehicle position prediction result and the prediction error of the corresponding vehicle position prediction result.
4. The traffic scene risk identification method considering dynamic and static information fusion according to claim 1 is characterized in that: In step 2, an obstacle map for each vehicle is obtained based on the high-precision time-series trajectory map of each vehicle obtained in step 1, and a visibility heat map of the corresponding vehicle is obtained based on the obstacle map of each vehicle; The visibility heat map of each vehicle is fused with the high-precision map of the corresponding vehicle's time-series trajectory to form a fused map corresponding to each vehicle. The fused map corresponding to each vehicle is then input separately into a pre-trained CNN network, and the trained CNN network extracts the static blind spot spatial features of the corresponding vehicle's time-series trajectory.
5. The traffic scene risk identification method considering dynamic and static information fusion according to claim 4 is characterized in that: The process of obtaining a visibility heat map is as follows: First, the semantic layer of each vehicle's time-series trajectory high-precision map is rasterized to obtain a raster map; Obstacles are determined in the grid map, and a plurality of target points are sampled in a traversable area of a corresponding vehicle in the grid map, thereby obtaining an obstacle map of the corresponding vehicle; Next, the time series trajectory of the corresponding vehicle is divided into multiple segments in the obstacle map. The visibility of each target point in each segment is detected by the ray method. Thus, the visibility of each target point is detected in each segment. Finally, the visibility of multiple target points in each segment of the corresponding vehicle time series trajectory is color mapped to obtain the visibility heat map of the corresponding vehicle.
6. The traffic scene risk identification method considering dynamic and static information fusion according to claim 1 is characterized in that: In step 2, the dynamic temporal behavior features obtained include time features, position and trajectory features, trajectory shape features, motion features, context features, and sequence features, wherein: time features include absolute time features, relative time features, and time interval features; position and trajectory features include absolute position features, relative position features, offset features, and relative distance features; trajectory shape features include curvature features and velocity change rate features; motion features include velocity features, acceleration features, angular velocity features, and heading angle consistency features; context features include neighboring target relationship features, environmental features, and time features; sequence features include sliding window sequence features and multi-target association features; The obtained static blind spot spatial features include blind spot edge shape features, local occlusion relationship features, blind spot area ratio features, and blind spot and lane centerline distance features.
7. The traffic scene risk identification method considering dynamic and static information fusion according to claim 1 is characterized in that: In step 3, the risk factor prediction model is a trained fully connected neural network.
8. The traffic scene risk identification method considering dynamic and static information fusion according to claim 1 is characterized in that: In step 3, the dynamic risk factors predicted by the risk factor prediction model include the collision risk index TTC, and the static risk factors include blind spot density, lane proximity, and maximum invisible ratio.
9. An electronic device comprising a processor and a memory, characterized in that: When the program instructions in the memory are read and executed by the processor, the traffic scene risk identification method considering dynamic and static information fusion as described in any one of claims 1 to 8 is executed.
10. A storage medium storing program instructions, characterized in that: When the program instructions are read and executed, the traffic scene risk identification method considering dynamic and static information fusion as described in any one of claims 1 to 8 is executed.
Citation Information
Cited By
Mobile object collision detection method and system based on real-time coordinate position
CN120748250A
Intelligent automobile operation high-risk scene identification method and device based on deep learning, and storage medium
CN121211213A
Vehicle dynamic state evaluation method and system fusing multi-sensing information
CN121686600A