Emergency rescue autonomous driving vehicle collision risk prediction method and system based on HM-TRW and HAGENN structure search
By constructing dynamic heterogeneous diagrams and hierarchical attention diagrams embedded in neural networks, the problem of collision risk prediction accuracy and efficiency of emergency rescue autonomous driving vehicles in complex environments is solved, and efficient and accurate collision risk prediction is achieved, adapting to various dynamic interactive relationships.
Patent Information
- Application Number
- CN202410283548.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-13
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-03-13
AI Technical Summary
The prior art is difficult to capture the complex heterogeneous characteristics and dynamic relationships of emergency rescue autonomous vehicles in real-time environments, resulting in limited prediction accuracy of collision risk and low computing efficiency, which cannot meet the real-time prediction needs.
Using the method based on HM-TRW and HAGENN structure search, a dynamic heterogeneous graph is constructed, and a high-order memory stochastic walk algorithm is used to learn heterogeneous characteristics and dynamic change laws, and embed neural networks in combination with hierarchical attention graphs to predict collision risk, and optimize computing efficiency through attention positioning and parameterized space.
It realizes efficient and accurate collision risk prediction between emergency rescue vehicles and surrounding moving objects in complex traffic environments, improves the calculation efficiency and generalization capabilities of the model, and can automatically adapt to various dynamic interactions.
Smart Images

Figure CN118013856B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of autonomous driving collision risk prediction. Specifically, it relates to a method and system for predicting the collision risk of an emergency rescue autonomous vehicle based on high-order memory-guided temporal random walk (HM-TRW) and hierarchical attention graph embedding neural network (HAGENN) structure search. Background Art
[0002] When an emergency rescue autonomous vehicle is performing rescue operations, it needs to reach the target location in a timely and safe manner to efficiently complete the rescue task. During the journey to the target location, various uncontrollable factors and possible accidents may occur, thus delaying the rescue task. Therefore, an emergency rescue autonomous vehicle must have the ability to make reasonable decisions based on the actual situation. In the process of decision-making, the prediction of collision risk is one of the essential links. The collision risk prediction first uses vehicle sensors and communication technologies to collect data of surrounding moving objects and predict their future trajectories. By using the data of surrounding moving objects and the data of the vehicle itself (the emergency rescue autonomous vehicle), the collision risk of the autonomous vehicle can be predicted, providing a basis for the vehicle's trajectory planning and action decision-making. Finally, the chassis drives the vehicle to travel safely and efficiently according to the output decision action parameters.
[0003] In the prior art, the methods for predicting the collision risk of emergency rescue autonomous vehicles are mostly machine learning and deep learning. However, existing models are difficult to capture complex heterogeneous characteristics and dynamic relationships in the real vehicle environment, resulting in limited prediction accuracy. In addition, some deep learning models have low computational efficiency when dealing with large-scale real-time data, which is not conducive to real-time prediction. Summary of the Invention
[0004] Aiming at the deficiencies in the prior art, the present invention provides a method and system for predicting the collision risk of an emergency rescue autonomous vehicle based on HM-TRW and HAGENN structure search.
[0005] The present invention achieves the above technical objectives through the following technical means.
[0006] Method for predicting the collision risk of an emergency rescue autonomous vehicle based on HM-TRW and HAGENN structure search:
[0007] Dataset construction: In a real vehicle experiment, vehicle data and data of surrounding moving objects are collected through data acquisition equipment, and the data is preprocessed and stored in a database.
[0008] Collision risk prediction model construction: A dynamic heterogeneous graph is constructed using the collected data of the ego vehicle and the data of surrounding moving objects. The heterogeneous characteristics and dynamic change rules of the dynamic heterogeneous graph are learned by the high-order memory-guided temporal random walk algorithm. The node representations of each surrounding moving object obtained are then fed into a hierarchical attention graph embedding neural network for collision risk prediction;
[0009] Model training: The training set and validation set in the database are respectively input into the collision risk prediction model to optimize the model parameters;
[0010] Model testing: The trained collision risk prediction model is tested using the test set in the database, and the model prediction performance is evaluated and analyzed according to the test results;
[0011] Collision risk prediction: The data of the emergency rescue autonomous vehicle and the data of surrounding moving objects are collected in real time and input into the tested collision risk prediction model to achieve collision risk prediction.
[0012] Furthermore, the learning of the heterogeneous characteristics of the dynamic heterogeneous graph includes:
[0013] (1) Transition vector
[0014] Set the initial high-order memory queue to be empty. The type transition vector accesses surrounding moving objects of each type with equal probability. When accessing the surrounding moving object v j , its type transition vector is updated according to the following formula:
[0015]
[0016] In the formula, φ(v j ) represents the type of the surrounding moving object v j , Q represents the first-in-first-out queue, and Norm() represents the two-norm of the returned vector;
[0017] (2) Type conversion
[0018] According to the probability distribution, determine the type of the next surrounding moving object to be accessed, and adopt a search mechanism with a search factor α∈[0,1] to solve the type trap problem, specifically as follows:
[0019]
[0020] In the formula, represents the surrounding moving object initially accessed by the type transition vector, and Pr(h n+1 ) represents the probability that the type of the next surrounding moving object to be accessed is h n+1 ;
[0021] (3) High-order memory recording
[0022] After accessing the surrounding moving object v j the transfer vector is stored in the next-in-first-out queue Q':
[0023]
[0024] where Put is a queue operator, indicating that when the first-in-first-out queue is full, the first transfer vector is popped and placed at the end of the queue.
[0025] Furthermore, the dynamic change rule of the learned dynamic heterogeneous graph is specifically as follows:
[0026] For the emergency rescue vehicle v i , use h n+1 to represent the type of its surrounding moving objects, so there is:
[0027]
[0028] where represents the set of surrounding moving objects of the emergency rescue vehicle v i at the future timestamp, t' represents the timestamp of the previous random walk, and e t is the edge set of the edge type;
[0029] Select the next moving object to be accessed from the set using an exponential decay distribution:
[0030]
[0031] where the timestamp t of the k-th step random walk in the future k ∈t, t represents the timestamp of the random walk; Pr(v n+1 ) represents the probability that the next moving object to be accessed is v n+1 , and v k represents the surrounding moving object accessed in the k-th step random walk; the discount rate δ ∈ [0, 1].[[]]
[0032] Furthermore, the node representation of each surrounding moving object is the fusion of the output of the high-order memory-guided time random walk algorithm and the original features of each surrounding moving object:
[0033]
[0034]
[0035]
[0036] where They are respectively the transfer vector, the original feature, and the one-hot vector of the surrounding moving object v j , the potential embedding of all moving objects, where represents the set of real numbers, N is the number of nodes, D is the feature dimension, and W and W represent the learnable parameters that are not shared between the surrounding moving object v f and other moving objects, j and x and x vj respectively represent the original feature, the recognition embedding feature, and the final node representation of the surrounding moving object.
[0037] Furthermore, the hierarchical attention graph embedding neural network predicts the collision risk by aggregating different information through node-level attention, edge-level attention, and time-level attention;
[0038] For node-level attention, at time stamp t, for interaction relationship type r, the importance between the emergency rescue vehicle v i and its surrounding moving object v j is calculated by the following formula: where σ is the activation function, x
[0039]
[0040] i and x j are respectively the input representations of the emergency rescue vehicle v i and the surrounding moving object v j , is a linear transformation matrix, || represents concatenation, represents all the surrounding moving objects of the emergency rescue vehicle v at time stamp t with interaction relationship type r, a i is the weight vector, r is the transpose of a r k i k represents the k-th moving object around the ego vehicle;
[0041] The node embedding of the emergency rescue vehicle v at time stamp t with interaction relationship type r i is represented as:
[0042]
[0043] For edge-level attention, the attention mechanism is used to learn the importance of different types of interaction relationships and is calculated through a multi-layer perceptron:
[0044]
[0045] Among them, w T is the edge-level attention vector, U el and b el are the single-layer parameters of the multi-layer perceptron, and R is the set of edge types;
[0046] Considering the importance of different interaction relationship types, the fusion embedding of the emergency rescue vehicle v i is expressed as: For the time-level attention, aggregate the fusion embeddings of the emergency rescue vehicles at all timestamps and package them as
[0047]
[0048] T represents the number of historical timestamps used to predict the collision risk; calculate the query-key-value vectors of the fusion embedding: P = G
[0049] ·U i ·U P
[0050] K = G i ·U K
[0051] V = G i ·U V
[0052] Among them, P, K, and V represent the query, key, and value vectors respectively, and U P , U K , U V represent the corresponding matrices for converting G i into query, key, and value vectors, D is the feature dimension, represents the set of real numbers;
[0053] Use the softmax function to calculate the time-level attention:
[0054]
[0055] Among them, Z i represents the time-level attention, is a mask matrix, and D' is the dimension of the query-key-value vector;
[0056] Take Z i T as the final fusion embedding and calculate the collision risk:
[0057] Y = softmax(W2·ReLU(W1·Z i T + b1)+ b2)
[0058] In the formula, softmax(·) is the output activation function, W1 and W2 are the weight matrices of the hierarchical attention map embedding neural network, ReLU(·) is the activation function, and b1 and b2 represent the bias terms.
[0059] Furthermore, the attention calculated by the hierarchical attention map embedding neural network in predicting the collision risk is sparsified using the attention localization space:
[0060] Through the matrix A LO Select the types of surrounding moving objects, interaction relationship types, and the number of timestamps to be concerned about. The specific calculation is as follows:
[0061]
[0062] When the timestamp is t, through the matrix Determine whether to pay attention to the surrounding moving objects with the interaction type r Represents all the surrounding moving objects of the emergency rescue vehicle v with the interaction relationship type r at the timestamp t′ i ;
[0063] By using the attention localization space, the time complexity is:
[0064]
[0065] Among them, Represents The number of non-zero values in, T represents the number of historical timestamps used to predict the collision risk, And Respectively represent the number of interaction relationships of type r at timestamps t′ and t, O(·) represents the time complexity, and |R| represents the number of edge types.
[0066] Furthermore, the attention calculated by the hierarchical attention map embedding neural network in predicting the collision risk is sparsified using the attention parameterization space:
[0067] Using the attention parameterization space Search for the attention function, and the expression is as follows:
[0068]
[0069] Among them, A N ={1,…,K N} T×|H| Is the parameterization matrix of the node mapping function F N (·), A R ={1,…,K R} 2T×|R| Is the parameterization matrix of the edge mapping matrix F R (·), KN and K R are two hyperparameters, and |H| represents the number of node types.
[0070] Furthermore, multi-stage differential search is used to reduce the parameter search complexity in the positioning space and the parameter space:
[0071] a) Spatial constraint
[0072] The following two constraints are introduced to reduce the search scope and limit the complexity: First, the emergency rescue vehicle can only receive information from surrounding moving objects in historical time. Second, is used to constrain the number of surrounding moving objects and the number of interaction relationships used for collision risk prediction at each timestamp, where is a hyperparameter, and 1 ≤ t ≤ T;
[0073] b) Hypernetwork construction: Use a hypernetwork to convert the parameter search in the positioning space and the parameterized space into a single neural architecture search problem. Specifically, the selection of operations is represented as a probability distribution:
[0074]
[0075] where x is the input, is the output, |A| represents the number of operations, and β i represents the mixing weight of the mapping function F i (·) corresponding to the i-th operation;
[0076] By adopting a hypernetwork, the mixing weight β and all parameters in the mapping function are jointly optimized in a differentiable manner:
[0077]
[0078] where η w and η β represent the learning rates of the structure weight and the model weight respectively, and represent the loss functions of the training set and the validation set respectively, w represents the structure weight, and β represents the model weight;
[0079] c) Multi-stage hypernetwork training: To stabilize the training of the hypernetwork, the training process is divided into three stages: moving object parameterization, interaction relationship parameterization, and attention positioning space search. Each stage focuses on a different parameterized space.
[0080] Furthermore, the construction of the dynamic heterogeneous graph is as follows: The emergency rescue vehicle and its surrounding moving objects are represented as nodes, and various interaction relationships between them are represented as edges. The nodes and edges change dynamically over time, forming a dynamic heterogeneous graph, and its expression is as follows:
[0081] G t =(v t ,e t ,u t )
[0082] where v t is a set of nodes with node type h, e t is a set of edges with edge type r, represents the feature set of all moving objects, h ∈ H, r ∈ R, and H and R are the node type set and edge type set respectively, represents the set of real numbers, N is the number of nodes, and D is the feature dimension.
[0083] An emergency rescue autonomous driving vehicle collision risk prediction system based on HM-TRW and HAGENN structure search, comprising:
[0084] Data acquisition equipment, including in-vehicle sensors, roadside devices, and communication technologies, for collecting self-vehicle data and surrounding moving object data;
[0085] A data preprocessing module for cleaning, normalizing, feature extracting, data dimension reducing, and dataset partitioning of the collected data;
[0086] A prediction model, including a high-order memory-guided temporal random walk algorithm, a hierarchical attention graph embedding neural network, and an optimal parameter search module, where the optimal parameter search module includes an attention localization space, an attention parameterization space, and a multi-stage differential search module;
[0087] A visualization module for displaying the predicted collision risk.
[0088] The beneficial effects of the present invention are:
[0089] (1) This application constructs a dynamic heterogeneous graph of the interaction between emergency rescue vehicles and surrounding moving objects, which can better learn different types of moving objects and the complex dynamic relationships between their interactions. In addition, the dynamic heterogeneous graph can simultaneously capture the static and dynamic features of surrounding moving objects, thereby providing more comprehensive and accurate collision risk prediction results. Specifically, it learns the position changes of each moving object at different time stamps based on the time series data of the self-vehicle and surrounding moving objects. Finally, the dynamic heterogeneous graph has strong reliability and generalization ability, can automatically adapt to various complex traffic environments, and can handle various dynamic interaction relationships.
[0090] (2) This application uses a high - order memory - guided temporal random walk algorithm and a hierarchical attention graph embedding neural network to further learn the importance of surrounding moving objects and their interaction relationships with emergency rescue vehicles. The high - order memory - guided temporal random walk algorithm can make full use of the historical data and non - decreasing time - constraint limitations of surrounding moving objects to consider the moving objects that have a greater impact on emergency rescue vehicles, thereby predicting collision risks more accurately and efficiently. At the same time, the hierarchical attention graph embedding neural network uses hierarchical attention layers to capture the importance of each surrounding moving object for the emergency rescue vehicle and the importance of various interaction relationships between the two, further improving the accuracy of prediction. Meanwhile, the future collision risk is calculated by fusing important features under historical timestamps.
[0091] (3) This application improves the model calculation efficiency by constructing an attention localization and parameterization space, and uses a multi - stage differentiable search algorithm to further accelerate the model's calculation. In actual collision risk prediction, the emergency rescue vehicle needs to consider the spatio - temporal features of all surrounding moving objects and calculate the importance of various interaction relationships to focus on the interaction features that have a greater impact on the collision risk, so the calculation cost is very high. To achieve efficient and accurate collision risk prediction, it is necessary to adopt a more lightweight and efficient model architecture. The attention localization space and parameterization space can flexibly determine the application location and calculation function of attention, while the multi - stage differentiable search algorithm can screen out invalid model architectures by adopting heuristic constraints and use a single - shot neural architecture search algorithm to determine the optimal architecture. BRIEF DESCRIPTION OF THE DRAWINGS
[0092] Figure 1 is a framework diagram of the HM - TRW and HAGENN - based structure search model according to the present invention;
[0093] Figure 2 is a flowchart of the method for predicting the collision risk of an emergency rescue autonomous vehicle based on HM - TRW and HAGENN structure search according to the present invention;
[0094] Figure 3 is a flowchart of the model training and verification according to the present invention;
[0095] Figure 4 is an example diagram of the interaction between an emergency rescue vehicle and other moving objects in the original traffic scenario according to the present invention;
[0096] Figure 5 is a dynamic heterogeneous graph of the interaction between an emergency rescue autonomous vehicle and surrounding moving objects according to the present invention;
[0097] Figure 6 is an example diagram of the scenario for predicting the collision risk of an emergency rescue autonomous vehicle according to the present invention;
[0098] Figure 7 is a display diagram of the collision risk prediction result of the emergency rescue autonomous vehicle according to the present invention. Specific Embodiments
[0099] To make the purpose and technical solutions of this application clearer and easier to understand, the following further describes this application with reference to the accompanying drawings. However, the protection scope of this application is not limited thereto.
[0100] Refer to Figure 1 、 Figure 2 The emergency rescue autonomous vehicle collision risk prediction system based on HM-TRW and HAGENN structure search described in this application includes a data acquisition device, a data preprocessing module, a prediction model, and a visualization module.
[0101] The data acquisition device includes in-vehicle sensors (lidar, acceleration sensor, speed sensor, steering angle sensor, GPS, camera, etc.) in real vehicle experiments, roadside devices (cameras and speed measurement radars, etc.), and communication technologies (vehicle-to-vehicle communication technology and vehicle-to-infrastructure communication technology), which are used to collect the data of the vehicle itself and the data of surrounding moving objects. The data of the vehicle itself collected mainly includes the speed, acceleration, steering angle, yaw rate, pedal force, and energy consumption of the vehicle itself; the data of surrounding moving objects mainly includes the positions, speeds, accelerations, and movement trajectories of the vehicles and vulnerable traffic groups around the emergency rescue autonomous vehicle.
[0102] The data preprocessing module is mainly used to clean, normalize, extract features, reduce the dimension of data, and divide the data set for the collected raw data, so that the prediction model can learn better; the preprocessed data is stored in the database. The specific steps of preprocessing are as follows:
[0103] (1) Data cleaning: Remove duplicate data, noise data, and irrelevant data, and at the same time fill in missing values and delete outliers.
[0104] (2) Data normalization: Convert the vehicle operation and pedestrian movement data with different dimensions into a unified format for subsequent processing and analysis.
[0105] (3) Feature extraction: Select important features through feature importance analysis for the prediction of collision risk.
[0106] (4) Data dimension reduction: Perform dimension reduction processing on high-dimensional data through principal component analysis, reduce the computational complexity and memory occupancy, and at the same time retain the main feature information to improve the efficiency and accuracy of collision risk prediction.
[0107] (5) Dataset division: Divide the data in the database into a training set, a validation set, and a test set. The training set is used to train the model; the validation set is used to optimize the model's prediction performance, calculate the prediction error, and continuously adjust the model parameters to achieve a specific accuracy in collision risk prediction; the test set is used to evaluate and analyze the model's accuracy and generalization ability. The former is evaluated through accuracy metrics and the AUC curve (area under the curve), and the latter is evaluated by inputting data in different traffic scenarios (such as signalized intersections, unsignalized intersections, roundabouts, merge areas, etc.) to assess the stability of the model's predictions. The training set, validation set, and test set are divided in the ratio of 70%, 20%, and 10%.
[0108] The prediction model consists of a random walk algorithm, a hierarchical attention graph embedding neural network, and an optimal parameter search module. The random walk algorithm adopts a high-order memory and non-decreasing time constraint strategy to capture the importance and dynamics of surrounding moving objects. The hierarchical attention graph embedding neural network includes three levels: node-level attention, edge-level attention, and time-level attention. The first two levels respectively learn the importance of each surrounding moving object for the emergency rescue vehicle and the importance of various interaction relationships between them. The last level aggregates the important features learned under historical timestamps to calculate the collision risk. The optimal parameter search module includes an attention localization space and an attention parameterization space, and also includes a multi-stage differential structure search module to accelerate the calculation of attention in hierarchical attention graph embedding and improve the efficiency of collision risk prediction.
[0109] The visualization module is used to display the predicted collision risk so that the driver can take appropriate risk avoidance measures in a timely manner.
[0110] Furthermore, the collision risk prediction method for emergency rescue autonomous vehicles based on high-order memory-guided temporal random walk and hierarchical attention graph embedding neural network structure search mainly relies on high-order memory-guided temporal random walk and hierarchical attention graph embedding neural network to achieve efficient and accurate collision risk prediction. The construction process of this model is as follows:
[0111] (1) Dynamic heterogeneous graph (DyHG) construction: All moving objects (including emergency rescue vehicles and their surrounding moving objects) are represented as nodes, and various interaction relationships between them (such as approaching trends, competitive relationships, cooperative relationships, etc.) are represented as edges (refer to Figure 4 ), and the nodes and edges change dynamically over time, forming a dynamic heterogeneous graph (refer to Figure 5 ), and its expression is as follows:
[0112] G t =(v t ,e t ,u t )
[0113] Among them, v t is the set of nodes with node type h ∈ H (including the types of surrounding moving objects), and e t is the set of edges with edge type r ∈ R, represents the feature set of all moving objects (such as speed, acceleration, motion trajectory, etc.). H and R are the node type set and edge type set respectively, |H| and |R| represent the number of node types and the number of edge types, represents the set of real numbers, N is the number of nodes, and D is the feature dimension.
[0114] (2) High-Order Memory-Guided Temporal Random Walk (HM-TRW): When an emergency rescue vehicle predicts collision risks, the importance of surrounding moving objects varies, and the interaction relationships with the emergency rescue vehicle also have different degrees of severity. Therefore, this heterogeneous characteristic needs to be considered. At the same time, the surrounding moving objects are constantly changing during the driving process of the emergency rescue vehicle, and the interaction relationships with the emergency rescue vehicle also change accordingly, which will greatly increase the complexity of the dynamic heterogeneous graph structure. Therefore, this dynamic characteristic also needs to be considered. Thus, the high-order memory-guided temporal random walk algorithm is introduced.
[0115] Specifically, first, the heterogeneous characteristic is learned through high-order memory guidance, where the high-order memory is a first-in-first-out queue storing different types of surrounding moving objects. The specific steps are as follows:
[0116] Step 1: Transition vector. Set the initial high-order memory queue to be empty, and the type transition vector accesses surrounding moving objects of each type with equal probability. When accessing the surrounding moving object v j , its type transition vector is updated according to the following formula:
[0117]
[0118] In the formula, φ(v j ) represents the type of the surrounding moving object v j , Q represents the first-in-first-out queue, and Norm() represents returning the two-norm of the vector.
[0119] Step 2: Type conversion. Determine the type of the next surrounding moving object to be accessed according to the probability distribution in step 1, and adopt a search mechanism with a search factor α ∈ [0, 1] to solve the type trap problem, as follows: In the formula,
[0120]
[0121] where represents the surrounding moving object initially accessed by the type transition vector, and Pr(hn+1 ) represents that the type of the next peripheral moving object to be accessed is h n+1 probability.
[0122] Step 3: High-order memory recording. After accessing the peripheral moving object v j , the transfer vector is stored in the first-in, first-out queue for the next step:
[0123]
[0124] where Put is a queue operator, indicating that when the first-in, first-out queue is full, the first transfer vector is popped and placed at the end of the queue.
[0125] Secondly, learn the dynamic change law of the dynamic heterogeneous graph through non-decreasing time constraints. For the emergency rescue vehicle v i (ego vehicle), use h n+1 to represent the type of its peripheral moving object, so there is:
[0126]
[0127] where, represents the set of peripheral moving objects of the emergency rescue vehicle v i at the future timestamp, t′ represents the timestamp of the previous random walk, and φ(v j ) represents the type of the peripheral moving object v j .
[0128] Select the next moving object to be accessed from the set using an exponential decay distribution:
[0129]
[0130] where, t k ∈t, t represents the timestamp, and t k represents the timestamp of the k-th step of the future random walk; Pr(v n+1 ) represents the probability that the next moving object to be accessed is v n+1 , v k represents the peripheral moving object accessed in the k-th step of the random walk; the discount rate δ ∈ [0, 1] is used to correct the time probability distribution.
[0131] Finally, fuse the output of the high-order memory-guided time random walk algorithm with the original features of each peripheral moving object to obtain their respective node representations:
[0132]
[0133]
[0134]
[0135] Among them, and represent the original feature, the recognition embedded feature, and the final node representation of the surrounding moving object respectively, are the transfer vector, the original feature, and the one-hot vector of the surrounding moving object v j respectively, is the latent embedding of all moving objects, W f and W represent the learnable parameters that are not shared between the surrounding moving object v j and other moving objects, represents the node representation of the surrounding moving object v j input to the hierarchical attention map embedding neural network.
[0136] (3) Hierarchical attention map embedding neural network (HAGENN): The node representations of each surrounding moving object obtained by the high-order memory-guided temporal random walk algorithm are then fed into the hierarchical attention map embedding neural network for collision risk prediction, and further capture the importance and temporal evolution trend of the surrounding moving objects and interaction relationships to improve the efficiency and accuracy of prediction. Specifically, different information is aggregated through node-level attention, edge-level attention, and time-level attention.
[0137] For node-level attention, at time stamp t, for interaction relationship type r, the importance between the emergency rescue vehicle v i and its surrounding moving object v j can be calculated by the following formula:
[0138]
[0139] where σ is the activation function, x i , x j are the input representations of the emergency rescue vehicle v i , the surrounding moving object v j respectively, is a linear transformation matrix, || represents concatenation, represents all the surrounding moving objects of the emergency rescue vehicle v i at time stamp t for interaction relationship type r, a r is a weight vector that parameterizes the attention function for interaction relationship type r, is the transpose of a r , x k represents the k-th moving object around the ego vehicle. Thus, the node embedding of the emergency rescue vehicle v i for interaction relationship type r at time stamp t can be obtained:
[0140]
[0141] For edge-level attention, the attention mechanism is adopted to learn the importance of different types of interaction relationships and calculate through a multi-layer perceptron:
[0142]
[0143] where σ is the activation function, w T is the edge-level attention vector, U el and b el are the single-layer parameters of the multi-layer perceptron, then the fusion embedding of the emergency rescue vehicle v i considering the importance of different interaction relationship types can be expressed as:
[0144]
[0145] For time-level attention, aggregate the fusion embeddings of the emergency rescue vehicle at all timestamps and package them as indicating the number of historical timestamps used to predict the collision risk. Then calculate the query-key-value vectors of the fusion embedding:
[0146] P = G i ·U P
[0147] K = G i ·U K
[0148] V = G i ·U V
[0149] where P, K, and V represent the query, key, and value vectors respectively, and U P , U K , U V represent the corresponding matrices for converting G i into query, key, and value vectors.
[0150] Calculate the time-level attention using the softmax function:
[0151]
[0152] where Z i represents the time-level attention, is a mask matrix, and D′ is the dimension of the query-key-value vectors. Taking Z i T as the final fusion embedding, the collision risk can be calculated:
[0153] Y = softmax(W2·ReLU(W1·Z i T +b1)+b2)
[0154] where softmax(·) is the output activation function, W1 and W2 are the weight matrices of the hierarchical attention map embedding neural network, ReLU(·) is the activation function, and b1 and b2 represent the bias terms.
[0155] (4) Attention localization and parameterized space: When predicting the collision risk, the hierarchical attention map embedding neural network needs to calculate the attention among different moving objects, different interaction relationships, and different timestamps, which has a large computational cost, resulting in low prediction efficiency and thus unable to predict the collision risk in a timely manner. To address this, attention localization and parameterized space are proposed to sparsify the attention to achieve a more efficient architecture.
[0156] a) Attention localization
[0157] Before adopting the attention localization space, the time complexity of the hierarchical attention map embedding is shown in the following formula:
[0158]
[0159] where T represents the number of historical timestamps used to predict the collision risk, and respectively represent the number of interaction relationships of type r at timestamps t' and t, and O(·) represents the time complexity.
[0160] By the matrix the types of surrounding moving objects, types of interaction relationships, and the number of timestamps to be concerned can be selected, and the specific calculation is as follows:
[0161]
[0162] Furthermore, when the timestamp is t, by the matrix it can be determined whether to pay attention to the surrounding moving objects of interaction type r (all the surrounding moving objects of the emergency rescue vehicle v of interaction relationship type r at timestamp t′ i ), It can be said that completely determines the application position of the attention function.
[0163] By using the attention localization space, the time complexity is greatly reduced:
[0164]
[0165] where represents The number of non-zero values. By controlling the total number, the time complexity of the hierarchical attention map embedding can be reduced to be independent of both T and |R|.
[0166] b) Parametric space
[0167] To further reduce the number of parameters, a parametric space is proposed to search for the calculation method of the attention function. The expression of the parametric space is as follows:
[0168] A Pa = A N ×A R
[0169] where A N = {1, …, K N} T×|H| is the parametric matrix of the node mapping function F N (·); A R = {1, …, K R} 2T×|R| is the parametric matrix of the edge mapping matrix F R (·); K N and K R are two hyperparameters. The mapping functions of the surrounding moving objects and interaction relationships are respectively selected from A N and A R .
[0170] By using the above parametric space, the shared parameters applicable to similar traffic scenarios and rescue tasks can be adaptively searched and learned. In addition, using the parametric space can also reduce the number of learnable parameters. The number of learnable parameters of the original hierarchical attention map embedding neural network can be expressed as O(T|H| + |R|). By adopting the parametric space, this number can be reduced to O(K N + K R ). When K N and K R are constrained to be constants, the number of learnable parameters also becomes a constant.
[0171] (5) Multi-stage differential search: To further reduce the parameter search complexity of the positioning space and the parametric space and achieve efficient collision risk prediction, heuristic constraints are proposed to remove invalid structures in the search space, and a one-time neural structure search algorithm is adopted to speed up the search process.
[0172] a) Space constraint: Two constraints are introduced to reduce the search range and limit the complexity: one is that the emergency rescue vehicle can only receive information from the surrounding moving objects in the historical time; the other is to constrain the number of surrounding moving objects and the number of interaction relationships used for collision risk prediction at each timestamp by 1 ≦ t ≦ T, where is a hyperparameter.
[0173] b) Hypernetwork construction: Use a hypernetwork to convert the parameter search in the positioning space and the parameterization space into a single neural architecture search problem. Specifically, represent the selection of operations as a probability distribution:
[0174]
[0175] where \(x\) is the input, is the output, \(|A|\) represents the number of operations, and \(\beta\) i represents the mixing weight of the mapping function \(F\) i (·) corresponding to the \(i\)-th operation. Here, the selection of operations represents the application of an attention function or the selection of moving objects and interaction relationships.
[0176] By adopting a hypernetwork, jointly optimize all the parameters in the mixing weight \(\beta\) and the mapping function in a differentiable manner:
[0177]
[0178] where \(\eta\) w and \(\eta\) β represent the learning rates of the structure weight and the model weight respectively, and represent the loss functions of the training set and the validation set respectively, \(w\) represents the structure weight, and \(\beta\) represents the model weight.
[0179] c) Multi-stage hypernetwork training: To stabilize the training of the hypernetwork, divide the training process into three stages: moving object parameterization, interaction relationship parameterization, and attention positioning space search, with each stage focusing on a different parameterization space.
[0180] Refer to Figure 3 , before performing actual collision risk prediction, it is necessary to train and validate the model, continuously adjust and optimize the model parameters to make it have the best prediction performance. Then, use the test set data to evaluate and analyze the prediction performance of the model. Finally, apply the prediction model to the actual traffic scenario and visualize the collision risk prediction results.
[0181] Furthermore, the process of model training and validation is as follows:
[0182] (1) Data input: Extract the feature vectors of the ego vehicle data and the surrounding moving object data from the divided training set as the input of the model.
[0183] (2) Model training: The high-order memory-guided temporal random walk algorithm and the hierarchical attention graph embedding neural network learn the importance of different moving objects and their interaction relationships, while the latter predicts the collision risk distribution at the current timestamp based on historical data.
[0184] (3) Model evaluation: Calculate the difference between the predicted collision risk and the actual collision risk.
[0185] (4) Parameter optimization: According to the model evaluation results, use the optimizer to optimize and update the model parameters, thereby reducing the prediction error.
[0186] (5) Iterative loop: Repeatedly execute the above steps (2) to (4), and use different training sample groups to iteratively train the model. Update the model parameters at the end of each iteration.
[0187] (6) Model validation: During the training process, periodically use the validation set to evaluate the model performance, monitor the generalization ability and prediction accuracy of the model, and adjust the model hyperparameters according to the validation results to obtain the best model performance.
[0188] Figures 7(a), (b), (c), and (d) are the display diagrams of the collision risk prediction results of the emergency rescue autonomous vehicle described in this application. It should be noted that this figure is only an interface example diagram for the collision risk prediction of the emergency rescue autonomous vehicle, only contains necessary functions, and can be improved according to specific needs in the future. It is not a constraint condition of this application here.
[0189] Referring to Figure 7(a), this interface consists of five parts. Frame ① shows that the function interface is the collision risk prediction interface. Other function interfaces such as driving model switching and emergency rescue information are not described here and can be further improved in the future; Frame ② shows the current operating state of the vehicle, specifically including starting, driving, braking, stopping, etc.; Frame ③ shows the current remaining battery power, signal strength, and time; Frame ④ shows the four display interface buttons of the collision risk prediction system: "Real-time road scene", "Vehicle operation data", "Collision risk prediction", and "Safe driving advice". Specific display information can be viewed by selecting different function buttons; Frame ⑤ corresponds to the specific information of the function button selected in Frame ④.
[0190] Specifically, the "Real-time road scene" interface (referring to Figure 7(a)) can comprehensively display the road environment where the emergency rescue vehicle is located at the current moment and the specific positions of the surrounding moving objects, enabling the driver to specifically and comprehensively perceive and understand the surrounding environment information; the "Vehicle operation data" interface (referring to Figure 7(b)) shows the motion state and driving route of the vehicle at the current moment, and these information can guide the driver's behavior operations and are crucial for collision risk prediction; the "Collision risk prediction" interface (referring to Figure 7(c)) shows the surrounding high collision risk objects and potential collision risk objects; the "Safe driving advice" interface (referring to Figure 7(d)) analyzes the predicted collision risk, provides a safe driving plan for the driver, and reminds the driver in the form of voice broadcast.
[0191] The graph embedding adopted by the emergency rescue autonomous driving vehicle collision risk prediction system based on HM-TRW and HAGENN structure search in this application is an algorithm for learning the complex relationships between nodes. It encodes all the nodes in the graph and maps them into equal-dimensional vectors that can be directly used by machine learning algorithms to achieve efficient and accurate prediction. To retain more effective information between nodes, this application further expands the graph embedding algorithm into a hierarchical attention graph embedding neural network algorithm, improving the capture ability for structural heterogeneity and dynamics. In the collision risk prediction of emergency rescue vehicles, the hierarchical attention graph embedding neural network can predict the potential conflict relationships between them by learning the complex relationships between emergency rescue vehicles and surrounding moving objects, and then predict the collision risk. Therefore, this application is expected to provide accurate and efficient collision risk prediction services for emergency rescue vehicles. On this basis, it can further guide the driver's decision-making and help autonomous driving vehicles plan safer and more efficient driving paths, thus better serving the emergency rescue cause.
[0192] The emergency rescue autonomous driving vehicle collision risk prediction method based on HM-TRW and HAGENN structure search described in this application, on the basis of perceiving the surrounding environment of the autonomous driving, combines the running data of the vehicle itself, and abstracts the complex interaction relationship between the emergency rescue vehicle and the surrounding moving objects into a dynamic heterogeneous graph; then, uses the random walk algorithm and the hierarchical attention graph embedding algorithm to model their complex interaction relationship, and learns the importance of the surrounding moving objects and various interaction relationships between them and the emergency rescue vehicle; then, inputs the vehicle itself data and the surrounding moving object data in the test set into the trained model to calculate the collision risk at future timestamps; finally, according to actual needs, visualizes the predicted collision risk to assist the driver in taking evasive actions safely and efficiently. This application can provide more efficient and accurate collision risk prediction for emergency rescue autonomous driving vehicles, ensuring the efficiency and safety of emergency rescue, and is conducive to building a safe traffic environment.
[0193] The above embodiments are the preferred embodiments of this application, but this application is not limited to the described embodiments. Any obvious modifications and substitutions that those skilled in the art can make to this application are within the protection scope of this application.
Claims
1. Emergency rescue autonomous vehicle collision risk prediction method based on HM-TRW and HAGENN structure search, characterized in that: Dataset construction: In real vehicle experiments, self-vehicle data and surrounding moving object data are collected by data acquisition equipment, and after preprocessing the data, it is stored in a database; Collision risk prediction model construction: Use the collected self-vehicle data and surrounding moving object data to construct a dynamic heterogeneous graph. The high-order memory-guided temporal random walk HM-TRW algorithm learns the heterogeneous characteristics and dynamic change rules of the dynamic heterogeneous graph, and the node representations of each surrounding moving object obtained are then fed into the hierarchical attention graph embedding neural network HAGENN for collision risk prediction; Model training: Input the training set and validation set in the database into the collision risk prediction model respectively to optimize the model parameters; Model testing: Use the test set in the database to test the trained collision risk prediction model, and evaluate and analyze the model prediction performance according to the test results; Collision risk prediction: Real-time collect the data of the emergency rescue autonomous vehicle and the surrounding moving object data, and input them into the tested collision risk prediction model to realize the prediction of collision risk; The construction of the dynamic heterogeneous graph is as follows: The emergency rescue vehicle and its surrounding moving objects represent nodes, and various interaction relationships between them are represented as edges. The nodes and edges change dynamically over time, forming a dynamic heterogeneous graph, and its expression is as follows: G t =(v t ,e t ,u t ) where, v t is the set of nodes with node type h, e t is the set of edges with edge type r, represents the feature set of all moving objects, h ∈ H, r ∈ R, where H and R are the node type set and the edge type set respectively, represents the set of real numbers, N is the number of nodes, and D is the feature dimension; The learning of the heterogeneous characteristics of the dynamic heterogeneous graph includes: (1) Transfer vector Set the initial high-order memory queue to be empty. The type transfer vector accesses the surrounding moving objects of each type with equal probability. When accessing the surrounding moving object v j its type transfer vector is updated according to the following formula: In the formula, represents the type of the surrounding moving object v j , Q represents a first-in-first-out queue, and Norm() represents returning the two-norm of a vector; (2) Type conversion According to the probability distribution, determine the type of the next peripheral moving object to be accessed, and adopt a search mechanism with a search factor α ∈ [0, 1] to solve the type trap problem, as follows: In the formula, represents the surrounding moving object initially accessed by the type transfer vector, and Pr(h n+1 ) represents the probability that the type of the next surrounding moving object to be accessed is h n+1 ; (3) High-order memory recording After accessing the surrounding moving object v j the type transfer vector is stored in the next-in, first-out queue Q': Among them, Put is a queue operator, indicating that when the first-in first-out queue is full, the first transfer vector is popped and placed at the end of the queue; The learning of the dynamic change rules of the dynamic heterogeneous graph is specifically: For the emergency rescue vehicle v i , use h n+1 to represent the type of moving objects around it. Thus, we have: Among them, represents the set of surrounding moving objects of the emergency rescue vehicle v i at the future timestamp, and t′ represents the timestamp of the previous random walk; Select the next motion object to be accessed from the set using an exponential decay distribution: where the time stamp \(t\) of the \(k\)-th future random walk k \(\in t\), \(t\) represents the time stamp of the random walk; \(\Pr(v\) n+1 ) represents the probability that the next moving object to be visited is \(v\) n+1 , \(v\) k represents the surrounding moving object visited in the \(k\)-th random walk; the discount rate \(\delta\in[0,1]\); The node representation of each surrounding moving object is the fusion of the output of the high-order memory-guided temporal random walk algorithm and the original features of each surrounding moving object; Among them, are the type transfer vector, original feature, and one-hot vector of the surrounding moving object v j respectively, is the potential embedding of all moving objects, represents the set of real numbers, N is the number of nodes, D is the feature dimension, and W f and W represent the learnable parameters that are not shared between the surrounding moving object v j and other moving objects, and x vj represent the recognition embedding feature and the final node representation of the surrounding moving object respectively.
2. The method for predicting the collision risk of an emergency rescue autonomous vehicle according to claim 1, wherein The hierarchical attention graph embedding neural network predicts collision risk by aggregating different information through node-level attention, edge-level attention, and time-level attention; For node-level attention, at time stamp t, for interaction relationship type r, for emergency rescue vehicle v i and its surrounding moving objects v j The importance between is calculated by the following formula: Among them, σ is the activation function, and x i and x j are the input representations of the emergency rescue vehicle v i and the surrounding moving object v j respectively, is a linear transformation matrix, || represents concatenation, represents all the surrounding moving objects of the emergency rescue vehicle v i at the interaction relationship type r at the time stamp t, a r is the weight vector, is the transpose of a r , and x k represents the input representation of the k-th surrounding moving object around the ego vehicle; The node embedding of the emergency rescue vehicle v with the interaction relationship type r at the timestamp t i is expressed as: denoted as: For edge-level attention, an attention mechanism is adopted to learn the importance of different types of interaction relationships and calculated through a multi-layer perceptron: where, w T is the edge-level attention vector, U el and b el are the single-layer parameters of the multi-layer perceptron, and |R| is the number of edge types; Emergency rescue vehicle v considering the importance of different types of interaction relationships i Fusion embedding Expressed as: For temporal attention, aggregate the fused embeddings of emergency rescue vehicles at all timestamps and package them as T represents the number of historical timestamps used to predict collision risk; calculate the query-key-value vectors of the fused embeddings: P = G i ·U P K = G i ·U K V = G i ·U V Among them, P, K, and V respectively represent query, key, and value vectors, and U P , U K , U V represent the corresponding matrices for converting G i into query, key, and value vectors. D is the feature dimension, represents the set of real numbers; Use the softmax function to calculate the time-level attention: Among them, Z i represents the temporal-level attention, is a mask matrix, D ′ is the dimension of the query-key-value vector; Take Z i T As the final fusion embedding, calculate the collision risk: Y = softmax(W2·ReLU(W1·Z i T + b1)+ b2) In the formula, softmax(·) is the output activation function, W1, W2 are the weight matrices of the hierarchical attention graph embedding neural network, ReLU(·) is the activation function, and b1, b2 represent the bias terms.
3. The emergency rescue autonomous vehicle collision risk prediction method according to claim 2, wherein Sparsify the attention calculated by the hierarchical attention graph embedding neural network when predicting collision risk using the attention localization space; Through a matrix Select the types of surrounding motion objects of interest, types of interaction relationships, and the number of timestamps. The specific calculation is as follows: When the timestamp is t, determine whether to pay attention to the surrounding moving objects with the interaction type r through the matrix All the surrounding moving objects of the emergency rescue vehicle v representing the interaction relationship type r at the timestamp t' ; i By using the attention localization space, the time complexity is: Among them, represents the number of non-zero values in A, LO T represents the number of historical timestamps used to predict the collision risk, and respectively represent the number of interaction relationships of type r at timestamps t' and t, O(·) represents the time complexity, and |R| represents the number of edge types.
4. The emergency rescue autonomous vehicle collision risk prediction method according to claim 3, wherein Sparsify the attention calculated by the hierarchical attention graph embedding neural network when predicting collision risk using the attention parameterization space; Utilize the attention parameterized space A Pa Search for the attention function, and the expression is as follows: A Pa = A N × A R Among them, A N ={1, …, K N} T×|H| is the parameterization matrix of the node mapping function F N (·), A R ={1, …, K R} 2T×|R| is the parameterization matrix of the edge mapping matrix F R (·), K N and K R are two hyperparameters, and |H| represents the number of node types.
5. The method for predicting the collision risk of an emergency rescue autonomous vehicle according to claim 4, wherein Use multi-stage differential search to reduce the parameter search complexity of the localization space and the parameter space: a) Space constraint The following two constraints are introduced to reduce the search scope and limit the complexity: First, the emergency rescue vehicle can only receive information about surrounding moving objects from historical time, and second, is used to constrain the number of surrounding moving objects and the number of interaction relationships for collision risk prediction at each timestamp, where is a hyperparameter, 1 ≤ t ≤ T; b) Supernet construction Use the supernet to convert the parameter search in the localization space and the parameterization space into a single neural architecture search problem, specifically representing the selection of operations as a probability distribution: where x is the input, is the output, |A| represents the number of operations, and β i represents the model weight of the mapping function F i (·) corresponding to the i-th operation; By adopting the supernet, jointly optimize all the parameters in the model weight β and the mapping function in a differentiable manner: Among them, η w and η β represent the learning rates of the structure weight and the model weight respectively, and represent the loss functions of the training set and the validation set respectively, w represents the structure weight, and β represents the model weight; c) Multi-stage supernet training To stabilize the training of the super network, the training process is divided into three stages: moving object parameterization, interaction relationship parameterization, and attention localization space search. Each stage focuses on a different parameterization space.
6. A system for implementing the emergency rescue autonomous driving vehicle collision risk prediction method based on HM-TRW and HAGENN structure search according to any one of claims 1-5, characterized in that, Including: Data acquisition equipment, including in-vehicle sensors, roadside devices, and communication technologies, for collecting vehicle data and surrounding moving object data; Data preprocessing module, for cleaning, normalizing, feature extracting, data dimension reduction, and dataset partitioning of the collected data; Prediction model, including a high-order memory-guided temporal random walk algorithm, a hierarchical attention map embedding neural network, and an optimal parameter search module. The optimal parameter search module includes an attention localization space, an attention parameterization space, and a multi-stage differential search module; Visualization module, for displaying the predicted collision risk.
Citation Information
Patent Citations
Driving safety monitoring device and method based on car networking BSM (Basic Safety Message) information fusion
CN109660967A
Complex scene driving risk prediction method based on multiple time-space diagrams
CN113762473A