A method for generating accident scenarios based on scene knowledge graphs and considering accident causes.

By constructing an accident scene generation method based on scene knowledge graphs and combining causal reinforcement learning and deep learning, a highly realistic and diverse accident scene database is generated, which solves the problem of scarce accident scene data in autonomous driving algorithms and improves the adaptability and safety of the algorithms.

CN120496333BActive Publication Date: 2025-10-31JILIN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510983114.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-10-31
Estimated Expiration
2045-07-17

AI Technical Summary

Technical Problem

Existing autonomous driving algorithms lack high-quality accident scenario data during training and verification, making it difficult to adapt to complex accident scenarios in the real world. In particular, the scarcity of accident scenario data and the insufficient realism of simulation scenarios make it difficult to achieve causal controllable generation of similar accident scenarios.

Method used

By employing a scenario-based knowledge graph approach, combined with causal reinforcement learning and deep learning, an accident scenario generation framework is constructed. Through scenario graphs and accident causation graphs, the diversity of scenarios and the consistency of causal relationships are enhanced, generating a highly realistic and diverse accident scenario database.

Benefits of technology

Under conditions of limited data sample size, this study significantly improves the adaptability of autonomous driving algorithms to similar accident scenarios, expands the number of accident scenarios in the closed-loop self-evolution process, and ensures the safe and reliable operation of autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496333B_ABST
    Figure CN120496333B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of autonomous driving testing technology, specifically a method for generating accident scenarios based on scene knowledge graphs and considering accident causes. It includes the following steps: Step 1, modeling the scene knowledge graph; Step 2, modeling a scene graph temporal prediction model; Step 3, modeling a temporal causal inference model; Step 4, modeling a scene graph temporal decision generation model, thereby generating accident scenarios. This invention can construct an accident scenario database with high realism, high diversity, and consistency in accident causes under conditions of limited accident scenario sample data. It efficiently expands the number of similar accident scenarios in the closed-loop self-evolving cloud database of autonomous driving algorithms. Using the constructed scenario database to train autonomous driving algorithms can effectively enhance the adaptability of autonomous vehicles to similar accident scenarios that have already occurred, ensuring the safe and reliable operation of autonomous vehicles in the real world.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of autonomous driving testing technology, specifically a method for generating accident scenarios based on scene knowledge graphs and considering the causes of accidents. Background Technology

[0002] In recent years, autonomous vehicles have developed rapidly, with significant progress in core technologies such as perception, decision-making, and control, greatly enhancing their ability to cope with complex scenarios. However, real-world driving scenarios are infinitely diverse, and long-tail scenarios with extremely low probability of occurrence have a significant impact on the safe operation of autonomous vehicles. Accident scenarios, as an important component of long-tail scenarios, are highly risky and involve complex interactive behaviors, making them invaluable for the training and validation of autonomous driving algorithms. Currently, autonomous driving algorithms generally adopt a data-driven architecture, and the closed-loop self-evolution process is an efficient solution for autonomous driving algorithms to cope with infinite driving environments using limited scenarios, significantly improving their adaptability to accident scenarios. However, due to the significant long-tail characteristics of accident scenarios, the limited number of accident scenarios that have occurred and been uploaded to the cloud are insufficient to meet the data sample size requirements for cloud model training. Inconsistent accident scenario types in the cloud database also increase the learning difficulty of the algorithm, making it difficult to truly adapt to real-world accident scenarios.

[0003] The closed-loop self-evolution process of autonomous driving algorithms aimed at improving adaptability to accident scenarios relies on a large amount of scenario data consistent with the types of accidents that have already occurred. However, acquiring a large amount of high-quality accident scenario data of the same types as those that have already occurred faces many challenges. On the one hand, the high-risk real-vehicle data collection process is inefficient, and cloud databases are extremely scarce for the same types of accident scenario data. On the other hand, simulated accident scenarios lack realism and are difficult to fully simulate the complex dynamic interactive behaviors in the real world.

[0004] Accident scene generation is an effective solution to the aforementioned problems, and typical accident scene generation methods include deep generative models and reinforcement learning models. However, existing deep generative models lack consideration of the essential causes of accidents and rely on a large amount of scene data as input. However, accident scene data is extremely scarce. These reasons make it difficult for deep generative models to achieve causal and controllable generation of similar accident scene data. Reinforcement learning methods have strong randomness in the generation of time-series scenes and lack consideration of the causal relationship during scene generation, making it difficult to achieve efficient and controllable generation of similar accident scenes. At the same time, the scene data may not conform to the objective laws of the real world during the generation process, and the authenticity of the scene data still needs to be verified.

[0005] To address the shortcomings of existing methods for constructing accident scene databases of the same type under limited data sample size, this invention proposes an accident scene generation framework enhanced by causal reinforcement learning and deep learning. This framework enhances the representation of driving scenarios in the form of scene graphs, and, in conjunction with deep learning models, achieves efficient extraction of realistic interaction characteristics and enhances scene diversity. By utilizing accident causation graphs in conjunction with reinforcement learning models to control the consistency of causal relationships, this framework can construct a database of accident scenes of the same type with high realism and diversity, even under limited data sample size, providing data support for the closed-loop self-evolution process of autonomous driving algorithms. Summary of the Invention

[0006] To address the aforementioned issues, this invention provides an accident scene generation method based on scene knowledge graphs and considering accident causes. This method can be used in the closed-loop self-evolution process of autonomous driving algorithms, significantly improving the adaptability of autonomous driving algorithms to similar accident scenarios.

[0007] The technical solution of this invention is described below in conjunction with the accompanying drawings:

[0008] This invention provides a method for generating accident scenarios based on scene knowledge graphs and considering accident causes, comprising the following steps:

[0009] Step 1: Model the scene knowledge graph;

[0010] A scene knowledge graph representation method is constructed, which includes scene graph SG and accident cause graph ACG. Scene graph SG is used to aggregate the interaction features of scene in spatial and temporal dimensions, and accident cause graph ACG is used to realize causal control in the accident scene generation process.

[0011] Step 2: Model the scene graph time series prediction model;

[0012] A scene graph SG is processed using a graph attention neural network to extract its spatial dimension information. A Transformer-based temporal feature extraction architecture is used to learn the long temporal dependencies of the temporal scene graph SG to extract its temporal dimension information. A variational graph autoencoder is used to construct a latent space for repeated sampling of latent variables and form a sampling layer. The latent variables are decoded by a decoder to obtain the scene graph sampling space, and a scene graph temporal prediction model is constructed.

[0013] Step 3: Model the time-series causal inference model;

[0014] Based on the temporal trajectory and related description of the input accident scene, a detailed analysis of the inherent causal relationship of the accident scene is conducted to obtain empirical knowledge. Based on this, an accident causal graph (ACG) and inference rules corresponding to each causal edge of the ACG graph are formulated, thereby obtaining a temporal causal inference model.

[0015] Step 4: Model the scene graph time-series decision generation model to generate accident scenarios;

[0016] By integrating key modules such as scene graph temporal prediction model and temporal causal inference model using reinforcement learning methods, a complete scene graph temporal decision generation model is constructed. Specifically, the scene graph temporal prediction model is used to generate the sampling space of the scene graph at the next time step in a guided manner, and the reinforcement learning agent is trained to sample in the sampling space. The temporal causal inference model is used to perform temporal causal inference on the generated temporal scene graph. A graph similarity measurement method is established between the inferred accident causation graph ACG0 and the real accident causation graph ACG, which serves as the main feedback reward when the reinforcement learning agent takes different actions, guiding the causal controllable generation of accident scenes.

[0017] Furthermore, the specific method for step one is as follows:

[0018] 11) Model the scene graph SG;

[0019] The scene information at each moment is modeled as a scene graph SG of one frame. t Each frame of the scene graph SG t It consists of several nodes and edges, where nodes represent different traffic participants and edges represent the interaction relationships between traffic participants; the scene graph SG corresponds to the traffic scenario at a certain time t. t Including node feature matrix N t Edge index matrix I t Edge feature matrix E t That is, SG t =(N t ,I t E t );

[0020] Node feature matrix Where m represents the number of nodes in the scene, i.e., the number of traffic participants, and n represents the node feature dimension. In this invention, n=4 node feature dimensions are selected, namely the position and velocity of the traffic participants in the x and y directions, i.e., (x, y, v). x ,v y );

[0021] Edge index matrix h represents the edge attribute dimension, the edge index matrix represents the interaction relationship between nodes, and a unidirectional edge N is defined. i →N j Two-way edge Self-circulating edge N i →N i Three edge types;

[0022] Edge feature matrix The edge carries a total of 5 types of feature information, namely △x, △y, △v. x ,△v y , , where △ x ,△ y For interactive nodes in x , y Relative position of direction; △ v x ,△ v y For interactive nodes in x , y Relative velocity in the direction; The interaction weighting coefficients are for vehicles at different distances in front of and behind the vehicle.

[0023] The Logistic function is used to model the interaction weights of vehicles at different distances in front of and behind the vehicle. The Logistic function is shown in equation (1);

[0024] (1)

[0025] In the formula, L is the maximum value of the interaction relationship weight coefficient; k is the rate at which the interaction relationship weight coefficient decays with distance; d0 is the distance when Logistic(d) = L / 2;

[0026] 12) Model the accident causation diagram (ACG);

[0027] The accident causation graph (ACG) consists of nodes and edges. In addition to the nodes representing traffic participants, a collision node C is introduced into the accident causation graph (ACG). The edges represent the causal relationships between nodes. Neither the nodes nor the edges carry any feature information. The essential causes of the accident scene are analyzed based on accident text data and video data. That is, the driving behavior of the vehicle that caused the accident is analyzed, and then the causal link of the accident scene is determined, thereby determining the specific content in the ACG graph.

[0028] Furthermore, the specific method for step two is as follows:

[0029] 21) Model the scene graph time series prediction model;

[0030] The scene graph temporal prediction model consists of three parts: a scene graph spatiotemporal encoder, a sampling layer, and a scene graph decoder. The scene graph spatiotemporal encoder is used to extract the spatiotemporal features of the input scene graph sequence. The sampling layer is used to sample in the latent space to obtain latent variables. The scene graph decoder decodes the latent variables to predict the scene graph sampling space at the next time step.

[0031] The scene graph spatiotemporal encoder consists of a graph attention neural network (GAT), a Transformer encoder, and fully connected layers. First, a multi-layer GAT is used to encode the node features at each time step. Then, a self-attention mechanism is used to aggregate node neighborhood information to extract spatial dimension features. For node features in the scene graph... Edge index features Edge features The output of the GAT network is shown in equation (2):

[0032] (2)

[0033] In the formula, N(i) is the set of neighboring nodes of node i; The features of node i at time step t; , The edge index features and edge features connecting node i and node j at time step t are respectively used to obtain the aggregated node feature matrix N.

[0034] The Transformer encoder is used to extract temporal features from the scene graph. The node feature sequence at each time step is assigned a position code (PE). The Transformer encoder outputs N. t As shown in equation (3):

[0035] (3)

[0036] In the formula, N represents the aggregated node features output by the GAT network; PE represents the positional encoding in a sine / cosine manner; N t Features of time-series nodes;

[0037] Design two fully connected layers to map the temporal coding information to the mean of the latent space according to equations (4) and (5). and variance :

[0038] (4)

[0039] (5)

[0040] In the formula, , This is the weight matrix of the fully connected layer. , This serves as a bias term when calculating latent variables;

[0041] Sampling layers are used to sample from the latent spatial distribution to obtain latent variables. As shown in equation (6);

[0042] (6)

[0043] In the formula, Standard deviation; The noise is a standard normally distributed noise, where I represents the identity matrix; d z Let z be the spatial dimension of the latent variable.

[0044] In the scene graph decoder, the latent variable z is expanded into a latent representation zt at each time step t. t Position encoding is used to process latent variables to recover temporal information, z t The data is fed into the Transformer decoder for processing, as shown in Equation (7), and the output is the node features at each time step. Together, they form the overall node feature N. d :

[0045] (7)

[0046] The features are mapped to the original node feature space through a fully connected layer, thus obtaining the predicted scene graph SG. p =( , , As shown in equation (8):

[0047] (8)

[0048] In the formula, W d This is the weight matrix of the fully connected layer of the decoder;

[0049] 22) Design the loss function;

[0050] The loss function includes the node feature reconstruction loss L. N Edge index reconstruction loss L I Edge feature reconstruction loss L E and KL divergence loss L KL Four parts;

[0051] Node feature reconstruction loss L N The mean squared error is used to measure the difference between the predicted node features and the actual node features, as shown in Equation (9):

[0052] (9)

[0053] In the formula, Let k be the value of the kth feature in the i-th node at time t+1. Let be the value of the k-th feature in the i-th node predicted at time t+1; m is the number of nodes in the scene, and n is the node feature dimension.

[0054] Edge index reconstruction loss L I The difference between the edge index features predicted by the model and the true index features is measured using cross-entropy loss, as shown in Equation (10):

[0055] (10)

[0056] In the formula, This is the actual edge index matrix, indicating whether there is an edge between node i and node j; if there is, it is 1, otherwise it is 0. The value is between 0 and 1, representing the probability that an edge exists between node i and node j.

[0057] Edge feature reconstruction loss L E The mean squared error (MSE) is used to measure the difference between the edge features predicted by the model and the actual edge features, as shown in Equation (11):

[0058] (11)

[0059] In the formula, e is the total number of edges in the scene graph at time t+1; h is the number of features carried by a single edge; Let be the value of the k-th feature in the edge formed by nodes i and j at time t+1; Let be the value of the k-th feature in the edge formed by nodes i and j predicted at time t+1;

[0060] KL divergence loss L KL Measuring distribution With prior distribution Differences between them, distribution yes The variational approximation is shown in equation (12):

[0061] (12)

[0062] The total loss L of the scene graph temporal prediction model is shown in equation (13):

[0063] (13)

[0064] In the formula, , , Representing L respectively I L E L KL The weight hyperparameters of the loss;

[0065] Furthermore, the specific method for step three is as follows:

[0066] 31) Based on the time-series trajectory information and text description of the accident scene, analyze what kind of driving behavior of the vehicle led to the accident, and then determine the causal link of the accident scene. In this way, the essential cause of the accident scene is obtained. Based on this, the ACG graph is modeled. The accident causation graph ACG includes vehicle nodes and edges representing the causal relationship between vehicle nodes.

[0067] 32) Based on the accident causation graph ACG, establish causal inference rules corresponding to each edge, take a scene graph sequence of arbitrary length as input, and infer the causal relationship graph corresponding to the scene graph sequence.

[0068] Furthermore, the specific method for step four is as follows:

[0069] 41) Define a graph similarity quantification method between the inferred accident causation graph ACG0 and the actual accident causation graph ACG, which serves as the main reward for the scene graph temporal decision generation model to control the accident scene generation type. The Jaccard similarity S is calculated by equation (14). J As a quantitative evaluation method between ACG0 diagrams and ACG diagrams:

[0070] (14)

[0071] In the formula, S J S represents the Jaccard graph similarity between ACG0 and ACG. J A larger value indicates that the ACG0 graph is more similar to the real ACG graph, meaning that the sequential scene has a more realistic cause of the accident; E0 and E are the set of oriented edges in the ACG0 and ACG graphs, respectively; TP=E0∩E are correctly identified causal edges with the correct direction; FP=E0\E are incorrectly identified causal edges that appear in E0 but not in E; FN=E\E0 are missing causal edges in the inference result that appear in E but not in E0.

[0072] 42) Assuming the temporal scene represented by the scene graph SG follows a Markov decision process, the reinforcement learning algorithm includes the state space S, the action space A, and the state transition function. Let s represent the state, a represent the action, and the reward function be R(s,a), where state s t ={SG t~t+n-1 This includes a fixed-length time sequence scene graph, and action a. t For a scenario graph SG containing multiple time points t+n t+n Sampling is performed in a discrete sampling space, and state transition is executed by the agent using action a. t Then a new state s is obtained.t+1 , that is s t+1 ={SG t+1~t+n}=f(s t ,a t ), where f is the scene graph update function considering trajectory smoothness and vehicle dynamics constraints, and the reward function R(s) is... t ,a t ) represents the state s in the current state. t Take action a t The reward obtained at that time consists of three parts, as shown in equation (15):

[0073] (15)

[0074] In the formula, R similarity The similarity reward between ACG0 and ACG obtained from the temporal scene graph inference is shown in Equation (16):

[0075] (16)

[0076] In the formula, a is the similarity reward coefficient;

[0077] R smootness The real-time negative reward for trajectory smoothness is used to measure the smoothness of the trajectory, as shown in Equation (17);

[0078] (17)

[0079] In the formula, t is the original trajectory; t* is the trajectory fitted by the SG algorithm; b is the smoothness reward coefficient;

[0080] R constraint The real-time negative reward for violating constraints is that actions that violate physical constraints or vehicle dynamics constraints are penalized, as shown in equation (18):

[0081] (18)

[0082] In the formula, This indicates that the sampling point violated the constraint, and the value is set to 1, where c is the penalty coefficient weight.

[0083] 43) Integrate traffic rule constraints and vehicle dynamics constraints into the reinforcement learning algorithm as the underlying hard constraints, and the upper layer is a reinforcement learning sampling strategy for inverse optimization. Traffic rule constraints mean that vehicles must not leave the road boundary and vehicles must not move in the opposite direction of the lane. Vehicle dynamics constraints specify the range of horizontal and vertical positions and speeds that the sampling points can reach.

[0084] If a sampling point meets one of the following conditions, it is considered to violate the constraint condition ViolationPenalty in equation (18), and the agent is given a constraint penalty reward: ① The sampling point violates the temporal motion trend of the trajectory, that is, it samples in the opposite direction of the temporal trajectory; ② The vehicle drives out of the road boundary; ③ The difference between the lateral position of the sampling point and the previous temporal point exceeds the vehicle dynamics limit; ④ The difference between the longitudinal position of the sampling point and the previous temporal point exceeds the vehicle dynamics limit.

[0085] 44) Select the PPO reinforcement learning algorithm to sample in the scene graph sampling space generated in step two, and set the loss of the PPO algorithm into three parts: policy loss, value loss, and entropy loss, which are used to optimize the policy, train the value network, and promote policy exploration, respectively, as shown in equation (19):

[0086] (19)

[0087] In the formula, c1 and c2 are the weight values ​​for different losses, respectively; , , These are strategy loss, value loss, and entropy loss, respectively.

[0088] The beneficial effects of this invention are as follows:

[0089] The accident scene generation method based on scene knowledge graph and considering accident causes provided by this invention can construct an accident scene database with high realism, high diversity and consistency of accident causes under the condition of limited accident scene sample data. It efficiently expands the number of similar accident scenes in the closed-loop self-evolutionary cloud database of autonomous driving algorithms. By using the constructed scene database to train autonomous driving algorithms, it can effectively enhance the adaptability of autonomous vehicles to similar accident scenes that have occurred, and ensure the safe and reliable operation of autonomous vehicles in the real world. Attached Figure Description

[0090] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0091] Figure 1 This is a flowchart of the present invention;

[0092] Figure 2 A schematic diagram of the traffic scene at time t;

[0093] Figure 3 Let SG be the scene graph corresponding to the traffic scene at time t.t Schematic diagram;

[0094] Figure 4 A schematic diagram of the spatial dimension features of the scene graph;

[0095] Figure 5 A schematic diagram illustrating the temporal dimension features of a scene graph;

[0096] Figure 6 This is a schematic diagram of accident scenario type 1;

[0097] Figure 7 This is a schematic diagram of the accident causation diagram ACG1 corresponding to accident scenario type 1.

[0098] Figure 8 This is a schematic diagram of accident scenario type 2;

[0099] Figure 9 This is a schematic diagram of the accident causation diagram (ACG2) corresponding to accident scenario type 2.

[0100] Figure 10 This is a diagram of the architecture of a scene graph time series prediction model;

[0101] Figure 11 This is a schematic diagram of the temporal motion process of the accident scene;

[0102] Figure 12 This is an ACG diagram corresponding to the temporal motion process of the accident scene;

[0103] Figure 13 This is a schematic diagram of the trajectory point density distribution in the accident scenario library 1.

[0104] Figure 14 A top view of the trajectory point density distribution in Accident Scene Library 1;

[0105] Figure 15 This is a schematic diagram of the trajectory point density distribution in the accident scenario database 2.

[0106] Figure 16 This is a top view of the density distribution of trajectory points in the accident scenario library 2. Detailed Implementation

[0107] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.

[0108] Example 1;

[0109] See Figure 1This embodiment provides a method for generating accident scenarios based on scene knowledge graphs and considering accident causes, including the following steps:

[0110] Step 1: Model the scene knowledge graph;

[0111] A scene knowledge graph representation method is constructed, comprising a scene graph (SG) and an accident cause graph (ACG). The scene graph (SG) is used to aggregate the interaction features of the scene in the spatial and temporal dimensions, and the accident cause graph (ACG) is used to implement causal control in the accident scene generation process, as detailed below:

[0112] 11) First, model the scene graph SG, refer to... Figure 2 and Figure 3 The scene information at each moment is modeled as a scene graph SG of one frame. t Each frame of the scene graph SG t It consists of several nodes and edges. Nodes represent different traffic participants, and edges represent the interaction relationships between traffic participants. See [link / reference] Figure 4 and Figure 5 A single-frame scene graph can represent the spatial dimension features of the scene at the current moment. It describes the interaction features between vehicles in a single-frame scene. The temporal changes of the scene graph represent the temporal dimension features of the scene. It describes the temporal interaction features between vehicles. The scene graph can effectively express the spatial and temporal dimensions of the scene.

[0113] Scene diagram SG corresponding to a traffic scenario at a certain time t t Including node feature matrix N t Edge index matrix I t Edge feature matrix E t That is, SG t =(N t ,I t E t Specifically:

[0114] Node feature matrix Where m represents the number of nodes in the scene, i.e., the number of traffic participants, and n represents the node feature dimension. In this invention, n=4 node feature dimensions are selected, namely the position and velocity of the traffic participants in the x and y directions, i.e., (x, y, v). x ,v y );

[0115] Edge index matrix h represents the edge attribute dimension, the edge index matrix represents the interaction relationship between nodes, and three edge types are defined: unidirectional edge, bidirectional edge, and self-looping edge, where unidirectional edge N i →N j Represents node N i The action will affect node N jBehavior, bidirectional edge N represents traffic participant i N j They will influence each other, with N self-circulating edges. i →N i Represents node N i Its actions are unaffected by other nodes, but may affect other nodes;

[0116] Edge feature matrix The edge carries a total of 5 types of feature information (△x, △y, △v). x ,△v y , ), where △ x ,△ y For interactive nodes in x , y Relative position of direction; △ v x ,△ v y For interactive nodes in x , y Relative velocity in the direction; The interaction weighting coefficients are used to represent the interaction relationships between vehicles at different distances in front of and behind the vehicle; they are used to describe the vehicle's attention to other vehicles in different positions around it, so as to achieve a more detailed and realistic expression of the interaction relationships between traffic participants.

[0117] Vehicles pay varying degrees of attention to other vehicles within a scene, tending to focus more on those nearby and in front of the vehicle, and less on those further away and behind. To incorporate realistic interaction relationships into the scene graph, a Logistic function is used to model the interaction relationship weights of vehicles at different distances in front of and behind the vehicle. The Logistic function is shown in equation (1);

[0118] (1)

[0119] In the formula, L is the maximum value of the interaction relationship weight coefficient; k is the rate at which the interaction relationship weight coefficient decays with distance; d0 is the distance when Logistic(d) = L / 2;

[0120] 12) Next, the accident causation graph (ACG) is modeled. The ACG also consists of nodes and edges. In addition to nodes representing traffic participants, this invention introduces collision nodes C into the ACG. Edges represent the causal relationships between nodes. Neither nodes nor edges carry any feature information. Based on accident text and video data, the essential causes of the accident scene are analyzed, i.e., what driving behaviors of the vehicle led to the accident, thereby determining the causal chain of the accident scene and thus the specific content in the ACG graph.

[0121] See Figure 6 , Figure 7 , Figure 8 and Figure 9 In the accident causation diagram ACG1 corresponding to accident scenario 1, the slow movement of vehicle 2 affected the behavior of vehicle 1, causing vehicle 1 to change lanes and overtake. At the same time, the slow movement of vehicle 4 also affected the behavior of vehicle 3, causing vehicle 3 to change lanes and overtake, resulting in a collision between vehicle 1 and vehicle 3. Both points to collision node C. In the accident causation diagram ACG2 corresponding to accident scenario 2, the slow movement of vehicle 2 motivated vehicle 1 to change lanes. At the same time, the slow movement of vehicle 3 also created favorable conditions for vehicle 1 to change lanes, leading to continuous lane-changing behavior by vehicle 1, ultimately colliding with vehicle 4, pointing to collision node C.

[0122] Step 2: Model the scene graph time series prediction model;

[0123] A Graph Attention Network (GAT) is used to process the scene graph SG to extract its spatial dimension information. A Transformer-based temporal feature extraction architecture is used to learn the long-term temporal dependencies of the temporal scene graph SG to extract its temporal dimension information. A Variational Graph Auto-Encoder (VGAE) is used to construct a latent space for repeated sampling of latent variables and form a sampling layer. After decoding the latent variables, the scene graph sampling space is obtained. A temporal prediction model for the scene graph is then constructed, as follows:

[0124] 21) See Figure 10 The scene graph temporal prediction model consists of three components: a scene graph spatiotemporal encoder, a sampling layer, and a scene graph decoder. The scene graph spatiotemporal encoder is used to extract the spatiotemporal features of the input scene graph sequence. The sampling layer is used to sample in the latent space to obtain latent variables. The scene graph decoder decodes the latent variables to predict the scene graph sampling space at the next moment.

[0125] The scene graph spatiotemporal encoder consists of a graph attention neural network (GAT), a Transformer encoder, and fully connected layers. Specifically, it first uses a multi-layer GAT to encode the node features at each time step, and then uses a self-attention mechanism to aggregate node neighborhood information to extract spatial dimension features. Edge index features Edge features The output of the GAT network is shown in equation (2):

[0126] (2)

[0127] In the formula, N(i) is the set of neighboring nodes of node i; This represents the characteristics of node i at time step t. , The edge index features and edge features connecting node i and node j at time step t are respectively used to obtain the aggregated node feature matrix N.

[0128] Then, a Transformer encoder is used to extract the temporal features of the scene graph, assigning a position code (PE) to the node feature sequence at each time step. The Transformer encoder outputs N. t As shown in equation (3):

[0129] (3)

[0130] In the formula, N represents the aggregated node features output by the GAT network, PE represents the positional encoding in sine and cosine manner, and N t Features of time-series nodes;

[0131] Design two fully connected layers to map the temporal coding information to the mean of the latent space according to equations (4) and (5). and variance :

[0132] (4)

[0133] (5)

[0134] In the formula, , This is the weight matrix of the fully connected layer. , This serves as a bias term when calculating latent variables;

[0135] Sampling layers are used to sample from the latent spatial distribution to obtain latent variables. As shown in equation (6):

[0136] (6)

[0137] In the formula, Standard deviation; The noise is a standard normally distributed noise, where I represents the identity matrix; d z Let z be the spatial dimension of the latent variable.

[0138] In the scene graph decoder, the latent variable z is expanded into a latent representation zt at each time step t. t Position encoding is used to process latent variables to recover temporal information, z t The data is fed into the Transformer decoder for processing, as shown in Equation (7), and the output is the node features at each time step. Together, they form the overall node feature N. d :

[0139] (7)

[0140] Then, by mapping the features to the original node feature space through a fully connected layer, the predicted scene graph SG can be obtained. p =( , , As shown in equation (8):

[0141] (8)

[0142] In the formula, W d This is the weight matrix of the fully connected layer of the decoder;

[0143] 22) Design the loss function;

[0144] The loss function includes the node feature reconstruction loss L. N Edge index reconstruction loss L I Edge feature reconstruction loss L E and KL divergence loss L KL Four parts;

[0145] Node feature reconstruction loss L N The mean squared error (MSE) is used to measure the difference between the predicted node features and the actual node features, as shown in equation (9):

[0146] (9)

[0147] In the formula, Let k be the value of the kth feature in the i-th node at time t+1. Let be the value of the k-th feature in the i-th node predicted at time t+1; m is the number of nodes in the scene, and n is the node feature dimension.

[0148] Edge index reconstruction loss L I To measure the difference between the edge index features predicted by the model and the true index features, cross-entropy loss is used for evaluation, as shown in Equation (10):

[0149] (10)

[0150] In the formula, This is the actual edge index matrix, indicating whether there is an edge between node i and node j; if there is, it is 1, otherwise it is 0. The value is between 0 and 1, representing the probability that an edge exists between node i and node j.

[0151] Edge feature reconstruction loss L E The mean squared error (MSE) is used to measure the difference between the edge features predicted by the model and the true edge features, as shown in Equation (11):

[0152] (11)

[0153] In the formula, e is the total number of edges in the scene graph at time t+1; h is the number of features carried by a single edge; Let be the value of the k-th feature in the edge formed by nodes i and j at time t+1; Let be the value of the k-th feature in the edge formed by nodes i and j predicted at time t+1;

[0154] KL divergence loss L KL Measuring distribution With prior distribution Differences between them, distribution yes The variational approximation is shown in equation (12):

[0155] (12)

[0156] The total loss L of the scene graph temporal prediction model is shown in equation (13):

[0157] (13)

[0158] In the formula, , , Representing L respectively I LE L KL The weight hyperparameters of the loss;

[0159] Step 3: Model the time-series causal inference model;

[0160] Based on the temporal trajectory and related description of the input accident scene, a detailed analysis of the inherent causal relationship of the accident scene is conducted to obtain empirical knowledge. Based on this, an accident causal graph (ACG) and inference rules corresponding to each causal edge of the ACG graph are formulated, thereby obtaining a temporal causal inference model.

[0161] 31) Based on the time-series trajectory information and text description of the accident scene, analyze what kind of driving behavior of the vehicle led to the accident, and then determine the causal link of the accident scene. In this way, the essential cause of the accident scene is obtained. Based on this, the ACG graph is modeled. The accident causation graph ACG includes vehicle nodes and edges representing the causal relationship between vehicle nodes.

[0162] For details, please refer to [link / reference]. Figure 11 and Figure 12 It depicts the temporal movement process of the accident scene and the corresponding ACG diagram. The scene takes place on a highway. Due to the slow speed of vehicles at nodes 2 and 4, vehicles at nodes 1 and 3 change lanes to overtake the vehicles in front, which leads to a collision and a serious traffic accident. The specific accident cause diagram ACG is shown as node 2 influencing the behavior of node 1 and node 4 influencing the behavior of node 3, which in turn causes nodes 1 and 3 to both point to the collision node C. The accident cause diagram ACG is thus constructed.

[0163] 32) Based on the accident causation graph ACG, establish causal inference rules corresponding to each edge, take a scene graph sequence of arbitrary length as input, and infer the causal relationship graph corresponding to the scene graph sequence.

[0164] See Figure 11 and Figure 12 In the accident scenario, a lane-changing causal rule can be constructed with the following conditions: ① The speed of the following vehicle is greater than the speed of the preceding vehicle; ② The longitudinal distance between the two vehicles is less than δx; ③ The lateral displacement of the following vehicle between two time steps is always less than δy1, and the interval between time steps is at least 1 second, specifically, the following vehicle does not produce obvious lane-crossing displacement and follows the preceding vehicle; ④ The lateral displacement of the following vehicle between two time steps is greater than δy2, specifically, the following vehicle produces obvious lane-changing behavior. If the above four conditions are met, a one-way causal edge is constructed from the preceding vehicle to the following vehicle, which is manifested as the slow movement of the preceding vehicle affecting the following vehicle to produce lane-changing behavior; for the collision causal rule, if the distance between the two vehicles is less than δ, then both point to the collision node C.

[0165] Step 4: Model the scene graph temporal decision generation model;

[0166] This paper utilizes reinforcement learning methods to integrate key modules such as a scene graph temporal prediction model and a temporal causal inference model to construct a complete scene graph temporal decision generation model. Specifically, the scene graph temporal prediction model is used to generate the sampling space of the scene graph at the next time step in a guided manner. A reinforcement learning agent is trained to sample in this sampling space. The temporal causal inference model is then used to perform temporal causal inference on the generated temporal scene graph. A graph similarity measurement method is established between the inferred accident causation graph ACG0 and the real accident causation graph ACG, serving as the primary feedback reward for the reinforcement learning agent when taking different actions. This guides the causal and controllable generation of accident scenes, as detailed below:

[0167] 41) Define a graph similarity metric between the inferred accident causation graph ACG0 and the actual accident causation graph ACG, where E0 and E represent the oriented edge sets in ACG0 and ACG, respectively, and TP d =E0∩E represents a correctly identified causal edge with the correct direction, FP=E0\E represents a misidentified causal edge that appears in E0 but not in E, and FN=E\E0 represents a missing causal edge in the inference result that appears in E but not in E0. Then, the Jaccard similarity S is calculated by equation (14). J This serves as a quantitative evaluation method between ACG0 diagrams and ACG diagrams:

[0168] (14)

[0169] In the formula, S J S represents the Jaccard graph similarity between ACG0 and ACG. J A larger value indicates that the ACG0 graph is more similar to the real ACG graph, meaning that the sequential scene has a more realistic cause of the accident; E0 and E are the set of oriented edges in the ACG0 and ACG graphs, respectively; TP=E0∩E are correctly identified causal edges with the correct direction; FP=E0\E are incorrectly identified causal edges that appear in E0 but not in E; FN=E\E0 are missing causal edges in the inference result that appear in E but not in E0.

[0170] 42) Assuming the temporal scene represented by the scene graph SG follows a Markov decision process, the reinforcement learning algorithm includes the state space S, the action space A, and the state transition function. Let s represent the state, a represent the action, and the reward function be R(s,a), where state s t ={SG t~t+n-1 This includes a fixed-length time sequence scene graph, and action a. t For a scenario graph SG containing multiple time points t+n t+nSampling is performed in a discrete sampling space, and state transition is executed by the agent using action a. t Then a new state s is obtained. t+1 , that is s t+1 ={SG t+1~t+n}=f(s t ,a t ), where f is the scene graph update function considering trajectory smoothness and vehicle dynamics constraints, and the reward function R(s) is... t ,a t ) represents the state s in the current state. t Take action a t The reward obtained at that time consists of three parts, as shown in equation (15):

[0171] (15)

[0172] In the formula, R similarity The similarity reward between ACG0 and ACG obtained from the temporal scene graph inference is shown in Equation (16):

[0173] (16)

[0174] In the formula, a is the similarity reward coefficient;

[0175] R smootness The real-time negative reward for trajectory smoothness is used to measure the smoothness of the trajectory, as shown in Equation (17):

[0176] (17)

[0177] In the formula, t is the original trajectory; t* is the trajectory fitted by the SG algorithm; b is the smoothness reward coefficient;

[0178] R constraint The real-time negative reward for violating constraints is that actions that violate physical constraints or vehicle dynamics constraints are penalized, as shown in equation (18):

[0179] (18)

[0180] In the formula, ViolationPenalty represents a sampling point violating the constraint, and its value is set to 1; c is the penalty coefficient weight;

[0181] 43) Traffic rule constraints and vehicle dynamics constraints are integrated into the reinforcement learning algorithm as the underlying hard constraints, and the upper layer is the reinforcement learning sampling strategy for inverse optimization. The traffic rule constraints mainly refer to the fact that vehicles must not drive out of the road boundary and vehicles must not move in the opposite direction of the lane. The vehicle dynamics constraints specify the range of horizontal and vertical positions and speeds that the sampling points can reach.

[0182] This invention defines a sampling point as violating the constraint condition ViolationPenalty in equation (18) if it meets one of the following conditions, and gives the agent a constraint penalty reward: ① The sampling point violates the temporal motion trend of the trajectory, that is, it samples in the opposite direction of the temporal trajectory; ② The vehicle drives out of the road boundary; ③ The difference between the lateral position of the sampling point and the previous temporal point exceeds the vehicle dynamics limit; ④ The difference between the longitudinal position of the sampling point and the previous temporal point exceeds the vehicle dynamics limit.

[0183] 44) This invention selects the PPO (Proximal Policy Optimization) reinforcement learning algorithm to complete the above tasks. The loss of the PPO algorithm is set into three parts: policy loss, value loss, and entropy loss, which are used to optimize the policy, train the value network, and promote policy exploration, respectively, as shown in Equation (19):

[0184] (19)

[0185] In the formula, c1 and c2 are the weight values ​​for different losses, respectively. , , These are strategy loss, value loss, and entropy loss, respectively.

[0186] Example 2;

[0187] This invention is aimed at Figure 6 , Figure 7 , Figure 8 and Figure 9 The two accident scenario types shown generate an accident scenario database, and the two databases are used in the closed-loop self-evolution process of the autonomous driving algorithm to improve the adaptability of the autonomous driving algorithm to such accident scenarios and verify the effectiveness of the method proposed in this invention.

[0188] This invention generates 1000 accident scene data points for each of two accident scenario types, forming two accident scenario databases. The density distribution of trajectory points in accident scenario database 1 and accident scenario database 2 are respectively referred to... Figure 13 , Figure 14 and Figure 15 and Figure 16 As can be seen from the figure, the density distribution of trajectory points fits into a clear parameter space, and the trajectory distribution range in the scene database is relatively wide, indicating the effectiveness of the method proposed in this invention in completing the task of constructing a database of similar accident scene scenarios.

[0189] Next, this invention uses one of the vehicles involved in the collision in the accident scenario as an autonomous vehicle, while the other vehicles use trajectories generated by the method presented in this paper. The autonomous vehicle is equipped with an autonomous driving algorithm that can predict traffic risks in real time. When the predicted traffic risk exceeds the accident scenario risk threshold, the autonomous vehicle can take corresponding driving actions to maintain a safe distance from other vehicles to avoid a collision. Poor risk prediction results in the autonomous vehicle lacking the ability to perceive accident risks, making it prone to traffic accidents. This invention uses a Long Short-Term Memory (LSTM) network as risk prediction algorithm 1, and an algorithm GTF constructed based on GAT and Transformer networks as risk prediction algorithm 2. The validation set consists of 1000 similar accident scenarios generated by the method presented in this paper, and the training set is divided into three groups: ① 1000 non-collision scenarios; ② 500 non-collision scenarios and 500 accident scenarios; ③ 1000 accident scenarios.

[0190] The non-collision scenarios in the training set are also generated by the method of this invention. They have the same basic elements as the accident scenarios and contain a large number of near-collision scenarios. It should be noted that the specific scenarios used in the training set and the validation set are not exactly the same. The true risk value is calculated based on prior accident scenario information, and historical traffic information is used to predict traffic risks in future moments during training. After training the algorithm on different training sets, the collision rates of the autonomous vehicle on the validation set are shown in Tables 1 and 2, respectively.

[0191] Table 1. Collision rates in the validation set for accident scenario type 1

[0192] Method Training set 1 Training set 2 Training set 3 LSTM 0.94 0.49 0.19 GTF 0.92 0.37 0.13

[0193] Table 2 Collision rates in the validation set for accident scenario type 2

[0194] Method Training set 1 Training set 2 Training set 3 LSTM 0.93 0.58 0.22 GTF 0.94 0.42 0.17

[0195] As can be seen from Tables 1 and 2, after training with more accident scenarios, the algorithm can accurately predict traffic risks under this type of accident scenario, reduce the average collision rate of autonomous vehicles by about 76%, and significantly enhance the risk perception ability of autonomous vehicles in accident scenarios. However, due to the relatively small amount of training data, there are inevitably some incorrect risk prediction results. Using richer scenario data can effectively improve this problem.

[0196] The results show that the accident scenario database constructed in this invention can effectively improve the adaptability of autonomous vehicles to similar accident scenarios, significantly reduce the accident rate when encountering similar driving scenarios, realize efficient closed-loop self-evolution of autonomous driving algorithms, and ensure the safe and reliable operation of autonomous vehicles in the real world.

[0197] In summary, the accident scene generation method based on scene knowledge graphs and considering accident causes provided by this invention can construct an accident scene database with high realism, high diversity, and consistency of accident causes under the condition of limited accident scene sample data. It efficiently expands the number of similar accident scenes in the closed-loop self-evolutionary cloud database of autonomous driving algorithms. By using the constructed scene database to train autonomous driving algorithms, it can effectively enhance the adaptability of autonomous vehicles to similar accident scenes that have already occurred, and ensure the safe and reliable operation of autonomous vehicles in the real world.

[0198] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for generating accident scenarios based on scene knowledge graphs and considering accident causes, characterized in that, Includes the following steps: Step 1: Model the scene knowledge graph; A scene knowledge graph representation method is constructed, which includes scene graph SG and accident cause graph ACG. Scene graph SG is used to aggregate the interaction features of scene in spatial and temporal dimensions, and accident cause graph ACG is used to realize causal control in the accident scene generation process. Step 2: Model the scene graph time series prediction model; A scene graph SG is processed using a graph attention neural network to extract its spatial dimension information. A Transformer-based temporal feature extraction architecture is used to learn the long temporal dependencies of the temporal scene graph SG to extract its temporal dimension information. A variational graph autoencoder is used to construct a latent space for repeated sampling of latent variables and form a sampling layer. The latent variables are decoded by a decoder to obtain the scene graph sampling space, and a scene graph temporal prediction model is constructed. Step 3: Model the time-series causal inference model; Based on the temporal trajectory and related description of the input accident scene, a detailed analysis of the inherent causal relationship of the accident scene is conducted to obtain empirical knowledge. Based on this, an accident causal graph (ACG) and inference rules corresponding to each causal edge of the ACG graph are formulated, thereby obtaining a temporal causal inference model. Step 4: Model the scene graph time-series decision generation model to generate accident scenarios; By integrating key modules such as scene graph temporal prediction model and temporal causal inference model using reinforcement learning methods, a complete scene graph temporal decision generation model is constructed. Specifically, the scene graph temporal prediction model is used to generate the sampling space of the scene graph at the next time step in a guided manner, and the reinforcement learning agent is trained to sample in the sampling space. The temporal causal inference model is used to perform temporal causal inference on the generated temporal scene graph. A graph similarity measurement method is established between the inferred accident causation graph ACG0 and the real accident causation graph ACG, which serves as the main feedback reward when the reinforcement learning agent takes different actions, guiding the causal controllable generation of accident scenes.

2. The method for generating accident scenarios based on scene knowledge graphs and considering accident causes according to claim 1, characterized in that, The specific method for step one is as follows: 11) Model the scene graph SG; The scene information at each moment is modeled as a scene graph SG of one frame. t Each frame of the scene graph SG t It consists of several nodes and edges, where nodes represent different traffic participants and edges represent the interaction relationships between traffic participants; the scene graph SG corresponds to the traffic scenario at a certain time t. t Including node feature matrix N t Edge index matrix I t Edge feature matrix E t That is, SG t =(N t ,I t E t ); Node feature matrix Where m represents the number of nodes in the scene, i.e., the number of traffic participants, and n represents the node feature dimension. We select n=4 node feature dimensions, which are the position and velocity of the traffic participants in the x and y directions, i.e., (x, y, v). x ,v y ); Edge index matrix h represents the edge attribute dimension, the edge index matrix represents the interaction relationship between nodes, and a unidirectional edge N is defined. i →N j Two-way edge Self-circulating edge N i →N i Three edge types; Edge feature matrix The edge carries a total of 5 types of feature information, namely △x, △y, △v. x ,△v y , , where △ x ,△ y For interactive nodes in x , y Relative position of direction; △ v x ,△ v y For interactive nodes in x , y Relative velocity in the direction; The interaction weighting coefficients are for vehicles at different distances in front of and behind the vehicle. The Logistic function is used to model the interaction weights of vehicles at different distances in front of and behind the vehicle. The Logistic function is shown in equation (1); (1) In the formula, L is the maximum value of the interaction relationship weight coefficient; k is the rate at which the interaction relationship weight coefficient decays with distance; d0 is the distance when Logistic(d) = L / 2; 12) Model the accident causation diagram (ACG); The accident causation graph (ACG) consists of nodes and edges. In addition to the nodes representing traffic participants, a collision node C is introduced into the accident causation graph (ACG). The edges represent the causal relationships between nodes. Neither the nodes nor the edges carry any feature information. The essential causes of the accident scene are analyzed based on accident text data and video data. That is, the driving behavior of the vehicle that caused the accident is analyzed, and then the causal link of the accident scene is determined, thereby determining the specific content in the ACG graph.

3. The method for generating accident scenarios based on scene knowledge graphs and considering accident causes according to claim 1, characterized in that, The specific method for step two is as follows: 21) Model the scene graph time series prediction model; The scene graph temporal prediction model consists of three parts: a scene graph spatiotemporal encoder, a sampling layer, and a scene graph decoder. The scene graph spatiotemporal encoder is used to extract the spatiotemporal features of the input scene graph sequence. The sampling layer is used to sample in the latent space to obtain latent variables. The scene graph decoder decodes the latent variables to predict the scene graph sampling space at the next time step. 22) Design the loss function; The loss function includes the node feature reconstruction loss L. N Edge index reconstruction loss L I Edge feature reconstruction loss L E and KL divergence loss L KL Four parts.

4. The method for generating accident scenarios based on scene knowledge graphs and considering accident causes according to claim 3, characterized in that, The specific method for step 21) is as follows: The scene graph spatiotemporal encoder consists of a graph attention neural network (GAT), a Transformer encoder, and fully connected layers. First, a multi-layer GAT is used to encode the node features at each time step. Then, a self-attention mechanism is used to aggregate node neighborhood information to extract spatial dimension features. For node features in the scene graph... Edge index features Edge features The output of the GAT network is shown in equation (2): (2) In the formula, N(i) is the set of neighboring nodes of node i; The features of node i at time step t; , The edge index features and edge features connecting node i and node j at time step t are respectively used to obtain the aggregated node feature matrix N. The Transformer encoder is used to extract temporal features from the scene graph. The node feature sequence at each time step is assigned a position code (PE). The Transformer encoder outputs N. t As shown in equation (3): (3) In the formula, N represents the aggregate node characteristics output by the GAT network; PE is the position code in sine / cosine mode; N t Features of time-series nodes; Design two fully connected layers to map the temporal coding information to the mean of the latent space according to equations (4) and (5). and variance : (4) (5) In the formula, , This is the weight matrix of the fully connected layer. , This serves as a bias term when calculating latent variables; Sampling layers are used to sample from the latent spatial distribution to obtain latent variables. As shown in equation (6); (6) In the formula, Standard deviation; The noise is a standard normally distributed noise, where I represents the identity matrix; d z Let z be the spatial dimension of the latent variable. In the scene graph decoder, the latent variable z is expanded into a latent representation zt at each time step t. t Position encoding is used to process latent variables to recover temporal information, z t The data is fed into the Transformer decoder for processing, as shown in Equation (7), and the output is the node features at each time step. Together, they form the overall node feature N. d : (7) The features are mapped to the original node feature space through a fully connected layer, thus obtaining the predicted scene graph SG. p =( , , As shown in equation (8): (8) In the formula, W d This is the weight matrix of the fully connected layer of the decoder.

5. The method for generating accident scenarios based on scene knowledge graphs and considering accident causes according to claim 3, characterized in that, The specific method for step 22) is as follows: Node feature reconstruction loss L N The mean squared error is used to measure the difference between the predicted node features and the actual node features, as shown in Equation (9): (9) In the formula, Let k be the value of the kth feature in the i-th node at time t+1. Let k be the value of the k-th feature in the i-th node predicted at time t+1. m is the number of nodes in the scene, and n is the node feature dimension; Edge index reconstruction loss L I The difference between the edge index features predicted by the model and the true index features is measured using cross-entropy loss, as shown in Equation (10): (10) In the formula, This is the actual edge index matrix, indicating whether there is an edge between node i and node j; if there is, it is 1, otherwise it is 0. The value is between 0 and 1, representing the probability that an edge exists between node i and node j. Edge feature reconstruction loss L E The mean squared error (MSE) is used to measure the difference between the edge features predicted by the model and the actual edge features, as shown in Equation (11): (11) In the formula, e is the total number of edges in the scene graph at time t+1; h is the number of features carried by a single edge; Let be the value of the k-th feature in the edge formed by nodes i and j at time t+1; Let be the value of the k-th feature in the edge formed by nodes i and j predicted at time t+1; KL divergence loss L KL Measuring distribution With prior distribution Differences between them, distribution yes The variational approximation is shown in equation (12): (12) The total loss L of the scene graph temporal prediction model is shown in equation (13): (13) In the formula, , , Representing L respectively I L E L KL The weight hyperparameter of the loss.

6. The method for generating accident scenarios based on scene knowledge graphs and considering accident causes according to claim 1, characterized in that, The specific method for step three is as follows: 31) Based on the time-series trajectory information and text description of the accident scene, analyze what kind of driving behavior of the vehicle led to the accident, and then determine the causal link of the accident scene. In this way, the essential cause of the accident scene is obtained. Based on this, the ACG graph is modeled. The accident causation graph ACG includes vehicle nodes and edges representing the causal relationship between vehicle nodes. 32) Based on the accident causation graph ACG, establish causal inference rules corresponding to each edge, take a scene graph sequence of arbitrary length as input, and infer the causal relationship graph corresponding to the scene graph sequence.

7. The method for generating accident scenarios based on scene knowledge graphs and considering accident causes according to claim 1, characterized in that, The specific method for step four is as follows: 41) Define a graph similarity quantification method between the inferred accident causation graph ACG0 and the actual accident causation graph ACG, which serves as the main reward for the scene graph temporal decision generation model to control the accident scene generation type. The Jaccard similarity S is calculated by equation (14). J As a quantitative evaluation method between ACG0 diagrams and ACG diagrams: (14) In the formula, S J S represents the Jaccard graph similarity between ACG0 and ACG. J The larger the value, the more similar the ACG0 diagram is to the real ACG diagram, that is, the more realistic the accident causes are in the sequential scene. E0 and E are the set of oriented edges in ACG0 and ACG graphs, respectively; TP=E0∩E are correctly identified causal edges with the correct orientation; FP=E0\E are incorrectly identified causal edges that appear in E0 but not in E; FN=E\E0 are missing causal edges in the inference results that appear in E but not in E0. 42) Assuming the temporal scene represented by the scene graph SG follows a Markov decision process, the reinforcement learning algorithm includes the state space S, the action space A, and the state transition function. Let s represent the state, a represent the action, and the reward function be R(s,a), where state s t ={SG t~t+n-1 This includes a fixed-length time sequence scene graph, and action a. t For a scenario graph SG containing multiple time points t+n t+n Sampling is performed in a discrete sampling space, and state transition is executed by the agent using action a. t Then a new state s is obtained. t+1 , that is s t+1 ={SG t +1~t+n }=f(s t ,a t ), where f is the scene graph update function considering trajectory smoothness and vehicle dynamics constraints, and the reward function R(s) is... t ,a t ) represents the state s in the current state. t Take action a t The reward obtained at that time consists of three parts, as shown in equation (15): (15) In the formula, R similarity The similarity reward between ACG0 and ACG obtained from the temporal scene graph inference is shown in Equation (16): (16) In the formula, a is the similarity reward coefficient; R smootness The real-time negative reward for trajectory smoothness is used to measure the smoothness of the trajectory, as shown in Equation (17); (17) In the formula, t is the original trajectory; t* is the trajectory fitted by the SG algorithm; b is the smoothness reward coefficient; R constraint The real-time negative reward for violating constraints is that actions that violate physical constraints or vehicle dynamics constraints are penalized, as shown in equation (18): (18) In the formula, ViolationPenalty represents that the sampling point violates the constraint, and its value is set to 1. c is the penalty coefficient weight. 43) Integrate traffic rule constraints and vehicle dynamics constraints into the reinforcement learning algorithm as the underlying hard constraints, and the upper layer is a reinforcement learning sampling strategy for inverse optimization. Traffic rule constraints mean that vehicles must not leave the road boundary and vehicles must not move in the opposite direction of the lane. Vehicle dynamics constraints specify the range of horizontal and vertical positions and speeds that the sampling points can reach. If a sampling point meets one of the following conditions, it is considered to violate the constraint condition ViolationPenalty in equation (18), and the agent is given a constraint penalty reward: ① The sampling point violates the temporal motion trend of the trajectory, that is, it samples in the opposite direction of the temporal trajectory; ② The vehicle drives out of the road boundary; ③ The difference between the lateral position of the sampling point and the previous temporal point exceeds the vehicle dynamics limit; ④ The difference between the longitudinal position of the sampling point and the previous temporal point exceeds the vehicle dynamics limit. 44) Select the PPO reinforcement learning algorithm to sample in the scene graph sampling space generated in step two, and set the loss of the PPO algorithm into three parts: policy loss, value loss, and entropy loss, which are used to optimize the policy, train the value network, and promote policy exploration, respectively, as shown in equation (19): (19) In the formula, c1 and c2 are the weight values ​​for different losses, respectively; , , These are strategy loss, value loss, and entropy loss, respectively.

Citation Information

Patent Citations

  • Intelligent network connection automobile test scene generation method based on knowledge graph

    CN117667699A

  • Causal reinforcement learning system for vehicles in safety-critical scene with causal confusion

    CN120012838A