A method for training a model for test scenario generation
By acquiring high-dimensional feature information and generating latent variable states using random functions, and using traffic prior models to train test scenarios for autonomous vehicles, the problem of generating safety-critical test scenarios in existing technologies is solved, thereby improving testing efficiency and safety.
Patent Information
- Application Number
- CN202310410081.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-17
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2043-04-17
AI Technical Summary
Existing methods for generating test scenarios for autonomous vehicles are insufficient to efficiently and effectively generate test scenarios for potential safety issues of vehicles under normal operating conditions, especially since safety hazards caused by the capability boundaries of each module are difficult to fully test.
By acquiring historical observation sequences and map information, high-dimensional feature information is obtained through dimensionality enhancement. Latent variable state information is generated by combining random functions, decoded using a traffic prior model, and the loss function is calculated to train the traffic prior model, thereby generating test scenarios for autonomous vehicles.
It enables the efficient and automatic generation of driving scenarios that demonstrate the safe operation of autonomous vehicles under normal working conditions, improving the efficiency of test scenario generation and testing, and providing support for the improvement of autonomous driving systems.
Smart Images

Figure CN116467599B_ABST
Abstract
Description
Technical Field
[0001] This article relates to, but is not limited to, autonomous driving technology, and in particular to a training method for a model that generates test scenarios. Background Technology
[0002] Currently, autonomous vehicles face serious safety issues. Autonomous vehicles developed by relevant companies have all encountered serious traffic accidents. These safety issues fundamentally hinder the large-scale application and commercialization of autonomous vehicles. Therefore, it is urgent to conduct safety tests on autonomous vehicles.
[0003] The basic process for testing and evaluating autonomous vehicles is as follows: Collect a series of real-world scenarios (referring to a comprehensive dynamic description of the interaction process between an autonomous vehicle and other vehicles, roads, traffic facilities, weather conditions, etc., within a certain time and space range; it is an organic combination of the autonomous vehicle's driving scenario and driving environment, including various entity elements, as well as the actions performed by entities and the connections between entities; for example, highway driving scenarios, following scenarios, lane-changing scenarios, and turning scenarios, etc.); generate a large number of virtual simulation test scenarios based on real-world scenarios using ontology or deep learning methods; test the autonomous vehicle according to the generated test scenarios and collect the test results; evaluate the safety of the autonomous vehicle based on the collected test results to obtain estimates of test indicators such as accident rate.
[0004] Currently, due to the low probability of real-world occurrences of critical test scenarios related to vehicle safety (such as vehicle collisions), testing using real-world road tests or virtual simulations based on reconstructions is generally inefficient. This makes it difficult to efficiently assess the average performance level of autonomous vehicles, and even more difficult to effectively test for potential safety issues, such as driving scenarios where an autonomous vehicle cannot drive safely under normal operating conditions. Furthermore, existing test scenario generation methods typically focus on whole-vehicle testing, neglecting potential safety issues arising from the limitations of individual modules' capabilities (such as a perception algorithm failing to recognize a specific object) even when each module is functioning normally. Therefore, how to efficiently and automatically generate such critical test scenarios has become an urgent problem to be solved. Summary of the Invention
[0005] The following is an overview of the subject matter described in detail herein. This overview is not intended to limit the scope of the claims.
[0006] This invention provides a training method for a model that generates test scenarios, which can efficiently and automatically generate driving scenarios for autonomous vehicles to drive safely under normal working conditions.
[0007] This invention provides a method for training a model to generate test scenarios, comprising:
[0008] Obtain the dataset for model training, which contains historical observation sequences and map information;
[0009] The historical observation sequences in the dataset are upgraded to obtain the first high-dimensional feature information;
[0010] Based on the correlation between discrete points contained in the map information in the dataset, the second high-dimensional feature information of the map information is obtained;
[0011] Based on the first high-dimensional feature information, the second high-dimensional feature information, and the pre-defined random function information representing the randomness of all vehicles in the scene, latent variable state information is generated.
[0012] The generated latent variable state information is decoded to obtain the output of the traffic prior model;
[0013] For all data in the dataset, calculate the pre-set first loss function based on the input of the traffic prior model and the output of the obtained traffic prior model to obtain the traffic prior model;
[0014] The first loss function is determined based on the distance between the predicted state information and the actual state information from the traffic prior model; the traffic prior model is used to generate test scenarios for autonomous vehicles.
[0015] On the other hand, embodiments of the present invention also provide a computer storage medium storing a computer program, which, when executed by a processor, implements the above-described training method for generating a model that realizes the test scenario.
[0016] Furthermore, embodiments of the present invention also provide a terminal, comprising: a memory and a processor, wherein the memory stores a computer program; wherein,
[0017] The processor is configured to execute computer programs in memory;
[0018] When the computer program is executed by the processor, it implements the training method for the model generated in the test scenario as described above.
[0019] The technical solution of this application includes: acquiring a dataset for model training, the dataset containing historical observation sequences and map information; increasing the dimensionality of the historical observation sequences in the dataset to obtain first high-dimensional feature information; obtaining second high-dimensional feature information of the map information based on the correlation between discrete points contained in the map information in the dataset; generating latent variable state information based on the first high-dimensional feature information, the second high-dimensional feature information, and pre-set random function information representing the randomness of all vehicles in the scene; decoding the generated latent variable state information to obtain the output of the traffic prior model; calculating a pre-set first loss function for all data in the dataset based on the input and output of the traffic prior model to obtain the traffic prior model; wherein, the first loss function is determined based on the distance between the predicted state information and the actual state information of the traffic prior model; the traffic prior model is used to generate test scenarios for autonomous vehicles. The embodiments of this invention train and obtain a traffic prior model, providing support for automatically generating driving scenarios for autonomous vehicles to drive safely under normal working conditions.
[0020] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the description, claims, and drawings. Attached Figure Description
[0021] The accompanying drawings are provided to further understand the technical solutions of the present invention and constitute a part of the specification. They are used together with the embodiments of this application to explain the technical solutions of the present invention and do not constitute a limitation on the technical solutions of the present invention.
[0022] Figure 1 This is a flowchart illustrating the training method for the model used to generate test scenarios according to an embodiment of the present invention.
[0023] Figure 2 This is a schematic diagram illustrating the application example of the vehicle three-circle approximation method of the present invention. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.
[0025] The steps illustrated in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases the steps shown or described may be performed in a different order than that presented here.
[0026] Figure 1 This is a flowchart illustrating the training method for the model used to generate test scenarios according to an embodiment of the present invention, as shown below. Figure 1 As shown, it includes:
[0027] Step 101: Obtain the dataset for model training. The dataset contains historical observation sequences and map information.
[0028] Step 102: Upscale the historical observation sequences in the dataset to obtain the first high-dimensional feature information;
[0029] Step 103: Based on the correlation between discrete points contained in the map information in the dataset, obtain the second high-dimensional feature information of the map information;
[0030] It should be noted that the high dimension of the first high-dimensional feature information and the second high-dimensional feature information in the embodiments of the present invention is the dimension determined by those skilled in the art based on the data structure of historical observation sequences and map information. The high dimension refers to the vector after linear transformation, whose dimension is higher than that of the vector before transformation.
[0031] Step 104: Generate latent variable state information based on the first high-dimensional feature information, the second high-dimensional feature information, and the pre-set random function information representing the randomness of all vehicles in the scene;
[0032] Step 105: Decode the generated latent variable state information to obtain the output of the traffic prior model;
[0033] Step 106: For all data in the dataset, calculate the pre-set first loss function based on the input of the traffic prior model and the output of the obtained traffic prior model to obtain the traffic prior model;
[0034] The first loss function is determined based on the distance between the predicted state information and the actual state information of the traffic prior model.
[0035] This invention addresses autonomous vehicles, considering potential safety issues arising from the decision-making layer under normal operating conditions. A deep learning model is trained to generate key test scenarios related to the decision-making layer, providing technical support for automated test scenario generation and laying the foundation for improving the efficiency of test scenario generation and testing. Furthermore, by understanding the potential safety hazards caused by the decision-making layer in autonomous vehicles, information support is provided for improving autonomous driving systems. In one exemplary instance, the dataset in this invention is:
[0036] Among them, (X) i ,Y i( ) represents a test scenario pair with a pre-set fixed duration T and a sampling interval of Δt. For map information; X i Represents a scenario used for training a traffic prior model. Historical observation sequence of time; Y i The actual sequence at scene TT′; s it ∈R N×d ,t∈{1,2,…,T}, represents the state of scene i at each time step as a vector of dimension N×d, where N is the total number of vehicles and d is the dimension of the state information; s it =(x it1 ,x it2 ,…,x itN ), representing the set of all vehicle states at time t in scenario i, x itj = (x, y, ...) ∈ R d This represents the state information of the j-th vehicle at time t in scenario i. This represents the status information of the decision-maker module of the vehicle under test.
[0037] In one exemplary instance, step 104 of this embodiment of the invention generates latent variable state information, including:
[0038] The first and second high-dimensional feature information are processed through a pre-defined self-attention mechanism module to obtain updated third high-dimensional feature information that represents the interaction relationship between vehicles.
[0039] The third high-dimensional feature information is processed through a pre-defined cross-attention mechanism module to obtain updated fourth high-dimensional feature information representing the interaction between vehicles and roads.
[0040] Latent variable state information is generated based on the third high-dimensional feature information, the fourth high-dimensional feature information, and the random function information.
[0041] In one exemplary instance, this embodiment of the invention generates latent variable state information based on third high-dimensional feature information, fourth high-dimensional feature information, and random function information, including:
[0042] Perform matrix multiplication on the updated third and fourth high-dimensional feature information;
[0043] The latent variable state information is determined based on the result of matrix multiplication and the information of the random function.
[0044] In one exemplary instance, the random function information in this embodiment of the invention is loaded through a preset fitter and regulator;
[0045] The fitter is used to predict the output state information of the decision-maker module based on the input state information of the decision-maker module of the test vehicle in the dataset; the regulator is used to adjust the behavior of background vehicles other than the test vehicle; here, background vehicles include: all vehicles in the scene other than the test vehicle.
[0046] It should be noted that, in this embodiment of the invention, when the tested vehicle changes, as long as the tested vehicle and the background vehicle in the regulator and fitter are adjusted, the above-described training of this embodiment of the invention can be performed to obtain a traffic prior model applicable to the adjusted vehicle. The adjustment of the background vehicle's behavior in this embodiment of the invention includes: adjusting the position information of the background vehicle at each moment within the scene.
[0047] In one exemplary instance, embodiments of the present invention decode latent variable state information, including:
[0048] The hidden variable state information is decoded by a pre-defined gated loop unit (GRU).
[0049] In one exemplary instance, the expression for the first loss function in this embodiment of the invention is:
[0050]
[0051]
[0052] Wherein, the subscript i1 is used to identify the decision-maker module of the test vehicle in scenario i, Y i1 This represents the actual state information of the decision-maker module in scenario i; This represents the predicted state information of the decision-maker module in scenario i, output by the traffic prior model; the subscript ij is used to identify vehicle j in scenario i, Y ij This represents the actual state information of vehicle j in scenario i; This represents the predicted state information of vehicle j in scenario i, output by the traffic prior model; N is the total number of vehicles in scenario i.
[0053] In one exemplary instance, the training method of this embodiment of the invention further includes adding the following first constraint to the traffic prior model:
[0054]
[0055] in, r ij Let r be the approximate radius of the circle of vehicle j in scene i. ik Let d be the approximate radius of the circle of vehicle k in scene i; itjk d represents the interval between vehicle j and vehicle k. itjk=min u,v dist(l itju ,l itkv ), l itju l represents the position of the center u of vehicle j in scene i at time t. itkv The center v of the circle representing vehicle k in scenario i is located at time t.
[0056] In one exemplary instance, the first constraint of this embodiment of the invention is used to ensure that no collision occurs between the predicted trajectories of the vehicles.
[0057] In one exemplary instance, the training method of this embodiment of the invention further includes:
[0058] The parameters of the multilayer perceptron (MLP) fitting are updated using the first gradient descent method, and the fitting is trained using a pre-defined second loss function.
[0059] After training the fitter using a pre-defined second loss function, the MLP parameters of the regulator are updated using a second gradient descent method, and the regulator is trained using a pre-defined third loss function.
[0060] In one exemplary instance, the second loss function in this embodiment of the invention is:
[0061]
[0062] Wherein, the subscript i1 is used to identify the decision-maker module of the test vehicle in scenario i, Y i1 This represents the actual state information of the decision-maker module in scenario i; This represents the predicted state information of the decision-maker module in scenario i, as output by the traffic prior model.
[0063] It should be noted that, in the embodiments of the present invention The parameters of the traffic prior model after training with the fitter are based on X. i The estimated future trajectory information of the decision-making module.
[0064] In one exemplary instance, the third loss function in this embodiment of the invention is:
[0065]
[0066] Here, the subscript i1 is used to identify the decision-maker module of the vehicle under test in scenario i. This represents the state information predicted by the decision-maker module in scenario i, as output by the traffic prior model. This represents the predicted state information of background vehicles in scenario i output by the traffic prior model, i.e., the predicted position value of background vehicles in scenario i output by the traffic prior model.
[0067] In one exemplary instance, the expression for the third loss function in this embodiment of the invention is:
[0068]
[0069] in, The trajectory information of the background vehicle. , r ij Let r be the approximate radius of the circle of vehicle j in scene i. ik Let d be the approximate radius of the circle of vehicle k in scene i; itjk d represents the interval between vehicle j and vehicle k. itjk =min u,v dist(l itju ,l itkv ), l itju l represents the position of the center u of vehicle j in scene i at time t. itkv The position of the center v of vehicle k in scenario i at time t; w tj For pre-defined weight items, This represents the distance between the background vehicle j and the vehicle being tested.
[0070] In one exemplary instance, the training method of this embodiment of the invention further includes: when training the fitter using a pre-defined second loss function, adding the following KL divergence (generally referring to relative entropy; relative entropy, also known as Kullback-Leibler divergence or information divergence, is an asymmetric measure of the difference between two probability distributions) loss function as a second constraint term to the traffic prior model:
[0071]
[0072] Where: w1 is a random variable representing the randomness of the decision-maker module of the tested vehicle before the MLP parameters of the fitter are updated, and w′1 is a random variable representing the randomness of the decision-maker module of the tested vehicle after the MLP parameters of the fitter are updated. w1 and w′1 respectively follow... d w Let σ be the dimension of random variables w1 and w′1; 1k σ′ represents the standard deviation of the future location distribution of background vehicles output by the traffic prior model. 1k μ represents the standard deviation of the background vehicle future position distribution in the regulator output. 1kμ′ represents the mean of the future location distribution of background vehicles output by the traffic prior model. 1k This represents the mean of the future location distribution of background vehicles output by the regulator.
[0073] In one exemplary instance, after generating the traffic prior model, the method of this embodiment of the invention further includes:
[0074] The generated traffic prior model is used to generate a test scenario for the vehicle under test. In other words, this embodiment of the invention provides a method for generating a test scenario for the vehicle under test based on the generated traffic prior model.
[0075] This invention also provides a computer storage medium storing a computer program, which, when executed by a processor, implements the training method for the model generated in the test scenario described above.
[0076] This invention also provides a terminal, comprising: a memory and a processor, wherein the memory stores a computer program; wherein,
[0077] The processor is configured to execute computer programs in memory;
[0078] When a computer program is executed by a processor, it implements the training method for the model generated in the test scenario as described above.
[0079] The following application examples briefly illustrate the embodiments of the present invention. These application examples are only used to describe the embodiments of the present invention and are not intended to limit the scope of protection of the present invention.
[0080] Application Examples
[0081] This application example uses a fixed time period denoted as T, with a sampling interval of Δt, in map information. The test scenarios for autonomous driving obtained from [the source] are In this context, each factor in S represents the state information of all vehicles at a sampling time step, s t ={x t1 ,x t2 ,…,x tN}, where N is the total number of vehicles in the scene, x tj =(x,y,v) x ,v y Let , ..., represent the state information of the j-th vehicle, including its horizontal and vertical coordinates and lateral and vertical velocities from a bird's-eye view. For ease of explanation, this application example denotes the first vehicle in any data point at any time in any scenario as the test vehicle (the autonomous driving vehicle under test), and denotes the decision-maker module of this vehicle as p(·). The decision-maker module p(·) uses the map information of the scene. Based on the observed state information of all vehicles over the past δ time steps (i.e., the historical observation sequence over the past δ time steps), the system plans its own state information for the next δ′ time steps; the decision-maker module runs the decision-maker algorithm from the relevant technology.
[0082] In one exemplary instance, this application example obtains test scenarios for autonomous driving by referring to relevant technologies, including: continuously driving on the road with a (autonomous driving) vehicle equipped with multiple types of sensors, collecting and storing surrounding environmental data, and obtaining test scenarios based on the collected surrounding environmental data; or obtaining vehicle data on a fixed road over a period of time through roadside sensor devices, and obtaining test scenarios based on the acquired vehicle data.
[0083] This application example generates safety-critical test scenarios based on the premise of generating reasonable traffic scenarios, i.e., scenarios that might occur in the real world. Therefore, it is necessary to first construct a traffic prior model. The problem of constructing the traffic prior model is modeled as a group trajectory prediction problem, i.e., predicting a group's trajectory based on a historical time period. The historical observation sequence predicts the future sequence TT′. This allows for end-to-end deep learning to learn traffic scenario operation rules from real data and implicitly store the corresponding traffic scenario operation relationships using neural networks. Simultaneously, the historical observation sequence of historical duration T′ is used as the scenario initialization condition to generate complex and diverse test scenarios. To ensure the decision-maker algorithm can efficiently generate safety-critical test scenarios, the constructed traffic prior model requires two additional functional modules: a fitter to approximate the decision-maker algorithm and a regulator to fine-tune the background vehicle behavior. In one exemplary instance, the decision-maker algorithm in this application example includes an autonomous driving path planning algorithm, which takes surrounding vehicle information and map information as input and outputs its own path information over a period of time. The decision-maker algorithm can also be other existing algorithms in related technologies that can perform the above processing; this application example will not elaborate further.
[0084] In one exemplary instance, this application example records a test scenario. Center front subset of sampling time steps Initialize data for a fixed scene, denoted as X. i X i ={s i1 ,s i2 ,…,s iT′};s it ∈R N×d ,t∈{1,2,…,T};S i -X i Let Y be the sequence of future real observations used by the traffic prior model for learning. iSimultaneously, the vehicle used for data collection was designated as the test vehicle, and its first position in the data at each moment was fixed, i.e.: s t ={x t1 ,x t2 ,…,x tN}, x t1 For the vehicle under test, Construct a dataset based on this. X i Indicates the input, Y i This indicates the output.
[0085] In one exemplary instance, this application example constructs a dataset with reference to relevant technologies, including: dividing the status data of all vehicles collected in the above process into data segments of a pre-set fixed duration (e.g., 20s, sampling time step), and constructing a dataset based on the divided data segments.
[0086] In one exemplary instance, this application example is based on input X i and output Y i The relationship between them will be defined from a mathematical perspective as a traffic prior model, and then processed through a deep neural network model. Indicates; among which, The function represented by the network, θ, is its set of parameters; using the dataset Based on input X i Predicted output Y i To address this problem, a corresponding first loss function is constructed, and the network parameters are updated using gradient descent, a technique in related technologies. When the training reaches a pre-set iteration limit, an approximate optimal solution θ is obtained. * In one exemplary instance, this application example sets an upper limit on the number of iterations based on experience. Assuming the upper limit is M, during the iteration process, if the ratio of the difference between the loss functions of two adjacent steps to the total loss function is less than a threshold within M iterations, the iteration stops (i.e.: Conversely, if the ratio of the difference between the loss functions of two adjacent steps to the loss function is greater than or equal to the threshold when the upper limit of the number of iterations is reached, it indicates that convergence has not yet occurred. In this application example, M is increased in accordance with relevant technologies, and iterative calculations are continued.
[0087] In one exemplary instance, the method for constructing a deep neural network module in this application example includes: determining the input X of the deep neural network. i and output Y i Based on input X i and output Y i The relationship between the input X and the network model structure is used to determine the network model structure; in an exemplary instance, this application demonstrates the structure of a deep neural network; wherein, based on the input X... iThe network incorporates GRUs from related technologies, taking into account the timing characteristics of the input X. i The network contains self-attention modules, a feature of related technologies, reflecting the characteristics of internal interactions.
[0088] In this embodiment of the invention, the test vehicle is equipped with a decision-maker module. The decision-maker module loads a decision-maker algorithm, and the data collected by the traffic prior model includes the state information of the vehicle equipped with the decision-maker module. Each vehicle (traffic subject) involved in the dataset is a node in the traffic prior model and has corresponding parameters used to map the original location information to high-dimensional feature information.
[0089] Secondly, for the decision-maker module, the aforementioned deep neural network model is used. Further optimization is performed on the network parameters of the nodes corresponding to the decision-maker module, so that the corresponding nodes can more accurately provide the possible outputs of the decision-maker module given the input. The fitter f(·) has its parameter set. Using datasets Consider the optimization problem of data fitting for the decision maker algorithm, and construct a second loss function accordingly, fixing the parameter set θ. * θ1 is updated using the gradient descent method in related techniques. When the upper limit of the number of iterations is reached, an approximate optimal solution is obtained. Corresponding Fitter And update Finally, consider the regulator g(·) that fine-tunes the background vehicle behavior, with its parameter set... Consider the fitter for the decision-maker algorithm The adversarial scenario is used as the basis for generating security-critical test scenarios, and a third loss function is constructed based on this, with a fixed parameter set θ. * θ2 is updated using the gradient descent method. When the upper limit of iteration is reached, an approximate optimal solution is obtained. Corresponding background vehicle behavior regulator And update The parameter is θ * The deep neural network model is the model ultimately used to generate the test scenario.
[0090] After completing the optimization of the above three stages in sequence, for any initialization scenario X... i ={s i1 ,s i2 ,…,s iT′}, s it ∈R N×d ,t∈{1,2,…,T}, through a deep neural network model Then it can be based on X iGenerate safety-critical test scenarios for the decision-maker module to disrupt its performance as much as possible and test its performance weaknesses; when changing a specific decision-maker module, only the parameters corresponding to the fitter and regulator in the above processing need to be adjusted.
[0091] This application example focuses on autonomous vehicles. Considering the potential safety issues caused by the decision-making layer under normal vehicle conditions, it proposes a test scenario generation method based on a deep learning model and a multi-stage optimization method. This method automates the generation of test scenarios, improving the efficiency of both generation and testing. Furthermore, by understanding the potential safety hazards caused by the decision-making layer in autonomous vehicles, it provides information support for the improvement of autonomous driving systems.
[0092] The following example dataset illustrates the above processing:
[0093] This application example demonstrates how to obtain a dataset. in,
[0094] X i ={s i1 ,s i2 ,…,s iT′}∈R T′×N×d ;
[0095] Y i =(s i,T ′ +1 ,s i,T ′ +2 ,…,s iT );
[0096] (X i ,Y i ( ) represents a test scenario pair with a pre-set fixed duration T and a sampling interval of Δt. For the corresponding high-precision map information, i is the data index, representing the data of the i-th scene; X i Represents a scenario used for training a traffic prior model. Historical observation sequence of time; Y i The actual sequence at scene TT′; s it ∈R N×d ,t∈{1,2,…,T}, represents the state at each time step as a vector of dimension N×d; s it =(x it1 ,x it2 ,…,x itN Let x represent the set of all vehicle states in scenario i at time t, where N is the total number of vehicles in the scenario, and x is the total number of vehicles in the scenario. itj = (x, y, ...) ∈ R dThis represents the state information of the j-th vehicle at time t in scenario i, where d is the data dimension. Also: Represents the status information of the decision-maker module;
[0097] This application example assumes scenario n, then the input to the traffic prior model is: X n Y is the initial data for the scene. n This is the output of the traffic prior model;
[0098] Historical observation sequence X i ={s i1 ,s i2 ,…,s iT′} Perform dimensionality increase to obtain the first high-dimensional feature information H0∈R N×D In one exemplary example, the historical observation sequence is upscaled using a pre-defined gated recurrent unit (GRU); where N represents the number of vehicles in the scene and N and D represent the dimensions of the scene's feature vectors; based on the correlation between discrete points in the map information, the second high-dimensional feature information of the map information is obtained. Assuming map information There are N M There are discrete points, each represented by a two-dimensional bird's-eye view coordinate system. Among them, discrete points have four types of relationships: predecessor, successor, left neighbor, and right neighbor, based on which a graph G can be constructed. u =(V,E) u ), u = 1, 2, 3, 4; where V is used to identify N M The index of a discrete point Used to represent the correlation between discrete points, where u represents the u-th type of correlation, if E u element e in uij =1, which means that i has the u-th type of correlation with respect to j. Using a graph convolutional neural network to analyze E... u The data is processed to obtain the second high-dimensional feature information of the map information. Where, N M D M These represent the number of discrete points in the map and the dimension of the feature vector, respectively.
[0099] The obtained first high-dimensional feature information H0 and second high-dimensional feature information M i The system processes the data through a pre-defined self-attention module to obtain a third high-dimensional feature that represents the interaction between vehicles.
[0100]
[0101] Where q = H i W Q k = H i W K v = H i W V For H i The linear transformations of H, respectively, represent H i The first query vector, the first key-value vector, and the first value vector are network structure parameters known to those skilled in the art; ← represents update;
[0102] The obtained third high-dimensional feature information is processed through a pre-defined cross-attention mechanism module to obtain a fourth high-dimensional feature information representing the interaction between vehicles and roads.
[0103]
[0104] Where q′=H′ i W Q , k′=H′ i W k v′=H′ i W v H′ i The linear transformations of H′ and H′ respectively represent the linear transformations of H′ and H′ respectively. i The second query vector, the second key-value vector, and the second value vector are network structure parameters known to those skilled in the art; ← represents update.
[0105] Get updated H i and H′ i Then, by multiplying the matrices using Formula 1 and Formula 2, we obtain... Considering that in real-world scenarios, a vehicle's behavior is not only related to its own historical trajectories and those of other vehicles, but also possesses a degree of randomness, a fitter and a regulator are introduced to characterize this randomness. Their inputs are z ~ N(0, I), and their outputs are... The fitter and regulator are structured as a multi-layer perceptron (MLP), which samples N z-values. i ~N(0,I), obtain the corresponding random function information.
[0106] According to the updated H I and W i Generate latent variable state information, decode the latent variable state information, and obtain the output function of the traffic prior model. In one exemplary example, embodiments of the present invention can use a GRU module to process latent variable state information (H... I Wi Decode the code to obtain the output.
[0107] For all data in the dataset, based on the determined inputs and outputs, calculate the pre-defined first loss function to obtain the traffic prior model.
[0108] For the decoding process, based on the distance between the predicted trajectory and the actual trajectory, and referring to relevant principles such as the L1 and L2, the first loss function can be designed as follows:
[0109]
[0110]
[0111] Wherein, the subscript i1 is used to identify the decision-maker module of the test vehicle in scenario i, Y i1 This represents the actual state information of the decision-maker module in scenario i; This represents the predicted state information of the decision-maker module in scenario i, output by the traffic prior model; the subscript ij is used to identify vehicle j in scenario i, Y ij This represents the actual state information of vehicle j in scenario i; This represents the predicted state information of vehicle j in scenario i, output by the traffic prior model; N is the total number of vehicles in scenario i.
[0112] Considering the limited number of vehicle collisions in real-world datasets, a first constraint term is added to ensure the reasonableness of the prediction results. To ensure that no collisions occur between predicted trajectories; in one exemplary example, this application example approximates the vehicles in the scene using a three-circle or five-circle method. Figure 2 This is a schematic diagram illustrating the application of the three-circle approximation method for vehicles in this invention, as shown below. Figure 2 As shown, the three circles have the same radius, which is The width of the vehicle is such that the centers of the circles are equidistant from the vehicle length; let r be the approximate radius of the circle for vehicle j in scene i. ij In the scenario at time t, the interval between vehicle j and vehicle k is d. itjk =min u,v dist(l itju ,l itkv ); where l itju l represents the position of the center u of vehicle j in scene i at time t. itkv Let v represent the position of the center of the circle containing vehicle k in scenario i at time t. Then, the first constraint term is transformed as follows:
[0113]
[0114] in,
[0115] The example method in this application also includes: updating the MLP parameters of the fitter using the first gradient descent method, and training the fitter using a pre-defined second loss function; after training the fitter using the pre-defined second loss function, updating the MLP parameters of the regulator using the second gradient descent method, and training the regulator using a pre-defined third loss function.
[0116] In this application example, to test the decision-maker module, the traffic prior model needs to be fitted first to determine behavioral trends for any scenario, thereby generating reasonable and effective safety-critical test scenarios. To complete the fitting task, based on the obtained traffic prior model, keeping the other parameters of the traffic prior model unchanged, only the MLP parameters of the fitter are updated using the first gradient descent method, and the fitter is trained using the second loss function:
[0117]
[0118] Wherein, the subscript i1 is used to identify the decision-maker module of the test vehicle in scenario i, Y i1 This represents the actual state information of the decision-maker module in scenario i; This represents the predicted state information of the decision-maker module in scenario i, as output by the traffic prior model.
[0119] Meanwhile, in order to make the most of the parameters of the traffic prior model obtained through training, and to ensure that the behavior of the fitter conforms to the interaction relationships in the real scene (the basic concept describing the differences between different distributions, which is a domain consensus), an additional KL divergence loss function is introduced as a second constraint:
[0120]
[0121] Where: w1 is a random variable representing the randomness of the decision-maker module of the tested vehicle before the MLP parameters of the fitter are updated, and w′1 is a random variable representing the randomness of the decision-maker module of the tested vehicle after the MLP parameters of the fitter are updated. w1 and w′1 respectively follow... d w Let σ be the dimension of random variables w1 and w′1; 1k σ′ represents the standard deviation of the future location distribution of background vehicles output by the traffic prior model. 1k μ represents the standard deviation of the background vehicle future position distribution in the regulator output. 1k μ′ represents the mean of the future location distribution of background vehicles output by the traffic prior model. 1k This represents the mean of the future location distribution of background vehicles output by the regulator.
[0122] This application example generates test scenarios by using a traffic prior model trained with a pre-defined second loss function. For each scenario in the dataset, the regulator is optimized to obtain the test scenario. At this point, the remaining parameters of the traffic prior model remain unchanged and are only updated as MLP parameters of the regulator. The third loss function is then changed to... in, The model is based on the optimized parameters in 3), based on X. i The estimated future trajectory information of the decision module; in one exemplary instance, the third loss function in this application example consists of two terms:
[0123]
[0124] in, The trajectory information of the background vehicle.
[0125] in, r ij Let r be the approximate radius of the circle of vehicle j in scene i. ik Let d be the approximate radius of the circle of vehicle k in scene i; itjk d represents the interval between vehicle j and vehicle k. itjk =min u,v dist(l itju ,l itkv ), l itju l represents the position of the center u of vehicle j in scene i at time t. itkv The position of the center v of vehicle k in scenario i at time t; w tj For pre-defined weight items, This represents the distance between the background vehicle j and the vehicle being tested.
[0126] This application example uses a third loss function to cause background vehicles in the traffic prior model to spontaneously "collide" with the test vehicle, thereby generating a safety-critical test scenario.
[0127] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
Claims
1. A method for training a model to generate test scenarios, comprising: Obtain the dataset for model training, which contains historical observation sequences and map information; The historical observation sequences in the dataset are upgraded to obtain the first high-dimensional feature information; Based on the correlation between discrete points contained in the map information in the dataset, the second high-dimensional feature information of the map information is obtained; Based on the first high-dimensional feature information, the second high-dimensional feature information, and the pre-defined random function information representing the randomness of all vehicles in the scene, latent variable state information is generated. The generated latent variable state information is decoded to obtain the output of the traffic prior model; For all data in the dataset, calculate the pre-set first loss function based on the input of the traffic prior model and the output of the obtained traffic prior model to obtain the traffic prior model; Wherein, the first loss function is determined based on the distance between the predicted state information and the actual state information of the traffic prior model; the traffic prior model is used to generate test scenarios for autonomous vehicles; the generation of latent variable state information includes: processing the first high-dimensional feature information and the second high-dimensional feature information through a pre-set self-attention mechanism module to obtain updated third high-dimensional feature information representing the interaction relationship between vehicles; processing the third high-dimensional feature information through a pre-set cross-attention mechanism module to obtain updated fourth high-dimensional feature information representing the interaction relationship between vehicles and roads; and based on the third high-dimensional feature information, the fourth high-dimensional feature information, and the random function... The process of generating latent variable state information based on the third high-dimensional feature information, the fourth high-dimensional feature information, and the random function information includes: performing matrix multiplication on the updated third high-dimensional feature information and the fourth high-dimensional feature information; determining the latent variable state information based on the result of the matrix multiplication and the random function information; loading the random function information through a preset fitter and regulator; the fitter is used to predict the output state information of the decision-maker module based on the state information input to the decision-maker module of the test vehicle in the dataset; the regulator is used to adjust the behavior of background vehicles other than the test vehicle.
2. The training method according to claim 1, characterized in that, The dataset is as follows: ; in, For a pre-set fixed duration The sampling interval is The test scenario is correct. The map information; Represents a scenario used for training a traffic prior model. The historical observation sequence at that time, ; Scene The actual sequence of time, ; , indicating that the state at each moment in scene i is of dimension i. The vector, The total number of vehicles. For the dimension of state information; , representing a scene middle The set of all vehicle states at any given moment. Represents the time t of scenario i. Vehicle status information, This represents the status information of the decision-maker module of the vehicle under test.
3. The training method according to claim 1, characterized in that, Decoding the generated latent variable state information includes: The hidden variable state information is decoded by a pre-defined gated loop unit (GRU).
4. The training method according to any one of claims 1-3, characterized in that, The expression for the first loss function is: (1); (2); Here, the subscript i1 is used to identify the decision-maker module of the vehicle under test in scenario i. This represents the actual state information of the decision-maker module in scenario i; This represents the predicted state information of the decision-maker module in scenario i, output by the traffic prior model; the subscript ij is used to identify vehicle j in scenario i. This represents the actual state information of vehicle j in scenario i; This represents the predicted state information of vehicle j in scenario i, output by the traffic prior model; N is the total number of vehicles in scenario i.
5. The training method according to any one of claims 1-3, characterized in that, The training method further includes adding the following first constraint to the traffic prior model: (3); in, , For vehicles in scenario i The approximate circle radius, For vehicles in scenario i The approximate radius of the circle; Indicates vehicle With vehicles The interval, , Representative scenarios Vehicles in exist The center of time Location Representative scenarios Vehicles in exist The center of time Location.
6. The training method according to any one of claims 1-3, characterized in that, The training method also includes: Multilayer perceptron of the fitter The parameters are updated using the first gradient descent method, and the fitter is trained using a pre-defined second loss function; After training the fitter using a pre-defined second loss function, the regulator... The parameters are updated using the second gradient descent method, and the regulator is trained using a pre-defined third loss function.
7. The training method according to claim 6, characterized in that, The second loss function is: (4); Here, the subscript i1 is used to identify the decision-maker module of the vehicle under test in scenario i. This represents the actual state information of the decision-maker module in scenario i; This represents the state information predicted by the decision-maker module in scenario i, as output by the traffic prior model.
8. The training method according to claim 6, characterized in that, The third loss function is: (5); Here, the subscript i1 is used to identify the decision-maker module of the vehicle under test in scenario i. This represents the predicted state information of the decision-maker module in scenario i, as output by the traffic prior model. This represents the predicted state information of background vehicles in scene i, as output by the traffic prior model.
9. The training method according to claim 8, characterized in that, The expression for the third loss function is: ; (6); in, The trajectory information of the background vehicle. , , For vehicles in scenario i The approximate circle radius, For vehicles in scenario i The approximate radius of the circle; Indicates vehicle With vehicles The interval, , Representative scenarios Vehicles in exist The center of time Location Representative scenarios Vehicles in exist The center of time Location; , For pre-defined weight items, Representative background vehicle The distance between the vehicle being tested and the vehicle being tested.
10. The training method according to claim 6, characterized in that, The training method further includes: when training the fitter using a pre-defined second loss function, adding the following to the traffic prior model. The divergence loss function is used as the second constraint term: (7); in: The random variable representing the randomness of the decision-maker module of the tested vehicle before updating the MLP parameters of the fitter. The randomness of the decision-maker module of the tested vehicle after updating the MLP parameters of the fitter is a random variable. Obey each , For random variables dimensionality; This represents the standard deviation of the future location distribution of background vehicles output by the traffic prior model. This represents the standard deviation of the future position distribution of background vehicles output by the regulator. This represents the mean of the future location distribution of background vehicles output by the traffic prior model. This indicates the mean of the future location distribution of background vehicles output by the regulator.
11. A computer storage medium storing a computer program, wherein the computer program, when executed by a processor, implements a training method for generating a model for a test scenario as described in any one of claims 1-10.
12. A terminal, comprising: A memory and a processor, wherein the memory stores a computer program; wherein, The processor is configured to execute computer programs in memory; When the computer program is executed by the processor, it implements the training method for the model that generates the test scenario as described in any one of claims 1-10.
Citation Information
Patent Citations
Self-driving automobile reinforcement learning method, system and device and storage medium
CN114357860A
Vehicle scene early warning decision-making method and device, computer equipment and storage medium
CN114863342A