Method and device for training automatic driving control parameter model, and parameter acquisition method

By generating training samples through the interaction of reinforcement learning algorithms and dynamic simulation environment, the autonomous driving control parameter model is updated, which solves the problem of poor control effect caused by single parameters in traditional autonomous driving control schemes, and realizes refined control and efficient parameter adaptation for different scenarios.

CN114861318BActive Publication Date: 2026-02-10APOLLO INTELLIGENT DRIVING (BEIJING) TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210547436.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-18
Publication Date
2026-02-10
Estimated Expiration
2042-05-18

AI Technical Summary

Technical Problem

Traditional autonomous driving control schemes use a single control parameter, which cannot be adaptively adjusted according to changes in different scenarios, resulting in poor control performance.

Method used

By using an autonomous driving control parameter model training method, reinforcement learning algorithms and dynamic simulation environments are interactively generated to generate training samples, update the control parameter model, and achieve adaptive control for different scenarios.

Benefits of technology

It improves the efficiency and speed of model training sample generation, enables fine-grained control over different scenarios, reduces model distortion, and improves the efficiency and effectiveness of control parameter adaptation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114861318B_ABST
    Figure CN114861318B_ABST
Patent Text Reader

Abstract

The present disclosure provides a training method and device of an automatic driving control parameter model, and a parameter acquisition method. The present disclosure relates to the technical field of computers, and in particular to the field of automatic driving and the field of automatic control. The specific implementation is as follows: scene information is input into a first control parameter model to obtain a control parameter output by the first control parameter model; the control parameter is interacted with a dynamics simulation environment to obtain a training sample; and the first control parameter model is updated according to the training sample to obtain a second control parameter model after training. According to the present disclosure, the first control parameter model automatically generates a control parameter according to scene information, and then automatically obtains a training sample, thereby improving the generation efficiency of the training sample and the model training speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and more particularly to the fields of autonomous driving and automatic control. Background Technology

[0002] Lateral and longitudinal control is fundamental to autonomous driving, and its quality directly impacts the effectiveness of the system. Traditional control schemes typically employ single control parameters that do not adapt to changes in the scenario. In reality, control objectives often differ across scenarios, and achieving optimal control requires employing appropriate parameters tailored to each scenario. Summary of the Invention

[0003] This disclosure proposes a training method for an autonomous driving control parameter model, a method for acquiring control parameters, and an apparatus.

[0004] According to one aspect of this disclosure, a method for training an autonomous driving control parameter model is provided, comprising:

[0005] The autonomous driving scenario information is input into the first control parameter model to obtain the control parameters output by the first control parameter model;

[0006] The system interacts with the dynamic simulation environment based on the control parameters to obtain training samples; and

[0007] The first control parameter model is updated based on the training samples to obtain the trained second control parameter model.

[0008] According to another aspect of this disclosure, a method for obtaining control parameters is provided, comprising:

[0009] The target scene information is input into the autonomous driving control parameter model for processing to obtain the target control parameters output by the autonomous driving control parameter model.

[0010] The autonomous driving control parameter model is a second control parameter model trained using any of the training methods in the embodiments of this disclosure.

[0011] According to another aspect of this disclosure, a training apparatus for a control parameter model is provided, comprising:

[0012] The input module is used to input scene information into the first control parameter model and obtain the control parameters output by the first control parameter model;

[0013] The acquisition module is used to interact with the dynamic simulation environment based on the control parameters to obtain training samples; and

[0014] The training module is used to train the first control parameter model based on the training samples to obtain the trained second control parameter model.

[0015] According to another aspect of this disclosure, a control parameter acquisition device is provided, comprising:

[0016] The acquisition module is used to input target scene information into the autonomous driving control parameter model for processing, and obtain the target control parameters output by the autonomous driving control parameter model.

[0017] The autonomous driving control parameter model is a second control parameter model obtained by training using any of the training devices described above.

[0018] According to another aspect of this disclosure, an electronic device is provided, comprising:

[0019] One or more processors;

[0020] A memory that is communicatively connected to the one or more processors;

[0021] One or more computer programs, wherein the one or more computer programs are stored in the memory, and when the one or more computer programs are executed by the electronic device, cause the electronic device to perform the method provided by any of the preceding claims.

[0022] According to another aspect of this disclosure, a computer-readable storage medium is provided that stores computer instructions that, when executed on a computer, cause the computer to perform the method provided in any of the preceding claims.

[0023] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method described in any of the above.

[0024] This embodiment of the disclosure automatically generates control parameters based on scene information through a first control parameter model, thereby automatically obtaining training samples, improving the efficiency of training sample generation, and thus improving the model training speed.

[0025] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0026] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0027] Figure 1This is a flowchart of a training method for an autonomous driving control parameter model according to an embodiment of the present disclosure;

[0028] Figure 2 A flowchart of a method for training an autonomous driving control parameter model according to another embodiment of the present disclosure;

[0029] Figure 3 A flowchart of a method for training an autonomous driving control parameter model according to another embodiment of the present disclosure;

[0030] Figure 4 A flowchart of a method for training an autonomous driving control parameter model according to another embodiment of the present disclosure;

[0031] Figure 5 This is a flowchart of a method for obtaining control parameters according to an embodiment of the present disclosure;

[0032] Figure 6 This is a schematic diagram of the structure of a training device for an autonomous driving control parameter model according to an embodiment of the present disclosure;

[0033] Figure 7 This is a schematic diagram of the structure of a training device for an autonomous driving control parameter model according to another embodiment of the present disclosure;

[0034] Figure 8 This is a schematic diagram of a control parameter acquisition device according to an embodiment of the present disclosure;

[0035] Figure 9 This is a schematic diagram of the structure of an automated parameter self-adjustment framework according to an embodiment of the present disclosure;

[0036] Figure 10 A schematic block diagram of an example electronic device used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0037] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0038] Figure 1 This is a flowchart of a method for training an autonomous driving control parameter model according to an embodiment of the present disclosure. The method may include:

[0039] S101. Input the autonomous driving scenario information into the first control parameter model to obtain the control parameters output by the first control parameter model;

[0040] S102. Interact with the dynamic simulation environment based on the control parameters to obtain training samples; and

[0041] S103. The first control parameter model is trained based on the training samples to obtain the trained second control parameter model.

[0042] In this embodiment, scene information may include various state parameters corresponding to a controlled object in a certain state. Different states may correspond to different scene information, which can be determined based on the characteristics of the actual state. Control parameters may include various control parameters corresponding to the controlled object performing a certain action in that state. The input of the first control parameter model is scene information, and the output is control parameters. The first control parameter model can process the input scene information and output the control parameters corresponding to that scene information. Then, it can interact with the dynamic simulation environment based on the control parameters to obtain the scene information corresponding to the next state, the reward function value, etc., thereby obtaining training samples. The first control parameter model is trained using the training samples to update the parameters of the first control parameter model, resulting in a second control parameter model. The first control parameter model automatically generates control parameters based on scene information and automatically obtains training samples, improving the efficiency of training sample generation and thus improving the model training speed.

[0043] In one possible implementation, the autonomous driving scenario information includes at least one of the following: speed, curvature, lateral position error, lateral heading angle error, longitudinal position error, longitudinal velocity error, longitudinal acceleration error, desired velocity, desired acceleration, desired lateral displacement, and desired heading angle.

[0044] In this embodiment, different states may correspond to different scene information. For example, the scene information corresponding to state 1 includes velocity, curvature, lateral position error, lateral heading angle error, desired velocity, and desired acceleration. The scene information corresponding to state 2 includes velocity, curvature, longitudinal position error, longitudinal velocity error, longitudinal acceleration error, desired lateral displacement, and desired heading angle. The scene information listed in this embodiment is merely an example and not a limitation. In practical applications, other scene information, such as jerk, may be included. With rich scene information, more refined scene segmentation can be achieved, resulting in a wider and more accurate range of applications.

[0045] In one possible implementation, the control parameters include at least one of the following: a lateral advance Q value, a lateral advance R1 value, a lateral advance R2 value, a lateral decay value, a longitudinal Q value, a longitudinal R1 value, a longitudinal R2 value, and a longitudinal decay value. Wherein, the Q value represents the penalty weight of the error term in the objective function of the Model Predictive Control (MPC) algorithm, the R1 value represents the penalty weight of the control quantity in the objective function, and the R2 value represents the penalty weight of the control quantity increment in the objective function.

[0046] In this embodiment of the disclosure, the control model may include a lateral model and a longitudinal model. The Q-value for lateral advancement is related to the state variables included in the lateral model. If the lateral model includes multiple state variables, the Q-value for lateral advancement may include the error term penalty weight value corresponding to each state variable. The Q-value for longitudinal advancement is related to the state variables included in the longitudinal model. If the longitudinal model includes multiple state variables, the Q-value for longitudinal advancement may include the error term penalty weight value corresponding to each state variable.

[0047] For example, if the state variables of the lateral model include lateral displacement, lateral velocity, heading angle, yaw rate, and actual front wheel steering angle, then the Q value of lateral forward movement can include a diagonal matrix of error term penalty weight values ​​q1, q2, q3, q4, and q5 for these five state variables.

[0048] For example, if the longitudinal model includes displacement, velocity, and actual torque, then the Q value of the longitudinal direction can be a diagonal matrix containing the error term penalty weights q1, q2, and q3 of these three state variables.

[0049] In the embodiments disclosed herein, the use of more multi-dimensional control parameters enables applicability to a wider range of control scenarios.

[0050] In one possible implementation, such as Figure 2 As shown, in S101, the autonomous driving scenario information is input into the first control parameter model to obtain the control parameters output by the first control parameter model, including: S201, the first scenario information is input into the first control parameter model and processed by the first policy network to obtain the first control parameters output by the first policy network.

[0051] In this embodiment, the control parameter model can be a model based on reinforcement learning algorithms, such as the DDPG (Deep Deterministic Policy Gradient) model. For example, the control parameter model can include an online network and a target network, which helps accelerate network convergence. The online network is a first control parameter model, which can include a first policy network and a first value network. The target network is a second control parameter model, which can include a second policy network and a second value network. During the acquisition of training samples, the first policy network of the online network can be mainly utilized. The first policy network processes the first scene information to obtain the first control parameters for execution in the first state. Thus, the first policy network of the first control parameter model can automatically generate the first control parameters based on the first scene information, improving the efficiency of control parameter generation. Then, the first policy network can send the first control parameters to the control module. The control module can use the first control parameters and the objective function to calculate the control quantity. Then, based on the control quantity, it interacts with the dynamic simulation environment to obtain information for transitioning from the first state to the second scene. The second scene information is used as the first scene information for the next iteration and input into the first policy network to iteratively generate the scene information corresponding to the next state. This process continues, without further elaboration. The control model of the control module can include a horizontal model and a vertical model. The objective function can be constructed based on the horizontal model, the vertical model, and penalty weight parameters.

[0052] In this embodiment, the control module may not directly use the first control parameter to calculate the objective function value, but instead attenuate the first control parameter before calculating the objective function value. In one approach, the control module may perform a penalty weight attenuation process on the first control parameter, then use the attenuated control parameter to calculate the objective function value, and interact with the dynamic simulation environment based on the objective function value and the attenuated control parameter to obtain second scene information.

[0053] In one possible implementation, such as Figure 2 As shown, in S102, the system interacts with the dynamic simulation environment based on the control parameters to obtain training samples, including:

[0054] S202. The Q value in the first control parameter is attenuated according to the attenuation coefficient, and the attenuated Q value is substituted into the objective function of the control model to calculate the control quantity, wherein the Q value represents the error term penalty weight in the objective function of the control model.

[0055] S203. Using the control quantity to interact with the dynamic simulation environment, obtain the simulation interaction result output by the dynamic simulation environment, wherein the simulation interaction result includes second scene information;

[0056] S204. Calculate the reward function value based on the first scene information and the simulation interaction result; and

[0057] S205. Obtain sample data based on the first scenario information, the first control parameters, the return function value, and the second scenario information.

[0058] In this embodiment of the disclosure, the control module can construct an objective function for the optimization problem using the horizontal model, vertical model, and penalty weight parameters in the control model, and use the objective function to attenuate the first control parameter. For example, the penalty weight parameters in the objective function may include the penalty weight Q of the error term, the penalty weight R1 of the control quantity, and the penalty weight R2 of the control quantity increment, etc.

[0059] In this embodiment of the disclosure, based on the first scene information s i Obtain the first control parameter a i Then, the first control parameter is subjected to penalty weight decay processing to obtain the decayed control parameter a. id Using the attenuation control parameter a id Substituting the objective function, the control variable is calculated. By interacting with the dynamic simulation environment using the control variable, information including the second scene s can be obtained. i+1 The simulation interaction results. Furthermore, based on the first scene information s i The reward function value r can be calculated from the simulation interaction results. i Then, (s) i ,a i ,r i ,s i+1 This can be used as a sample data point. This sample data can be used directly as a training sample, or it can be obtained by sampling from multiple sample data points.

[0060] In this embodiment of the disclosure, the first control parameter is subjected to penalty weight decay processing, which enables the model's sample data to be obtained based on the decayed control parameter, thus helping to reduce model distortion.

[0061] In one possible implementation, the reward function value is determined based on the error reward value, the error rate of change reward value, the control quantity change reward value, and the simulation metric reward value.

[0062] For example, the reward function r(i) = b1r1(i) + b2r2(i) + b3r3(i) + b4r4(i). Here, r1(i) represents the error reward, r2(i) represents the error rate of change reward, r3(i) represents the control quantity change reward, and r4(i) represents the simulation metric reward. b1 to b4 represent the coefficients of each term, which can be empirical values. The simulation metric reward can be calculated based on one or more of the following: collision reward, sudden braking reward, sudden steering reward, and trajectory replanning reward. Calculating the reward function value from multiple perspectives, including error, error rate of change, control quantity change, and simulation metric, helps to obtain a more suitable reward function value.

[0063] In one possible implementation, in S202, the first control parameter is subjected to penalty weight attenuation processing to obtain attenuated control parameters corresponding to the first action. This includes: attenuating the Q value in the first control parameter according to an attenuation coefficient to obtain a attenuated Q value, wherein the attenuated Q value is less than the Q value before attenuation processing. The Q value represents the error term penalty weight in the objective function of the Model Predictive Control (MPC) algorithm. For example, if the attenuation coefficient d < 1, the error term penalty weight Q at prediction step i... i =Q*d i i can be a natural number. For example, Q0 = Q in the first step, Q1 = Q*d in the second step, and so on. In this way, the weight Q can be penalized as the number of prediction steps i increases. i The smaller the value, the better to solve the problem of model distortion.

[0064] In one possible implementation, in S205, sample data is obtained based on the first scene information, the first control parameter, the reward function value, and the second scene information, including:

[0065] Based on the first scenario information, the first control parameters, the reward function value, and the second scenario information, the sample data for this test is generated; and

[0066] The second scene information of the sample data generated each time is used as the first scene information for the next input of the first control parameter model. The steps of obtaining sample data are executed iteratively, for example, S201 to S205, to obtain multiple sample data.

[0067] In this embodiment of the disclosure, a sample data set can also be referred to as a group of sample data, which is a dataset. For example, the sample data includes first scene information, first control parameters, reward function values, and second scene information. The second scene information s of the sample data is obtained in this instance. i+1 This can be used as the input to the first control parameter model for the next first scene information, and the first control parameter model will output the first control parameter a again. i+1Then, the first control parameter is attenuated to obtain the attenuated control parameter. Using the attenuated control parameter to interact with the dynamic simulation environment, the second scene information s can be obtained. i+2 Furthermore, sample data (s) can be obtained. i+1 ,a i+1 ,r i+1 ,s i+2 In this way, multiple sample data can be obtained step by step, thereby establishing a suitable training sample set for subsequent model training.

[0068] In one possible implementation, in S102, interacting with the dynamic simulation environment based on the control parameters to obtain training samples further includes:

[0069] S206. Add the multiple sample data sets to the playback pool; and

[0070] S207. Sample multiple sample data in the playback pool to obtain training samples.

[0071] In this embodiment, the reinforcement learning algorithm may include a replay pool, also known as a memory replay pool or experience replay pool. The replay pool may contain multiple sample data points, and training samples can be obtained by randomly sampling from these multiple sample data points. For example, a minimum batch size can be set, and a minimum batch of training samples can be randomly sampled from the replay pool according to this minimum batch size, thus forming the training sample set. By sampling, such as random sampling, the correlation between selected samples can be eliminated, resulting in a better training sample set and a more accurate trained model.

[0072] In one possible implementation, the training samples include first scene information, first control parameters, reward function values, and second scene information. The first control parameter model includes a first policy network, a first value network, a second policy network, and a second value network.

[0073] In one possible implementation, such as Figure 3 As shown, in S103, the first control parameter model is trained based on the training samples to obtain the trained second control parameter model, including:

[0074] S301. Input the first scene information into the first policy network and the second policy network respectively for processing to obtain the first control parameter output by the first policy network and the second control parameter output by the second policy network.

[0075] S302. Input the first scene information and the first control parameters into the first value network for processing to obtain the first value function output by the first value network.

[0076] S303. Input the first scene information and the second control parameters into the second value network for processing to obtain the second value function output by the second value network; and

[0077] S304. Based on the training samples, the first value function, and the second value function, train the first control parameter model to obtain the trained second control parameter model.

[0078] In this embodiment, the first policy network and the first value network are online networks, while the second policy network and the second value network are target networks. In reinforcement learning, updating model parameters through multiple networks helps to accelerate model convergence.

[0079] In one possible implementation, such as Figure 4 As shown, in S304, the first control parameter model is trained based on the training samples, the first value function, and the second value function to obtain the trained second control parameter model, including:

[0080] S401. Update the parameters of the first value network using a loss function, wherein the loss function is determined based on the first scene information, the first control parameters, the second scene information, the second control parameters, the first value function, and the second value function;

[0081] S402. Update the parameters of the first policy network using gradient descent, wherein the gradient descent method requires at least the first scene information, the first control parameters, and the first value function.

[0082] S403. Update the parameters of the second policy network according to the parameters of the first policy network and the parameters of the second policy network; and

[0083] S404. Update the parameters of the second value network according to the parameters of the first value network and the parameters of the second value network.

[0084] In this embodiment of the disclosure, the first policy network, the first value network, the second policy network, and the second value network may each have a corresponding update formula.

[0085] For example, the update formula for a first-valued network is: Where L is the loss function and N is the prediction time domain. i =r i +γV'(s i+1 ,μ'(s i+1 |θ μ’ )|θ V’ V is the output of the first-valued network, si For the first scene information, a i Represents the first control parameter, s i+1 This is information for the second scene. θ V Let θ be the parameter of the first-valued network. V′ These are the parameters of the second-valued network.

[0086] For example, the update formula for the first-policy network is:

[0087]

[0088] in, Let θ be the policy gradient. μ Here are the parameters for policy network 1, where s represents scene information, a represents control parameters, and θ represents the control parameters. V The network parameters are denoted by μ(s|θ). μ V(s,a|θ) represents the output of the first policy network. V ) represents the output of the first-policy network. Based on this formula, the first-policy network can be updated using gradient descent.

[0089] For example, the update formula for a second-valued network is: θ V′ ←τθ V +(1τ)θ V′ The parameters θ based on the first-valued network V The parameters θ of the second-valued network 2 V′ The parameters of value network 2 can be updated. τ represents the proportion, which can be an empirical value.

[0090] For example, the update formula for the second policy network is: θ μ′ ←τθ μ +(1τ)θ μ′ Based on the parameters θ of policy network 1 μ The parameters θ of policy network 2 μ′ The parameters of policy network 2 can be updated. τ represents the proportion, which can be an empirical value. The value of τ in the update formula of the second policy network can be the same as or different from the value of τ in the update formula of the second value network.

[0091] The timing of steps S401 to S404 is not restricted and can be adjusted arbitrarily according to requirements. For example, the first policy network and the first value network can be updated in parallel, the second policy network and the second value network can be updated in parallel, or the second policy network can be updated first and then the first policy network. By using loss functions and gradient descent methods to update the first policy network, the first value network, the second policy network, and the second value network, the control parameter model can converge quickly.

[0092] Figure 5This is a flowchart of a method for obtaining control parameters according to an embodiment of the present disclosure. The method may include: S501, inputting target scene information into an autonomous driving control parameter model for processing to obtain target control parameters output by the autonomous driving control parameter model; wherein, the autonomous driving control parameter model is a second control parameter model trained using any of the training methods described in the above embodiments. Based on the trained second control parameter model, corresponding target control parameters can be generated quickly and accurately for different target scene information without manual parameter tuning, thus improving tuning efficiency.

[0093] Figure 6 This is a schematic diagram of a training device for an autonomous driving control parameter model according to an embodiment of the present disclosure. The device may include:

[0094] Input module 601 is used to input autonomous driving scenario information into the first control parameter model and obtain the control parameters output by the first control parameter model;

[0095] Acquisition module 602 is used to interact with the dynamic simulation environment based on the control parameters to acquire training samples; and

[0096] The training module 603 is used to train the first control parameter model based on the training samples to obtain the trained second control parameter model.

[0097] In one possible implementation, the scene information includes at least one of the following: velocity, curvature, lateral position error, lateral heading angle error, longitudinal position error, longitudinal velocity error, longitudinal acceleration error, desired velocity, desired acceleration, desired lateral displacement, and desired heading angle.

[0098] In one possible implementation, the control parameters include at least one of the following: a lateral advance Q value, a lateral advance R1 value, a lateral advance R2 value, a lateral decay value, a longitudinal Q value, a longitudinal R1 value, a longitudinal R2 value, and a longitudinal decay value. Wherein, the Q value represents the penalty weight of the error term in the objective function of the Model Predictive Control (MPC) algorithm, the R1 value represents the penalty weight of the control quantity in the objective function, and the R2 value represents the penalty weight of the control quantity increment in the objective function.

[0099] Figure 7 This is a schematic diagram of a training apparatus for an autonomous driving control parameter model according to an embodiment of the present disclosure. The apparatus of this embodiment includes one or more features of the above-described training apparatus embodiment for an autonomous driving control parameter model. In one possible implementation, the input module 601 includes:

[0100] The input submodule 701 is used to input the first scene information into the first policy network of the first control parameter model for processing, and obtain the first control parameters output by the first policy network.

[0101] In one possible implementation, the acquisition module 602 includes:

[0102] The processing submodule 702 is used to attenuate the Q value in the first control parameter according to the attenuation coefficient, and to substitute the attenuated Q value into the objective function of the control model to calculate the control quantity, wherein the Q value represents the error term penalty weight in the objective function of the control model;

[0103] The interaction submodule 703 is used to interact with the dynamic simulation environment using the control quantity to obtain the simulation interaction result output by the dynamic simulation environment, and the simulation interaction result includes second scene information;

[0104] Calculation submodule 704 is used to calculate the reward function value based on the first scene information and the simulation interaction result; and

[0105] The acquisition submodule 705 is used to acquire sample data based on the first scene information, the first control parameters, the reward function value, and the second scene information.

[0106] In one possible implementation, the reward function value is determined based on the error reward value, the error rate of change reward value, the control quantity change reward value, and the simulation metric reward value.

[0107] In one possible implementation, the acquisition submodule 705 is further configured to:

[0108] Based on the first scenario information, the first control parameters, the reward function value, and the second scenario information, the sample data for this test is generated; and

[0109] The second scene information of each generated sample data is used as the first scene information for the next input of the first control parameter model. The steps of obtaining sample data are iteratively executed to obtain multiple sample data.

[0110] In one possible implementation, the acquisition module 602 further includes:

[0111] Playback submodule 706 is used to add multiple sample data sets to the playback pool; and

[0112] The sampling submodule 707 is used to sample multiple sample data in the playback pool to obtain training samples.

[0113] In one possible implementation, the training samples include first scene information, first control parameters, reward function values, and second scene information. The first control parameter model includes a first policy network, a first value network, a second policy network, and a second value network.

[0114] In one possible implementation, the training module 603 includes:

[0115] The strategy network submodule 708 is used to input the first scene information into the first strategy network and the second strategy network respectively for processing, and obtain the first control parameter output by the first strategy network and the second control parameter output by the second strategy network.

[0116] Value network submodule 709 is used to input the first scene information and the first control parameters into the first value network for processing to obtain a first value function output by the first value network; input the first scene information and the second control parameters into the second value network for processing to obtain a second value function output by the second value network; and

[0117] The update submodule 710 is used to train the first control parameter model based on the training samples, the first value function, and the second value function to obtain the trained second control parameter model.

[0118] In one possible implementation, the update submodule 303 is used to:

[0119] The parameters of the first value network are updated using a loss function, which is determined based on the first scene information, the first control parameters, the second scene information, the second control parameters, the first value function, and the second value function.

[0120] The parameters of the first policy network are updated using gradient descent, and the information required by the gradient descent method includes at least the first scene information, the first control parameters, and the first value function.

[0121] Update the parameters of the second policy network based on the parameters of the first policy network and the parameters of the second policy network; and

[0122] Update the parameters of the second value network based on the parameters of the first value network and the parameters of the second value network.

[0123] Figure 8 This is a schematic diagram of a control parameter acquisition device according to an embodiment of the present disclosure. The device may include:

[0124] The acquisition module 801 is used to input the target scene information into the autonomous driving control parameter model for processing, and obtain the target control parameters output by the autonomous driving control parameter model.

[0125] The autonomous driving control parameter model is a second control parameter model obtained by training using any of the training devices in the embodiments.

[0126] The functions of each module and / or sub-module in the device embodiments of this disclosure can be found in the relevant descriptions in the above method embodiments of this disclosure, and will not be repeated here.

[0127] This disclosure provides a method for full-scale parameter self-adjustment in vehicle lateral and longitudinal control, which can significantly improve the efficiency of parameter adaptation and the control effect. This disclosure is applicable to, but not limited to, applications such as autonomous driving lateral and longitudinal control and robot control.

[0128] There are various methods for optimizing autonomous driving control. For example, optimizing the dynamic model involves considering more complex models, adding more reasonable constraints, and modeling environmental uncertainties and noise. Another approach is to add adaptive features to a simple model, adjusting control parameters or outputs for different scenarios to achieve better control performance. The first method is difficult to implement due to its large computational latency, and the second method is currently the mainstream choice in the industry. However, the design of related technologies in terms of adaptation is relatively simple. For example, they only distinguish scenarios based on speed, curvature (lateral), and acceleration (longitudinal), providing corresponding MPC control parameters (e.g., QR matrix), without performing fine-grained parameter optimization for each scenario.

[0129] The related technologies have the following problems:

[0130] 1. The scene division is rather coarse, generally only considering velocity, curvature, and acceleration information. It does not consider the impact of control errors on control parameters, nor does it take into account the mutual influence between lateral and longitudinal control.

[0131] 2. The control parameters for different scenarios are not adapted in a differentiated manner. Often, a few scenarios are roughly selected to determine their control parameters, while the parameters for the remaining scenarios are determined by linear interpolation, which cannot achieve better control results.

[0132] 3. The parameters of different vehicles are often different, and there are many parameters that need to be adjusted. Manual parameter adjustment is costly and has a low degree of automation.

[0133] This disclosure provides a model predictive control method, which mainly includes a predictive model, rolling optimization, and feedback correction. At each sampling time, based on the current state information, the future state is predicted using the system model, and the error between the predicted trajectory and the desired trajectory is calculated as the cost function. Furthermore, a constrained control problem is constructed based on the constraint information. Then, the optimal control sequence is solved, and the first control variable in the optimal control sequence is applied to the system. A new optimal control sequence is calculated at the next sampling time. This model predictive control method requires setting the objective function of the optimization problem.

[0134] For example, the objective function for an MPC control model with prediction time domain N is:

[0135]

[0136] Y = DX, Y r =DX r

[0137] The objective function consists of three parts: a penalty for the error term (Y). i -Y ri ) T Q(Y i -Y ri ), penalty for control quantity U i T R1U i And the penalty ΔU for the increment of the control quantity i T R2ΔU i Among them, (Y) i -Y ri ) represents the error term, Y i Y is the current state value. ri U represents the desired state value. i For control quantity, ΔU i The increment is the control quantity. In this objective function, i is the discretized time value, which can be understood as one i corresponding to each execution step. X is the state equation of the control model.

[0138] For general control problems, D is an identity matrix, and the penalty weight parameters include the penalty weight Q for the error term, the penalty weight R1 for the control quantity, and the penalty weight R2 for the control quantity increment. The penalty weight parameters are a diagonal matrix. If the values ​​on the diagonal do not change with i in MPC control, it means that the penalty weights are equal for different prediction steps. The parameters Q, R1, and R2 directly affect the control effect, and the determination of the magnitudes of these three parameters is the result of balancing the control error and the control quantity. For example, increasing Q can reduce the error; increasing R1 can reduce lateral sway (e.g., causing severe lateral rolling), making the longitudinal movement more energy-efficient; increasing R2 can prevent excessive lateral steering and excessive sudden braking and acceleration in the longitudinal direction.

[0139] The following section provides examples of control models. Control models can include longitudinal and lateral models.

[0140] The longitudinal model can include state variables, Q, and control variables. For example, state variables can include displacement x, velocity v, and actual torque T. The penalty weight for the error term is Q = blockdiag{q1,q2,q3}. Control variables can include the applied torque. The state equations of the longitudinal model are as follows. Then we can obtain the following formula:

[0141]

[0142] Among them, v r Let m be the vehicle's desired speed, m be the vehicle's mass, and m be the vehicle's mass. e Let k be the vehicle's equivalent mass, k be the drag coefficient, R be the tire radius, g be the gravitational constant, and C be the weight of the vehicle. r τ is the rolling resistance coefficient, θ is the pitch angle, and τ is the time delay coefficient. des This is to apply downward torque.

[0143] The lateral model can include state variables, Q-values, and control variables. For example, state variables can include lateral displacement y. e lateral velocity heading angle θ e yaw rate And the actual front wheel steering angle δ. The penalty weight of the error term Q = blockdiag{q1,q2,q3,q4,q5}. The control quantity can include the issued front wheel steering angle.

[0144] State equations of the longitudinal model Then we can obtain the following formula:

[0145]

[0146] c f For front wheel lateral stiffness, c r For rear wheel lateral stiffness, v xLet m be the vehicle's longitudinal velocity, m be the vehicle's mass, and l be the vehicle's longitudinal velocity. f The distance from the center of gravity to the front axle, l r The distance from the center of mass to the back axis, I z Let δ be the moment of inertia, τ be the time delay coefficient, and δ be the time delay coefficient. des To issue the front wheel steering angle.

[0147] In this embodiment, a weight decay strategy is adopted to reduce model prediction distortion; a neural network is used as a carrier to establish a mapping relationship between scene information and optimal parameters to achieve fine-grained scene segmentation; and an automated parameter self-adjustment framework is used to train the neural network to obtain the optimal control parameters (also referred to as control information, control parameter information, etc.).

[0148] I. Weight decay strategy in the prediction time domain

[0149] On the one hand, the simplified model used in MPC has certain errors, resulting in poor prediction accuracy for distant states. On the other hand, control is actually more concerned with nearby errors. Traditional MPC control uses the same penalty weight for different prediction steps. The scheme in this embodiment adds a decay coefficient (d < 1). This decay coefficient allows the penalty weight Q to decrease as the prediction step number i increases, thus addressing the model distortion problem. In this embodiment, instead of using a constant penalty weight Q for the error term, the penalty weight Q decreases as the prediction step number increases. Let Q... i To determine the penalty weight for predicting step i, then Q... i =Q*d i-1 .

[0150] II. Refined Scene Division Methods

[0151] Compared with the traditional linear interpolation method for scene segmentation, the solution of this disclosure uses a more refined neural network model as a carrier to describe the mapping relationship between scenes and parameters, thereby segmenting scenes more precisely.

[0152] The input to the neural network model can be scene information, such as autonomous driving scene information. For example, scene information can include: speed, curvature, lateral position error, lateral heading angle error, longitudinal position error, longitudinal velocity error, longitudinal acceleration error, desired velocity, desired acceleration, desired lateral displacement, and desired heading angle (system state), totaling 11 dimensions.

[0153] The output of the neural network model is control parameter information. For example, the control parameter information includes: horizontal forward Q-value (5 dimensions), horizontal forward R1 and R2 values, horizontal decay value, vertical Q-value (3 dimensions), vertical R1 and R2 values, and vertical decay value, for a total of 14 dimensions. Since the penalty weights are relative, the horizontal and vertical R1 values ​​can be fixed at 1. In this case, the R values ​​that need to be trained can only include the penalty weight R2 for Δu, thus reducing the dimensionality of the control parameter information to 12 dimensions.

[0154] The specific content and dimensions of the scenario information and control parameters mentioned above are merely examples and not limitations. In practical applications, they can be added or reduced according to specific needs.

[0155] III. Automated Parameter Self-Tuning Framework

[0156] Sample data is obtained by interacting with the simulation environment using reinforcement learning methods. Then, the neural network model is trained using the sample data to obtain the optimal control parameters of MPC, such as Q, R2, and decay coefficient.

[0157] See Figure 9 The following is an example of the specific algorithm steps based on this parameter self-adjustment framework:

[0158] 1. The policy network 1 of the online network can generate control parameters a (e.g., including Q, R2, and decay) based on the scenario information s, and transmit the control parameters a to the control module.

[0159] 2. The control module interacts with the simulation environment to generate a series of sample data, which is then stored in the memory playback pool. An example of the sample data can be found in the following formula.

[0160]

[0161] Among them, s i For the scene information corresponding to the first state, a i The control parameter information corresponding to the first action taken in the first state, r i The reward function value obtained after taking the first action in the first state, s i+1 This refers to the control parameter information corresponding to the transition to the second state after taking the first action in the first state. Here, the subscript i ranges from 0 to n, where n is a natural number. For example, the scene information s corresponding to the first state is input into policy network 1. n Policy network 1 can output the control parameter information a corresponding to taking the first action in the first state. n(For example, including Q, R2, and decay). In the control module, a penalty weight decay strategy can be used to calculate the weight decay of Q in the control parameter information corresponding to the first action, causing Q to change. The calculated decayed a... n Send it to the simulation environment. The simulation environment can be based on a n Output the scene information s corresponding to the second state. n+1 Furthermore, according to s n+1 The return function value r can be calculated. n If n is 0, the first set of sample data includes (s0, a0, r0, s1). If n is 1, the second set of sample data includes (s1, a1, r1, s2). Multiple sets of sample data can be generated iteratively through the policy network.

[0162] 3. By randomly sampling, the correlation between samples is eliminated, and the minimum batch of sample data is selected from the sample data in the memory replay pool.

[0163] 4. Using a reinforcement learning algorithm such as DDPG, train the policy network and value network offline with a minimum batch of sample data. Update the policy using a reinforcement learning algorithm. Repeat steps 1 to 4 until the network converges (both the policy network and value function converge).

[0164] Examples of the input and output of the policy network and value network are as follows:

[0165] (1) Policy Network:

[0166] Input: Scenario information s corresponding to the system state n For example, it includes speed, curvature, error, and desired speed.

[0167] Output: Control parameter information a corresponding to the action n Examples include Q, R2, and decay values.

[0168] (2) Value network:

[0169] Input: Scenario information s corresponding to the system state n and the control parameter information a corresponding to the action. n

[0170] Output: Value function V(s) n ) = r n +V(s n+1 )*γ, where γ is the discount factor. The value function can evaluate the quality of a state and action.

[0171] The neural network model based on the DDPG algorithm can include an online network and a target network. The online network can include policy network 1 and value network 1, and the target network can include policy network 2 and value network 2. V represents the output of the value network, s represents the scene information corresponding to the state, and a represents the control parameter information corresponding to the action. During training, the scene information s corresponding to the system state in the samples can be used... i Input policy network 1 and policy network 2. Policy network 1 can output the control parameter information a corresponding to the action. i Policy Network 2 can output control parameter information a corresponding to the action. i '. Then you can put s i and a i Input value network 1, s i and a i Input value network 2. Value network 1 and value network 2 can be based on V(s) i ) = r i +V(s i+1 )*γ calculates value function 1 as V and value function 2 as V'.

[0172] To address the overestimation problem, the DDPG algorithm employs two sets of value networks and policy networks. The second set of value networks and policy networks consists of a target value network and a target policy network, with the same initial parameters as the first set of networks.

[0173] The following section presents examples of update formulas for the policy network and value network.

[0174] The output of value network 1 can be represented as V(s,a|θ) V ), where V represents the value network, s represents the state, a represents the action, and θ represents the action. V These are the network parameters.

[0175] The output of the policy network can be represented as μ(s|θ) μ Where μ represents the policy network, θ μ These are the policy network parameters.

[0176] The target y of Value Network 2 is the target value, the policy gradient includes the update formula of the policy network, and the value loss function includes the update formula of the value network.

[0177] The update formula for value network 1 is: Where L is the loss function and N is the prediction time domain. i =r i +γV'(s i+1 ,μ'(s i+1 |θ μ’ )|θ V’V is the output of value network 1, s i For the first scene information, a i Represents the first control parameter, s i+1 For the second scene information, 'i' can represent time. θ V Let θ be the parameter of the first-valued network. V′ Let L be the parameters of the second-valued network. Based on this formula, the valued network 1 can be updated based on the minimized loss L.

[0178] The update formula for Policy Network 1 is:

[0179]

[0180] in, Let θ be the policy gradient. μ Here are the parameters for Policy Network 1, where s represents the state, a represents the action, and θ represents the action. V For the network parameters, μ(s|θ) μ V(s,a|θ) represents the output of policy network 1. V The value represents the output of network 1. Based on this formula, policy network 1 can be updated using gradient descent.

[0181] The update formula for Value Network 2 is: θ V′ ←τθ V +(1-τ)θ V′ Where θ V Let θ be the parameter of value network 1. V′ These are the parameters of value network 2. The parameters of value network 2 can be updated based on the parameters of value network 1 and value network 2.

[0182] The update formula for Policy Network 2 is: θ μ′ ←τθ μ +(1τ)θ μ′ Where θ μ Let θ be the parameter of policy network 1. μ′ These are the parameters of Policy Network 2. The parameters of Policy Network 2 can be updated based on the parameters of Policy Network 1 and Policy Network 2.

[0183] The following is an example of a reward function r(i).

[0184] r(i) = b1r1(i) + b2r2(i) + b3r3(i) + b4r4(i), where the returns for each term in the return function are as follows:

[0185] 1. Error Reporting

[0186] e(i) represents the error at the current time i. g(e) represents a linear function of e(i). ∈ represents the allowable error range. A and -A represent the specific error reward values. If e(i) is less than ∈ this range, it means the error is already very small, no adjustment is needed, and a large positive reward can be given. If the error e(i) exceeds ∈ but is less than ∈ max Then a continuous negative reward can be given. (Exceeding ∈) max That would result in a large negative return.

[0187] 2. Error Change Rate Report

[0188] If the error increases at the next time step i+1, a negative reward is given. B and -C represent the specific error reward values.

[0189] 3. The return on the change in the control quantity r3(i) = Δu

[0190] Changes in the control quantity can also be used as a penalty, giving a negative reward.

[0191] 4. Simulation metric return r4(i) = ∑r 4-j (i), where j ranges from 1 to 4, and the reward function represents the sum of the following terms: collision reward r 4-1 (i) Sudden braking response r 4-2 (i) A sharp turn in the direction will result in a return r 4-3 (i) and the reward r for trajectory replanning 4-4 (i). These are just examples; other rewards may also be included.

[0192] Traditional solutions typically employ linear interpolation, simply dividing the scene into different intervals for velocity, curvature, and acceleration, followed by manual adjustment of parameters based on experience. However, this approach requires adjusting numerous parameters, resulting in low adaptation efficiency. This solution, based on reinforcement learning, eliminates the need for manual parameter tuning and automatically finds the optimal control parameters, significantly improving adaptation efficiency. Furthermore, the predictive temporal weight decay strategy and refined scene segmentation method further optimize the control performance.

[0193] This solution is based on reinforcement learning and can automatically find the optimal control parameters without manual parameter tuning, which greatly improves the efficiency of adaptation. The prediction time-domain weight decay strategy and the refined scene segmentation method in the solution further optimize the control effect.

[0194] It should be noted that the division of functional units in this embodiment is illustrative and represents only one logical functional division. In actual implementation, other division methods may be used. The functional units in this embodiment can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0195] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods provided in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.

[0196] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0197] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0198] According to embodiments of this disclosure, this disclosure also provides an autonomous driving vehicle, which may include electronic devices for training a method for an autonomous driving control parameter model or for acquiring control parameters for implementing embodiments of this disclosure.

[0199] Figure 10 A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0200] like Figure 10As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1002 or a computer program loaded from storage unit 1008 into random access memory (RAM) 1003. The RAM 1003 may also store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.

[0201] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0202] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as methods for training autonomous driving control parameter models or methods for acquiring control parameters. For example, in some embodiments, the methods for training autonomous driving control parameter models or methods for acquiring control parameters can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by the computing unit 1001, one or more steps of the methods for training autonomous driving control parameter models or methods for acquiring control parameters described above can be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured by any other suitable means (e.g., by means of firmware) to perform a method for training an autonomous driving control parameter model or a method for acquiring control parameters.

[0203] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0204] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0205] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0206] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0207] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0208] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0209] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0210] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for training an autonomous driving control parameter model, comprising: The process of inputting autonomous driving scenario information into a first control parameter model to obtain control parameters output by the first control parameter model includes: inputting the first scenario information into a first policy network of the first control parameter model for processing to obtain the first control parameters output by the first policy network; wherein, the policy network is a reinforcement learning policy network. The process of obtaining training samples by interacting with the control parameters and the dynamic simulation environment includes: attenuating the Q value in the first control parameter according to the attenuation coefficient, and substituting the attenuated Q value into the objective function of the control model to calculate the control quantity, wherein the Q value represents the error term penalty weight in the objective function of the model predictive control (MPC) algorithm; the dynamic simulation environment is a dynamic simulation environment for an autonomous vehicle; interacting with the dynamic simulation environment using the control quantity to obtain the simulation interaction result output by the dynamic simulation environment, wherein the simulation interaction result includes second scene information; calculating a reward function value based on the first scene information and the simulation interaction result; the reward function value is determined based on the error reward value, the error rate of change reward value, the control quantity change reward value, and the simulation metric reward value; and obtaining sample data based on the first scene information, the first control parameter, the reward function value, and the second scene information; and The first control parameter model is updated based on the training samples to obtain the trained second control parameter model.

2. The method according to claim 1, wherein, The autonomous driving scenario information includes at least one of the following: Velocity, curvature, lateral position error, lateral heading angle error, longitudinal position error, longitudinal velocity error, longitudinal acceleration error, desired velocity, desired acceleration, desired lateral displacement, and desired heading angle.

3. The method according to claim 1 or 2, wherein, The control parameters include at least one of the following: The Q value for lateral advance, the R1 value for lateral advance, the R2 value for lateral advance, the lateral attenuation value, the Q value for longitudinal advance, the R1 value for longitudinal advance, the R2 value for longitudinal advance, and the longitudinal attenuation value. Wherein, the Q value represents the penalty weight of the error term in the objective function of the Model Predictive Control (MPC) algorithm, the R1 value represents the penalty weight of the control quantity in the objective function, and the R2 value represents the penalty weight of the control quantity increment in the objective function.

4. The method according to claim 1, wherein, Based on the first scenario information, the first control parameters, the reward function value, and the second scenario information, sample data is obtained, including: Based on the first scenario information, the first control parameters, the reward function value, and the second scenario information, the sample data for this test is generated; and The second scene information of each generated sample data is used as the first scene information for the next input of the first control parameter model. The steps of obtaining sample data are iteratively executed to obtain multiple sample data.

5. The method according to claim 4, wherein, The method of interacting with the dynamic simulation environment based on the control parameters to obtain training samples also includes: Add multiple sets of the aforementioned sample data to the playback pool; and Multiple sample data in the playback pool are sampled to obtain training samples.

6. The method according to claim 1 or 2, wherein, The training samples include first scene information, first control parameters, reward function values, and second scene information; the first control parameter model includes a first policy network, a first value network, a second policy network, and a second value network. The first control parameter model is updated based on the training samples to obtain the trained second control parameter model, including: The first scene information is input into the first policy network and the second policy network respectively for processing to obtain the first control parameter output by the first policy network and the second control parameter output by the second policy network. The first scene information and the first control parameters are input into the first value network for processing to obtain the first value function output by the first value network. The first scene information and the second control parameters are input into the second value network for processing to obtain the second value function output by the second value network; and Based on the training samples, the first value function, and the second value function, the first control parameter model is updated to obtain the trained second control parameter model.

7. The method according to claim 6, wherein, Based on the training samples, the first value function, and the second value function, the first control parameter model is updated to obtain the trained second control parameter model, including: The parameters of the first value network are updated using a loss function, which is determined based on the first scene information, the first control parameters, the second scene information, the second control parameters, the first value function, and the second value function. The parameters of the first policy network are updated using gradient descent, and the information required by the gradient descent method includes at least the first scene information, the first control parameters, and the first value function. Update the parameters of the second policy network based on the parameters of the first policy network and the parameters of the second policy network; and Update the parameters of the second value network based on the parameters of the first value network and the parameters of the second value network.

8. A method for obtaining control parameters, comprising: The target scene information is input into the autonomous driving control parameter model for processing to obtain the target control parameters output by the autonomous driving control parameter model. The autonomous driving control parameter model is a second control parameter model trained using the training method described in any one of claims 1 to 7.

9. A training device for an autonomous driving control parameter model, comprising: The input module is used to input autonomous driving scenario information into the first control parameter model and obtain the control parameters output by the first control parameter model; wherein, the policy network is a reinforcement learning policy network; The acquisition module is used to interact with the dynamics simulation environment based on the control parameters to acquire training samples; wherein, the dynamics simulation environment is the dynamics simulation environment of an autonomous vehicle; and The training module is used to update the first control parameter model based on the training samples to obtain the trained second control parameter model. The input module includes an input submodule, which is used to input the first scene information into the first policy network of the first control parameter model for processing, and obtain the first control parameters output by the first policy network. The acquisition module includes: The processing submodule is used to attenuate the Q value in the first control parameter according to the attenuation coefficient, and to substitute the attenuated Q value into the objective function of the control model to calculate the control quantity, wherein the Q value represents the error term penalty weight in the objective function of the model predictive control (MPC) algorithm. An interaction submodule is used to interact with the dynamic simulation environment using the control quantity to obtain the simulation interaction result output by the dynamic simulation environment, wherein the simulation interaction result includes second scene information; A calculation submodule is used to calculate a reward function value based on the first scene information and the simulation interaction result; wherein the reward function value is determined based on the error reward value, the error change rate reward value, the control quantity change reward value, and the simulation metric reward value; and The acquisition submodule is used to acquire sample data based on the first scene information, the first control parameters, the reward function value, and the second scene information.

10. The apparatus according to claim 9, wherein, The autonomous driving scenario information includes at least one of the following: Velocity, curvature, lateral position error, lateral heading angle error, longitudinal position error, longitudinal velocity error, longitudinal acceleration error, desired velocity, desired acceleration, desired lateral displacement, and desired heading angle.

11. The apparatus according to claim 9 or 10, wherein, The control parameters include at least one of the following: The Q value for lateral advance, the R1 value for lateral advance, the R2 value for lateral advance, the lateral attenuation value, the Q value for longitudinal advance, the R1 value for longitudinal advance, the R2 value for longitudinal advance, and the longitudinal attenuation value. Wherein, the Q value represents the penalty weight of the error term in the objective function of the Model Predictive Control (MPC) algorithm, the R1 value represents the penalty weight of the control quantity in the objective function, and the R2 value represents the penalty weight of the control quantity increment in the objective function.

12. The apparatus according to claim 9, wherein, The acquisition submodule is also used for: Based on the first scenario information, the first control parameters, the reward function value, and the second scenario information, the sample data for this test is generated; and The second scene information of each generated sample data is used as the first scene information for the next input of the first control parameter model. The steps of obtaining sample data are iteratively executed to obtain multiple sample data.

13. The apparatus according to claim 12, wherein, The acquisition module further includes: The playback submodule is used to add multiple sample data sets to the playback pool; and The sampling submodule is used to sample multiple sample data in the playback pool to obtain training samples.

14. The apparatus according to claim 9 or 10, wherein, The training samples include first scene information, first control parameters, reward function values, and second scene information. The first control parameter model includes a first policy network, a first value network, a second policy network, and a second value network. The training module includes: The strategy network submodule is used to input the first scenario information into the first strategy network and the second strategy network respectively for processing, and obtain the first control parameter output by the first strategy network and the second control parameter output by the second strategy network. The value network submodule is used to input the first scene information and the first control parameters into the first value network for processing, to obtain a first value function output by the first value network; and to input the first scene information and the second control parameters into the second value network for processing, to obtain a second value function output by the second value network; and The update submodule is used to update the first control parameter model based on the training samples, the first value function, and the second value function to obtain the trained second control parameter model.

15. The apparatus according to claim 14, wherein, The update submodule is used for: The parameters of the first value network are updated using a loss function, which is determined based on the first scene information, the first control parameters, the second scene information, the second control parameters, the first value function, and the second value function. The parameters of the first policy network are updated using gradient descent, and the information required by the gradient descent method includes at least the first scene information, the first control parameters, and the first value function. Update the parameters of the second policy network based on the parameters of the first policy network and the parameters of the second policy network; and Update the parameters of the second value network based on the parameters of the first value network and the parameters of the second value network.

16. A device for acquiring control parameters, comprising: The acquisition module is used to input target scene information into the autonomous driving control parameter model for processing, and obtain the target control parameters output by the autonomous driving control parameter model. The autonomous driving control parameter model is a second control parameter model obtained by training using the training device described in any one of claims 10 to 15.

17. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-8.

19. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-8.

20. An autonomous vehicle, including the electronic equipment as claimed in claim 17.

Citation Information

Patent Citations

  • Automatic driving training method and system based on combination of imitation learning and reinforcement learning

    CN114282433A