Vehicle control method, vehicle control device, electronic device, and readable storage medium
By using the target neural network to optimize the nonlinear control quantity in autonomous driving, the problem of inefficient trajectory planning is solved, fast and effective vehicle trajectory tracking control is achieved, and the real-time and accuracy of autonomous driving are improved.
Patent Information
- Application Number
- CN202410688714.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-29
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2044-05-29
AI Technical Summary
In autonomous driving, the complex form of nonlinear functions leads to inefficient trajectory planning, difficult optimization process and long calculation time, making it difficult to improve efficiency while ensuring the accuracy of trajectory planning.
By obtaining the nonlinear control quantity and inputting it into the target neural network based on the target error model, the target function in the form of a norm is used for optimization to obtain the control increment for vehicle motion control, and the target neural network is used to quickly solve the nonlinear problem.
It achieves the rapid and effective solution of nonlinear problems under the requirements of real-time and accuracy, and improves the dynamic performance of vehicle trajectory tracking automatic control.
Smart Images

Figure CN118722714B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of computer technology, and specifically relates to a vehicle control method, a vehicle control device, an electronic device, and a readable storage medium. Background Art
[0002] In autonomous driving technologies, the complex nature of nonlinear functions can lead to multiple local optimal solutions and even vanishing or exploding gradients, making the optimization process extremely difficult. Furthermore, even if a usable solution is achieved, it can take a long time to optimize. Therefore, improving trajectory planning efficiency and reducing computation time while maintaining accuracy have become pressing challenges for those skilled in the art. Summary of the Invention
[0003] The embodiments of the present application provide a vehicle control method, a vehicle control device, an electronic device, and a readable storage medium, which can solve the problem of low vehicle control efficiency in related technologies.
[0004] In a first aspect, an embodiment of the present application provides a vehicle control method, the method comprising: obtaining a nonlinear control quantity corresponding to a target vehicle at a current moment, wherein the nonlinear control quantity comprises an actual state quantity, an expected state quantity, an actual control input quantity, and an expected control input quantity of the target vehicle; inputting the nonlinear control quantity into a target neural network to obtain a control increment, wherein the target neural network is an equivalent representation of an objective function in a norm form constructed based on a target error model, the target error model is determined based on a target dynamics model, and the control increment is the first term of a target control sequence that minimizes the objective function; at the next moment, performing motion control on the target vehicle according to the control increment.
[0005] In a second aspect, an embodiment of the present application provides a vehicle control device, which includes: an acquisition module for acquiring a nonlinear control quantity corresponding to a target vehicle at a current moment, wherein the nonlinear control quantity includes an actual state quantity, an expected state quantity, an actual control input quantity, and an expected control input quantity of the target vehicle; a processing module for inputting the nonlinear control quantity into a target neural network to obtain a control increment, wherein the target neural network is an equivalent representation of an objective function in a norm form constructed based on a target error model, the target error model is determined based on a target dynamics model, and the control increment is the first term of a target control sequence that minimizes the objective function; a control module for performing motion control on the target vehicle at the next moment according to the control increment.
[0006] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the method described in the first aspect.
[0007] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0008] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the steps of the method described in the first aspect.
[0009] In an embodiment of the present application, a nonlinear control quantity corresponding to a target vehicle at a current moment is obtained, wherein the nonlinear control quantity includes an actual state quantity, an expected state quantity, an actual control input quantity, and an expected control input quantity of the target vehicle; the nonlinear control quantity is input into a target neural network to obtain a control increment, wherein the target neural network is an equivalent representation of an objective function in a one-norm form constructed based on a target error model, the target error model is determined based on a target dynamics model, and the control increment is the first term of a target control sequence that minimizes the objective function; at the next moment, the target vehicle is motion-controlled according to the control increment, and nonlinear problems, especially one-norm nonlinear optimization problems, can be solved quickly and effectively, thereby meeting the real-time and accuracy requirements for controlling the target vehicle, and effectively improving the dynamic performance of vehicle trajectory tracking automatic control. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 This is a flow chart of a vehicle control method provided in an embodiment of the present application;
[0011] Figure 2a is a schematic structural diagram of a target neural network provided in an embodiment of the present application;
[0012] Figure 2b is a schematic structural diagram of another target neural network provided in an embodiment of the present application;
[0013] Figure 3a This is a schematic diagram of the structure of another target neural network provided in an embodiment of the present application;
[0014] Figure 3b is a schematic diagram of the structure of another target neural network provided in an embodiment of the present application;
[0015] Figure 4aThis is a schematic diagram of a network expression of a function provided in an embodiment of the present application;
[0016] Figure 4b This is a schematic diagram of a network expression of a function provided in an embodiment of the present application;
[0017] Figure 4c This is a schematic diagram of a network expression of a function provided in an embodiment of the present application;
[0018] Figure 4d This is a schematic diagram of a network expression of a function provided in an embodiment of the present application;
[0019] Figure 5a This is a schematic diagram of a network expression of an input incremental penalty provided in an embodiment of the present application;
[0020] Figure 5b This is a schematic diagram of another network expression of input incremental penalty provided in an embodiment of the present application;
[0021] Figure 6a This is a schematic diagram of a network expression of a state penalty provided in an embodiment of the present application;
[0022] Figure 6b This is a schematic diagram of a network expression of another state penalty term provided in an embodiment of the present application;
[0023] Figure 7a This is a schematic diagram of a network expression of an input target penalty provided in an embodiment of the present application;
[0024] Figure 7b This is another network expression diagram of input target penalty provided in an embodiment of the present application;
[0025] Figure 8a This is a schematic diagram of a network expression of an input quantity restriction constraint provided in an embodiment of the present application;
[0026] Figure 8b This is another network expression diagram of input quantity restriction provided by an embodiment of the present application;
[0027] Figure 8c This is a schematic diagram of a network expression of another restriction function provided in an embodiment of the present application;
[0028] Figure 9 is a flow chart of another vehicle control method provided in an embodiment of the present application;
[0029] Figure 10 is a structural diagram of a vehicle control device provided in an embodiment of the present application;
[0030] Figure 11 Schematic diagram of the structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0031] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0032] In order to better understand the technical solution provided by this application, the mathematical symbols involved in this application are first introduced.
[0033] (1) Represented as an m-dimensional vector.
[0034] (2) For vector |a| represents the sum of the absolute values of the included elements |a|=|a1|+|a2|+|a3|, |a| b =b1|a1|+b2|a2|+b3|a3|.
[0035] (3) Activation functions can include: linear function (purelin), radial basis function (radbas), symmetric sigmoid function (tansig), positive linear function (poslin), and their expressions are:
[0036] purelin(x)=x;
[0037] radbas(x)=exp(-x 2 );
[0038]
[0039] Below, in conjunction with the accompanying drawings, a vehicle control method, a vehicle control device, an electronic device and a readable storage medium provided by the embodiments of the present application are described in detail through specific embodiments and their application scenarios.
[0040] Figure 1 An embodiment of the present application provides a vehicle control method, which can be executed by an electronic device, which may include a server and / or a terminal device. In other words, the method can be executed by software or hardware installed in the electronic device, and the method includes the following steps:
[0041] S110: Obtain the nonlinear control variable corresponding to the target vehicle at the current moment.
[0042] The nonlinear control variable includes the actual state variable, expected state variable, actual control input variable, and expected control input variable of the target vehicle. The actual state variable is the actual state parameter of the target vehicle at the current moment, the expected state variable is the expected state parameter preset at the next moment, the actual control input variable is the motion parameter applied to the target vehicle based on the control increment determined at the previous moment, and the expected control input variable is the expected state parameter preset at the current moment. The motion parameter may include the angle control variable and speed control variable at the previous moment.
[0043] S120: Inputting the nonlinear control variable into the target neural network to obtain a control increment.
[0044] Among them, the target neural network is an equivalent representation of an objective function in a norm form constructed based on a target error model, the target error model is determined based on a target dynamics model, and the control increment is the first item of the target control sequence that minimizes the objective function.
[0045] Among them, the L1 norm is also called the L1 norm or the Manhattan norm, which refers to the sum of the absolute values of all elements in the vector. In this application, the objective function in the form of the L1 norm can improve the vehicle control performance.
[0046] The control increment represents the change in the control input from one moment to the next, assuming that the vehicle motion can be described by a nonlinear function expressed as follows:
[0047] x k+1 =f(x k ,u k )#
[0048] in, is the system state at time k, is the system input at time k.
[0049] In model predictive control (MPC), the control input is typically solved via an optimization problem that considers a sequence of control increments Δu, rather than directly optimizing the control input u (the actual control input quantity). For example, at time k, the optimization problem might solve for a sequence of control increments of the form:
[0050] Δu(k|k), Δu(k+1|k), Δu(k+2|k),…
[0051] Then, the control input sequence is obtained by accumulating these control increments:
[0052] u(k|k)=u(k-1)+Δu(k|k);
[0053] u(k+1|k)=u(k|k)+Δu(k+1|k);
[0054] u(k+2|k)=u(k+1|k)+Δu(k+2|k); ......
[0056] It can be understood that due to the expressive power of neural networks and their good optimization strategies, 1-norm nonlinear optimization problems can be solved quickly. At the same time, the neural network optimization algorithm has good convergence. The 1-norm objective function is optimized based on the target neural network. Such a 1-norm objective function is also convergent, so the final control increment is the first item of the target control sequence that minimizes the objective function.
[0057] Exemplarily, the equivalent representation is as follows Figure 2a and 2b As shown, Figure 2b for Figure 2a The target neural network can be trained based on the input and output data of the nonlinear model or constructed based on the target error model.
[0058] In this step, the target dynamics model is represented by constructing a corresponding neural network, while the target error model is obtained by fitting the target dynamics model using input and output data. Since the target dynamics model describes the main laws of vehicle motion, expressing it using a neural network through construction can avoid safety issues. Furthermore, since the target error model is the compensation component of the entire dynamics model, it is used to improve prediction accuracy and, therefore, control performance. However, since the analytical expression of this target error model is relatively complex, expressing it through a neural network in a constructed form would significantly reduce optimization speed. Therefore, fitting the target dynamics model based on input and output data allows for a more compact expression, thereby improving output efficiency and the accuracy of control increments.
[0059] S130: At the next moment, the target vehicle is motion controlled according to the control increment.
[0060] Exemplarily, at time k, the target control sequence is minimized in the prediction time domain based on the nonlinear control amount at the current time. At time k+1, the target vehicle is controlled based on the first item in the target control sequence, i.e., the control increment, and the target vehicle will produce a new state. In addition, when reaching time (k+1), when the time domain conditions are met, the above steps S110-S130 are repeated. Therefore, the vehicle control process is a process that repeats the above steps at each moment, so that the dynamic changes and external disturbances of the target vehicle can be grasped in real time through the target neural network, thereby maintaining effective control of the target vehicle.
[0061] Furthermore, in one implementation, the time domain condition includes one of the following:
[0062] (1) The control time domain is not equal to the prediction time domain.
[0063] (2) Both the control time domain and the prediction time domain are 1.
[0064] It is understandable that when N c =N p =1, that is, when the control time domain is equal to the prediction time domain and is 1, the equivalent representation of the objective function is the construction of the target neural network as follows: Figure 3a As shown, the neurons with black and gray backgrounds belong to the input layer and output layer respectively. The input data are the state x0 at time 1 and k, the input u0 at time k-1, and the expected state x at time k+1. ref (k+1|k); and all output data are set to 0, which means minimizing the objective function. The activation function of all neurons in the figure is a linear function (purelin) (excluding the neurons in the rectangular module). The weights of the dotted connections are undetermined, while the weights of the solid connections are all 1. When controlling the vehicle through the target neural network, only the weights corresponding to the dotted connections are optimized. This weight corresponds to the optimization variable, that is, the control increment Δu(k|k) in the objective function. Therefore, the weight of the dotted connection obtained by the output of the target neural network is the optimized solution of the objective function (when N c =N p =1).
[0065] Furthermore, when the difference between the prediction time domain and the control time domain is large or N c ≠N p , the equivalent representation of the objective function is the construction of the target neural network as follows Figure 3b As shown, similarly, when controlling the vehicle through the target neural network, only the weights corresponding to the dotted connections are optimized, and the result obtained is the solution of the objective function, that is, the control increment sequence.
[0066] In the embodiment of the present application, the nonlinear control quantity corresponding to the target vehicle at the current time is obtained, wherein the nonlinear control quantity comprises actual state quantity, expected state quantity, actual control input quantity and expected control input quantity of the target vehicle; the nonlinear control quantity is input into a target neural network to obtain a control increment, wherein the target neural network is an equivalent representation of a target function in a norm form based on a target error model, the target error model is determined based on a target dynamics model, and the control increment is the first term of a target control sequence for minimizing the target function; at the next time, the target vehicle is controlled according to the control increment, which can quickly and effectively solve nonlinear problems, especially 1-norm nonlinear optimization problems, thereby meeting the real-time and accuracy requirements of controlling the target vehicle and effectively improving the dynamic performance of vehicle trajectory tracking automatic control.
[0067] In an implementation manner, before the nonlinear control quantity is input into the target neural network to obtain the control increment, the method further comprises: constructing the target error model; wherein the constructing the target error model comprises: obtaining an initial dynamics model; obtaining a target dynamics model by discretizing the initial dynamics model based on the association between the position of the target vehicle and the geodetic coordinate system; and obtaining the target error model by fitting the target dynamics model based on first training data output by the target dynamics model.
[0068] Exemplarily, it is assumed that the longitudinal velocity is constant, i.e. and it is assumed that the longitudinal force F xf experienced by the vehicle is 0, the vehicle dynamics model can be obtained as:
[0069]
[0070] wherein, δ is the front wheel steering angle, v x and v y are the longitudinal velocity and lateral velocity at the center of mass of the vehicle, r is the yaw rate of the vehicle, l f and l r are the distances from the center of mass of the vehicle to the front wheel and the rear wheel, m and I z are the mass of the vehicle and the moment of inertia of the vehicle around the z axis respectively, c f and c r are the cornering stiffness of the front wheel and the rear wheel respectively.
[0071] When tracking the expected path, since the reference trajectory of the controller is usually established on the geodetic coordinate system, the relationship between the vehicle position and the geodetic coordinate system needs to be established:
[0072]
[0073] Where L is the vehicle wheelbase, is the heading angle of the vehicle.
[0074] Discretizing the above equation into small time intervals, we get:
[0075]
[0076] Where T is the time interval.
[0077] Similarly, the discretized initial dynamic model is as follows:
[0078]
[0079]
[0080] It can be seen that the target dynamics model is based on the vehicle kinematic model with compensation terms added.
[0081] Furthermore, an error model is trained based on the target dynamics model to obtain a target error model. In one implementation, the target error model is obtained by fitting the target dynamics model based on the first training data output by the target dynamics model, including:
[0082] Acquire multiple sets of second training data according to a preset control range, wherein each set of the second training data includes a vehicle longitudinal speed, a vehicle lateral speed, a vehicle yaw rate, a vehicle steering angle, a horizontal coordinate of a vehicle position, a vertical coordinate of a vehicle position, and a vehicle heading angle at a first moment;
[0083] Inputting multiple sets of the second training data into the target dynamics model at preset time intervals to obtain multiple sets of the first training data, wherein the first training data includes a vehicle longitudinal speed, a vehicle lateral speed, a vehicle yaw rate, a vehicle steering angle, a horizontal coordinate of a vehicle position, a vertical coordinate of a vehicle position, and a vehicle heading angle at a second moment, where the second moment is a moment subsequent to the first moment;
[0084] The target dynamics model is fitted based on the second training data and the first training data to obtain the target error model.
[0085] It is understandable that training a lateral speed Yaw angular velocity r k and the front wheel steering angle δ k As input, the new yaw rate r k+1 and lateral velocity The target dynamic model is output. First, use the above two equations to collect input and output data with a certain accuracy and control range. For example, the range of longitudinal velocity of data collection is set to lateral speed Yaw angular velocity Front wheel angle The data interval is 0.01, and then all combinations are input into the above two equations to obtain the corresponding outputs, and multiple sets of data are obtained. These data are then used to train the target dynamics model to fit the two formulas and obtain the target error model, such as Figure 4a As shown. Among them, the function
[0086] The neural network expression is as follows Figure 4b As shown; function The neural network expression is as follows Figure 4c As shown; function The neural network expression is as follows Figure 4d shown.
[0087] In the neural network expressions of the above functions, the neurons with black and gray backgrounds belong to the input layer and output layer respectively. The input data are the longitudinal speed of the vehicle at time k Vehicle lateral speed Vehicle yaw rate r k , vehicle steering angle δ k , X coordinate of vehicle position X k , Y coordinate of the vehicle position k , vehicle heading angle The output data are the vehicle lateral speed at time k+1 Vehicle yaw rate r k+1 , X coordinate of vehicle position X k+1 , Y coordinate of the vehicle position k+1 and vehicle heading angle The vehicle longitudinal speed and the vehicle steering angle δ k The input variable of the model is the X coordinate of the vehicle position X k+1 , Y coordinate of the vehicle position k+1 and vehicle heading angle is the output variable of the model, and the others are state variables. The activation functions of the neurons in the figure are sine function (sin(x)), cosine function (cos(x)), tangent function (tan(x)), inverse tangent function (tan -1 (x)) and the inverse proportional function The corresponding neurons are marked, and the activation functions of other neurons are all linear functions (purelin) (excluding the neurons in the rectangular module). The weights of the dotted connections are marked at the corresponding positions in the figure, and the weights of the solid line connections are all 1.
[0088] In one implementation, before inputting the nonlinear control variable into the target neural network to obtain the control increment, the method further includes the following steps:
[0089] Step 1: Based on model predictive control, construct the initial function according to the first constraint.
[0090] In one implementation, the first constraint includes an input increment penalty, an input target penalty, and a state penalty, wherein the input increment penalty represents a weight corresponding to the control increment, the state penalty represents a weight corresponding to a difference between the actual state quantity and the expected state quantity, and the input target penalty represents a weight corresponding to a difference between the actual control input quantity and the expected control input quantity.
[0091] Among them, the network expression of input incremental penalty can be referred to Figure 5a and 5b As shown, Figure 5b for Figure 5a A simplified representation of the network. The black and gray neurons belong to the input and output layers, respectively. The activation function of the neurons is a linear function (purelin). The weights connected by dashed lines are the weights r(i) of the corresponding input increments, and the weights connected by solid lines are all 1. The network takes the input increment as input, multiplies it by the weight r(i), and then outputs it. This network corresponds to the objective function |Δu(k+i-1|k)| r(i) .
[0092] The network expression of state penalty can be referred to Figure 6a and 6b As shown, Figure 6b for Figure 6a The black and gray neurons belong to the input layer and output layer respectively, and the activation function of the neurons is a linear function (purelin); the weights of the dotted lines are -1 and q(i), and the weights of the solid lines are all 1. The network is in state x and the desired state x ref The input is then subtracted, and the difference is multiplied by the weight q(i) and output. This network corresponds to the objective function |x(k+i|k)-x ref (k+i|k)| q(i) .
[0093] The network expression of input target penalty can be referred to Figure 7a and 7b As shown, Figure 7b for Figure 7aThe black and gray neurons belong to the input layer and output layer respectively, and the activation function of the neurons is a linear function (purelin); the weights of the dotted lines are -1 and p(i), and the weights of the solid lines are all 1. The network is based on the input u and the expected input u ref The input is then subtracted, and the difference is multiplied by the weight p(i) and output. This network corresponds to the objective function |u(k+i-1|k)-u ref (k+i-1|k)| p(i) .
[0094] It is understandable that assuming the vehicle motion can be expressed as x k+1 =f(x k ,u k ) description, in order to make the actual state sequence of the vehicle motion as close as possible to the given reference state sequence, at time k, the control increment sequence can be obtained by optimizing the following initial function containing the first constraint condition:
[0095]
[0096] st
[0097] x k+1 =f(x k ,u k )
[0098] x k =x0,u k-1 =u0
[0099]
[0100] Where Δu is the control increment; u(k+i-1|k)=u(k+i-2|k)+Δu(k+i-1|k), i=1,2,…,N c ;u(k+i-1|k)=u(k+i-2|k), i=N c +1,N c +2,…,N p ; N p and N c are respectively the prediction time domain and the control time domain, and N p ≥N c ;x ref (k+i|k) is the expected state of the vehicle at time k+i, i.e., the expected state quantity in the nonlinear control quantity; u ref (k+i|k) is the expected input at time k+i, i.e., the expected control input in the nonlinear control quantity; x0 is the state at time k, i.e., the actual state in the nonlinear control quantity; u0 is the input at time k-1, i.e., the actual control input; p(i), q(i), and r(i) are u respectively.k 、x k and Δu k The weight vector of u k 、x k and Δu k The constraint set.
[0101] Step 2: For the initial function, construct the objective function according to the second constraint condition.
[0102] In one implementation, the second constraint condition includes an input quantity limitation constraint and an input increment limitation constraint, the input increment limitation constraint is a constraint on the control increment, and the input quantity limitation constraint is a constraint on the actual control input quantity.
[0103] Optionally, the input limit constraint can be used Figure 8a and 8b The network shown is expressed as follows, Figure 8b for Figure 8a A simplified representation of the network. The black and gray neurons belong to the input and output layers, respectively. The activation function of the neurons below the dashed line is the corresponding input restriction function, or the first input restriction function, restriction1(u). The activation functions of the other neurons can be linear functions (purelin). The weight of the dotted connection is M1, and the weight of the solid connection is 1. The network takes the state x as input, then the restriction function outputs the corresponding value, which is finally multiplied by the weight M1 and output. This network corresponds to M2*restriction(u(k+i-1|k)) in the objective function. The neural network representation of the input increment restriction function is the same as this and will not be repeated here.
[0104] Furthermore, based on the initial function in S112, the non-equality constraint is converted into a soft constraint using a penalty function through the second constraint condition, and the objective function is obtained:
[0105]
[0106] st
[0107] x k+1 =f(x k ,u k )
[0108] x k =x0,u k-1 =u0
[0109] in,
[0110]
[0111] a>0,b>0,c>0,g(u)>0,i(Δu k )>0;
[0112] restriction1(u) and restriction2(Δu k ) is a penalty function, also called a restriction function, M1 and M2 are weight constants and are large enough.
[0113] The following describes the specific forms of the first and second restriction functions and their neural network expressions, taking into account the input increment constraint and input amount constraint:
[0114]
[0115] c≤Δδ i ≤d,i=1,2,…N c
[0116]
[0117] C≤δ i ≤D,i=1,2,…N p
[0118] The above constraints are formally the same, so we will use the longitudinal velocity input constraint as an example to illustrate the constraint and convert it into the following penalty function:
[0119]
[0120] Among them, when the variable satisfies the constraint When , the function output is 0, otherwise the function will output a positive value as a penalty, and the farther the variable is from the constraint range, the greater the value of the function output will be. The neural network expression corresponding to this function is as follows Figure 8c As shown in the figure, the neurons with black and gray backgrounds belong to the input layer and output layer respectively. The circles represent neurons, and the hexagons express the bias values of the corresponding neurons. The activation functions of neurons 1 and 3 are linear functions (purelin), while the activation function of neuron 2 is a positive linear function (poslin); the weights of the dotted connections are marked at the dotted line positions, M is a sufficiently large constant, and the weights of the solid line connections are all 1.
[0121] In one implementation, the control increment includes a longitudinal velocity increment and a steering angle increment, the actual control input includes an actual longitudinal velocity input and an actual steering angle input, the input constraint includes a dynamic constraint on the longitudinal velocity increment and a dynamic constraint on the steering angle increment, and the input constraint includes a dynamic constraint on the actual longitudinal velocity input and a dynamic constraint on the actual steering angle input. In another implementation, the dynamic constraint on the longitudinal velocity increment is: Δa < Δv x (k+i-1|k) <b;转向角度增量的动态约束为:c<Δδ(k+i-1|k)<d;纵向速度实际输入量的动态约束为:A<v x (k+j-1|k) <B;转向角度实际输入量的动态约束为:C<δ(k+j-1|k)<D。其中,Δv x (k+i-1|k) represents the longitudinal velocity increment, a represents the minimum value of the longitudinal velocity increment, and b represents the maximum value of the longitudinal velocity increment; Δδ(k+i-1|k) represents the steering angle increment, c represents the minimum value of the steering angle increment, and d represents the maximum value of the steering angle increment; v x (k+j-1|k) represents the actual input of longitudinal velocity, A represents the minimum value of the actual input of longitudinal velocity, and B represents the maximum value of the actual input of longitudinal velocity; δ(k+j-1|k) represents the actual input of steering angle, C represents the minimum value of the actual input of steering angle, and D represents the maximum value of the actual input of steering angle, k represents the current time, i and j represent the time step index, i = 1, 2, ... N c , j=1,2,…N p , N c To control the time domain, N p For the prediction time domain.
[0122] Based on this implementation, the solution of the objective function can be determined as:
[0123]
[0124] a<Δv x (k+i-1|k) <b,i=1,2,…N c
[0125] c<Δδ(k+i-1|k) <d,i=1,2,…N c
[0126] A <v x (k+i-1|k) <B,i=1,2,…N p
[0127] C<δ(k+i-1|k) <D,i=1,2,…N p
[0128] Where Δv x and Δδ are the longitudinal velocity increment and steering angle increment of the target vehicle, respectively, and are also optimization variables. In addition:
[0129] v x (k+i-1|k)=v x (k+i-2|k)+Δv x (k+i-1|k),i=1,2,…,N c ;
[0130] v x (k+i-1|k)=v x (k+i-2|k),i=N c +1,N c +2,…,N p ;
[0131] δ(k+i-1|k)=δ(k+i-2|k)+Δδ(k+i-1|k),i=1,2,…,N c ;
[0132] δ(k+i-1|k)=δ(k+i-2|k),i=N c +1,N c +2,…,N p ;
[0133] Among them, N p and N c are respectively the prediction time domain and the control time domain, and N p ≥N c q, p and r are state penalty, input target penalty and input increment penalty respectively. ref (k+i|k),y ref (k+i|k) and are the expected values of X, Y coordinates and heading angle at time k+i respectively. f is the dynamic model. and are the system state at time k and the input at time k-1 respectively. a, b, c, d, A, B, C, D are parameters related to the input increment and input constraint.
[0134] In the objective function, and Used to make the longitudinal speed and vehicle steering angle close to a given speed and steering; and Used to make the vehicle's driving trajectory and heading angle track the reference trajectory and desired heading as much as possible; and It is used to make the control increment of the target vehicle as small as possible to make the control smoother.
[0135] In one implementation, at the next moment, the target vehicle is motion controlled according to the control increment, including: determining a desired speed based on the longitudinal speed increment, and determining a desired steering amount based on the steering angle increment; determining a turning angle control amount based on the desired steering amount and the measured actual steering amount, and determining a speed control amount based on the desired speed and the measured actual speed; at the next moment, controlling the driving state of the target vehicle based on the turning angle control amount and the speed control amount.
[0136] In this implementation, the nonlinear control quantity corresponding to the target vehicle at the current moment is first obtained, and then the nonlinear control quantity is input into the target neural network to obtain the expected steering amount and expected speed corresponding to the target vehicle. Finally, based on the expected steering amount and the expected speed, high-frequency real-time control of the target vehicle can be achieved, which can effectively improve the dynamic performance of the vehicle trajectory tracking automatic control system.
[0137] The following is a specific example Figure 1 The vehicle control method shown in FIG is specifically described. Figure 9 As shown in the figure, the model predictive controller is responsible for calculating the control sequence based on the neural network constructed by the dynamic model introduced above, and outputting the control quantity at time k (i.e. v x (k|k) and δ(k|k)) are the expected values of the steering controller and speed controller in the figure. At each control moment, the model predictive controller calculates the control quantity (i.e., v) at time k based on the expected path, expected heading sequence, and vehicle state information measured and estimated by the signal filtering and fusion processor, such as position, longitudinal velocity, lateral velocity, roll angular velocity, heading angle, and steering angle. x (k|k) and δ(k|k)). Then the steering controller estimates the steering output of the steering system based on the desired steering δ(k|k) and real-time measurement, and calculates the steering angle control amount to act on the steering system. Similarly, the speed controller is based on the desired speed v x (k|k), the speed output of the speed system is estimated by real-time measurement, and the speed control variable is calculated and applied to the speed system. The above process is then repeated at each control moment. In this embodiment, the use of this cascaded hierarchical structure facilitates the simplification of the model predictive controller design and improves the dynamic performance of the entire system.
[0138] Figure 10A vehicle control device provided in an embodiment of the present application is shown. The vehicle control device 1100 may include: an acquisition module 1010 , an input module 1020 , and a control module 1030 .
[0139] In this embodiment, the acquisition module 1010 is used to obtain the nonlinear control quantity corresponding to the target vehicle at the current moment, wherein the nonlinear control quantity includes the actual state quantity, expected state quantity, actual control input quantity and expected control input quantity of the target vehicle; the processing module 1020 is used to input the nonlinear control quantity into the target neural network to obtain a control increment, wherein the target neural network is an equivalent representation of an objective function in a norm form constructed based on a target error model, the target error model is determined based on a target dynamics model, and the control increment is the first item of the target control sequence that minimizes the objective function; the control module 1030 is used to perform motion control on the target vehicle at the next moment according to the control increment.
[0140] In one implementation, the device also includes a first construction module for constructing the target error model; wherein, constructing the target error model includes: obtaining an initial dynamic model; based on the relationship between the position of the target vehicle and the geodetic coordinate system, obtaining a target dynamic model by discretizing the initial dynamic model; and obtaining the target error model by fitting the target dynamic model according to the first training data output by the target dynamic model.
[0141] In one implementation, obtaining the target error model by fitting the target dynamics model based on the first training data output by the target dynamics model includes:
[0142] According to a preset control range, multiple sets of second training data are obtained, wherein each set of the second training data includes the vehicle longitudinal speed, vehicle lateral speed, vehicle yaw angular velocity, vehicle steering angle, vehicle horizontal coordinate of vehicle position, vehicle vertical coordinate of vehicle position, and vehicle heading angle at a first moment; the multiple sets of the second training data are input into the target dynamics model at preset time intervals to obtain multiple sets of the first training data, wherein the first training data includes the vehicle longitudinal speed, vehicle lateral speed, vehicle yaw angular velocity, vehicle steering angle, vehicle horizontal coordinate of vehicle position, vehicle vertical coordinate of vehicle position, and vehicle heading angle at a second moment, the second moment being the next moment after the first moment; the target dynamics model is fitted based on the second training data and the first training data to obtain the target error model.
[0143] In one implementation, the device also includes a second construction module, which is used to construct an initial function based on model predictive control and the first constraint condition; for the initial function, the objective function is constructed according to the second constraint condition; wherein the first constraint condition includes an input increment penalty, an input target penalty and a state penalty, the input increment penalty represents the weight corresponding to the control increment, the state penalty represents the weight corresponding to the difference between the actual state quantity and the expected state quantity, and the input target penalty represents the weight corresponding to the difference between the actual control input quantity and the expected control input quantity; the second constraint condition includes an input quantity limitation constraint and an input increment limitation constraint, the input increment limitation constraint is a constraint on the control increment, and the input quantity limitation constraint is a constraint on the actual control input quantity.
[0144] In one implementation, the control increment includes a longitudinal speed increment and a steering angle increment, the actual control input includes: an actual longitudinal speed input and an actual steering angle input, the input limit constraint includes a dynamic constraint on the longitudinal speed increment and a dynamic constraint on the steering angle increment, and the input limit constraint includes a dynamic constraint on the actual longitudinal speed input and a dynamic constraint on the actual steering angle input.
[0145] In one implementation, the dynamic constraint of the longitudinal velocity increment is: Δa<Δv x (k+i-1|k) <b;转向角度增量的动态约束为:c<Δδ(k+i-1|k)<d;纵向速度实际输入量的动态约束为:A<v x (k+j-1|k) <B;转向角度实际输入量的动态约束为:C<δ(k+j-1|k)<D;其中,Δv x (k+i-1|k) represents the longitudinal velocity increment, a represents the minimum value of the longitudinal velocity increment, and b represents the maximum value of the longitudinal velocity increment; Δδ(k+i-1|k) represents the steering angle increment, c represents the minimum value of the steering angle increment, and d represents the maximum value of the steering angle increment; v x (k+j-1|k) represents the actual input of longitudinal velocity, A represents the minimum value of the actual input of longitudinal velocity, and B represents the maximum value of the actual input of longitudinal velocity; δ(k+j-1|k) represents the actual input of steering angle, C represents the minimum value of the actual input of steering angle, and D represents the maximum value of the actual input of steering angle, k represents the current time, i and j represent the time step index, i = 1, 2, ... N c , j=1,2,…N p , N c To control the time domain, N p For the prediction time domain.
[0146] In one implementation, at the next moment, the target vehicle is motion controlled according to the control increment, including: determining a desired speed based on the longitudinal speed increment, and determining a desired steering amount based on the steering angle increment; determining a turning angle control amount based on the desired steering amount and the measured actual steering amount, and determining a speed control amount based on the desired speed and the measured actual speed; at the next moment, controlling the driving state of the target vehicle based on the turning angle control amount and the speed control amount.
[0147] The vehicle control device provided in the embodiment of the present application can realize Figure 1 To avoid repetition, the various processes implemented in the illustrated method embodiment will not be described again here.
[0148] The model training device and vehicle control device in the embodiments of the present application can be devices, or components, integrated circuits, or chips in electronic devices. The embodiments of the present application are not specifically limited.
[0149] The model training device and the vehicle control device in the embodiments of the present application may be devices having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiments of the present application.
[0150] Optional, such as Figure 11 As shown, an embodiment of the present application also provides an electronic device 1100, including a processor 1110, a memory 1120, and a program or instruction stored in the memory 1120 and executable on the processor 1110. When the program or instruction is executed by the processor 1110, the various processes of the above-mentioned vehicle control method embodiment are implemented, or the various processes of the above-mentioned vehicle control method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0151] An embodiment of the present application also provides a computer-readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned vehicle control method embodiment are implemented, or the various processes of the above-mentioned vehicle control method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0152] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk.
[0153] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned vehicle control method embodiment, or to implement the various processes of the above-mentioned vehicle control method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0154] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0155] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0156] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0157] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.
Claims
1. A vehicle control method, characterized in that: include: Obtaining a nonlinear control variable corresponding to a target vehicle at a current moment, wherein the nonlinear control variable includes an actual state variable, an expected state variable, an actual control input variable, and an expected control input variable of the target vehicle; Inputting the nonlinear control variable into a target neural network to obtain a control increment, wherein the target neural network is an equivalent representation of a target function in a norm form constructed based on a target error model, the target error model is determined based on a target dynamics model, and the control increment is the first term of a target control sequence that minimizes the target function; At a next moment, performing motion control on the target vehicle according to the control increment; Before inputting the nonlinear control variable into the target neural network to obtain the control increment, the method further includes: Constructing the target error model; Wherein, constructing the target error model includes: Obtaining an initial kinetic model; Based on the relationship between the position of the target vehicle and the earth coordinate system, the target dynamic model is obtained by discretizing the initial dynamic model; Obtaining the target error model by fitting the target dynamics model according to the first training data output by the target dynamics model; The target error model is obtained by fitting the target dynamics model according to the first training data output by the target dynamics model, including: Acquire multiple sets of second training data according to a preset control range, wherein each set of the second training data includes the vehicle longitudinal velocity, the vehicle lateral velocity, the vehicle yaw angular velocity, the vehicle steering angle, the abscissa of the vehicle position, the ordinate of the vehicle position, and the vehicle heading angle at the first moment; Inputting multiple sets of the second training data into the target dynamics model at preset time intervals to obtain multiple sets of the first training data, wherein the first training data includes a vehicle longitudinal velocity, a vehicle lateral velocity, a vehicle yaw rate, a vehicle steering angle, a horizontal coordinate of a vehicle position, a vertical coordinate of a vehicle position, and a vehicle heading angle at a second moment, where the second moment is a moment subsequent to the first moment; The target dynamics model is fitted based on the second training data and the first training data to obtain the target error model.
2. The method according to claim 1, characterized in that Before inputting the nonlinear control variable into the target neural network to obtain the control increment, the method further includes: Based on the model predictive control, an initial function is constructed according to the first constraint condition; For the initial function, construct the objective function according to the second constraint condition; The first constraint condition includes an input increment penalty, an input target penalty, and a state penalty. The input increment penalty represents a weight corresponding to the control increment, the state penalty represents a weight corresponding to a difference between the actual state quantity and the desired state quantity, and the input target penalty represents a weight corresponding to a difference between the actual control input quantity and the desired control input quantity. The second constraint condition includes an input quantity limitation constraint and an input increment limitation constraint. The input increment limitation constraint is a constraint on the control increment, and the input quantity limitation constraint is a constraint on the actual control input quantity.
3. The method according to claim 2, characterized in that The control increment includes a longitudinal speed increment and a steering angle increment, the actual control input includes: an actual longitudinal speed input and an actual steering angle input, the input limit constraint includes a dynamic constraint on the longitudinal speed increment and a dynamic constraint on the steering angle increment, and the input limit constraint includes a dynamic constraint on the actual longitudinal speed input and a dynamic constraint on the actual steering angle input.
4. The method according to claim 3, characterized in that The dynamic constraint of the longitudinal velocity increment is: ; The dynamic constraints on the steering angle increment are: ; The dynamic constraint of the actual input of longitudinal velocity is: ;The dynamic constraint of the actual input of the steering angle is: ; in, represents the longitudinal velocity increment, a represents the minimum value of the longitudinal velocity increment, Indicates the maximum value of the longitudinal velocity increment; represents the steering angle increment, c represents the minimum value of the steering angle increment, Indicates the maximum value of the steering angle increment; Indicates the actual input of longitudinal speed, A indicates the minimum value of the actual input of longitudinal speed, and B indicates the maximum value of the actual input of longitudinal speed; represents the actual input value of the steering angle, C represents the minimum value of the actual input value of the steering angle, D represents the maximum value of the actual input value of the steering angle, k represents the current moment, i and j represent the time step index, ,j , To control the time domain, For the prediction time domain.
5. The method according to claim 3, characterized in that The step of controlling the motion of the target vehicle at the next moment according to the control increment includes: determining a desired speed amount based on the longitudinal speed increment, and determining a desired steering amount based on the steering angle increment; Determining a steering angle control amount based on the desired steering amount and the measured actual steering amount, and determining a speed control amount based on the desired speed amount and the measured actual speed amount; At the next moment, the running state of the target vehicle is controlled based on the turning angle control amount and the speed control amount.
6. A vehicle control device, characterized in that: include: an acquisition module, configured to acquire a nonlinear control variable corresponding to a target vehicle at a current moment, wherein the nonlinear control variable includes an actual state variable, an expected state variable, an actual control input variable, and an expected control input variable of the target vehicle; a processing module, configured to input the nonlinear control variable into a target neural network to obtain a control increment, wherein the target neural network is an equivalent representation of an objective function in a norm form constructed based on a target error model, the target error model is determined based on a target dynamics model, and the control increment is the first term of a target control sequence that minimizes the objective function; A control module, configured to control the motion of the target vehicle according to the control increment at a next moment; The method further includes a first construction module for constructing the target error model; wherein constructing the target error model includes: obtaining an initial dynamic model; obtaining a target dynamic model by discretizing the initial dynamic model based on the relationship between the position of the target vehicle and the earth coordinate system; and obtaining the target error model by fitting the target dynamic model according to first training data output by the target dynamic model; The target error model is obtained by fitting the target dynamics model according to the first training data output by the target dynamics model, including: According to a preset control range, multiple sets of second training data are obtained, wherein each set of the second training data includes the vehicle longitudinal speed, the vehicle lateral speed, the vehicle yaw angular velocity, the vehicle steering angle, the abscissa of the vehicle position, the ordinate of the vehicle position, and the vehicle heading angle at a first moment; the multiple sets of the second training data are input into the target dynamics model at preset time intervals to obtain multiple sets of the first training data, wherein the first training data includes the vehicle longitudinal speed, the vehicle lateral speed, the vehicle yaw angular velocity, the vehicle steering angle, the abscissa of the vehicle position, the ordinate of the vehicle position, and the vehicle heading angle at a second moment, the second moment being the next moment after the first moment; the target dynamics model is fitted based on the second training data and the first training data to obtain the target error model.
7. An electronic device, characterized in that: The method comprises a processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program, and when the program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Intelligent vehicle trajectory tracking model prediction control method based on model compensation
CN107561942A
Intelligent fleet longitudinal following control method based on fuzzy model predictive control
CN112148001A