Model-free self-learning dynamic optimization control method and system for autonomous driving vehicles
Through the model-free self-learning dynamic optimization control method, a dynamic feedback control strategy for autonomous driving is constructed for driving comfort and safety, which solves the problem that it is difficult to achieve comfortable driving by optimizing vehicle acceleration indicators in the existing technology, realizes automatic driving control that does not rely on accurate models, and improves vehicle comfort and safety.
Patent Information
- Application Number
- CN202510144031.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-10
AI Technical Summary
Existing methods that optimize vehicle acceleration indicators cannot achieve comfortable driving well, and the model-based autonomous driving control method is highly dependent on vehicle dynamic models.
The dynamic optimization control method of model-free self-learning is adopted to construct an augmented vehicle dynamic model and establish a tracking error system to construct a dynamic feedback control strategy for driving comfort and safety, and train the strategy through reinforcement learning.
It realizes autonomous driving control that does not rely on precise vehicle dynamics models, improves the comfort and safety of autonomous driving vehicles, adapts to various complex road environments and traffic conditions, and improves the robustness and adaptability of autonomous driving systems.
Smart Images

Figure CN119620619B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving vehicle motion control technology, and in particular to a model-free self-learning dynamic optimization control method and system for autonomous driving vehicles based on dynamic feedback for improving the comfort and safety of autonomous driving. Background Art
[0002] The continuous development of artificial intelligence and information technology has brought autonomous driving to a new level. At the same time, as a key link in smart cities, autonomous driving can effectively improve road traffic safety, reduce urban road traffic congestion, and bring great changes to the current transportation system. Autonomous driving vehicles are usually composed of the following modules: perception and positioning, high-level and low-level path planning, and path tracking. Path tracking plays a vital role in these modules, which not only directly determines the safety of autonomous driving vehicles, but also directly determines the driving experience of passengers. Therefore, it is of great application value to improve the comfort and safety of autonomous driving in the path tracking control module.
[0003] When dealing with the problem of autonomous driving path tracking control, traditional reinforcement learning can already train the controller very well. However, when considering driving comfort and safety issues, relying solely on traditional reinforcement learning methods may not be very good, and additional modules need to be introduced to assist it. For example, when discussing the safety of autonomous driving, the time to collision indicator is often introduced to regulate the autonomous driving vehicle, that is, when the collision time is less than a given threshold, additional strategies will be triggered to ensure safety. When discussing the comfort of autonomous driving, the vehicle acceleration indicator is introduced, and comfortable autonomous driving is achieved by optimizing the vehicle's acceleration. However, this method still has some problems, that is, the acceleration is generally determined by the rate of change of the vehicle control quantity, and only optimizing the vehicle acceleration indicator may not be able to achieve comfortable driving well. The method based on the vehicle acceleration indicator has a certain hysteresis, that is, it has already generated a large acceleration and then suppressed. More importantly, this optimization indicator does not conform to the driving habits of human drivers, that is, human drivers generally do not determine driving behavior based on the real-time acceleration of the vehicle. On the contrary, human drivers will optimize driving behavior from the control side, such as lightly stepping on the accelerator and lightly turning the steering wheel. In addition to safety and comfort issues, the dynamic characteristics of autonomous vehicles have problems such as being unable to be accurately modeled or gradually changing with use. Therefore, studying autonomous driving control methods that do not completely rely on the dynamics model of autonomous driving vehicles has broad practical value. Summary of the invention
[0004] The present invention provides a model-free self-learning dynamic optimization control method and system for an autonomous driving vehicle, so as to solve the problems that the existing optimization of vehicle acceleration indicators cannot well achieve comfortable driving and the model-based autonomous driving control method has a high dependence on the vehicle dynamics model.
[0005] According to a first aspect of the present invention, an embodiment provides a model-free self-learning dynamic optimization control method for an autonomous driving vehicle, comprising the following steps: obtaining a vehicle state quantity and a vehicle control quantity of a target autonomous driving vehicle according to a pre-constructed vehicle dynamics model of the target autonomous driving vehicle; constructing an augmented vehicle dynamics model of the target autonomous driving vehicle, obtaining tracking trajectory information of the target autonomous driving vehicle, and establishing a tracking error system according to the augmented vehicle dynamics model and the tracking trajectory information; constructing an autonomous driving dynamic feedback control strategy for driving comfort and safety according to the augmented vehicle dynamics model and the tracking error system; and training the autonomous driving dynamic feedback control strategy for driving comfort and safety until a preset maximum training time is reached to obtain a final autonomous driving strategy.
[0006] Optionally, the vehicle state quantities include the vehicle's lateral and longitudinal positions, the vehicle's lateral and longitudinal velocities, and the vehicle's heading angle and yaw rate; the vehicle control quantities include the front and rear axle propulsion or braking torque and the sideslip angles of the front and rear tires.
[0007] Optionally, constructing an augmented vehicle dynamics model of the target autonomous driving vehicle, acquiring tracking trajectory information of the target autonomous driving vehicle, and establishing a tracking error system according to the augmented vehicle dynamics model and the tracking trajectory information includes:
[0008] Performing a differentiation operation on the vehicle control amount to obtain a virtual control amount; constructing an augmented vehicle dynamics model based on the virtual control amount and the vehicle dynamics model; acquiring tracking trajectory information of the target autonomous driving vehicle based on the augmented vehicle dynamics model; defining an augmented tracking error vector based on the tracking trajectory information, so as to establish the tracking error system using the augmented tracking error vector and the augmented vehicle dynamics model.
[0009] Optionally, constructing an automatic driving dynamic feedback control strategy for driving comfort and safety according to the augmented vehicle dynamics model and the tracking error system includes:
[0010] Design the original performance indicators of the target autonomous driving vehicle; design the augmented performance indicators based on the augmented vehicle dynamics model and the original performance indicators; construct the autonomous driving dynamic feedback control strategy for driving comfort and safety based on the tracking error system, with the goal of minimizing the augmented performance indicators.
[0011] Optionally, the training of the autonomous driving dynamic feedback control strategy for driving comfort and safety until a preset maximum training time is reached to obtain a final autonomous driving strategy includes:
[0012] A safe operating state range and a backup safe operating driving strategy are set; a feedforward control strategy is constructed according to the tracking trajectory information or the vehicle dynamics model; based on the safe operating state range, the autonomous driving dynamic feedback control strategy for driving comfort and safety is trained using a reinforcement learning algorithm based on an execution-evaluation structure to obtain an autonomous driving dynamic feedback control strategy; the autonomous driving dynamic feedback control strategy is used to control the operation of the target autonomous driving vehicle, and the training process is iteratively executed until the preset maximum training time is reached to obtain the final autonomous driving strategy.
[0013] Optionally, when the target autonomous driving vehicle exceeds the safe operating state range during operation, the autonomous driving dynamic feedback control strategy for driving comfort and safety is replaced with the backup safe operation driving strategy to control the operation of the target autonomous driving vehicle until it runs within the safe operating state range, and then the backup safe operation driving strategy is switched to the autonomous driving dynamic feedback control strategy for driving comfort and safety to continue training.
[0014] The second aspect of the present invention provides a model-free self-learning dynamic optimization control system for an autonomous driving vehicle, including: an acquisition module, used to acquire the vehicle state quantity and vehicle control quantity of the target autonomous driving vehicle according to a pre-constructed vehicle dynamics model of the target autonomous driving vehicle; a construction module, used to construct an augmented vehicle dynamics model of the target autonomous driving vehicle, acquire the tracking trajectory information of the target autonomous driving vehicle, and establish a tracking error system according to the augmented vehicle dynamics model and the tracking trajectory information; a construction module, used to construct an autonomous driving dynamic feedback control strategy for driving comfort and safety according to the augmented vehicle dynamics model and the tracking error system; an iterative training module, used to train the autonomous driving dynamic feedback control strategy for driving comfort and safety until a preset maximum training time is reached to obtain a final autonomous driving strategy.
[0015] A third aspect of the present invention provides an autonomous driving vehicle, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the model-free self-learning dynamic optimization control method for the autonomous driving vehicle as described in the above embodiment.
[0016] A fourth aspect of the present invention provides a computer program product, which, when executed by a processor, implements the above-mentioned model-free self-learning dynamic optimization control method for an autonomous driving vehicle.
[0017] A fifth aspect of the present invention provides a computer-readable storage medium, which stores a computer program that, when executed by a processor, implements the above-mentioned model-free self-learning dynamic optimization control method for an autonomous driving vehicle.
[0018] The model-free self-learning dynamic optimization control method and system for autonomous driving vehicles proposed in the embodiments of the present invention construct an autonomous driving dynamic feedback control strategy that does not rely on an accurate vehicle dynamics model, and train it through a reinforcement learning method to obtain a final autonomous driving strategy, thereby achieving autonomous driving control that coexists with comfort and safety, thereby effectively improving the experience of passengers in autonomous driving vehicles and ensuring driving safety, and can better adapt to various complex road environments and traffic conditions, improve the robustness and adaptability of the autonomous driving system, and can promote the construction and development of autonomous driving and smart cities, with a wide range of applications.
[0019] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0021] Figure 1 A flow chart of a model-free self-learning dynamic optimization control method for an autonomous driving vehicle provided by an embodiment of the present invention;
[0022] Figure 2 A schematic diagram of a specific flow chart of constructing an automatic driving dynamic feedback control strategy for driving comfort and safety provided by an embodiment of the present invention;
[0023] Figure 3 A schematic diagram of the overall control flow provided by an embodiment of the present invention;
[0024] Figure 4 A block diagram of a model-free self-learning dynamic optimization control system for an autonomous driving vehicle provided by an embodiment of the present invention;
[0025] Figure 5 A schematic diagram of the structure of an autonomous driving vehicle provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0026] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.
[0027] The following describes the model-free self-learning dynamic optimization control method and system for an autonomous driving vehicle according to an embodiment of the present invention with reference to the accompanying drawings.
[0028] Figure 1 A flow chart of a model-free self-learning dynamic optimization control method for an autonomous driving vehicle provided in an embodiment of the present invention.
[0029] like Figure 1 As shown, the model-free self-learning dynamic optimization control method of the autonomous driving vehicle includes the following steps:
[0030] In step S101, the vehicle state quantity and vehicle control quantity of the target autonomous driving vehicle are obtained according to a pre-constructed vehicle dynamics model of the target autonomous driving vehicle.
[0031] In some embodiments, the vehicle state quantities include the vehicle's lateral and longitudinal positions, the vehicle's lateral and longitudinal velocities, and the vehicle's heading angle and yaw rate; the vehicle control quantities include the front and rear axle propulsion or braking torques and the sideslip angles of the front and rear tires.
[0032] In the actual implementation process, Figure 2 As shown, the vehicle dynamics model of the target autonomous driving vehicle is established in advance, and the specific expression is:
[0033]
[0034] In the formula, and are the lateral and longitudinal positions of the autonomous vehicle, and are the lateral and longitudinal speeds of the autonomous vehicle, is the heading angle of the autonomous driving vehicle, is the yaw rate of the autonomous driving vehicle, , are the tire forces on the front and rear axles of the autonomous driving vehicle in the lateral direction, , are the tire forces of the front and rear axles in the longitudinal direction of the autonomous driving vehicle, and are the distances from the center of gravity of the autonomous vehicle to the center of the front axle and the center of the rear axle, and are the mass and moment of inertia of the autonomous vehicle, is the air resistance. Usually, when the influence of wind speed is ignored, the air resistance can be expressed as ,in, is the air mass density, is the air resistance coefficient, The effective front area.
[0035] According to the vehicle state equation, the tire force can be calculated as follows:
[0036]
[0037] In the formula, Indicates the side slip angle of the front or rear tire. and is the tire force under the wheel frame, and It can be expressed as:
[0038]
[0039] In the formula, Indicates the propulsion or braking torque of the front and rear axles, is the tire radius, is the tire cornering stiffness, Used to describe the road surface. is the slip angle, is the normal force and can be calculated as:
[0040]
[0041] in, is the acceleration due to gravity.
[0042] Autonomous driving control task adjustment The automatic driving vehicle can be driven according to a predetermined trajectory. Therefore, the vehicle control amount selected in the embodiment of the present invention includes the front and rear axle propulsion or braking torque and the side slip angle of the front and rear tires, that is, At the same time, the state quantity of the autonomous driving vehicle is selected as: , we can further obtain the expression of the following dynamic nonlinear equation:
[0043] .
[0044] In step S102, an augmented vehicle dynamics model of the target autonomous driving vehicle is constructed, tracking trajectory information of the target autonomous driving vehicle is obtained, and a tracking error system is established based on the augmented vehicle dynamics model and the tracking trajectory information.
[0045] In some embodiments, constructing an augmented vehicle dynamics model of a target autonomous driving vehicle, obtaining tracking trajectory information of the target autonomous driving vehicle, and establishing a tracking error system based on the augmented vehicle dynamics model and the tracking trajectory information include:
[0046] Performing a differential operation on the vehicle control quantity to obtain a virtual control quantity;
[0047] An augmented vehicle dynamics model is constructed according to the virtual control quantity and the vehicle dynamics model;
[0048] Obtaining tracking trajectory information of the target autonomous driving vehicle according to the augmented vehicle dynamics model;
[0049] An augmented tracking error vector is defined according to the tracking trajectory information, so as to establish a tracking error system by using the augmented tracking error vector and an augmented vehicle dynamics model.
[0050] In the actual implementation process, Figure 2 As shown in the figure, by performing differential operation on the control amount of the autonomous driving vehicle, the virtual control amount can be obtained:
[0051]
[0052] Furthermore, we define the augmented state , the following augmented vehicle dynamics model can be obtained:
[0053]
[0054] In the formula, , .
[0055] Note that in the original vehicle dynamics model, the system dynamic characteristics are completely unknown, i.e. In the augmented vehicle dynamics model, the input dynamic is known and is set to a constant matrix.
[0056] In the actual implementation process, the preset tracking trajectory information is generally given by the path planning module of the autonomous driving vehicle. Since its setting must follow the constraints of the vehicle's dynamic characteristics (i.e., the constraints of the dynamic characteristics that are not set by humans), the tracking trajectory information can usually be expressed as:
[0057]
[0058] In the formula, is the desired augmented state trajectory, that is , is the desired control quantity, that is, the feedforward control strategy.
[0059] After obtaining the tracking trajectory, the tracking error state information can be determined by the following formula:
[0060]
[0061] in, is the lateral tracking error, is the desired lateral displacement, is the lateral velocity error, is the desired lateral velocity, is the longitudinal tracking error, is the desired longitudinal displacement, is the longitudinal velocity error, is the desired longitudinal velocity, is the heading angle error, is the desired heading angle, is the yaw rate error, is the desired yaw rate, is the rear axle propulsion or braking torque error, is the desired rear axle propulsion or braking torque, is the side slip angle error of the rear tire, is the desired rear tire slip angle, is the front axle propulsion or braking torque error, is the desired front axle propulsion or braking torque, is the side slip angle error of the front tire, is the desired front tire slip angle.
[0062] Further, the augmented tracking error vector is defined according to the tracking error state information as: , , , and then the tracking error system can be constructed using the augmented tracking error vector and the augmented vehicle dynamics model:
[0063]
[0064] In the formula, and .
[0065] In step S103, an automatic driving dynamic feedback control strategy for driving comfort and safety is constructed based on the augmented vehicle dynamics model and the tracking error system.
[0066] In some embodiments, an automatic driving dynamic feedback control strategy for driving comfort and safety is constructed based on an augmented vehicle dynamics model and a tracking error system, including:
[0067] Design the original performance indicators of the target autonomous driving vehicle;
[0068] Designing an augmented performance index based on the augmented vehicle dynamics model and the original performance index;
[0069] An automatic driving dynamic feedback control strategy for driving comfort and safety is constructed based on the tracking error system, with the goal of minimizing the augmented performance index.
[0070] In actual implementation, if there is no augmented vehicle dynamics model, the general optimization control problem of autonomous driving vehicles is to find a static optimization control strategy. Making autonomous vehicles track errors As small as possible, that is, to track the preset tracking trajectory while minimizing the following raw performance indicators:
[0071]
[0072] In the formula, It is usually called the utility function, which is generally a positive definite function. and When it converges to 0, the performance index Also convergent.
[0073] Furthermore, the augmented performance index is designed according to the augmented vehicle dynamics model and the original performance index as follows:
[0074]
[0075] In the formula, , About is a positive definite function of .
[0076] Furthermore, in order to achieve a comfortable autonomous driving control strategy, the utility function is designed as:
[0077]
[0078] In the formula, and is a positive definite matrix of suitable dimension, used to optimize the tracking error system and .
[0079] The above performance indicators mainly include two parts, one for autonomous driving performance and the other for autonomous driving comfort. The former mainly ensures driving performance and safety by optimizing tracking errors. The latter achieves comfortable and smooth autonomous driving control by imitating human driving behavior, that is, making appropriate small changes to driving actions. It can be seen that the original vehicle dynamics model cannot introduce driving comfort indicators into the indicators. At this point, the goal is to design the following automatic driving dynamic feedback control strategy for driving comfort and safety based on the tracking error system to minimize the augmented performance index and realize the path tracking control of automatic driving. Among them, the automatic driving dynamic feedback control strategy for driving comfort and safety is:
[0080]
[0081] At this point, the original control input of the autonomous driving vehicle can be expressed as: , which is consistent with the general optimization control and reinforcement learning used There is a clear difference, that is, the control strategy is a memory function of the state and has a differential form, so it is called dynamic optimization control.
[0082] In step S104, the automatic driving dynamic feedback control strategy for driving comfort and safety is trained until the preset maximum training time is reached to obtain the final automatic driving strategy.
[0083] In some embodiments, the automatic driving dynamic feedback control strategy for driving comfort and safety is trained until a preset maximum training time is reached to obtain a final automatic driving strategy, including:
[0084] Setting safe operating state ranges and backup safe operating driving strategies;
[0085] Construct a feedforward control strategy based on tracking trajectory information or vehicle dynamics model;
[0086] Based on the safe operating state range, the reinforcement learning algorithm based on the execution-evaluation structure is used to train the automatic driving dynamic feedback control strategy for driving comfort and safety, and the automatic driving dynamic feedback control strategy is obtained;
[0087] The target autonomous driving vehicle is controlled by using the autonomous driving dynamic feedback control strategy, and the training process is iteratively executed until the target autonomous driving vehicle reaches the preset maximum training time to obtain the final autonomous driving strategy.
[0088] Among them, when the target autonomous driving vehicle exceeds the safe operating state range during operation, the autonomous driving dynamic feedback control strategy for driving comfort and safety is replaced with the backup safe operation driving strategy to control the operation of the target autonomous driving vehicle until it runs within the safe operating state range, and then the backup safe operation driving strategy is switched to the autonomous driving dynamic feedback control strategy for driving comfort and safety to continue training.
[0089] In the actual implementation process, Figure 2 and 3As shown, the integral reinforcement learning technology is used to train the dynamic feedback control strategy of autonomous driving for driving comfort and safety until the preset conditions are met, and the final autonomous driving strategy is obtained, including:
[0090] First, determine the safe operating state range (the maximum value allowed for each state in the operation of the autonomous driving vehicle), that is,
[0091]
[0092]
[0093]
[0094]
[0095]
[0096]
[0097]
[0098]
[0099] in, is the operator norm, is the maximum lateral tracking error allowed in the operation of the autonomous vehicle, is the maximum lateral velocity error allowed in the operation of the autonomous vehicle, is the maximum longitudinal tracking error allowed in the operation of the autonomous vehicle, is the maximum longitudinal speed error allowed in the operation of the autonomous vehicle, is the maximum heading angle error allowed in the operation of the autonomous vehicle, is the maximum yaw rate error allowed in the operation of the autonomous vehicle, is the maximum propulsion or braking torque error allowed on the front axle during the operation of the autonomous vehicle, It is the maximum side slip angle error of the front tire allowed in the operation of the autonomous vehicle. It should be pointed out that the safe operating state range is determined by the hardware configuration of the autonomous vehicle and the designer's requirements for driving performance.
[0100] Furthermore, given a backup safe operating driving strategy, the strategy can be any control strategy, without having to be a dynamic control strategy, and is intended to ensure that the autonomous driving vehicle drives in a safe area when exceeding the safe operating state.
[0101] Before training the dynamic feedback optimization control strategy, build a feedforward control strategy The feedforward control strategy is related to the tracking trajectory information. The embodiments of the present invention provide two methods for determining the feedforward control strategy.
[0102] The first method: select different control quantities Stimulate the autonomous vehicle and measure the state quantities when the autonomous vehicle reaches a steady state The data table can also be made into a neural network, which is a feedforward network in the specific implementation.
[0103] The second method is to use a preset PID controller to make the autonomous vehicle track the preset , collect status data to achieve steady-state conditions and , the control input at this time That is feedforward control, so feedforward control is obtained.
[0104] It should be pointed out that the feedforward control strategy does not need to be precise and accurate. Common control indicators such as accuracy and robustness of autonomous driving vehicles can be adjusted by the feedback control part.
[0105] The execution-evaluation structure is used to train the autonomous driving strategy. As mentioned above, the training goal is to minimize the augmented performance index, that is, the evaluation network and the execution network are used respectively, where the evaluation network is:
[0106]
[0107] The execution network is:
[0108]
[0109] In the formula, and They are the neural network weights of the evaluation network and the neural network weights of the execution network respectively.
[0110] Approximate performance indicators and control strategies, where represents the activation function vector, . Take the initial network weights , and the initial network weights must ensure that the control strategy is an admissible control strategy. It should be noted that the execution network uses the input dynamics in the augmented vehicle dynamics model , so there is no need to know the information of the original vehicle dynamics model, so it is called a model-free training method, which is also impossible to achieve with static optimization control methods. At the same time, without knowing the information of the original vehicle dynamics model, an analytical form of strategy improvement is given based on the first-order necessary conditions, which can be regarded as a physical information neural network (Physics-Informed Neural Networks, referred to as PINN).
[0111] Furthermore, the autonomous driving dynamic feedback control strategy is used to control the operation of the autonomous driving vehicle. The following steps are followed to train and obtain the final autonomous driving strategy:
[0112] Initialize evaluation interval ,counter , Experience playback length , Experience Replay Cache , Network update trigger threshold , Operation end time And other parameters.
[0113] Using the automatic driving dynamic feedback control strategy to operate the automatic driving vehicle, In, traverse the running time From 0 to If the running time , then terminate the training and save the learned parameters; if the running time , then continue to judge the following conditions: If it meets , then execute the subsequent steps, where is a natural number, otherwise continue to execute the autonomous driving vehicle operation steps.
[0114] Update experience replay cache If the counter satisfy , then add a tuple To the experience replay cache In, order , and return to the autonomous vehicle operation step; otherwise, add a tuple To the experience replay cache and release the experience replay cache Middle group , in the tuple and They are:
[0115]
[0116] .
[0117] Replay Cache Based on Experience Update the evaluation network regulation law. Update the evaluation network regulation law as follows:
[0118]
[0119] in, To evaluate the network learning rate, , .
[0120] Replay Cache Based on Experience Update the execution network. When the evaluation network change rate (derivative) approaches zero, that is, Satisfy the conditions , then let , and update the execution network according to the following regulation law:
[0121]
[0122] in, To implement the network learning rate, Represents the projection operator. According to the definition of the projection operator, the learning law can be further written as the following jump form:
[0123]
[0124] Then clear the experience replay cache and counter, , , and then return to the autonomous driving vehicle operation step; otherwise, directly return to the autonomous driving vehicle operation step.
[0125] The autonomous driving vehicle stops running, and the evaluation and execution networks are trained according to the online data during the running time. The optimized performance indicators and optimized control strategies are obtained through the evaluation and execution networks respectively, and the final autonomous driving strategy is obtained.
[0126] Furthermore, if the state of the autonomous driving vehicle during operation exceeds the safe operating state, it switches to the backup safe operating driving strategy, and uses the backup safe operating driving strategy to control the operation of the target autonomous driving vehicle until it runs within the safe operating state range. All learned parameters are saved, and then the backup safe operating driving strategy is switched to the autonomous driving dynamic feedback control strategy for driving comfort and safety to re-train or terminate the training.
[0127] In summary, according to the model-free self-learning dynamic optimization control method of the autonomous driving vehicle proposed in the embodiment of the present invention, an autonomous driving dynamic feedback control strategy that does not rely on an accurate vehicle dynamics model is constructed, and it is trained through a reinforcement learning method to obtain the final autonomous driving strategy, thereby achieving autonomous driving control with coexistence of comfort and safety, thereby effectively improving the experience of passengers in the autonomous driving vehicle and ensuring driving safety, and can better adapt to various complex road environments and traffic conditions, improve the robustness and adaptability of the autonomous driving system, and can promote the construction and development of autonomous driving and smart cities, and has a wide range of applications.
[0128] Next, the model-free self-learning dynamic optimization control system for an autonomous driving vehicle proposed in accordance with an embodiment of the present invention is described with reference to the accompanying drawings.
[0129] Figure 4It is a block diagram of a model-free self-learning dynamic optimization control system for an autonomous driving vehicle according to an embodiment of the present invention.
[0130] like Figure 4 As shown, the model-free self-learning dynamic optimization control system 40 of the autonomous driving vehicle includes: an acquisition module 401, a construction module 402, a construction module 403 and an iterative training module 404.
[0131] Among them, the acquisition module 401 is used to obtain the vehicle state quantity and vehicle control quantity of the target autonomous driving vehicle according to the pre-constructed vehicle dynamics model of the target autonomous driving vehicle. The construction module 402 is used to construct the augmented vehicle dynamics model of the target autonomous driving vehicle, obtain the tracking trajectory information of the target autonomous driving vehicle, and establish a tracking error system based on the augmented vehicle dynamics model and the tracking trajectory information. The construction module 403 is used to construct an autonomous driving dynamic feedback control strategy for driving comfort and safety based on the augmented vehicle dynamics model and the tracking error system. The iterative training module 404 is used to train the autonomous driving dynamic feedback control strategy for driving comfort and safety until the preset maximum training time is reached to obtain the final autonomous driving strategy.
[0132] In some embodiments, the vehicle state quantities include the vehicle's lateral and longitudinal positions, the vehicle's lateral and longitudinal velocities, and the vehicle's heading angle and yaw rate; the vehicle control quantities include the front and rear axle propulsion or braking torques and the sideslip angles of the front and rear tires.
[0133] In some embodiments, building block 402 includes:
[0134] Performing a differential operation on the vehicle control quantity to obtain a virtual control quantity;
[0135] An augmented vehicle dynamics model is constructed according to the virtual control quantity and the vehicle dynamics model;
[0136] Obtaining tracking trajectory information of the target autonomous driving vehicle according to the augmented vehicle dynamics model;
[0137] An augmented tracking error vector is defined according to the tracking trajectory information, so as to establish a tracking error system by using the augmented tracking error vector and an augmented vehicle dynamics model.
[0138] In some embodiments, the construction module 403 includes:
[0139] Design the original performance indicators of the target autonomous driving vehicle;
[0140] Designing an augmented performance index based on the augmented vehicle dynamics model and the original performance index;
[0141] An automatic driving dynamic feedback control strategy for driving comfort and safety is constructed based on the tracking error system, with the goal of minimizing the augmented performance index.
[0142] In some embodiments, the iterative training module 404 includes:
[0143] Setting safe operating state ranges, safe operating driving strategies and backup safe operating driving strategies;
[0144] Construct a feedforward control strategy based on tracking trajectory information or vehicle dynamics model;
[0145] Based on the safe operation state range and safe operation driving strategy, the reinforcement learning algorithm based on the execution-evaluation structure is used to train the automatic driving dynamic feedback control strategy for driving comfort and safety, and the automatic driving dynamic feedback control strategy is obtained;
[0146] The target autonomous driving vehicle is controlled by using the autonomous driving dynamic feedback control strategy, and the training process is iteratively executed until the preset maximum training time is reached to obtain the final autonomous driving strategy.
[0147] Among them, when the target autonomous driving vehicle exceeds the safe operating state range during operation, the autonomous driving dynamic feedback control strategy for driving comfort and safety is replaced with the backup safe operation driving strategy to control the operation of the target autonomous driving vehicle until it runs within the safe operating state range, and then the backup safe operation driving strategy is switched to the autonomous driving dynamic feedback control strategy for driving comfort and safety to continue training.
[0148] It should be noted that the aforementioned explanation of the embodiment of the model-free self-learning dynamic optimization control method for the autonomous driving vehicle is also applicable to the model-free self-learning dynamic optimization control system for the autonomous driving vehicle of this embodiment, and will not be repeated here.
[0149] According to the model-free self-learning dynamic optimization control system of the autonomous driving vehicle proposed in the embodiment of the present invention, an autonomous driving dynamic feedback control strategy that does not rely on an accurate vehicle dynamics model is constructed, and is trained through a reinforcement learning method to obtain a final autonomous driving strategy, thereby achieving autonomous driving control that combines comfort and safety, thereby effectively improving the experience of passengers in the autonomous driving vehicle and ensuring driving safety. It can also better adapt to various complex road environments and traffic conditions, improve the robustness and adaptability of the autonomous driving system, promote the construction and development of autonomous driving and smart cities, and has a wide range of applications.
[0150] Figure 5 A schematic diagram of the structure of an autonomous driving vehicle provided by an embodiment of the present invention. The autonomous driving vehicle may include:
[0151] A memory 501 , a processor 502 , and a computer program stored in the memory 501 and executable on the processor 502 .
[0152] When the processor 502 executes the program, the model-free self-learning dynamic optimization control method for the autonomous driving vehicle provided in the above embodiment is implemented.
[0153] Furthermore, the autonomous driving vehicle also includes:
[0154] The communication interface 503 is used for communication between the memory 501 and the processor 502 .
[0155] The memory 501 is used to store computer programs that can be executed on the processor 502 .
[0156] The memory 501 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0157] If the memory 501, the processor 502 and the communication interface 503 are implemented independently, the communication interface 503, the memory 501 and the processor 502 can be connected to each other through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0158] Optionally, in a specific implementation, if the memory 501, the processor 502 and the communication interface 503 are integrated on a chip, the memory 501, the processor 502 and the communication interface 503 can communicate with each other through an internal interface.
[0159] The processor 502 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.
[0160] An embodiment of the present invention also provides a computer program product, which, when executed by a processor, implements the above-mentioned model-free self-learning dynamic optimization control method for an autonomous driving vehicle.
[0161] An embodiment of the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned model-free self-learning dynamic optimization control method for an autonomous driving vehicle.
[0162] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0163] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "N" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0164] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present invention belong.
[0165] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection with one or N wirings (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways as necessary and then storing it in a computer memory.
[0166] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiment, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0167] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0168] In addition, each functional unit in each embodiment of the present invention may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0169] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present invention. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.
Claims
1. A model-free self-learning dynamic optimization control method for an autonomous driving vehicle, characterized in that: The following steps are involved: Acquire a vehicle state quantity and a vehicle control quantity of the target autonomous driving vehicle according to a pre-constructed vehicle dynamics model of the target autonomous driving vehicle; Constructing an augmented vehicle dynamics model of the target autonomous driving vehicle, obtaining tracking trajectory information of the target autonomous driving vehicle, and establishing a tracking error system according to the augmented vehicle dynamics model and the tracking trajectory information, specifically comprising: Performing a differential operation on the vehicle control amount to obtain a virtual control amount; constructing an augmented vehicle dynamics model according to the virtual control amount and the vehicle dynamics model; Acquiring tracking trajectory information of the target autonomous driving vehicle according to the augmented vehicle dynamics model; defining an augmented tracking error vector according to the tracking trajectory information, so as to establish the tracking error system by using the augmented tracking error vector and the augmented vehicle dynamics model; Constructing an automatic driving dynamic feedback control strategy for driving comfort and safety according to the augmented vehicle dynamics model and the tracking error system; The automatic driving dynamic feedback control strategy for driving comfort and safety is trained until a preset maximum training time is reached to obtain a final automatic driving strategy.
2. The model-free self-learning dynamic optimization control method for an autonomous driving vehicle according to claim 1, characterized in that: The vehicle state includes the lateral and longitudinal positions of the vehicle, the lateral and longitudinal velocities of the vehicle, and the heading angle and yaw rate of the vehicle; The vehicle control quantities include front and rear axle propulsion or braking torques and the sideslip angles of the front and rear tires.
3. The model-free self-learning dynamic optimization control method for an autonomous driving vehicle according to claim 1, characterized in that: The method of constructing an automatic driving dynamic feedback control strategy for driving comfort and safety based on the augmented vehicle dynamics model and the tracking error system includes: Designing original performance indicators of the target autonomous driving vehicle; designing an augmented performance index according to the augmented vehicle dynamics model and the original performance index; The automatic driving dynamic feedback control strategy for driving comfort and safety is constructed based on the tracking error system, with the goal of minimizing the augmented performance index.
4. The model-free self-learning dynamic optimization control method for an autonomous driving vehicle according to claim 1, characterized in that: The step of training the automatic driving dynamic feedback control strategy for driving comfort and safety until a preset maximum training time is reached to obtain a final automatic driving strategy includes: Setting safe operating state ranges and backup safe operating driving strategies; Constructing a feedforward control strategy according to the tracking trajectory information or the vehicle dynamics model; Based on the safe operating state range, the automatic driving dynamic feedback control strategy for driving comfort and safety is trained using a reinforcement learning algorithm based on an execution-evaluation structure to obtain an automatic driving dynamic feedback control strategy; The target autonomous driving vehicle is controlled to operate using the autonomous driving dynamic feedback control strategy, and the training process is iteratively executed until the preset maximum training time is reached to obtain the final autonomous driving strategy.
5. The model-free self-learning dynamic optimization control method for an autonomous driving vehicle according to claim 4, characterized in that: When the target autonomous driving vehicle exceeds the safe operating state range during operation, the autonomous driving dynamic feedback control strategy for driving comfort and safety is replaced with the backup safe operation driving strategy to control the operation of the target autonomous driving vehicle until it runs within the safe operating state range, and then the backup safe operation driving strategy is switched to the autonomous driving dynamic feedback control strategy for driving comfort and safety to continue training.
6. A model-free self-learning dynamic optimization control system for an autonomous driving vehicle, characterized in that: include: An acquisition module, used to acquire a vehicle state quantity and a vehicle control quantity of the target autonomous driving vehicle according to a pre-built vehicle dynamics model of the target autonomous driving vehicle; A construction module is used to construct an augmented vehicle dynamics model of the target autonomous driving vehicle, obtain tracking trajectory information of the target autonomous driving vehicle, and establish a tracking error system according to the augmented vehicle dynamics model and the tracking trajectory information, specifically comprising: Performing a differential operation on the vehicle control quantity to obtain a virtual control quantity; An augmented vehicle dynamics model is constructed according to the virtual control quantity and the vehicle dynamics model; Obtaining tracking trajectory information of the target autonomous driving vehicle according to the augmented vehicle dynamics model; An augmented tracking error vector is defined according to the tracking trajectory information, so as to establish a tracking error system by using the augmented tracking error vector and an augmented vehicle dynamics model; A construction module, used for constructing an automatic driving dynamic feedback control strategy for driving comfort and safety according to the augmented vehicle dynamics model and the tracking error system; The iterative training module is used to train the autonomous driving dynamic feedback control strategy for driving comfort and safety until a preset maximum training time is reached to obtain a final autonomous driving strategy.
7. An autonomous driving vehicle, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the model-free self-learning dynamic optimization control method for an autonomous driving vehicle as described in any one of claims 1 to 5.
8. A computer program product, characterized in that When the computer program / instructions are executed by a processor, the model-free self-learning dynamic optimization control method for the autonomous driving vehicle described in any one of claims 1 to 5 is implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement a model-free self-learning dynamic optimization control method for an autonomous driving vehicle as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Trace tracking method of autonomous vehicle
CN109407677A
Automatic tracking control method for intelligent commercial vehicle based on output feedback gain planning
CN110471277A