Iterative learning control evolutionary method for autonomous vehicle in recurrent scenarios

By using an iterative learning control method and optimizing the control law of autonomous vehicles using preceding iterative data, the performance gap of online controllers in cyclic scenarios is solved, and adaptive evolution and performance improvement of autonomous vehicle control are achieved.

CN121541472BActive Publication Date: 2026-05-01BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIHANG UNIV
Filing Date
2025-11-24
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing online controllers for autonomous vehicles perform poorly in real-world scenarios. There is a performance gap between offline design and online application, leading to degradation or failure of control algorithms, especially in cyclical scenarios where it is difficult to effectively improve performance.

Method used

An iterative learning control method is adopted to optimize the control solution of subsequent iterations by leveraging the control experience of previous iterations. This constructs an iterative learning control evolution framework for unmanned vehicles, which automatically upgrades the control law using historical data, reducing error accumulation and improving control performance.

Benefits of technology

It enables the adaptive evolution of autonomous vehicle controllers in cyclical scenarios, improving control performance, reducing the risk of error accumulation, and enhancing system safety and control accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121541472B_ABST
    Figure CN121541472B_ABST
Patent Text Reader

Abstract

The application discloses an iterative learning control evolution method for an automatic driving vehicle in a circulation scenario and relates to the field of automatic driving vehicle control. First, an offline controller is designed for the automatic driving vehicle in the circulation scenario, and reference states in the corresponding scenario are obtained. Then, based on the reference states and considering the influence factors that are difficult to traverse in the offline controller design, an online controller is designed. Finally, the online controller is continuously evolved through iterative learning until the learning process converges, and the optimal control effect is achieved. Based on the optimal control effect, the motion state of the automatic driving vehicle converges to the expected state in the circulation scenario, and the task in the corresponding scenario is executed. The application guarantees the effectiveness and real-time performance of the online controller, avoids the error accumulation in the iterative process, and improves the safety of the unmanned system.
Need to check novelty before this filing date? Find Prior Art

Description

Iterative learning control evolution method for autonomous vehicles in recurring scenarios Technical Field

[0001] This invention relates to the field of autonomous vehicle control, and in particular to a control evolution method based on iterative learning aimed at improving the control performance of existing autonomous vehicles in cyclic scenarios. Background Technology

[0002] Autonomous vehicles have significant advantages in expanding traffic flow, optimizing energy consumption and emissions, and enhancing driving safety. In particular, autonomous driving technology has shown great application potential in cyclical scenarios such as transportation, patrol, and guidance. Among these, vehicle control issues directly determine the feasibility and effectiveness of autonomous driving.

[0003] There is a significant contradiction between current offline controller design and online controller application. Offline designs cannot fully cover the countless influencing factors such as vehicle type, road gradient, and trajectory curvature, and typically employ a fixed configuration scheme that performs well in most situations for online applications. On the one hand, with the development of autonomous driving technology, practical problems such as high-speed cruising and trajectory tracking require superior performance, but the high costs of repeated trial and error in online applications are unbearable. On the other hand, the performance gap between the ideal performance in the design phase and the actual effectiveness in the application phase accumulates over time, leading to degradation or even failure of the control algorithm.

[0004] Given the above discussion, scholars have begun to introduce iterative learning control (ILC), which improves the current control law by leveraging the experience of previous iterations to achieve performance enhancement. A typical application scenario for iterative learning control is in a cyclical scenario, where autonomous vehicles are periodically assigned to patrol, guide, or transport tasks along fixed routes; each task process is called an "iteration." Under similar operating conditions and control objectives, the experience of previous iterations can support performance improvements in subsequent iterations, thereby enabling the evolution of autonomous vehicle control. Summary of the Invention

[0005] To address the issue of poor control algorithm performance when online controllers for autonomous vehicles are applied in real-world scenarios, this invention proposes an iterative learning control evolution method for autonomous vehicles in cyclic scenarios. For existing control laws designed offline, this method provides an iterative learning-based control evolution approach, incorporating the control results of previous iterations into the control calculations of subsequent iterations. This optimizes the manual parameter tuning process between iterations, achieving automatic control evolution and improving the performance of existing control laws.

[0006] An iterative learning control evolution method for autonomous vehicles in recurring scenarios includes the following steps:

[0007] Step 1: Design an offline controller for autonomous vehicles in a loop scenario to obtain the vehicle's initial control system and reference state;

[0008] The design process of the offline controller is as follows:

[0009] First, for the cyclic scenario, based on the Frenet coordinate system, in the... In the next iteration, at the corresponding time The vehicle closed-loop control system, composed of the vehicle's state vector and control vector, is represented as follows:

[0010]

[0011] Vehicle state vector This includes the vehicle's coordinates along the road direction. Perpendicular to the road direction coordinates The deviation between the road heading angle and the vehicle heading angle longitudinal speed Lateral speed yaw rate Control vector Including the steering angle of the vehicle's front wheels Vehicle driving torque ;

[0012] Then, define the desired state of the vehicle. , For the longitudinal desired vehicle speed, For slowly changing road curvature. The sampling time step is used to further obtain the vehicle's reference state. .

[0013] Finally, the design goal of the offline controller is expressed as: minimizing the control error under the constraints of the vehicle's closed-loop control system. .

[0014] Step 2: Based on the reference state, the autonomous vehicle adopts the MPC controller as the optimal controller and designs the controller to construct an iterative learning control problem.

[0015] The optimal control law obtained based on the reference state is:

[0016] (1)

[0017] in, For vectors A reference state vector of the same dimension. Represents a 6-dimensional real vector space. for A positive definite symmetric matrix. express 3D real matrix space, For optimal control law The implicit function expression satisfies .

[0018] Based on the optimal control law, and considering the influencing factors that are difficult to traverse in offline controller design, the following iterative learning control problem is constructed:

[0019] (2)

[0020] in, The dataset recorded in the previous iteration, For iterative learning loss function, used to... During the iteration, the first... Secondary cycle data evolution control law.

[0021] That is, in the recurrent scenario, the iterative learning control problem is constructed as follows:

[0022] (3)

[0023] in This is the implicit function expression for optimal control. This is the implicit function expression for the iterative learning loss.

[0024] Step 3: In the cyclical scenario, the autonomous vehicle performs tasks and iterates through historical data in each task execution. Through data-driven iterative learning, the control problem is evolved online until the learning process converges and the optimal control effect is achieved.

[0025] The process of online evolution in the iterative learning control problem is as follows:

[0026] In the first mission execution, i.e., the 0th iteration, the autonomous vehicle applies the control law designed offline. and obtain basic data .

[0027] In the first iteration, the control law is updated to This caused the unconsidered factors in the design of the offline controller to be transformed into... Feedback to In the process, the control performance of the first iteration was improved and data was obtained. .

[0028] Based on the Historical data from the next iteration, in any given iteration The online evolution process of the controller is repeated in each iteration. Introduction Deploy iterative learning control laws The control law continues to evolve online until the learning process converges, achieving the optimal control effect.

[0029] Step four: Based on the optimal control effect, the autonomous vehicle's motion state converges to the desired state in the cyclical scenario, and it executes the task in the corresponding scenario.

[0030] The present invention has the following advancements and advantages:

[0031] (1) In view of the unavoidable performance gap between offline design and online application of controller, the present invention provides a data-driven online evolution framework for controller, which ensures the effectiveness of the designed controller when applied online.

[0032] (2) For cyclic scenarios, this invention designs an adaptive control scheme that does not require human intervention, feeds back the error out-of-bounds problem in the previous iteration to the subsequent iteration, and punishes improper control signals through a cost function to avoid the risk of error accumulation and improve the safety of unmanned systems.

[0033] (3) In repetitive tasks, make full use of historical data to automatically upgrade the deployed control law and make up for the deficiencies in the robustness, conservatism and accuracy of the initial control law. Attached Figure Description

[0034] Figure 1 is a flowchart of the iterative learning control evolution algorithm for autonomous vehicles in cyclic scenarios provided by the present invention.

[0035] Figure 2 is a schematic diagram of the variable definition and coordinate relationship of the controlled object provided by the present invention;

[0036] Figure 3 shows the application effect of the algorithm provided by the present invention in the embodiment of iterative learning. With the intervention of the iterative learning algorithm, the trajectory tracking control effect gradually improves.

[0037] Figure 4 shows the control effect of the algorithm provided by the present invention on the speed of the controlled vehicle in an embodiment. As the iterative learning algorithm is introduced, the speed control effect gradually improves.

[0038] Figure 5 shows the application effect of the algorithm provided by the present invention on the tracking performance of the controlled vehicle in the embodiment. With the intervention of the iterative learning algorithm, the lateral error of path tracking gradually decreases. Detailed Implementation

[0039] The embodiments of the present invention will be further described in detail below with reference to the accompanying drawings and specific examples:

[0040] The iterative learning control evolution method for autonomous vehicles in cyclic scenarios provided by this invention is an unmanned vehicle control scheme based on iterative learning and possessing self-evolution capabilities. As shown in Figure 1, the specific implementation steps include:

[0041] Step 1: For autonomous vehicles in cyclical scenarios, construct a closed-loop control system for the vehicle;

[0042] As shown in Figure 2, the definition is in the... In the next iteration, at the corresponding time Below, in the Frenet coordinate system, the coordinates of the vehicle along the road direction Perpendicular to the road direction coordinates The deviation between the road heading angle and the vehicle heading angle longitudinal speed Lateral speed yaw rate The vehicle state vector consists of 6 dimensions. superscript Represents the transpose of a matrix or vector; determined by the steering angle of the vehicle's front wheels. and vehicle drive torque Composition of control vector The vehicle closed-loop control system can be described by formula (1):

[0043] (1)

[0044] in, For the overall vehicle quality, Let the vehicle's yaw moment of inertia be... The gear ratio of the vehicle drive system. For the transmission efficiency of the vehicle drive system, The rolling resistance coefficient, This is the equivalent air drag coefficient. For the front axle equivalent lateral stiffness, For the equivalent lateral stiffness of the rear axle, This is the longitudinal distance from the center of gravity to the front axle. This is the longitudinal distance from the center of mass to the rear axle. This is the sampling time step.

[0045] Without loss of generality, system (1) is constructed as an implicit expression (2);

[0046] (2)

[0047] Step 2: In the loop scenario, design the offline controller for the autonomous vehicle to obtain the reference state of the vehicle;

[0048] The controlled vehicles need to perform tasks such as transportation and patrol along a predetermined route without human intervention, and each task execution is called an iteration.

[0049] In the In each iteration, as the task execution time... The control objective for the scrolling can be expressed as:

[0050] (3)

[0051] in, For the longitudinal desired vehicle speed, The curvature of the road changes slowly.

[0052] With the number of iterations With the increase of , the overall iterative control objective of the system can be expressed as:

[0053] (4)

[0054] Define the desired state The system control objective can be expressed as:

[0055] (5)

[0056] in, This represents the negative limit at 0 (i.e., from the negative half-axis closer to 0, indicating the first...). The control error of the second iteration is less than that of the third iteration. (next iteration)

[0057] Therefore, the offline design goal of the controller can be decomposed into the principle of applying the laws of system dynamics. Minimize control error under constraints .

[0058] Meanwhile, in repetitive task scenarios, the reference state in the control problem is considered... In any of the first It remains unchanged in each iteration.

[0059] Step 3: Based on the reference state, the autonomous vehicle adopts the MPC controller as the optimal controller, and designs the controller to construct an iterative learning control problem;

[0060] The specific process of controller design is as follows:

[0061] S301. In each iteration of the cyclic scenario, the controlled vehicle follows a similar desired trajectory. Perform repetitive tasks where the task scenario is given by a predetermined route and road curvature. The task requires a desired vehicle speed. At the same time, vehicles are affected by lane width. Speed ​​limits on road sections The constraints satisfy .

[0062] S302. In the offline design of the controller, the vehicle kinematics model (1) is introduced to predict the future vehicle state and calculate the cost function; inevitably, the actual driverless car roughly follows the vehicle kinematics (1), but there are model disturbances, such as:

[0063] (9)

[0064] in, This represents the unknown deviation between the actual unmanned vehicle dynamics and the vehicle dynamics model (1).

[0065] The dynamics of actual driverless vehicles (9) are difficult to predict, but real-time feedback is crucial. In addition, in the driverless car (9), the wheel turning angle and driving torque Subject to the mechanical limit constraints of the vehicle's steering system and drive system respectively, satisfying , .

[0066] S303. Considering the control problem corresponding to the multi-input multi-output system (1), the robustness requirements of the system (9), and the constraints of environmental and mechanical limits, the sampled model predictive control (MPC) controller is used as the optimal controller for controller design.

[0067] Define prediction domain Within the prediction domain, combining Figure 2 and the vehicle dynamics model shown in (1), the state sequence is predicted. ,satisfy:

[0068] (10)

[0069] Considering the above constraints, construct the MPC problem (11):

[0070] (11)

[0071] in, , Indicates that the diagonal elements are A diagonal matrix.

[0072] Based on the optimization problem (11), the implicit expression of the optimal control law can be written as:

[0073] (12)

[0074] That is, the optimal solution to the MPC problem (11) The first item Used to output control signals, it is an optimal control law and can be expressed as state feedback. and expected trajectory function ;

[0075] S304. Based on the MPC controller (12), an iterative learning loss is introduced. This allows for the optimization of the control law in subsequent iterations using historical data generated from previous iterations. For optimal control;

[0076] In this embodiment, Designed as a piecewise function (13)

[0077] (13)

[0078] in, For the constant used for replacement, To determine the state and Tiny constants of similarity.

[0079] Therefore, iterative learning control law This can be expressed as:

[0080] (14)

[0081] Similarly to (12), the optimal solution to problem (14) is... The first item Used to output control signals, it is an iterative learning control law, which can be expressed as state feedback. Expected trajectory Historical data and decision variables its own function ;

[0082] Step 4: In the cyclical scenario, the autonomous vehicle performs tasks and iterates through historical data in each task execution. Through data-driven iterative learning, the control problem evolves online until the learning process converges and the optimal control effect is achieved.

[0083] In the first task execution (i.e., the 0th iteration), the control law designed offline is applied. and obtain basic data .

[0084] In the first iteration, the control law is updated to This caused the unconsidered factors in the design of the offline controller to be transformed into... Feedback to In the process, the control performance of the first iteration was improved and data was obtained. This will improve the control precision of unmanned vehicles, stabilize speed tracking, reduce path tracking errors, and enhance task completion efficiency.

[0085] Based on the Historical data from the next iteration, in any given iteration Repeat step 3 in the next iteration, Introduction Deploy iterative learning control laws It continuously evolves the control law online, gradually improving control efficiency.

[0086] As shown in Figure 3, in each iteration, the controlled unmanned vehicle starts from the origin of the route, accelerates to the desired speed, and maintains high speed to track the desired route. Compared with the initial optimization control, the first iteration of learning control significantly shortens the acceleration process and reduces the path tracking error. The difference between the second and first iterations is relatively small, indicating that the iterative learning control gradually converges in the second learning process and reaches the optimal control performance.

[0087] As shown in Figure 4, under the desired vehicle speed setting, the controlled vehicle gradually fluctuates and converges to the desired vehicle speed; with the progress of iterative learning, the vehicle speed tracking oscillation gradually decreases, and the performance is improved.

[0088] As shown in Figure 5, in the Frenet coordinate system Coordinates are used as indicators to evaluate vehicle path tracking performance; as the learning process progresses, the average tracking error level decreases and the error curve oscillates less, thus improving performance.

[0089] The present invention has been described according to specific embodiments, but is not limited to the above-described embodiments. Equivalent modifications or substitutions made by those skilled in the art without departing from the scope of the claims are all within the scope of protection of the present invention, and the scope of protection of the present invention should be determined by the appended claims.

Claims

1. An iterative learning control evolution method for autonomous vehicles in recurring scenarios, characterized in that, The specific steps are as follows: Step 1, for autonomous vehicles in cyclic scenarios, design an offline controller to obtain the vehicle's initial control system and reference state; the design process of the offline controller is as follows: First, for cyclic scenarios, based on the Frenet coordinate system, in the... In the next iteration, at the corresponding time The vehicle closed-loop control system, composed of the vehicle's state vector and control vector, is represented as follows: This is the vehicle state vector. The vehicle control vector is defined; then, the desired state of the vehicle is defined. Then the vehicle's reference state is Finally, the design objective of the offline controller is expressed as: minimizing the control error under the constraints of the vehicle's closed-loop control system. Step two: Based on the reference state, the autonomous vehicle adopts the MPC controller as the optimal controller and designs the controller to construct an iterative learning control problem. The construction process of the iterative learning control problem is as follows: Based on the reference state, the optimal control law is obtained as follows: (1) Among them, For vectors A reference state vector of the same dimension. Represents a 6-dimensional real vector space. for A positive definite symmetric matrix. express 3D real matrix space, For optimal control law The implicit function expression satisfies Based on the optimal control law, and considering the influencing factors that are difficult to traverse in offline controller design, the following iterative learning control problem is constructed: (2) Among them, The dataset recorded in the previous iteration, For iterative learning loss function, used to... During the iteration, the first... The data evolution control law for the next iteration; in the recurring scenario, the iterative learning control problem is constructed as follows: in The implicit function expression for the iterative learning loss is given. Step 3: In the cyclic scenario, the autonomous vehicle performs a task and iterates through historical data in each task execution. The iterative learning control problem is evolved online through data-driven learning until the learning process converges and the optimal control effect is achieved. Step 4: Based on the optimal control effect, the autonomous vehicle's motion state in the cyclic scenario converges to the desired state, and the corresponding task in the scenario is executed.

2. The iterative learning control evolution method for autonomous vehicles oriented towards cyclic scenarios according to claim 1, characterized in that, The online evolution process of the iterative learning control problem is as follows: In the first task execution, i.e., the 0th iteration, the control law designed offline is applied to the autonomous vehicle. and obtain basic data In the first iteration, the control law is updated to... This caused the unconsidered factors in the design of the offline controller to be transformed into... Feedback to In the process, the control performance of the first iteration was improved and data was obtained. Based on the first Historical data from the next iteration, in any given iteration In the next iteration, return to step two and repeat the online evolution process of the controller. Introduction Deploy iterative learning control laws The control law continues to evolve online until the learning process converges, achieving the optimal control effect.

Citation Information

Patent Citations

  • Predictive learning tracking control method and device for self-driving automobile

    CN119535959A

  • Method and system for controlling an autonomous vehicle device to repeatedly follow a same predetermined trajectory

    US20210208596A1