Adaptive Cruise Control Method, Device and Medium Based on Online Adaptive Learning

By constructing the mathematical model of the adaptive cruise control system and the online strategy learning algorithm, the adaptability and robustness of the adaptive cruise control system in complex environments is solved, and stable tracking and efficient energy saving are achieved.

CN119821393BActive Publication Date: 2025-07-04SHANDONG UNIV OF SCI & TECH +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510273705.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-07-04
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

The existing adaptive cruise control system is not adaptable and robust in complex dynamic environments, making it difficult to deal with changes in road conditions and external interference, affecting driving safety and ride comfort.

Method used

Build a mathematical model of the adaptive cruise control system, define the H infinite control problem, and determine the optimal control strategy through online strategy learning algorithms to improve the adaptability and robustness of the system.

Benefits of technology

It realizes stable tracking and efficient energy saving of the adaptive cruise system in complex environments, improves driving comfort and safety, and enhances the robustness and stability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119821393B_ABST
    Figure CN119821393B_ABST
Patent Text Reader

Abstract

The present application discloses an adaptive cruise control method, device and medium based on online adaptive learning, belonging to the technical field of adaptive control. The method includes: constructing a mathematical model of the adaptive cruise control system and defining an H-infinity control problem related to the adaptive cruise control system; defining the state of the adaptive cruise control system based on a preset performance function algorithm and constructing an online policy learning algorithm; solving the H-infinity control problem based on the online policy learning algorithm to determine an optimal control strategy. The present application achieves the technical effect of improving the adaptability and robustness of the ACC system in a complex dynamic environment through the above method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of adaptive control, and particularly to an adaptive cruise control method, device, and medium based on online adaptive learning. Background Art

[0002] An Adaptive Cruise Control (ACC) system is an advanced driver assistance system that adds a function of maintaining a safe distance from the vehicle ahead on the basis of traditional cruise control. This system can monitor in real time information such as the acceleration and speed of the host vehicle (i.e., the vehicle equipped with the ACC system), as well as the distance and relative speed from the vehicle ahead, and precisely control the throttle and braking systems of the host vehicle through a built-in controller. The ACC system consists of two core controllers: an upstream controller and a downstream controller. The upstream controller is responsible for calculating the ideal host vehicle acceleration signal based on the collected real-time data, while the downstream controller is responsible for converting this ideal acceleration signal into the actual vehicle acceleration and achieving it by precisely adjusting the throttle and brakes.

[0003] Current adaptive cruise control systems have limitations in dealing with dynamic uncertainties (such as changes in road conditions and difficulties in predicting the behavior of other vehicles) and external disturbances (such as sudden deceleration by the driver or passengers). These uncertainties and disturbances may cause the system to fail to adjust the vehicle state in a timely and accurate manner, thus affecting driving safety and riding comfort.

[0004] Therefore, how to improve the adaptability and robustness of the ACC system in a complex dynamic environment has become a technical problem to be solved urgently. Summary of the Invention

[0005] Embodiments of this application provide an adaptive cruise control method, device, and medium based on online adaptive learning to solve the following technical problem: how to improve the adaptability and robustness of the ACC system in a complex dynamic environment.

[0006] In a first aspect, embodiments of this application provide an adaptive cruise control method based on online adaptive learning. The method includes: constructing a mathematical model of the adaptive cruise control system and defining an H-infinity control problem related to the adaptive cruise control system; defining the state of the adaptive cruise control system based on a preset performance function algorithm and constructing an online policy learning algorithm; solving the H-infinity control problem based on the online policy learning algorithm to determine the optimal control strategy.

[0007] In an implementation manner of this application, establishing a mathematical model of the adaptive cruise control system and defining an H-infinity control problem related to the adaptive cruise control system specifically includes:

[0008] The formula for the mathematical model of the adaptive cruise control system is as follows:

[0009] (1)

[0010] Among them, matrix A is the state matrix, matrix B is the input matrix, and matrix D is the disturbance matrix. is the state vector, is the control input, is the external disturbance, represents the headway time difference, that is, the time interval maintained between the host vehicle and the target vehicle. represents the time constant related to the dynamic response of the vehicle. , , represents the ideal stopping distance of the host vehicle. represents the actual stopping distance of the host vehicle. , represents the vehicle speed of the host vehicle. is the speed of the target vehicle. , is the acceleration of the host vehicle. , is the input desired acceleration, w is the disturbance of the adaptive cruise control system.

[0011] The formula for the H-infinity control problem of the adaptive cruise control system is as follows:

[0012] (2)

[0013] Among them, is the positive definite solution of this formula, * represents the number of iterations, Q and R are both positive definite weight matrices. is a small positive real number. is used to limit the output control of the external disturbance on the adaptive cruise control system within a controllable range.

[0014] In an implementation manner of this application, the state of the adaptive cruise control system is defined based on a preset performance function algorithm, and an online policy learning algorithm is constructed, specifically including:

[0015] Construct an offline policy learning algorithm;

[0016] Construct a planning performance function to constrain the offline policy learning algorithm, and process the offline policy learning algorithm based on the planning performance function to determine the initial state of the online policy learning algorithm;

[0017] Construct an online policy learning algorithm based on the initial state.

[0018] In an implementation manner of this application, constructing an offline policy learning algorithm specifically includes:

[0019] The offline policy learning algorithm is represented by the following formula:

[0020] Set an initial stable matrix ;

[0021] Based on , solve according to formula (3) :

[0022] (3)

[0023] Wherein, is a matrix describing the dynamic changes of the adaptive cruise control system, is the control gain matrix for calculating the control input according to the state of the adaptive cruise control system, is the disturbance gain matrix for calculating the disturbance input according to the state of the adaptive cruise control system, and are the state and control right matrices respectively;

[0024] Solve the initial optimal solution according to formula (4) :

[0025] Let , , ;

[0026] Take the limit: , , ;

[0027] (4)

[0028] Wherein, is the identity matrix, represents the control gain matrix of the th iteration, represents the parameter for adjusting the matrix during the iteration process, represents the positive definite solution of the th iteration, ;

[0029] For other cases of :

[0030] (5)

[0031] Wherein, is the eigenvalue of the matrix , is the real part of the eigenvalue of the matrix , is a matrix for adjusting and greater than real numbers;

[0032] For each matrix satisfies:

[0033]

[0034] Thus, an offline policy learning algorithm is constructed.

[0035] In an implementation manner of the present application, a planning performance function is constructed to constrain the offline policy learning algorithm, and the offline policy learning algorithm is processed based on the planning performance function to determine the initial state of the online policy learning algorithm, specifically including:

[0036] Construct a smooth function with a positive decreasing property , the smooth function has the formula:

[0037] (7)

[0038] where , , satisfies the condition , and respectively serve as the lower bound and the upper bound to limit the state of the adaptive cruise control system , is a positive constant greater than 0;

[0039] Based on and limit the state of the adaptive cruise control system to obtain the transformed state of the adaptive cruise control system .

[0040] In an implementation manner of the present application, an online policy learning algorithm is constructed, specifically including:

[0041] Based on formula (2) and the transformed state of the adaptive cruise control system , define , define , and define the cost function as ;

[0042] Incorporate two exploration signals and for collecting data into the control input and external disturbance respectively, then rewrite formula (1) as:

[0043] (8)

[0044] Deriving the cost function gives:

[0045] (9)

[0046] where is the cost function of the adaptive cruise control system, representing the target performance that the adaptive cruise control system aims to achieve, represents the instantaneous cost of the adaptive cruise control system under the combined action of the state , control input and disturbance input ;

[0047] To collect data of the adaptive cruise control system, the following matrices are defined:

[0048]

[0049]

[0050]

[0051] (10)

[0052] where, ;

[0053] Define a with full column rank. According to formula (11), we get , and ;

[0054] (11)

[0055] where, represents the vector containing the estimated solution of the equation, and respectively represent the vectorized forms of the control gain matrix and the disturbance gain matrix ;

[0056] Define as a positive constant greater than 0, and , then:

[0057] (12)

[0058] where, are the dimensions of the state , control input and disturbance input of the adaptive cruise control system respectively;

[0059] Construct a preliminary online policy learning algorithm based on the above steps; among them, the preliminary online policy learning algorithm needs to determine an initial matrix ;

[0060] Remove the initial matrix to rewrite formula (1) to determine the online policy learning algorithm.

[0061] In an implementation manner of this application, rewrite formula (1) to determine the online policy learning algorithm, which specifically includes:

[0062] Rewrite formula (1) to generate formula (13):

[0063] (13)

[0064] Among them, , is the adjusted adaptive cruise control system state matrix to describe the dynamic change of the system state;

[0065] Define and , take the derivative of formula (9) and rewrite it to obtain formula (14):

[0066] (14)

[0067] Solve the problem of containing and and information through the matrix , and information, use to replace , and then select to obtain formula (15):

[0068] (15)

[0069] Among them, , is a positive number used to adjust the weight matrix , and respectively represent the maximum eigenvalue and the minimum eigenvalue of the matrix;

[0070] Based on the above steps, use metric to sort the adaptive cruise control data and import the processed data into the matrix tool; among them, , , and After standardization in [context], the formula (16) is obtained:

[0071]

[0072]

[0073]

[0074]

[0075] (16)

[0076] Through the full column rank matrix and , solve :

[0077] (17)

[0078] Define the full column rank matrix:

[0079] , , to calculate , and :

[0080] (18)

[0081] Based on the above steps to obtain the online policy learning algorithm.

[0082] In one implementation of this application, based on the online policy learning algorithm, solve the H-infinity control problem to determine the optimal control strategy, specifically including:

[0083] Initialize the online policy learning algorithm to make , , and ;

[0084] Repeat the control input and the disturbance to collect the data of the adaptive cruise control system until ;

[0085] When satisfies , then solve according to formula (17);

[0086] Solve , and according to formula (18);

[0087] Select Meet the conditions of formula (15) and let Until And ;

[0088] Suppose , , And ;

[0089] Solve according to formula (3) ;

[0090] Set , , collect the information of the adaptive cruise control system until the condition of formula (12) is met;

[0091] Solve according to formula (11) , And , until And Return to obtain , , , to determine the optimal control strategy.

[0092] In a second aspect, an embodiment of the present application further provides an adaptive cruise control device based on online adaptive learning, characterized in that the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: construct a mathematical model of the adaptive cruise control system and define an H-infinity control problem related to the adaptive cruise control system; define the state of the adaptive cruise control system based on a preset performance function algorithm and construct an online policy learning algorithm; solve the H-infinity control problem based on the online policy learning algorithm to determine the optimal control strategy.

[0093] In a third aspect, an embodiment of the present application further provides a non-volatile computer storage medium for adaptive cruise control based on online adaptive learning, storing computer-executable instructions, characterized in that the computer-executable instructions are set to: construct a mathematical model of the adaptive cruise control system and define an H-infinity control problem related to the adaptive cruise control system; define the state of the adaptive cruise control system based on a preset performance function algorithm and construct an online policy learning algorithm; solve the H-infinity control problem based on the online policy learning algorithm to determine the optimal control strategy.

[0094] An adaptive cruise control method, device and medium based on online adaptive learning provided by an embodiment of the present application have at least the following technical effects:

[0095] A mathematical model of the adaptive cruise control system is constructed, and the relevant H-infinity control problem is clearly defined, laying a solid theoretical foundation for the performance optimization of the system. An online policy learning algorithm with real-time learning and optimization capabilities is constructed. By continuously analyzing real-time data, the algorithm can dynamically adjust the control strategy to adapt to the changing traffic environment. This adaptability and flexibility enable the system to track the vehicle ahead more stably and accurately, effectively avoiding potential safety hazards caused by sudden changes in traffic conditions. By applying control theory and optimization algorithms, the optimal performance of the adaptive cruise system can be determined, ensuring that the vehicle maintains a safe distance during cruising while achieving high energy efficiency. It not only improves the comfort and convenience of driving but also enhances the robustness and stability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0096] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0097] Figure 1 It is a flowchart of an adaptive cruise control method based on online adaptive learning provided by an embodiment of the present application;

[0098] Figure 2 It is a structural diagram of the adaptive cruise control system of the present application;

[0099] Figure 3 It is a convergence effect diagram of the adaptive cruise control system of the present application when there is PPF;

[0100] Figure 4 It is a convergence effect diagram of the adaptive cruise control system of the present application when there is no PPF;

[0101] Figure 5 It is the maximum change rate in the present application and the numerical value of the policy calculation;

[0102] Figure 6 It is in the online policy learning algorithm of the present application the policy error;

[0103] Figure 7 It is in the online policy learning algorithm of the present application the policy error;

[0104] Figure 8 It is in the online policy learning algorithm of the present application the policy error;

[0105] Figure 9Effect diagram of the embodiment of the present application;

[0106] Figure 9 Among them, (a) is the comparison diagram of the actual distance and the ideal distance in the embodiment of the present application;

[0107] Figure 9 Among them, (b) is the comparison diagram of the speed of the host vehicle and the target vehicle in the embodiment of the present application;

[0108] Figure 10 Effect diagram of the prior art;

[0109] Figure 10 Among them, (a) is the comparison diagram of the actual distance and the ideal distance of the prior art;

[0110] Figure 10 Among them, (b) is the comparison diagram of the speed of the host vehicle and the target vehicle of the prior art;

[0111] Figure 11 Schematic diagram of the trajectory error between the embodiment of the present application and the prior art under PPF limitation;

[0112] Figure 11 Among them, (a) is the schematic diagram of the trajectory error of the embodiment of the present application under PPF limitation;

[0113] Figure 11 Among them, (b) is the schematic diagram of the trajectory error of the prior art under PPF limitation;

[0114] Figure 12 Comparison diagram of the integral absolute error of the error between an adaptive cruise control method based on online adaptive learning provided by the embodiment of the present application and the prior art;

[0115] Figure 13 Schematic diagram of the internal structure of an adaptive cruise control device based on online adaptive learning provided by the embodiment of the present application. Detailed implementation manners

[0116] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0117] The embodiment of the present application provides an adaptive cruise control method, device and medium based on online adaptive learning to solve the following technical problems: how to improve the adaptability and robustness of the ACC system in a complex dynamic environment.

[0118] The technical solutions proposed in the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0119] Figure 1 It is a flowchart of an adaptive cruise control based on online adaptive learning provided for the embodiments of the present application. As Figure 1 shown, an adaptive cruise control method based on online adaptive learning provided for the embodiments of the present application specifically includes the following steps:

[0120] Step 1: Construct a mathematical model of the adaptive cruise control system and define the H-infinity control problem related to the adaptive cruise control system.

[0121] Figure 2 It is the structural diagram of the adaptive cruise control system.

[0122] First of all, it is necessary to construct a mathematical model of the adaptive cruise control system.

[0123] The formula of the mathematical model of the adaptive cruise control system is:

[0124] (1)

[0125] Among them, matrix A is the state matrix, matrix B is the input matrix, matrix D is the disturbance matrix, is the state vector, is the control input, is the external disturbance, represents the headway time difference, that is, the time interval maintained between the host vehicle and the target vehicle, represents the time constant related to the dynamic response of the vehicle, , , represents the ideal stopping distance of the host vehicle, represents the actual stopping distance of the host vehicle, , represents the vehicle speed of the host vehicle, is the speed of the target vehicle, , is the acceleration of the host vehicle, , is the input desired acceleration, w is the disturbance of the adaptive cruise control system.

[0126] Furthermore, construct the H-infinity control problem related to the mathematical model of the adaptive cruise control system.

[0127] The formula of the H-infinity control problem is:

[0128] (2)

[0129] Among them, is the positive definite solution of this formula, and both Q and R are positive definite weight matrices. is a small positive real number. It is used to limit the output control of the adaptive cruise control system by external disturbances within a controllable range.

[0130] Step 2: Define the state of the adaptive cruise control system based on the preset performance function algorithm, and construct an online policy learning algorithm.

[0131] First, construct an offline policy learning algorithm.

[0132] The offline policy learning algorithm is represented by the following formula:

[0133] Set an initial stable matrix ;

[0134] Based on , solve according to formula (3):

[0135] (3)

[0136] where is the matrix describing the dynamic changes of the adaptive cruise control system, is the control gain matrix for calculating the control input according to the state of the adaptive cruise control system, is the disturbance gain matrix for calculating the disturbance input according to the state of the adaptive cruise control system, and are the state and control right matrices respectively;

[0137] Solve the initial optimal solution according to formula (4):

[0138] Let , , ;

[0139] Take the limit: , , ;

[0140] (4)

[0141] where is the identity matrix, represents the control gain matrix of the th iteration, represents the parameter for adjusting the matrix during the iteration process, represents the positive definite solution of the th iteration, ;

[0142] For other cases of:

[0143] (5)

[0144] where is the eigenvalue of the matrix , is the real part of the eigenvalue of the matrix , is a real number used to adjust the matrix and greater than ;

[0145] For each matrix satisfies:

[0146]

[0147] Thus, an offline policy learning algorithm is constructed.

[0148] Furthermore, a planning performance function is constructed to constrain the offline policy learning algorithm, and the offline policy learning algorithm is processed based on the planning performance function to determine the initial state of the online policy learning algorithm;

[0149] Construct a smooth function with a positive decreasing property . The formula of the smooth function

[0150] (7)

[0151] where , , satisfying the condition , and respectively serve as the lower bound and upper bound to limit the state of the adaptive cruise control system , is a positive constant greater than 0;

[0152] Based on and limit the state of the adaptive cruise control system to obtain the transformed state of the adaptive cruise control system .

[0153] In a specific example, set the parameters , , . When there are and restrictions, for the state of the adaptive cruise control system The limiting effect is as Figure 3 shown below.

[0154] When there is no and limiting the adaptive cruise control system state the adaptive cruise control system state 1, 2, 3's tracking error is as Figure 4 shown below.

[0155] Furthermore, an online policy learning algorithm is constructed based on the initial state.

[0156] Based on formula (2) and the transformed adaptive cruise control system state , define , define , define the cost function as ;

[0157] Incorporate the two exploration signals and used to collect data into the control input and external interference respectively, then rewrite formula (1) as:

[0158] (8)

[0159] Take the derivative of the cost function to get:

[0160] (9)

[0161] where is the cost function of the adaptive cruise control system, representing the target performance that the adaptive cruise control system is required to achieve, represents the instantaneous cost of the adaptive cruise control system under the combined action of the state , control input and interference input ;

[0162] To collect data of the adaptive cruise control system, define the following matrices:

[0163]

[0164]

[0165]

[0166] (10)

[0167] where, ;

[0168] Define full column rank , according to formula (11), we get , and ;

[0169] (11)

[0170] where represents the vector containing the solutions of the equations and the estimated values, and represent the vectorized forms of the control gain matrix and the disturbance gain matrix respectively;

[0171] Define as a positive constant greater than 0, and , then:

[0172] (12)

[0173] where are the dimensions of the state , the control input and the disturbance input of the adaptive cruise control system respectively;

[0174] Based on the above steps, a preliminary online policy learning algorithm is constructed; among them, the preliminary online policy learning algorithm needs to determine the initial matrix ;

[0175] Remove the initial matrix to rewrite formula (1) to generate formula (13):

[0176] (13)

[0177] where , is the adjusted state matrix of the adaptive cruise control system to describe the dynamic changes of the system state;

[0178] Define and , take the derivative and rewrite formula (9) to get formula (14):

[0179] (14)

[0180] Solve the middle through the matrix and contains , and For the problem of information, use to replace , and then select to obtain formula (15):

[0181] (15)

[0182] where , is a positive number used to adjust the weight matrix , and represent the maximum eigenvalue and the minimum eigenvalue of the matrix respectively;

[0183] Based on the above steps, use the metric to sort the adaptive cruise control data and import the processed data into the matrix tool; where , , and After normalization in

[0184]

[0185]

[0186]

[0187]

[0188] (16)

[0189] Through the column full rank matrices and , solve :

[0190] (17)

[0191] Define the column full rank matrix:

[0192] , , to calculate , and :

[0193] (18)

[0194] Based on the above steps to obtain the online policy learning algorithm.

[0195] Step 3: Solve the H-infinity control problem based on the online policy learning algorithm to determine the optimal performance of the adaptive cruise control system.

[0196] First, initialize the online policy learning algorithm to make , , and ;

[0197] Furthermore, repeat the control input and the disturbance to collect the data of the adaptive cruise control system until ;

[0198] Furthermore, when satisfies , then solve according to Equation (17);

[0199] Furthermore, solve , and according to Equation (18);

[0200] Furthermore, select that satisfies the condition of Equation (15), and make until and ;

[0201] Furthermore, let , , and ;

[0202] Furthermore, solve according to Equation (3);

[0203] Furthermore, set , , and collect the information of the adaptive cruise control system until it satisfies Equation (12);

[0204] Furthermore, solve , and , until and Return to obtain , , to determine the optimal control strategy.

[0205] In a specific example, the experimental vehicle is equipped with sensors such as millimeter-wave radar, lidar, and cameras to collect perception information of the surrounding environment. The on-vehicle computing unit is equipped with a multi-modal environment perception fusion algorithm to accurately measure the vehicle's position, speed, and acceleration information. The time constant of the lower controller of the adaptive cruise control system is set to , the time headway . Set a non-zero initial system state , and the goal of the control problem in adaptive cruise control is to drive the state vector to zero by implementing the designed control law. Select the identity matrix as and , set , and the solution of formula (3) is obtained as:

[0206]

[0207] P* is the solution of the offline policy algorithm, and the main role of P* is to determine whether the result of the online policy learning algorithm is consistent with the solution of the offline policy learning algorithm.

[0208] Obtain the initial matrix through the online policy learning algorithm, set , , and introduce the exploration signals and into the online policy learning algorithm. Increase by 0.1 in each iteration to make the calculation result of greater than 0. Once the condition is satisfied, continue to use the online policy learning algorithm to solve . Figure 5 is the maximum rate of change and .

[0209] Select as the steady-state feedback matrix of the online policy learning algorithm . Figures 6 to 8 The calculation result in shows the convergence of , and in the algorithm.

[0210] The learning result of the tenth iteration is:

[0211]

[0212] Based on the above iteration results, the approximation error The solution obtained from the online policy learning algorithm is close to the solution derived from Equation (3), which indicates that it is feasible to use the online policy learning-based method to solve the model-free H∞ control problem in the adaptive cruise control system.

[0213] Based on the learning results of the proposed online policy learning algorithm, the tracking trajectory of the host vehicle is as Figure 9 and Figure 10 shown. The superiority of the tracking effect of the method designed in this application applied to the ACC system compared with the tracking effect of the reinforcement learning method adopted by the existing invention applied to the ACC system is demonstrated through the trajectory diagrams of the distance and speed between the host vehicle and the target vehicle.

[0214] To further demonstrate the superiority of the proposed control method in terms of accuracy, it is compared with the reinforcement learning algorithm in the prior art. The distance error x1, speed error x2, and acceleration x3 of the host vehicle are as Figure 11 and Figure 12 shown:

[0215] It can be seen from the above results that the algorithm designed in the embodiments of this application has excellent tracking accuracy. The designed and adopted model-free control method can learn from real-time input and output data, providing higher robustness and adaptability, especially for systems with uncertainties and dynamic changes.

[0216] The above is the method embodiment proposed in this application. Based on the same inventive concept, the embodiments of this application also provide an adaptive cruise control device based on online adaptive learning, and its structure is as Figure 13 shown.

[0217] Figure 13 FIG. is a schematic internal structure diagram of an adaptive cruise control device based on online adaptive learning provided by an embodiment of this application. As Figure 13 shown, the device includes:

[0218] At least one processor 201;

[0219] And a memory 202 communicatively connected to the at least one processor;

[0220] Wherein, the memory 202 stores instructions executable by the at least one processor. The instructions are executed by the at least one processor 201 so that the at least one processor 201 can:

[0221] Construct a mathematical model of an adaptive cruise control system and define the H-infinity control problem related to the adaptive cruise control system; construct a mathematical model of the adaptive cruise control system and define the H-infinity control problem related to the adaptive cruise control system; define the state of the adaptive cruise control system based on a preset performance function algorithm and construct an online policy learning algorithm; solve the H-infinity control problem based on the online policy learning algorithm to determine the optimal control strategy.

[0222] Some embodiments of the present application provide a non-volatile computer storage medium corresponding to Figure 1 an adaptive cruise control based on online adaptive learning, storing computer-executable instructions, and the computer-executable instructions are set as:

[0223] Construct a mathematical model of an adaptive cruise control system and define the H-infinity control problem related to the adaptive cruise control system; define the state of the adaptive cruise control system based on a preset performance function algorithm and construct an online policy learning algorithm; solve the H-infinity control problem based on the online policy learning algorithm to determine the optimal control strategy.

[0224] The various embodiments in the present application are described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the Internet of Things devices and media, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.

[0225] The systems and media provided by the embodiments of the present application correspond one-to-one with the methods. Therefore, the systems and media also have beneficial technical effects similar to the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the systems and media will not be repeated here.

[0226] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0227] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0228] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction means that implements the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0229] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0230] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0231] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.

[0232] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0233] It should also be noted that the terms "comprises", "comprising", and any other variants are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0234] The above description is only for the embodiments of the present application and is not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. An adaptive cruise control method based on online adaptive learning, characterized in that, The method includes: Construct an adaptive cruise control system mathematical model and define the H-infinity control problem for the adaptive cruise control system; Define the state of the adaptive cruise control system based on a preset performance function algorithm and construct an online policy learning algorithm; Solve the H-infinity control problem based on the online policy learning algorithm to determine the optimal control strategy; Define the state of the adaptive cruise control system based on a preset performance function algorithm and construct an online policy learning algorithm, specifically including: Construct an offline policy learning algorithm; Construct a planning performance function to constrain the offline policy learning algorithm and process the offline policy learning algorithm based on the planning performance function to determine the initial state of the online policy learning algorithm; Construct the online policy learning algorithm based on the initial state; Construct a planning performance function to constrain the offline policy learning algorithm and process the offline policy learning algorithm based on the planning performance function to determine the initial state of the online policy learning algorithm, specifically including: Construct a smooth function with the property of positive decreasing , the smooth function has the formula: (7) Among them, is the time, , , satisfying the condition that ; and respectively serve as the lower bound and the upper bound to restrict the state of the adaptive cruise control system , is a positive constant greater than 0; Based on and restrict the state of the adaptive cruise control system to obtain the transformed state of the adaptive cruise control system .

2. The adaptive cruise control method based on online adaptive learning according to claim 1, wherein, Establish an adaptive cruise control system mathematical model and define the H-infinity control problem for the adaptive cruise control system, specifically including: The formula of the adaptive cruise control system mathematical model is: (1) Among them, matrix A is the state matrix, matrix B is the input matrix, and matrix D is the disturbance matrix. is the state vector, is the control input, is the external disturbance, represents the headway time difference, that is, the time interval maintained between the host vehicle and the target vehicle. represents the time constant related to the dynamic response of the vehicle. , , represents the ideal stopping distance of the host vehicle. represents the actual stopping distance of the host vehicle. , represents the vehicle speed of the host vehicle. is the speed of the target vehicle. , is the acceleration of the host vehicle. , is the input desired acceleration, and w is the disturbance of the adaptive cruise control system. The formula of the H-infinity control problem for the adaptive cruise control system is: (2) wherein, is the positive definite solution of this formula, * represents the number of iterations, both Q and R are positive definite weight matrices, is a small positive real number, and the is used to limit the output control of the adaptive cruise control system by external interference within a controllable range.

3. An adaptive cruise control method based on online adaptive learning according to claim 2, characterized in that, Construct an offline policy learning algorithm, specifically including: The offline policy learning algorithm is represented by the following formula: Set an initial stable matrix ; Based on , solve according to formula (3) : (3) Among them, is a matrix describing the dynamic changes of the adaptive cruise control system, is a control gain matrix for calculating the control input according to the state of the adaptive cruise control system, is a disturbance gain matrix for calculating the disturbance input according to the state of the adaptive cruise control system, and are the state and control right matrices respectively; Solve the initial optimal solution according to formula (4). : Set , , ; Take the limit: , , ; (4) Among them, is the identity matrix, represents the control gain matrix of the -th iteration, represents the parameter of the adjustment matrix during the iteration process, represents the positive definite solution of the -th iteration, ; For other cases of: (5) wherein is the matrix 's eigenvalue, is the real part of the eigenvalue of the matrix , and is a real number used to adjust the matrix and greater than . For each matrix satisfies: Thus, construct an offline policy learning algorithm.

4. An adaptive cruise control method based on online adaptive learning according to claim 3, characterized in that, Construct an online policy learning algorithm, specifically including: Based on formula (2) and the transformed state of the adaptive cruise control system , define , define , define the cost function as ; Two exploration signals for collecting data and are respectively incorporated into the control input and external interference, and then Equation (1) is rewritten as: (8) Take the derivative of the cost function to obtain: (9) wherein is the cost function of the adaptive cruise control system, representing the target performance to be achieved by the adaptive cruise control system, represents the instantaneous cost of the adaptive cruise control system under the combined action of the state , control input and disturbance input ; To collect data of the adaptive cruise control system, define the following matrices: (10) Among them, ; Define column full rank , according to formula (11), obtain and ; (11) Among them, denotes the inclusion of the solution of the equation the vector of estimated values, and respectively represent the vectorized forms of the control gain matrix and the disturbance gain matrix ; Definition is a positive constant greater than 0, and , then: (12) Among them, are the states of the transformed adaptive cruise control system , the control input and the disturbance input dimensions; Based on the above steps to construct a preliminary online policy learning algorithm; wherein, the preliminary online policy learning algorithm needs to determine an initial matrix ; Remove the initial matrix , to rewrite formula (1) to determine the online policy learning algorithm.

5. An adaptive cruise control method based on online adaptive learning according to claim 4, characterized in that, Rewrite formula (1) to determine the online policy learning algorithm, specifically including: Rewrite formula (1) to generate formula (13): (13) Among them, , is the adjusted state matrix of the adaptive cruise control system to describe the dynamic changes of the system state; Definition and , by taking the derivative and rewriting formula (9), formula (14) is obtained: (14) Through the matrix In progress and Containing , and For the problem of information, use Instead of , and then select to obtain formula (15): (15) Among them, , is a positive number used to adjust the weight matrix , and represent the maximum eigenvalue and the minimum eigenvalue of the matrix respectively; Based on the above steps, use metrics to organize the adaptive cruise control data and import the processed data into a matrix tool; among them, , , and After standardization in, the formula (16) is obtained: (16) Through a column full-rank matrix and , solve : (17) Define a column full-rank matrix: , , , to calculate , and : (18) Based on the above steps, obtain the online policy learning algorithm.

6. The adaptive cruise control method based on online adaptive learning according to claim 5, characterized in that, Solve the H-infinity control problem based on the online policy learning algorithm to determine the optimal control strategy, specifically including: Initialize the online policy learning algorithm to make , , and ; Repeat control input and interference Collect adaptive cruise control system data so that ; When is satisfied , solve according to formula (17) ; Solve according to formula (18) , and ; Select Meet the conditions of formula (15) and let Until And ; Let , , and ; Solve according to formula (3) ; Set , , collect the information of the adaptive cruise control system until the formula (12) is satisfied; Solve according to formula (11) , and until and return and obtain , , to determine the optimal control strategy.

7. An adaptive cruise control device based on online adaptive learning, characterized in that, The device includes: At least one processor; And a memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can: Construct an adaptive cruise control system mathematical model and define the H-infinity control problem for the adaptive cruise control system; Define the state of the adaptive cruise control system based on a preset performance function algorithm and construct an online policy learning algorithm; Solve the H-infinity control problem based on the online policy learning algorithm to determine the optimal control strategy; Define the state of the adaptive cruise control system based on a preset performance function algorithm and construct an online policy learning algorithm, specifically including: Construct an offline policy learning algorithm; Construct a planning performance function to constrain the offline policy learning algorithm and process the offline policy learning algorithm based on the planning performance function to determine the initial state of the online policy learning algorithm; Construct the online policy learning algorithm based on the initial state; Construct a planning performance function to constrain the offline policy learning algorithm, and process the offline policy learning algorithm based on the planning performance function to determine the initial state of the online policy learning algorithm, specifically including: Construct a smooth function with the property of positive decreasing , where the smooth function has the formula: (7) wherein, is time, , , satisfying the condition of ; and respectively serve as the lower bound and the upper bound to restrict the state of the adaptive cruise control system , is a positive constant greater than 0; Based on and restrict the state of the adaptive cruise control system to obtain the transformed state of the adaptive cruise control system .

8. A non - volatile computer storage medium for adaptive cruise control based on online adaptive learning, storing computer - executable instructions, characterized in that, The computer-executable instructions are set as: Construct a mathematical model of the adaptive cruise control system and define the H-infinity control problem related to the adaptive cruise control system; Define the state of the adaptive cruise control system based on a preset performance function algorithm and construct an online policy learning algorithm; Solve the H-infinity control problem based on the online policy learning algorithm to determine the optimal control strategy; Define the state of the adaptive cruise control system based on a preset performance function algorithm and construct an online policy learning algorithm, specifically including: Construct an offline policy learning algorithm; Construct a planning performance function to constrain the offline policy learning algorithm, and process the offline policy learning algorithm based on the planning performance function to determine the initial state of the online policy learning algorithm; Construct the online policy learning algorithm based on the initial state; Construct a planning performance function to constrain the offline policy learning algorithm, and process the offline policy learning algorithm based on the planning performance function to determine the initial state of the online policy learning algorithm, specifically including: Construct a smooth function with a positive decreasing property , the smooth function has the formula:[[]] (7) wherein, is time, , , satisfying condition, and respectively serve as the lower bound and the upper bound to restrict the state of the adaptive cruise control system , is a positive constant greater than 0; Based on and limit the state of the adaptive cruise control system to obtain the transformed state of the adaptive cruise control system .

Citation Information

Patent Citations

  • Adaptive vehicle following algorithm based on improved model prediction control

    CN107808027A

  • Adaptive dynamic programming method of aircraft engine in optimal acceleration tracking control

    CN110821683A

  • Automatic driving speed control framework based on spatio-temporal data reinforcement learning

    CN113741464A

  • Self-adaptive cruise vehicle speed control method, system and device and storage medium

    CN117774971A