A Multi-Agent Following Control Method and System Based on Adaptive Dynamic Programming

Through adaptive dynamic programming method and neural network fitting technology, the problem of optimal control in multi-agent systems is solved, and the optimal follow-up control in linear and nonlinear systems is realized.

CN115755615BActive Publication Date: 2025-07-11GUANGZHOU INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211501456.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-28
Publication Date
2025-07-11
Estimated Expiration
2042-11-28

AI Technical Summary

Technical Problem

The conventional control methods of existing multi-agent systems cannot achieve optimal control, especially in nonlinear systems, and the optimal control strategy cannot be solved. The conventional methods mainly use algebraic calculation rather than data-driven methods.

Method used

Adaptive dynamic programming is adopted to solve the optimal control strategy by defining utility functions and cost functions, using iterative methods and neural network fitting technology, which is suitable for linear and nonlinear systems.

Benefits of technology

It realizes that followers follow with minimum trajectory error and control energy in multi-agent systems, and is suitable for isomorphic and heterogeneous systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115755615B_ABST
    Figure CN115755615B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a multi-agent following control method and system based on adaptive dynamic programming. According to the differences in the states and control quantities of the followers and the leader, the following state error and control quantity error are obtained. A utility function is defined with the goal of minimizing the following state error and the consumed energy, and a cost function is obtained based on the utility function. The optimal control strategy is solved with the idea of dynamic programming. Since both the cost function and the control strategy are non-explicitly expressed, an action neural network and an evaluation neural network are used to respectively fit the control strategy and the cost function, and the optimal control strategy is solved by means of iterative calculation. The action neural network and the evaluation neural network are trained with the collected state values and control quantity values of the leader and the followers, which can enable the followers to achieve following motion of the leader with the minimum trajectory error and control energy, and is applicable not only to the followers of linear systems but also to the followers of nonlinear systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the technical field of multi-agent following, and in particular, to a multi-agent following control method and system based on adaptive dynamic programming. Background Art

[0002] With the development of technology, agents have attracted great attention from researchers in recent decades, and they have potential research value in various aspects such as communication, computer technology, biology, and social behavior. With the gradual arrival of the intelligent era, agents are also widely used in all aspects of life. However, as the application functions become increasingly complex, a multi-agent system composed of multiple simple agents has greater advantages than a single agent. Multi-agent systems are widely used in various fields such as military, aerospace, and industry. For example, the formation flight of unmanned aerial vehicles, the coordinated operation of multiple satellites, and the formation transportation of intelligent vehicles. Therefore, the coordinated control of multi-agent systems has received extensive research.

[0003] Almost all existing conventional control methods for multi-agents do not consider the cost function and cannot achieve optimal control. The currently used optimal control methods mainly use algebraic calculation methods to solve the optimal control strategy, rather than data-driven methods. The currently used optimal control methods mainly target linear systems and cannot solve the optimal control strategy of non-linear systems. Summary of the Invention

[0004] The embodiments of the present invention provide a multi-agent following control method and system based on adaptive dynamic programming, which uses a data-based adaptive dynamic programming method rather than an analytical formula method to solve the optimal control strategy, enabling the follower to achieve following motion of the leader with the smallest trajectory error and control energy.

[0005] In a first aspect, the embodiments of the present invention provide a multi-agent following control method based on adaptive dynamic programming. The multi-agents include a leader and at least one follower. The following control method includes:

[0006] Step S1: Determine the state equation and control quantity of the follower based on the state quantity of the leader, determine the following state error between the follower and the leader based on the state equation of the follower, and the control quantity error of the follower.

[0007] Step S2: Determine the utility function with the goal of minimizing the following state error and energy consumption of the follower, determine the cost function that minimizes the energy of the follower based on the utility function, and determine the control error function based on the cost function.

[0008] Step S3: Solve the control error function based on an iterative method, fit the control error function and the cost function based on a preset neural network, so as to gradually fit out the optimal control error function during the iteration process, and determine the optimal control strategy based on the optimal control error function.

[0009] Preferably, step S1 specifically includes:

[0010] Determine the state quantity ξ of the leader at the current moment and the state quantity ξ′ at the next moment; the state quantity λ of the follower at the current moment and the state quantity λ′ at the next moment; determine the state equation of the follower:

[0011] λ′ = f(λ) + g(λ)v(λ)

[0012] In the above formula, f(·) is the state coupling function, g(·) is the input coupling function, v(·) is the control strategy function of the follower, and v(λ) is the control quantity under the state quantity of the follower at the current moment;

[0013] Determine the desired control quantity v e as:

[0014] v e = g -1 (ξ)(ξ′ - f(ξ))

[0015] In the above formula, v e is the desired control quantity of the follower, and g -1 (ξ) is the transpose of the input coupling function value;

[0016] The following state error x between the follower and the leader at the current moment is:

[0017] x = λ - ξ

[0018] The control quantity error u(x) between the control quantity v(λ) of the follower under the state quantity at the current moment and the desired control quantity v e is:

[0019] u(x) = v(λ) - v e

[0020] = v(x + ξ) - v e

[0021] The following state error x′ of the follower at the next moment is:

[0022] x′ = λ′ - ξ′

[0023] = f(λ) + g(λ)v - ξ′

[0024] = f(x + ξ) + g(x + ξ)(u(x) + v e ) - ξ′

[0025] The following state error x' of the follower at the next moment is a function of the following state error x at the current moment and the control quantity error u(x), denoted as:

[0026] x' = F(x, u(x))

[0027] In the above formula, F(·) represents the mapping from x and u(x) to x'.

[0028] Preferably, the step S2 specifically includes:

[0029] Determine the utility function of the following state error as:

[0030] U(x, u(x)) = x T Qx + u(x) T Ru(x)

[0031] In the above formula, U(x, u(x)) represents the utility function when the following state error is x and the control quantity error is u(x), and Q and R are both positive definite matrices;

[0032] Determine the cost function of the follower based on the utility function:

[0033] V(x, u(x)) = ∑U(x, u(x)) = (x T Qx + u(x) T Ru(x)) + (x' T Qx' + u' T (x')Ru'(x')) + …

[0034] In the above formula, V(·) is the cost function, U(·) is the utility function; x' is the following state error of the follower at the next moment, and u'(x') is the control quantity error of the follower at the next moment;

[0035] Based on the Bellman optimality principle:

[0036]

[0037] In the above formula, V * (x) is the optimal cost function under the following state error x at the current moment, and V * (x) is the optimal cost function under the following state error x' at the next moment, and min{·} represents finding the minimum value of the function in the curly brackets;

[0038] The optimal control error function u * (x) is:

[0039]

[0040] In the formula, u* $(x)$ represents the control amount when the following state error $x$ is at the current moment, represents the value of $u(x)$ when the function in the curly brackets is minimized; $R$ -1 represents the inverse matrix of the positive definite matrix $R$, $g$ T $(x)$ represents the transpose of $g(x)$.

[0041] Preferably, in the step S3, when solving the control error function based on the iterative method, assume that the value of the initialization value function $V_0(·)$ is 0, and the initial control error function is solved as:

[0042]

[0043] The cost function $V$ at the $i$th ($i = 1, 2, 3, \cdots$) iteration i $(x)$ is:

[0044]

[0045] The control error function $u$ at the $i$th ($i = 1, 2, 3, \cdots$) iteration i $(x)$ is:

[0046]

[0047] Preferably, in the step S3, the control error function and the cost function are fitted based on a preset neural network to gradually fit the optimal control error function during the iteration process, specifically including:

[0048] Set the structures of the evaluation neural network and the action neural network. Both the evaluation neural network and the action neural network include an input layer, a hidden layer, and an output layer;

[0049] Take the following state error $x$ at the current moment as the input of the action neural network, and take the control amount error amount $u(x)$ as the output of the action neural network; Based on the output value of the action neural network at the $i$th iteration and the control amount error amount $u$ at the $i$th iteration i (x) to determine the first loss function of the action neural network;

[0050] Take the following state error $x$ at the previous moment as the input of the evaluation neural network, and take the cost function $V(x)$ as the output of the evaluation neural network; Based on the output value of the evaluation neural network at the $i$th iteration and $V$ at the $i$th iteration i (x) to determine the second loss function of the action neural network;

[0051] Train the action neural network and the evaluation neural network respectively based on the first loss function and the second loss function, obtain the optimal control error function by fitting based on the action neural network, obtain the optimal cost function by fitting based on the evaluation neural network, and determine the optimal control strategy based on the optimal control error function and the desired control quantity.

[0052] Preferably, the first loss function is:

[0053]

[0054]

[0055]

[0056] In the above formula, u i (x) is the output value of the action neural network at the i-th iteration, W1 a is the weight matrix of the hidden layer of the action neural network, b a is the bias matrix of the hidden layer of the action neural network, σ(·) is the activation function, is the weight matrix of the output layer of the action neural network; is the objective function of u i (x) at the i-th iteration; V i (·) is the cost function at the i-th iteration, represents the partial derivative of the cost function V i (x′) with respect to x′ at the i-th iteration;

[0057] The second loss function is:

[0058]

[0059]

[0060]

[0061] In the above formula, V i (·) represents the mapping function of the evaluation neural network at the i-th iteration, V i (x) represents the output value of the evaluation neural network at the i-th iteration, W1 c is the weight matrix of the hidden layer of the evaluation neural network, b c is the bias matrix of the hidden layer of the evaluation neural network, σ(·) is the activation function, is the weight matrix of the output layer of the evaluation neural network; V i o (x) is V i(x) objective function; u i-1 (x) represents the output value of the evaluation neural network at the (i - 1)-th iteration, x T represents the transpose of x, represents u i-1 transpose of (x), V i-1 (·) represents the mapping relationship of the evaluation neural network at the (i - 1)-th iteration, V i-1 (x′) represents the output value of the evaluation neural network with x′ input at the (i - 1)-th iteration.

[0062] Preferably, the optimal control strategy is:

[0063] v * (λ) = u * (x) + v e

[0064] In the above formula, v * (·) is the optimal control strategy of the follower, v * (λ) represents the control quantity of the follower under the state quantity λ at the current moment, v e represents the desired control quantity of the follower, u * (·) represents the optimal control error function of the follower, u * (x) represents the optimal control quantity error when the following state error of the follower is x at the current moment.

[0065] In a second aspect, an embodiment of the present invention provides a multi-agent following control system based on adaptive dynamic programming. The multi-agent includes a leader and at least one follower. The following control system includes:

[0066] A state quantity calculation module, which determines the state equation and control quantity of the follower based on the state quantity of the leader, determines the following state error between the follower and the leader based on the state equation of the follower, and the control quantity error of the follower;

[0067] A control error determination module, which determines the utility function with the following state error of the follower and the minimum energy consumption as the goal, determines the cost function that minimizes the energy of the follower based on the utility function, and determines the control error function based on the cost function;

[0068] An optimal control strategy module, which solves the control error function based on an iterative method, fits the control error function and the cost function based on a preset neural network to gradually fit out the optimal control error function during the iterative process, and determines the optimal control strategy based on the optimal control error function.

[0069] In a third aspect, an embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the multi-agent following control method based on adaptive dynamic programming as described in the embodiment of the first aspect of the present invention are implemented.

[0070] In a fourth aspect, an embodiment of the present invention provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the multi-agent following control method based on adaptive dynamic programming as described in the embodiment of the first aspect of the present invention are implemented.

[0071] A multi-agent following control method and system based on adaptive dynamic programming provided by an embodiment of the present invention obtain a following state error and a control quantity error according to the differences between the states and control quantities of followers and leaders. A utility function is defined with the goal of minimizing the following state error and the consumed energy, and a cost function is obtained according to the utility function. The optimal control strategy is solved with the idea of dynamic programming. Since both the cost function and the control strategy are non-explicitly expressed, an action neural network and a critic neural network are used to fit the control strategy and the cost function respectively, and the optimal control strategy is solved by an iterative calculation method. The action neural network and the critic neural network are trained with the collected state values and control quantity values of the leader and the followers, so that the followers can achieve following motion of the leader with the minimum trajectory error and control energy, which is applicable not only to followers of linear systems but also to followers of nonlinear systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0073] Figure 1 It is a flowchart of a multi-agent following control method based on adaptive dynamic programming according to an embodiment of the present invention;

[0074] Figure 2 It is an example diagram of multi-agent following motion according to an embodiment of the present invention;

[0075] Figure 3 It is a schematic diagram of the network structures of the critic neural network and the action neural network according to an embodiment of the present invention;

[0076] Figure 4 It is an adaptive dynamic programming block diagram according to an embodiment of the present invention;

[0077] Figure 5 Schematic diagram of an entity structure according to an embodiment of the present invention. Detailed implementation manners

[0078] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0079] In the embodiments of the present application, the term "and / or" merely describes an association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent three cases: A exists alone, both A and B exist simultaneously, and B exists alone.

[0080] The terms "first" and "second" in the embodiments of the present application are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present application, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a system, product, or device that includes a series of components or units is not limited to the listed components or units, but may optionally further include components or units not listed, or may optionally further include other components or units inherent to these products or devices. In the description of the present application, the meaning of "a plurality of" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0081] Referring to "embodiments" herein means that a specific feature, structure, or characteristic described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0082] Existing multi-agent following control technologies have the following disadvantages: (1) Conventional control methods do not consider the cost function and cannot achieve optimal control. (2) The currently used optimal control methods mainly use algebraic calculation methods to solve the optimal control strategy, rather than data-driven methods. (3) The currently used optimal control methods mainly target linear systems and cannot solve the optimal control strategy of nonlinear systems.

[0083] Therefore, the embodiments of the present invention provide a multi-agent following control method and system based on adaptive dynamic programming, which can enable the follower to achieve following motion of the leader with the minimum trajectory error and control energy. By using a data-based adaptive dynamic programming method instead of an analytical formula method to solve the optimal control strategy, it is applicable not only to followers of linear systems but also to followers of nonlinear systems; it is applicable not only to homogeneous multi-agent systems but also to heterogeneous multi-agent systems. The following will be elaborated and introduced through multiple embodiments.

[0084] Figure 1 and Figure 2 A multi-agent following control method based on adaptive dynamic programming provided by an embodiment of the present invention, wherein the multi-agent includes a leader and at least one follower, and the following control method includes:

[0085] Step S1: Determine the state equation and control quantity of the follower based on the state quantity of the leader, determine the following state error between the follower and the leader based on the state equation of the follower, and the control quantity error of the follower.

[0086] As Figure 2 shown, in a multi-agent system composed of one leader and multiple followers, wherein the leader moves autonomously, and followers a, b, and c follow the leader to move. The follower can obtain the state value of the leader at any time. In the figure, ξ represents the state quantity of the leader at the current moment, ξ′ represents the state quantity of the leader at the next moment; λ represents the state quantity of the follower at the current moment, λ′ represents the state quantity of the follower at the next moment, and v(λ) represents the control quantity of the follower at the current moment. x represents the following state error between the follower and the leader at the current moment, x′ represents the following state error between the follower and the leader at the next moment, and u(x) represents the control quantity error of the follower at the current moment.

[0087] The state quantity of the leader at the current moment is ξ, and the state quantity at the next moment is ξ′; the state quantity of the follower at the current moment is λ, and the state quantity of the follower at the next moment is λ′; at any time, the follower can obtain the state value of the leader, and determine the state equation of the follower:

[0088] λ′ = f(λ) + g(λ)v(λ) (1)

[0089] In the above formula, f(·) is the state coupling function, g(·) is the input coupling function, v(·) is the control strategy function of the follower, and v(λ) is the control quantity under the state quantity of the follower at the current moment; f(·) and g(·) are functions with known expressions, and v(·) is an unknown function.

[0090] If the follower is to achieve following motion of the leader, the desired control quantity v of the follower needs to be determinede is:

[0091] v e = g -1 (ξ)(ξ′ - f(ξ)) (2)

[0092] In the above formula, v e is the expected control quantity of the follower, and g -1 (ξ) is the transpose of the input coupling function value.

[0093] Let x represent the following state error between the current follower and the leader, and let u(x) represent the error between the control quantity v(λ) at the current moment state of the follower and the expected control quantity v e Then the calculation method of x is:

[0094] x = λ - ξ (3)

[0095] The control quantity error u(x) between the control quantity v(λ) at the current moment state of the follower and the expected control quantity v e is:

[0096]

[0097] The following state error x′ of the follower at the next moment is:

[0098]

[0099] In the above formula, λ′ is the state quantity of the follower at the next moment, ξ′ is the state quantity of the leader at the next moment, x represents the following state error of the follower at the current moment, and x′ represents the following state error of the follower at the next moment.

[0100] The following state error x′ of the follower at the next moment is a function of the following state error x and the control quantity error u(x) at the current moment, denoted as:

[0101] x′ = F(x, u(x)) (6)

[0102] In the above formula, F(·) represents the mapping from x and u(x) to x′.

[0103] Step S2: Determine the utility function with the following state error of the follower and the minimum energy consumption as the goal, determine the cost function that minimizes the energy of the follower based on the utility function, and determine the control error function based on the cost function.

[0104] Determine the utility function of the following state error as:

[0105] U(x, u(x)) = x T Qx + u(x) T Ru(x) (7)

[0106] In the above formula, U(x, u(x)) represents the utility function when the following state error is x and the control quantity error is u(x), and Q and R are both positive definite matrices;

[0107] Determine the cost function of the follower based on the utility function:

[0108] V(x, u(x)) = ∑U(x, u(x)) = (x T Qx + u(x) T Ru(x)) + (x′ T Qx′ + u′ T (x′)Ru′(x′)) + … (8)

[0109] In the above formula, V(·) is the cost function, and U(·) is the utility function; x′ is the following state error of the follower at the next moment, and u′(x′) is the control quantity error of the follower at the next moment;

[0110] The purpose of this embodiment is to find the optimal tracking control strategy u(·) of the follower, so that the follower can minimize the cost function V(x, u). Based on the Bellman optimality principle, there is:

[0111]

[0112] In the above formula, V * (x) is the optimal cost function under the following state error x at the current moment, and V * (x) is the optimal cost function under the following state error x′ at the next moment, and min{·} represents finding the minimum value of the function in the curly brackets;

[0113] The optimal control error function u * (x) is:

[0114]

[0115] In the formula, u * (x) represents the control quantity when the following state error x is at the current moment, represents the value of u(x) when the function in the curly brackets is minimized; R -1 represents the inverse matrix of the positive definite matrix R in formula (7), and g T (x) represents the transpose of g(x) in formula (1).

[0116] Step S3, solve the control error function based on the iterative method, fit the control error function and the cost function based on a preset neural network, so as to gradually fit the optimal control error function during the iteration process, and determine the optimal control strategy based on the optimal control error function.

[0117] Solve for V using an iterative method * (·) and u * When (·) and u(·), set the initial value of the function V0(·) to 0, and solve for the initial control error function as follows:

[0118]

[0119] The cost function V at the i-th (i = 1, 2, 3, …) iteration i (x) is:

[0120]

[0121] The control error function u at the i-th (i = 1, 2, 3, …) iteration i (x) is:

[0122]

[0123] When the iteration number i approaches infinity, the cost function V i (·) approaches its optimal function V i * (·), and the control policy function u i (·) approaches its optimal function u * (·). That is, when i → ∞, V i → V * , π i → π * .

[0124] Since in equations (12) and (13), the value function V(·) and the policy function u(·) are both non-analytical, the functional expressions of V(·) and u(·) cannot be explicitly expressed. Neural networks have very good non-linear fitting capabilities. To achieve an explicit expression of V i (·) in equation (12) and u i (·) in equation (13), an evaluation neural network and an action neural network are respectively used to fit these two functions. By updating the weight parameters in the evaluation neural network and the action neural network, the gradual improvement of V i (·) and u i (·) during the iteration process is achieved. Both the evaluation neural network and the action neural network adopt the structure of input layer - hidden layer - output layer (as Figure 3 shown). The data required for their training are the randomly sampled state value ξ of the leader at the current moment, the state value λ of the follower at the current moment, the current control quantity v(λ) of the follower, and the state value ξ′ of the leader at the next moment, the state value λ′ of the follower at the next moment, and are transformed into the current state error x, the next moment state error x′, and the current control quantity error u(x) according to equations (3) and (4).

[0125] Based on the above embodiments, as a preferred implementation manner, in step S3, the control error function and the cost function are fitted based on a preset neural network to gradually fit the optimal control error function during the iteration process. Specifically, it includes:

[0126] Set the structures of the evaluation neural network and the action neural network. Both the evaluation neural network and the action neural network include an input layer, a hidden layer, and an output layer;

[0127] Take the following state error x at the current moment as the input of the action neural network, and take the control quantity error u(x) as the output of the action neural network; based on the output value of the action neural network at the i-th iteration and the objective function of the control quantity error u i (x), determine the first loss function of the action neural network;

[0128] Take the following state error x at the previous moment as the input of the evaluation neural network, and take the cost function V(x) as the output of the evaluation neural network; based on the output value of the evaluation neural network at the i-th iteration and the objective function of V i (x), determine the second loss function of the action neural network;

[0129] Based on the first loss function and the second loss function, train the action neural network and the evaluation neural network respectively. Based on the action neural network, fit the optimal control error function. Based on the evaluation neural network, fit the optimal cost function. Based on the optimal control error function and the desired control quantity, determine the optimal control strategy.

[0130] Among them, the input of the action network is the state error x of the follower at the current moment, and the output is the control quantity error u(x) of the follower. Its mapping relationship is expressed as:

[0131]

[0132] In the above formula, u i (x) is the output value of the action neural network at the i-th iteration, W1 a is the weight matrix of the hidden layer of the action neural network, b a is the bias matrix of the hidden layer of the action neural network, σ(·) is the activation function, is the weight matrix of the output layer of the action neural network.

[0133] Set the loss function of the action neural network, that is, the first loss function as:

[0134]

[0135] In the above formula, is the objective function of u i (x) at the i-th iteration; it can be calculated according to formula (10):

[0136]

[0137] In the above formula, V i (·) is the cost function at the i-th iteration, represents the partial derivative of the cost function V i (x′) with respect to x′ at the i-th iteration;

[0138] The input of the evaluation neural network is x, and the output is the cost function V(x). Its mapping relationship is expressed as:

[0139]

[0140] In the above formula, V i (·) represents the mapping function of the evaluation neural network at the i-th iteration, and V i (x) represents the output value of the evaluation neural network at the i-th iteration. W1 c is the weight matrix of the hidden layer of the evaluation neural network, and b c is the bias matrix of the hidden layer of the evaluation neural network. σ(·) is the activation function, is the weight matrix of the output layer of the evaluation neural network; set the loss function of the evaluation neural network, that is, the second loss function is:

[0141]

[0142] Among them, V i o (x) is the objective function of V i (x) at the i-th iteration, and it can be calculated according to formula (12):

[0143]

[0144] In the above formula, u i-1 (x) represents the output value of the evaluation neural network at the (i - 1)-th iteration, and x T represents the transpose of x, represents the transpose of u i-1 (x), and V i-1 (·) represents the mapping relationship of the evaluation neural network at the (i - 1)-th iteration, and V i-1 (x′) represents the output value of the evaluation neural network with x′ as the input at the (i - 1)-th iteration.

[0145] Such as Figure 4As shown, the current state error x is input into the action neural network (action network in the figure), and the current control quantity error u(x) is obtained. x and u(x) are input into the state equation to obtain the state error x' at the next moment. x' is input into the evaluation network at the (i - 1)-th iteration (evaluation network in the figure) to obtain the output V i-1 (x′) of the evaluation neural network. Through formula (16), the target value of the action neural network is obtained and the action neural network is trained to make u i (x) and as close as possible. Through formula (19), the target value V i o (x) of the evaluation neural network is obtained. The evaluation neural network is trained to make the value of V i (x) and V i o (x) as close as possible.

[0146] In each iteration, formula (15) and formula (18) are used as the loss functions to train the action network and the evaluation network respectively. As the iteration number i increases, the action network gradually fits the optimal control strategy function u * (x), and the evaluation network gradually fits the optimal cost function V * (x). Finally, the optimal control strategy under the following error can be obtained according to formula (4):

[0147] v * (λ) = u * (x) + v e (20)

[0148] In the above formula, v * (·) is the optimal control strategy of the follower, v * (λ) represents the control quantity of the follower at the current moment state quantity λ, v e represents the expected control quantity of the follower, u * (·) represents the optimal control error function of the follower, and u * (x) represents the optimal control quantity error of the follower when the following state error at the current moment is x. After obtaining the optimal control strategy of each follower through formula (20), each follower can achieve the optimal following motion of the leader under the control of its respective optimal control strategy.

[0149] The embodiment of the present invention also provides a multi-agent following control system based on adaptive dynamic programming. The multi-agent includes a leader and at least one follower. Based on the multi-agent following control method based on adaptive dynamic programming in the above embodiments, it includes:

[0150] The state variable calculation module determines the state equation and control quantity of the follower based on the state quantity of the leader, determines the following state error between the follower and the leader based on the state equation of the follower, and the control quantity error of the follower;

[0151] The control error determination module determines the utility function with the goal of minimizing the following state error and energy consumption of the follower, determines the cost function that minimizes the energy of the follower based on the utility function, and determines the control error function based on the cost function;

[0152] The optimal control strategy module solves the control error function based on the iterative method, fits the control error function and the cost function based on a preset neural network, gradually fits the optimal control error function during the iteration process, and determines the optimal control strategy based on the optimal control error function.

[0153] Based on the same concept, the embodiment of the present invention also provides a schematic diagram of an entity structure, as Figure 5 shown. The server may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 complete communication with each other through the communication bus 840. The processor 810 may call the logical instructions in the memory 830 to execute the steps of the multi-agent following control method based on adaptive dynamic programming as described in the above embodiments. For example, it includes:

[0154] Step S1: Determine the state equation and control quantity of the follower based on the state quantity of the leader, determine the following state error between the follower and the leader based on the state equation of the follower, and the control quantity error of the follower;

[0155] Step S2: Determine the utility function with the goal of minimizing the following state error and energy consumption of the follower, determine the cost function that minimizes the energy of the follower based on the utility function, and determine the control error function based on the cost function;

[0156] Step S3: Solve the control error function based on the iterative method, fit the control error function and the cost function based on a preset neural network, gradually fit the optimal control error function during the iteration process, and determine the optimal control strategy based on the optimal control error function.

[0157] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0158] Based on the same concept, an embodiment of the present invention further provides a non-transitory computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program includes at least one segment of code. The at least one segment of code can be executed by a master device to control the master device to implement the steps of the multi-agent following control method based on adaptive dynamic programming as described in the above embodiments. For example, it includes:

[0159] Step S1: Determine the state equation and control quantity of the follower based on the state quantity of the leader, determine the following state error between the follower and the leader based on the state equation of the follower, and the control quantity error of the follower.

[0160] Step S2: Determine the utility function with the goal of minimizing the following state error and energy consumption of the follower, determine the cost function that minimizes the energy of the follower based on the utility function, and determine the control error function based on the cost function.

[0161] Step S3: Solve the control error function based on an iterative method, fit the control error function and the cost function based on a preset neural network, so as to gradually fit out the optimal control error function during the iterative process, and determine the optimal control strategy based on the optimal control error function.

[0162] Based on the same technical concept, an embodiment of the present application further provides a computer program. When the computer program is executed by a master device, it is used to implement the above method embodiments.

[0163] The program can be stored in whole or in part on a storage medium packaged together with the processor, or can be stored in whole or in part on a memory not packaged together with the processor.

[0164] Based on the same inventive concept, an embodiment of the present application further provides a processor for implementing the above method embodiment. The above processor may be a chip.

[0165] In summary, a multi-agent following control method and system based on adaptive dynamic programming provided by an embodiment of the present invention obtains a following state error and a control quantity error according to the differences in the states and control quantities of the followers and the leader. A utility function is defined with the goal of minimizing the following state error and the consumed energy, and a cost function is obtained according to the utility function. The optimal control strategy is solved with the idea of dynamic programming. Since both the cost function and the control strategy are non-explicitly expressed, an actor neural network and a critic neural network are used to fit the control strategy and the cost function respectively, and the optimal control strategy is solved by means of iterative calculation. The actor neural network and the critic neural network are trained with the collected state values and control quantity values of the leader and the followers, so that the followers can achieve following motion of the leader with the minimum trajectory error and control energy. It is applicable not only to the followers of linear systems but also to the followers of non-linear systems.

[0166] The various embodiments of the present invention can be combined arbitrarily to achieve different technical effects.

[0167] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. that incorporates one or more available media. The available media may be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media (such as solid state disks, SSDs), etc.

[0168] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by relevant hardware instructed by a computer program. This program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The foregoing storage media include: various media such as ROM or random access memory RAM, magnetic disks, or optical discs that can store program codes.

[0169] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-agent following control method based on adaptive dynamic programming, the multi-agent includes a leader and at least one follower, characterized in that, The following control method includes: Step S1: Determine the state equation and control quantity of the follower based on the state quantity of the leader, determine the following state error between the follower and the leader based on the state equation of the follower, and the control quantity error of the follower; Step S2: Determine the utility function with the goal of minimizing the following state error and energy consumption of the follower, determine the cost function that minimizes the energy of the follower based on the utility function, and determine the control error function based on the cost function; Step S3: Solve the control error function based on the iterative method, fit the control error function and the cost function based on a preset neural network, gradually fit the optimal control error function during the iteration process, and determine the optimal control strategy based on the optimal control error function In step S3, fitting the control error function and the cost function based on a preset neural network to gradually fit the optimal control error function during the iteration process specifically includes: Set the structures of the evaluation neural network and the action neural network, both of which include an input layer, a hidden layer, and an output layer; Using the following state error x at the current moment as the input of the action neural network, and using the control quantity error u(x) as the output of the action neural network; determining the first loss function of the action neural network based on the output value of the action neural network at the i-th iteration and the objective function of the control quantity error u i (x) The following state error x at a previous moment is the input of the evaluation neural network, and the cost function V(x) is the output of the evaluation neural network; based on the output value of the evaluation neural network at the i-th iteration and the objective function of V i (x) at the i-th iteration, determine the second loss function of the action neural network; Train the action neural network and the evaluation neural network based on the first loss function and the second loss function respectively, fit the optimal control error function based on the action neural network, fit the optimal cost function based on the evaluation neural network, and determine the optimal control strategy based on the optimal control error function and the desired control quantity; The first loss function is as follows: In the above formula, u i (x) is the output value of the action neural network at the i-th iteration, W1 a is the weight matrix of the hidden layer of the action neural network, b a is the bias matrix of the hidden layer of the action neural network, σ(·) is the activation function, is the weight matrix of the output layer of the action neural network; is the objective function of u i (x) at the i-th iteration; V i (·) is the cost function at the i-th iteration, represents the partial derivative of the cost function V i (x′) with respect to x′ at the i-th iteration; g(·) is the input coupling function; x′ is the following state error of the follower at the next moment; The second loss function is as follows: In the above formula, V i (x) represents the output value of the evaluation neural network at the i-th iteration. W1 c is the weight matrix of the hidden layer of the evaluation neural network, b c is the bias matrix of the hidden layer of the evaluation neural network, σ(·) is the activation function, is the weight matrix of the output layer of the evaluation neural network; V i o (x) is the objective function of V i (x) at the i-th iteration; u i-1 (x) represents the output value of the action neural network at the (i - 1)-th iteration, x T represents the transpose of x, represents the transpose of u i-1 (x), V i-1 (x′) represents the output value of the evaluation neural network at the (i - 1)-th iteration with x′ as the input; Q and R are both positive definite matrices.

2. The multi-agent following control method based on adaptive dynamic programming according to claim 1, characterized in that The specific content of step S1 includes: Determine the state quantity ξ at the current moment and the state quantity ξ' at the next moment of the leader; the state quantity λ at the current moment and the state quantity λ' at the next moment of the follower; determine the state equation of the follower: λ′ = f(λ) + g(λ)v(λ) In the above formula, f(·) is the state coupling function, v(·) is the control strategy function of the follower, and v(λ) is the control quantity under the current moment state quantity of the follower; Determine the desired control quantity v of the follower e It is: v e = g -1 (ξ)(ξ′ - f(ξ)) In the above formula, v e is the expected control quantity of the follower, and g -1 (ξ) is the inverse matrix of the input coupling function value; The following state error x at the current moment between the follower and the leader is: x = λ - ξ The control quantity error \(u(x)\) between the control quantity \(v(\lambda)\) of the follower under the state quantity at the current moment and the desired control quantity \(v\) is as follows: e ​ u(x) = v(λ) - v e = v(x + ξ) - v e The following state error x' of the follower at the next moment is: x' = λ' - ξ' = f(λ) + g(λ)v - ξ' = f(x + ξ) + g(x + ξ)(u(x) + v e ) - ξ' The following state error x' of the follower at the next moment is a function of the following state error x at the current moment and the control quantity error u(x), denoted as: x' = F(x, u(x)) In the above formula, F(·) represents the mapping from x and u(x) to x'.

3. The multi-agent following control method based on adaptive dynamic programming according to claim 2, wherein The specific content of step S2 includes: Determine the utility function of the following state error as: U(x, u(x)) = x T Qx + u(x) T Ru(x) In the above formula, U(x, u(x)) represents the utility function when the following state error is x and the control quantity error is u(x); Determine the cost function of the follower based on the utility function: V(x, u(x)) = ∑U(x, u(x)) = (x T Qx + u(x) T Ru(x)) + (x′ T Qx′ + u(x′) T Ru′(x′)) + … In the above formula, V(·) is the cost function, U(·) is the utility function; x' is the following state error of the follower at the next moment, and u'(x') is the control quantity error of the follower at the next moment; Based on the Bellman optimality principle: In the above formula, V * (x) is the optimal cost function under the following state error x at the current moment, and V * (x’) is the optimal cost function under the following state error x′ at the next moment, and min{·} represents the minimum value of the function in the curly brackets; The optimal control error function u * (x) is as follows: where u * (x) represents the control quantity at the current moment when following the state error x, represents the value of u(x) when minimizing the function in the curly brackets; R -1 represents the inverse matrix of the positive definite matrix R, g T (x) represents the transpose of g(x).

4. The multi-agent following control method based on adaptive dynamic programming according to claim 3, characterized in that In step S3, when solving the control error function based on the iterative method, set the value of the initial value function V0(·) to 0, and solve the initial control error function as: The cost function $V_{(x)}$ at the $i$-th ($i = 1, 2, 3, \cdots$) iteration is as follows: i (x) is: The control error function u at the i-th (i = 1, 2, 3,...) iteration i (x) is as follows:

5. The multi-agent following control method based on adaptive dynamic programming according to claim 1, wherein The optimal control strategy is: v * (λ) = u * (x) + v e In the above formula, v * (·) is the optimal control strategy of the follower, v * (λ) represents the control quantity of the follower under the state quantity λ at the current moment, v e represents the desired control quantity of the follower, u * (·) represents the optimal control error function of the follower, u * (x) represents the optimal control quantity error of the follower when the following state error at the current moment is x.

6. A multi-agent following control system based on adaptive dynamic programming, which is used to execute the multi-agent following control method based on adaptive dynamic programming according to any one of claims 1-5, characterized in that, The multi-agent includes a leader and at least one follower, characterized in that the following control system includes: A state quantity calculation module determines the state equation and control quantity of the follower based on the state quantity of the leader, determines the following state error between the follower and the leader based on the state equation of the follower, and the control quantity error of the follower; A control error determination module determines a utility function with the goal of minimizing the following state error and energy consumption of the follower, determines a cost function that minimizes the energy of the follower based on the utility function, and determines a control error function based on the cost function; An optimal control strategy module solves the control error function based on an iterative method, fits the control error function and the cost function based on a preset neural network to gradually fit out the optimal control error function during the iteration process, and determines the optimal control strategy based on the optimal control error function.

7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the multi-agent following control method based on adaptive dynamic programming according to any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the multi-agent following control method based on adaptive dynamic programming according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Target following and dynamic obstacle avoidance control method for speed difference slip steering vehicle

    CN110989576A

  • Hybrid power system energy management strategy based on reverse deep reinforcement learning

    CN111367172A