Multi-agent encircling elasticity control method based on human experience feedback

By adopting a closed elastic control method based on human experience feedback in a multi-agent system, combined with reinforcement learning and inverse step method, the system's robustness and optimization control problems in a dynamic environment are solved, and higher adaptability and task execution security are achieved.

CN120044784APending Publication Date: 2025-05-27BEIHANG UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510119972.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-25
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing multi-agent systems lack sufficient robustness in handling external interference, actuator failures and input saturation, especially in dynamic environments, making it difficult to achieve optimized control.

Method used

The multi-agent encirclement elastic control method based on human experience feedback is adopted, combined with the operator's manual input obstacle avoidance or emergency stop signals, reinforcement learning algorithms and inverse steps to realize the online robust control of the multi-agent system.

Benefits of technology

In the presence of continuous and rapid change of actuator disturbance environment, the online robust control of multi-agent systems is realized, which improves the system's ability to adapt to environmental changes and ensures the safety of task execution and the consistency of formation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120044784A_ABST
    Figure CN120044784A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-agent encirclement elasticity control method based on human experience feedback, belongs to the technical field of multi-agent system safety consistency control, and realizes multi-agent system encirclement elasticity online control by combining manual input obstacle avoidance and sudden stop signals of an operator with a reinforcement learning algorithm and a backstepping method. On-line robust control can be realized in an environment with continuously and rapidly changing perturbation of the actuator, the learning process is further accelerated, and the adaptive capacity of a multi-agent system to environment change is improved; besides, the backstepping method can realize anti-interference control on a nonlinear system, so that followers in the multi-agent system can effectively surround leaders in an actuator disturbance environment with continuous and rapid change, and formation consistency and task execution safety are kept; according to the method, the MASs actuator fault of the multi-agent system can be effectively handled, and the selected performance index is minimized when non-autonomous human interactive input exists in the multi-agent system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-agent system security consistency control, and in particular to a multi-agent encirclement elastic control method based on human experience feedback. Background Art

[0002] In recent decades, multi-agent systems (MASs) have rapidly developed into a major research hotspot in the control field, gaining recognition due to their widespread practical applications. MASs can collect environmental data or information from neighboring nodes and transmit this data to a central base station via a communication network. In recent years, the coordinated control of MASs has become a hot topic, particularly in areas such as collective tracking with a dynamic leader and multi-leader control.

[0003] For applications in encirclement control and pure feedback systems with unknown nonlinearities, researchers have proposed fault-tolerant control methods. These methods design controllers by integrating the error surface over a performance boundary function, eliminating the need to compensate for uncertainties and faults through differential equations. The error limit is defined by a pre-set performance boundary function. Although these methods have shown some effectiveness in dealing with multi-agent systems with unknown and bounded dynamic leader inputs, they may lack robustness to external disturbances.

[0004] In addition, research has also shown that existing finite-time fault-tolerant control methods may lack sufficient robustness in the face of actuator failures and input saturation, especially when dealing with persistent or intermittent faults. In order to improve the robustness of multi-agent systems in dynamic environments with sensor attacks.

[0005] At the same time, in the field of optimal encirclement control of MASs, researchers face the challenge of achieving optimal control while ensuring robustness. Although optimal control theory provides a framework for minimizing long-term objectives under limited resources, in practical applications, it is often difficult to directly solve the Hamilton-Jacobi-Bellman (HJB) equation due to incomplete system information. Dynamic programming methods are used as a method to solve the optimal solution, but they are limited by the "curse of dimensionality." Although adaptive dynamic programming and reinforcement learning (RL) have made progress as methods for solving online optimal control problems, they still have problems such as complex update laws, poor robustness to environmental changes, and excessive dependence on system dynamics. Summary of the Invention

[0006] In view of the above problems, the present invention provides a multi-agent encirclement elastic control method based on human experience feedback. The present invention combines the operator's manual input of obstacle avoidance, emergency stop signals, reinforcement learning algorithm and backstepping method to realize the elastic online control of the multi-agent system encirclement, and can realize online robust control in an environment with continuous and rapid changes in actuator disturbances, further accelerating the learning process and improving the adaptability of the multi-agent system to environmental changes; in addition, the backstepping method can realize anti-disturbance control of nonlinear systems, so that the followers in the multi-agent system can effectively encircle the leader in an environment with continuous and rapid changes in actuator disturbances, maintaining the consistency of the formation and the safety of task execution; the present invention can effectively deal with actuator failures of the multi-agent system MASs. When there is a fault in the multi-agent system and the system cannot control itself and requires human input of obstacle avoidance or emergency stop signals, the selected performance indicators can be minimized.

[0007] The present invention provides a multi-agent encirclement elastic control method based on human experience feedback, comprising:

[0008] Step S1, let t = 1, when t = 1, it represents the initial time; let i = 1, when i = 1, it represents the first agent, i = 1, 2, 3.....M+N, M+N represents the total number of agents;

[0009] Step S2: Let q = 1. When q = 1, it represents the q-th dimension state variable of the i-th agent at time t;

[0010] Step S3: Determine whether q = 1. If so, proceed to the next step; if not, proceed to step S10;

[0011] Step S4: Obtain the distributed encirclement error of the q-th dimension state variable of the ith agent at time t, and further obtain the weight estimate of the q-th dimension state variable of the ith agent at time t;

[0012] Step S5: Based on the neural network S iq,t and feedback behavior M iq,t , obtain the state estimate of the q-th dimension state variable of the i-th agent at time t;

[0013] Step S6: Obtain the value function of the q-th dimension state variable of the ith agent at time t and solve the Hamilton-Jacobi-Bellman equation to obtain the expected virtual control law of the q-th dimension state variable of the ith agent at time t;

[0014] Step S7: using the reinforcement learning method and the distributed encirclement error of the q-th dimension state variable of the ith agent at time t in step S4, obtain the weight design value of the critic network of the q-th dimension state variable of the ith agent at time t;

[0015] Step S8, design the virtual control law for the q-th state variable of the i-th agent at time t;

[0016] Step S9, based on the virtual control law for the q-th state variable of the i-th agent at time t, obtain the weight design value of the orator network for the q-th state variable of the i-th agent at time t, and proceed to the next step;

[0017] Step S10, determine whether 1 < q ≤ Q - 1, where Q represents the total dimension of the state variables. If not, proceed to the next step; if so, obtain the updated distributed合围 error for the q-th state variable of the i-th agent at time t, let q = q + 1, return to Step S6, and obtain the weight design value of the orator network for the q-th state variable of the i-th agent at time t;

[0018] Step S11, determine whether q is equal to Q. If not, let q = q + 1 and return to Step S6. If so, obtain the distributed合围 error for the q-th state variable of the i-th agent at time t, return to Step S6, and obtain the control law for the q-th state variable of the i-th agent at time t, and proceed to the next step;

[0019] Step S12, determine whether q is greater than or equal to Q. If so, obtain the control law for the i-th agent at time t, and based on the control law for the i-th agent at time t, obtain the estimated value of the disturbance for the i-th agent at time t.

[0020] If not, let q = q + 1 and return to Step S4;

[0021] Step S13, traverse the I agents at time t, repeat Steps S1 - S12, and obtain the control laws, disturbance estimated values, and state estimated values for each agent at time t;

[0022] Step S14, determine whether t is greater than or equal to T, where T represents the total number of time steps. If so, obtain the control laws, disturbance values, and state estimated values for each agent at each time step, and perform multi-agent合围 elastic control based on the control laws and disturbance values for each agent at each time step; if not, let t = t + 1 and return to Step S1.

[0023] Optionally, the specific steps for obtaining the state estimated value of the q-th state variable of the i-th agent at time t in Step S5 are as follows:

[0024] Obtain the distributed合围 error for the q-th state variable of the i-th agent at time t, and based on the distributed合围 error for the q-th state variable of the i-th agent at time t, obtain the derivative weight estimated value for the q-th state variable of the i-th agent at time t;

[0025] Based on the derivative weight estimate of the q-dimensional state variable of the ith agent at time t, obtain the weight estimate of the q-dimensional state variable of the ith agent at time t;

[0026] The weight estimate of the q-dimensional state variable of the i-th agent at time t is input into the neural network, and based on the feedback behavior, the state estimate of the q-dimensional state variable of the i-th agent at time t is obtained.

[0027] Optionally, the specific steps of obtaining the weight design value of the critic network of the q-th dimension state variable of the i-th agent at time t include:

[0028] Obtain the weight estimate of the critic network for the q-th dimension state variable of the ith agent at time t using reinforcement learning and the distributed encirclement error of the q-th dimension state variable of the ith agent at time t;

[0029] The weight design value of the critic network for the q-th dimensional state variable of the ith agent at time t is obtained based on the weight estimate value of the critic network for the q-th dimensional state variable of the ith agent at time t.

[0030] Optionally, obtain the state estimate of the q-th dimension state variable of the i-th agent at time t, expressed as:

[0031]

[0032] in, is the estimated value of the q-dimensional state variable of the i-th agent at time t, represents the weight estimate of the q-dimensional state variable of the i-th agent at time t, S i,q (t) represents the radial basis function of the q-th dimension state variable of the i-th agent at time t, x i,q (t) is the true value of the q-dimensional state variable of the i-th agent at time t, u i,q (t) represents the control law to be designed for the q-dimensional state variable of the i-th agent at time t, u i,human,q (t) represents the input signal introduced by the q-dimensional state variable of the i-th feedback behavior at time t, b i,q (t) represents the qth dimension state variable of the i-th feedback behavior at time t. The neural network is used to identify the neuron bias of nonlinear dynamics.

[0033] Optionally, the weight design value of the critic network of the q-th dimension state variable of the i-th agent at time t is The expression is:

[0034]

[0035] Among them, S i,q(t) represents the radial basis function of the q-th dimension state variable of the i-th agent at time t, η ci,q (t) represents the positive parameter to be designed of the critic network of the q-th dimension state variable of the i-th agent, The weight estimate of the critic network representing the q-th dimension state variable of the i-th agent at time t, δ i,q (t) represents the distributed encirclement error of the q-th dimension state variable of the i-th agent at time t.

[0036] Optionally, the virtual control law a of the q-dimensional state variable of the i-th agent at time t is i,q (t), the expression is:

[0037]

[0038] Among them, α i,q (t) represents the virtual control law of the q-th dimension state variable of the i-th agent at time t, τ i,q (t) represents the constant term designed in the q-dimensional state variable of the i-th agent at time t, represents the estimated weight of the q-th dimension state variable of the i-th agent at time t, The estimated weight of the orator network representing the q-th dimension of the state variable of the ith agent at time t.

[0039] Optionally, the weight of the orator network of the q-th dimension state variable of the i-th agent at time t is The expression of the design value is:

[0040]

[0041] Among them, η ai,q The positive parameters to be designed of the orator actor network representing the q-th dimension state variable of the i-th agent, The estimated weight of the orator actor network representing the q-th dimension state variable of the i-th agent; The estimated weights of the critic network representing the q-th state variable of the ith agent at time t.

[0042] Compared with the prior art, the present invention has at least the following beneficial effects:

[0043] (1) The present invention uses manual input of obstacle avoidance or emergency stop signals for real-time feedback and reinforcement learning algorithms to perform elastic control of the multi-agent system. This enables the multi-agent system to estimate interference online and achieve robust control in a continuously and rapidly changing actuator interference environment. This allows the system to better adapt to environmental changes, exhibit higher robustness in tasks with high uncertainty, and reduce the risk of formation breakdown and path deviation.

[0044] (2) The present invention significantly improves the real-time response capability of the multi-agent system; the operator's real-time feedback can promptly influence the decision-making process of the reinforcement learning algorithm, enabling the multi-agent system to quickly adjust its strategy to cope with sudden changes and complex challenges in the environment; the control method of the present invention that combines human operation and machine control effectively compensates for the shortcomings of relying solely on automated algorithms, and at the same time, robust control can be achieved through the design of control laws;

[0045] (3) While ensuring the continuous and rapid change of the actuator interference environment, the present invention also greatly improves the efficiency of task execution; significantly improves the multi-agent system's ability to close the gap in the environment of continuous and rapid change of actuator disturbances, and provides strong support for the reliability and flexibility of the multi-agent system in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The drawings are only for purposes of illustrating particular embodiments and are not to be considered limiting of the invention.

[0047] Figure 1 is a schematic diagram of a multi-agent system according to an embodiment of the present invention;

[0048] Figure 2 (a)-(d) are schematic diagrams of simulation diagrams of online estimation of x-axis error interference according to an embodiment of the present invention;

[0049] Figure 3 (a)-(d) are schematic diagrams of simulation diagrams of online estimation of y-axis error interference in an embodiment of the present invention;

[0050] Figure 4 (a)-(d) are schematic diagrams of simulation diagrams of online estimation of z-axis error interference in an embodiment of the present invention;

[0051] Figure 5 (a)-(d) are schematic diagrams of radial basis function (RBF) network weight convergence simulation diagrams according to an embodiment of the present invention;

[0052] Figure 6 This is a schematic diagram of a comparative simulation diagram of reinforcement learning RL performed by multi-agent encirclement elastic control and disturbance observer DOB and state observer SOB in an embodiment of the present invention. DETAILED DESCRIPTION

[0053] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other. In addition, the present invention can also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present invention is not limited by the specific embodiments disclosed below.

[0054] A specific embodiment of the present invention, as Figure 1-6 , discloses a multi-agent enclosure elastic control method based on human experience feedback, the specific implementation steps include:

[0055] Step S1, let t = 1, when t = 1, it represents the initial time; let i = 1, when i = 1, it represents the first agent, i = 1, 2, 3.....M+N, M+N represents the total number of agents;

[0056] Step S2: Determine the distributed encirclement error δ of the i-th agent at time t i (t), the expression is:

[0057]

[0058] Among them, x r (t) represents the state of the rth leader at time t, M represents the number of followers, and N represents the number of leaders; p=1,2,3.....M+N, M+N represents the total number of agents, N and M are both integers; p, j∈M, r∈N, a p,j Indicates the communication status between the jth follower and the jth follower. If there is a communication connection between followers, then a p,j =1; if there is no communication connection between followers, then a p,j =0;x p (t) represents the state of the p-th follower at time t, x j (t) represents the state of the j-th follower at time t, a p,r Represents the communication connection between the pth follower and the rth leader, r∈N, if there is a communication connection between the pth follower and the rth leader, then a p,r =1; if there is no communication connection between the follower and the leader, then a p,r =0.

[0059] Step S3, let q = 1. When q = 1, it represents the q-th dimension state variable of the i-th agent at time t;

[0060] Step S4: Determine whether q = 1. If so, proceed to the next step; if not, proceed to step S12;

[0061] Step S5: Obtain the distributed encirclement error δ of the q-dimensional state variable of the i-th agent at time t i,q (t),

[0062] The distributed encirclement error δ based on the q-th dimension state variable of the i-th agent at time t i,q (t), get the derivative weight estimate of the q-dimensional state variable of the i-th agent at time t The expression is:

[0063]

[0064] Among them, S i,q-1 (t) represents the radial basis function of the q-1th dimension state variable of the i-th agent at time t, which is used to approximate the unknown nonlinear dynamics of the agent, x i,q-1 (t) represents the state variable of the q-1th dimension of the i-th agent at time t, q = 1, 2, 3...Q, Q represents the total dimension, Represents the estimated weight of the q-1th dimension state variable of the i-th agent at time t.

[0065] Step S6: Estimated value of derivative weight of the q-th dimension state variable of the i-th agent at time t Get the weight estimate of the q-th dimension state variable of the i-th agent at time t

[0066] The estimated weight of the q-th dimension state variable of the i-th agent at time t Input the neural network and obtain the state estimate of the q-dimensional state variable of the i-th agent at time t based on the feedback behavior The expression is:

[0067]

[0068] in, is the estimated value of the q-dimensional state variable of the i-th agent at time t, represents the weight estimate of the q-dimensional state variable of the i-th agent at time t, S i,q (t) represents the radial basis function of the q-th dimension state variable of the i-th agent at time t, which is used to approximate the unknown nonlinear dynamics of the agent, x i,q (t) is the true value of the q-dimensional state variable of the i-th agent at time t, u i,q (t) represents the control law to be designed for the q-dimensional state variable of the i-th agent at time i, u i,human,q (t) represents the input signal introduced by the q-th dimension feedback behavior of the i-th agent at time t, where the feedback behavior is human experience, and b i,q(t) represents the qth dimension state variable of the i-th feedback behavior at time t. The neural network is used to identify the neuron bias of nonlinear dynamics.

[0069] It is understandable that the feedback behavior includes emergency obstacle avoidance in the event of a sudden obstacle, and an additional manual auxiliary input of a human operator into the emergency stop signal of the multi-intelligent system when an interference fault occurs;

[0070] Exemplarily, the feedback behavior includes an emergency stop signal and an obstacle avoidance signal;

[0071] Furthermore, the leader and follower in the multi-agent system move toward the target point. There are static and dynamic circular obstacles in the motion space. The obstacle avoidance signal is input into the multi-agent system with the manual assistance of the human operator, thereby achieving safe movement of the agent.

[0072] Step S7: Obtain the value function of the q-th dimension state variable of the i-th agent at time t, expressed as:

[0073]

[0074] Among them, V i,q (·) represents the value function of the q-th dimension state variable of the i-th agent at time t, δ i,q (t) represents the distributed encirclement error of the q-th dimension state variable of the i-th agent at time t, α i,q (t) represents the virtual control quantity of the q-th dimension state variable of the i-th agent at time t.

[0075] Step S8: Solve the Hamilton-Jacobi-Bellman equation HJB equation based on the value function of the q-dimensional state variable of the i-th agent at time t to obtain the virtual control law expected by the q-dimensional state variable of the i-th agent at time t The expression is:

[0076]

[0077] in, represents the virtual control law of the q-dimensional state variable expectation of the i-th agent at time t, d i,q (t) represents the topological structure coefficient of the q-dimensional state variable of the i-th agent at time t, τ i,q (t) is the continuous variable of the q-th dimension state variable of the i-th agent at time t, τ i,q >0,f i,q (·) represents the nonlinear dynamics term of the unknown q-dimensional state variable of the i-th agent, is the expected weight value of the q-th dimension state variable of the i-th agent at time t;

[0078] Step S9: using the reinforcement learning method and the distributed encirclement error of the q-th dimension state variable of the ith agent at time t, obtain the weight estimate of the critic network of the q-th dimension state variable of the ith agent at time t;

[0079] The weight design value of the critic-critic network of the q-th dimension state variable of the ith agent at time t is obtained based on the estimated weight value of the critic-critic network of the q-th dimension state variable of the ith agent at time t. The expression is:

[0080]

[0081] Where ηci,q(t) represents the positive parameter to be designed of the critic network of the q-th dimension state variable of the i-th agent, The estimated weight of the critic network representing the q-th dimension of the state variable of the i-th agent at time t;

[0082] Step S10: Design the virtual control law α of the q-th dimension state variable of the i-th agent at time t i,q (t), the expression is:

[0083]

[0084] Among them, α i,q (t) represents the virtual control law of the q-th dimension state variable of the i-th agent at time i, τ i,q (t) represents the constant term designed in the q-dimensional state variable of the i-th agent at time t, represents the estimated weight of the q-th dimension state variable of the i-th agent at time t, The estimated weights of the orator actor network representing the q-th dimension of the state variable of the ith agent at time t.

[0085] Step S11: Based on the virtual control law of the q-th dimension state variable of the ith agent at time t, obtain the weight design value of the orator actor network of the q-th dimension state variable of the ith agent at time t Go to the next step;

[0086] Optionally, the weight design value of the speaker actor network of the qth state variable of the i-th agent at time t in step S11 is The expression is:

[0087]

[0088] Among them, η ai,q The positive parameters to be designed of the orator actor network representing the q-th dimension state variable of the i-th agent, The weight estimation value of the orator actor network representing the q - dimensional state variable of the i - th agent The weight estimation value of the critic network representing the q - dimensional state variable of the i - th agent at time t

[0089] Step S12: Determine whether 1 < qS ≤ Q - 1, where Q represents the total dimension of the state variables. If not, proceed to the next step; if so, obtain the updated distributed合围 error δ′ i,q (t) of the q - dimensional state variable of the i - th agent at time t, and use the updated distributed合围 error δ′ i,q (t) of the q - dimensional state variable of the i - th agent at time t as the distributed合围 error δ i,q+1 (t) of the (q + 1) - dimensional state variable of the i - th agent at time t. Let q = q + 1, return to step S7, and obtain the weight of the orator actor network of the q - dimensional state variable of the i - th agent at time t Design value

[0090] Optionally, the expression for the updated distributed合围 error of the q - dimensional state variable of the i - th agent at time t in step S12 is:

[0091]

[0092] where x i,q (t) represents the q - dimensional state variable of the i - th agent at time t, represents the virtual control law of the (q - 1) - dimensional state variable of the i - th agent at time t

[0093] Step S13: Determine whether q is equal to Q. If not, let q = q + 1 and return to step S7. If so, obtain the distributed合围 error of the q - dimensional state variable of the i - th agent at time t, return to step S7, and obtain the weight Estimation value of the orator actor network of the q - dimensional state variable of the i - th agent at time t

[0094] where the expression for the distributed合围 error of the q - dimensional state variable of the i - th agent at time t in step S13 is:

[0095]

[0096] where g i,q (t) represents the constant coefficient of the q - dimensional state variable of the i - th agent at time t, represents the virtual control law of the (q - 1) - dimensional state variable of the i - th agent at time t, u i,q (t) is the control law to be designed for the q - dimensional state variable of the i - th agent at time t, f i,q It should be noted that the term "合围误差" is not a common technical term in the context of patent translation. You may need to double - check if this is a correct expression in the original language or provide more context for a more accurate translation. If it's a made - up or misspelled term, it might affect the overall comprehensibility of the translation.(t) represents the nonlinear dynamics of the qth dimension state variable of the i-th agent at time t, represents the disturbance value of the q-dimensional state variable of the i-th agent at time t caused by the actuator attack;

[0097] The weight design value of the orator actor network based on the q-th dimension state variable of the i-th agent at time t Get the control law of the q-th dimension state variable of the i-th agent at time t Go to the next step;

[0098] Optionally, the control law of the q-th dimension state variable of the i-th agent at time t in step S13 is The expression is:

[0099]

[0100] in, It represents the estimated value of the disturbance of the q-th dimension state variable of the i-th agent at time t caused by the actuator attack.

[0101] Step S14: Determine whether q is greater than or equal to Q. If not, set q=q+1 and return to step S5.

[0102] If so, obtain the control law of the ith agent at time t, and obtain the derivative estimate of the disturbance of the ith agent at time t based on the control law of the ith agent at time t The expression is:

[0103]

[0104] Where β represents the design value of the disturbance input to be estimated, g i-1 (t) represents the constant coefficient of the state of the i-1th agent at time t, It represents the estimated disturbance value of the actuator attack suffered by the i-1th agent at time t.

[0105] The derivative estimate of the perturbation of the i-th agent at time t Get the derivative estimate of the perturbation of the i-th agent at time t

[0106] Step S15: traverse the I agents at time t and repeat steps S1-S14 to obtain the control law, disturbance estimate, and state estimate of each agent at time t;

[0107] Step S16: Determine whether t is greater than or equal to T, where T represents the total number of moments. If so, obtain the control law, disturbance value, and state estimation value of each agent at each moment, and perform multi-agent elastic control based on the control law, disturbance value, and state estimation value of each agent at each moment; if not, set t = t + 1 and return to step S1.

[0108] It can be understood that each agent performs combined elastic control through the control law, disturbance value and state estimation value of each agent at each moment.

[0109] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any technician familiar with this technical field within the technical scope disclosed by the present invention should be covered by the scope of protection of the present invention.

Claims

1. A multi-agent encirclement elastic control method based on human experience feedback, characterized in that: Including: Step S1: Let t = 1. When t = 1, it represents the initial moment; Let i = 1. When i = 1, it represents the first agent, and i = 1, 2, 3.....M + N, where M + N represents the total number of agents; Step S2: Let q = 1. When q = 1, it represents the q - th state variable of the i - th agent at time t; Step S3: Judge whether q = 1. If so, proceed to the next step; if not, enter Step S10; Step S4: Obtain the distributed合围 error of the q - th state variable of the i - th agent at time t, and further obtain the weight estimation value of the q - th state variable of the i - th agent at time t; Step S5: Based on the neural network S iq,t and feedback behavior iq,t , obtain the state estimate of the q-dimensional state variable of the i-th agent at time t; Step S6: Obtain the value function of the q - th state variable of the i - th agent at time t and solve the Hamilton - Jacobi - Bellman equation to obtain the expected virtual control law of the q - th state variable of the i - th agent at time t; Step S7: Using the reinforcement learning method and the distributed合围 error of the q - th state variable of the i - th agent at time t in Step S4, obtain the weight design value of the critic network of the q - th state variable of the t - th agent at time t; Step S8: Design the virtual control law of the q - th state variable of the i - th agent at time t; Step S9: Based on the virtual control law of the q - th state variable of the i - th agent at time t, obtain the weight design value of the speaker network of the q - th state variable of the i - th agent at time t, and proceed to the next step; Step S10: Judge whether 1 < q ≤ Q - 1, where Q represents the total dimension of the state variables. If not, proceed to the next step; if so, obtain the updated distributed合围 error of the q - th state variable of the i - th agent at time t, let q = q + 1, return to Step S6, and obtain the weight design value of the speaker network of the q - th state variable of the i - th agent at time t; Step S11: Judge whether q is equal to Q. If not, let q = q + 1, return to Step S6. If so, obtain the distributed合围 error of the q - th state variable of the i - th agent at time t, return to Step S6, obtain the control law of the q - th state variable of the i - th agent at time t, and proceed to the next step; Step S12: Judge whether q is greater than or equal to Q. If so, obtain the control law of the i - th agent at time t, and obtain the estimated value of the disturbance of the i - th agent at time t based on the control law of the i - th agent at time t, If not, let q = q + 1, return to Step S4; Step S13: Traverse the I agents at time t, repeat Steps S1 - S12, and obtain the control laws, disturbance estimation values, and state estimation values of each agent at time t; Step S14: Judge whether t is greater than or equal to T, where T represents the total number of time steps. If so, obtain the control laws, disturbance values, and state estimation values of each agent at each time step, and perform multi - agent合围 elastic control based on the control laws and disturbance values of each agent at each time step; if not, let t = t + 1, return to Step S1.

2. The multi - agent合围 elastic control method according to claim 1, characterized in that, Step S5 obtains the state estimate of the q-dimensional state variable of the i-th agent at time t The specific steps include: Obtain the distributed encirclement error of the q-dimensional state variable of the ith agent at time t, and obtain the derivative weight estimate of the q-dimensional state variable of the ith agent at time t based on the distributed encirclement error of the q-dimensional state variable of the ith agent at time t; Based on the derivative weight estimate of the q-dimensional state variable of the ith agent at time t, obtain the weight estimate of the q-dimensional state variable of the ith agent at time t; The weight estimate of the q-dimensional state variable of the ith agent at time t is input into the neural network, and based on the feedback behavior, the state estimate of the q-dimensional state variable of the ith agent at time i is obtained.

3. The multi-agent encirclement elastic control method according to claim 1, characterized in that: The specific steps of obtaining the weight design value of the critic network of the q-th dimension state variable of the i-th agent at time t include: Using the reinforcement learning method and the distributed encirclement error of the q-th dimension state variable of the ith agent at time t, the weight estimate of the critic network of the q-th dimension state variable of the ith agent at time t is obtained; The weight design value of the critic network of the q-th dimensional state variable of the ith agent at time t is obtained based on the weight estimate value of the critic network of the q-th dimensional state variable of the ith agent at time t.

4. The multi-agent encirclement elastic control method according to claim 2, characterized in that: Get the state estimate of the q-dimensional state variable of the i-th agent at time t, expressed as: in, is the estimated value of the q-dimensional state variable of the i-th agent at time t, represents the weight estimate of the q-dimensional state variable of the ith agent at time t, S i,q (t) represents the radial basis function of the q-dimensional state variable of the i-th agent at time t, x i,q (t) is the true value of the q-dimensional state variable of the i-th agent at time t, u i,q (t) represents the control law to be designed for the q-dimensional state variable of the ith agent at time t, u i,human,q (t) represents the input signal introduced by the q-dimensional state variable of the i-th feedback behavior at time t, b i,q (t) represents the qth dimensional state variable of the i-th feedback behavior at time t, and uses neural networks to identify the neuron bias of nonlinear dynamics.

5. The multi-agent encirclement elastic control method according to claim 3, characterized in that: The weight design value of the critic network of the q-th dimension state variable of the ith agent at the time t The expression is: Among them, S i,q (t) represents the radial basis function of the q-dimensional state variable of the i-th agent at time t, η ci,q (t) represents the positive parameter to be designed of the critic network of the q-th dimensional state variable of the i-th agent, The weight estimate of the critic network representing the q-th dimension state variable of the ith agent at time t, δ i,q (t) represents the distributed encirclement error of the q-th dimensional state variable of the i-th agent at time t.

6. The multi-agent encirclement elastic control method according to claim 5, characterized in that: The virtual control law α of the q-dimensional state variable of the i-th agent at the time t i,q (t), the expression is: Among them, α i,q (t) represents the virtual control law of the q-dimensional state variable of the i-th agent at time t, τ i,q (t) represents the constant term designed in the q-dimensional state variable of the i-th agent at time t, represents the estimated weight of the q-dimensional state variable of the ith agent at time t, The estimated value of the weight of the orator network representing the q-th dimension state variable of the ith agent at time t.

7. The multi-agent encirclement elastic control method according to claim 6, characterized in that: The weight of the orator network of the q-dimensional state variable of the ith agent at the time t The expression of the design value is: Among them, η ai,q The positive parameters to be designed of the orator actor network representing the q-th dimension state variable of the ith agent, The estimated weights of the orator actor network representing the q-th dimension state variable of the ith agent; The estimated value of the weight of the critic network representing the q-th dimension of the state variable of the ith agent at time t.

Citation Information

Patent Citations

  • Multi-agent system encirclement control method and system

    CN111176327A

  • High-order multi-agent reinforcement learning optimization controller construction method and system

    CN116500893A

  • Multi-agent system fixed time optimal control method with unknown lag

    CN118759843A

  • Multi-agent system human-in-loop elastic control method based on event trigger integral reinforcement learning

    CN118915466A

  • Validating and computing stability limits of human-in-the-loop adaptive control systems

    US20180148069A1