A distributed online output feedback control method for uncertain multi-agent systems
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING UNIV OF POSTS & TELECOMM
- Filing Date
- 2026-07-09
- Publication Date
- 2026-08-07
AI Technical Summary
[0006]本发明的主要目的在于提供一种不确定多智能体系统分布式在线输出反馈控制方法,以解决现有技术中分布式优化、纳什均衡搜索和多智能体控制方法难以同时适应时变代价函数、物理输出约束、系统参数未知、状态不可完全测量、邻居通信限制以及实时梯度反馈的问题
适用于具有物理动力学约束的智能体系统。本发明将智能体输出作为实际博弈决策变量,并通过控制输入
间接调节该输出,相比直接更新抽象决策变量的分布式优化方法,更适用于真实工程系统。
Smart Images

Figure CN122525952A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of multi-agent system control, distributed optimization control, online non-cooperative game control, output feedback control, and networked control, and particularly to a distributed online output feedback control method for uncertain multi-agent systems. Background Technology
[0002] With the development of smart grids, robot swarms, communication networks, intelligent transportation, and industrial collaborative systems, multiple intelligent agents with autonomous decision-making capabilities need to conduct collaborative or competitive decisions under local objectives and mutual influences. Non-cooperative game theory and Nash equilibrium search methods have been widely applied in the field of distributed decision-making and control of networked systems because they can effectively describe policy optimization problems under multi-agent coupling conditions.
[0003] Existing distributed Nash equilibrium search methods typically assume that the agent can directly update its own decision variables. However, in practical engineering, the agent's decision variables are often the physical outputs of the underlying dynamic system (such as the output power of distributed power sources in a microgrid, or the position or velocity output of a robot), which cannot be directly set by the algorithm and must be indirectly achieved through control inputs. Furthermore, the cost function of real-world systems often changes dynamically over time (such as fluctuations in power system load demand and electricity prices, or changes in the robot's task environment), causing the equilibrium point to evolve into a dynamic trajectory. Existing methods are mostly designed for static cost functions or fixed equilibrium points, making it difficult to generate online-updated reference signals and ensuring that the system output continuously tracks the dynamic equilibrium trajectory.
[0004] Furthermore, practical multi-agent systems commonly suffer from problems such as unknown parameters, high dynamic order, incomplete state measurement, and communication limited to neighboring nodes, limiting the applicability of existing control methods that rely on accurate models or complete state information. Moreover, existing methods often assume that the agent can obtain ideal gradient information (e.g., calculating the gradient of the cost function at the reference variable), but in real-world scenarios, the analytical expression of the cost function is often unavailable, and gradient information can only be obtained through measurement, estimation, or numerical approximation after the actual output, making it difficult for existing methods to adapt to engineering realities.
[0005] The existing technology is a "distributed Nash equilibrium search output feedback integral control method for uncertain linear multi-agent systems," which combines an upper-level virtual reference generator with a lower-level output feedback integral controller to bring the system output to the Nash equilibrium point of a static non-cooperative game. However, this type of existing technology still has significant limitations: First, it is only applicable to static games or fixed equilibrium points and cannot adapt to dynamic equilibrium trajectories under time-varying cost functions; second, it aims at asymptotic convergence and lacks characterization of online performance in time-varying environments; third, it relies on ideal gradient information, which is inconsistent with actual engineering scenarios; and fourth, it does not fully consider the underlying physical dynamic constraints, which may lead to inconsistencies between the reference decision variables and the actual output, making it difficult to meet the online control requirements in time-varying non-cooperative game scenarios. Summary of the Invention
[0006] The main objective of this invention is to provide a distributed online output feedback control method for uncertain multi-agent systems, addressing the challenges of existing distributed optimization, Nash equilibrium search, and multi-agent control methods in simultaneously adapting to time-varying cost functions, physical output constraints, unknown system parameters, incomplete state measurement, neighbor communication limitations, and real-time gradient feedback. This enables end-to-end models to perform efficient reinforcement learning and exploration through reward feedback on their behavior, significantly surpassing the performance achieved through imitation learning alone.
[0007] Another objective of this invention is to propose a distributed online output feedback control system for uncertain multi-agent systems.
[0008] A third objective of this invention is to provide a non-transitory computer-readable storage medium.
[0009] The fourth objective of this invention is to provide a computer device.
[0010] To achieve the above objectives, a first aspect of the present invention proposes a distributed online output feedback control method for an uncertain multi-agent system, comprising:
[0011] The multi-agent system comprises multiple agents, each agent's output variable serving as its actual decision variable in a time-varying non-cooperative game. The method includes: Obtain the actual output of each agent, local cost function information, and estimation information sent by neighboring agents; For each agent, a local estimation vector is constructed, which contains the agent's estimates of the decision variables of other agents and the agent's reference signal; Based on the local estimation vector, the estimation information sent by neighboring agents, and the local gradient feedback signal, a distributed online reference generation dynamic is constructed to generate the reference signal; The output tracking error is constructed based on the actual output and the reference signal, and the integral error state is constructed accordingly. Construct a dynamic output feedback auxiliary state based on the actual output; The control input is generated based on the integral error state, the output tracking error, and the dynamic output feedback auxiliary state. The control input is applied to the corresponding intelligent agent, so that the actual output of the intelligent agent tracks the reference signal in real time.
[0012] In one embodiment of the present invention, each agent satisfies an uncertain output dynamics model, which includes the agent's state variables, control inputs, output variables, unknown parameters or uncertain parameters, and a system matrix or vector associated with the unknown parameters or uncertain parameters.
[0013] In one embodiment of the present invention, each agent in the time-varying non-cooperative game has a time-varying cost function, which includes the agent's output decision variables, the set of output decision variables of other agents besides the agent, and a time variable.
[0014] In one embodiment of the present invention, the neighboring agents interact with each other through a communication graph, which includes a set of agent nodes, a set of communication edges, and an adjacency weight matrix. The elements in the adjacency weight matrix represent the communication weights for sending information between corresponding agents.
[0015] In one embodiment of the present invention, the construction of the distributed online reference generation dynamic includes: Based on the consistency adjustment parameters, communication weights, selection vectors, and the local gradient feedback signals of the corresponding agents, dynamic equations are established to generate reference signals.
[0016] In one embodiment of the present invention, the acquisition of the local gradient feedback signal includes: Calculate the partial derivative of the agent's cost function with respect to its own decision variables. The calculation of this partial derivative is based on the agent's estimation set of decision variables for other agents besides itself. The resulting partial derivative is the local gradient at the reference variable.
[0017] In one embodiment of the present invention, when the corresponding agent cannot obtain the local gradient at the reference variable, the acquisition of the local gradient feedback signal includes: The real-time gradient at the actual output of the agent is obtained through online measurement, numerical difference, sensor feedback, operation cost evaluation module or learning model estimation, and this real-time gradient is used as the local gradient feedback signal.
[0018] In one embodiment of the present invention, constructing the output tracking error and the integral error states includes: The output tracking error is constructed based on the difference between the agent's actual output and the reference signal; Establish a dynamic relationship between the integral error states, so that the integral error states change with time integration, and these integral error states are error state variables specific to the corresponding agent.
[0019] In one embodiment of the present invention, constructing the dynamic output feedback auxiliary state includes: Determine the relative order and output feedback gain of the corresponding agent; Based on the relative order, output feedback gain, and the actual output of the agent, a dynamic equation is established to generate a dynamic output feedback auxiliary state containing multiple components.
[0020] In one embodiment of the present invention, generating the control input includes: Based on preset controller gain parameters, combined with integral error state, output tracking error and dynamic output feedback auxiliary state, control input is generated through linear combination or preset operation relationship. The controller gain parameters include multiple preset adjustment coefficients.
[0021] In one embodiment of the present invention, the method further includes: Calculate the difference between the local gradient at the reference variable and the real-time gradient at the actual output, and define this difference as the gradient mismatch. Based on the gradient mismatch, adjust the parameter settings of the distributed online reference generation dynamics, or analyze the closed-loop interaction between the distributed online reference generation dynamics and the control input. The gradient mismatch is calculated by subtracting the real-time gradient at the actual output from the local gradient at the reference variable.
[0022] In one embodiment of the present invention, the multi-agent system is selected from microgrid distributed power systems, robot swarm systems, communication network resource allocation systems, intelligent transportation systems, or industrial multi-device collaborative control systems.
[0023] In one embodiment of the present invention, when the multi-agent system is a microgrid distributed power system, the specific implementation steps include: Each distributed generation unit is treated as a corresponding intelligent agent; The output power of each distributed generation unit is taken as the actual output of the intelligent agent; Construct a time-varying cost function that includes the local generation cost coefficient, the time-varying electricity price or incentive signal, the power imbalance penalty coefficient, and the time-varying load demand or capacity demand.
[0024] To achieve the above objectives, a second aspect of the present invention provides a distributed online output feedback control system for an uncertain multi-agent system, comprising: The output acquisition module is used to acquire the actual output of each agent; The local cost information acquisition module is used to acquire the local cost function information or local gradient feedback signal of each agent; The neighbor communication module is used to receive estimation information sent by neighboring intelligent agents; The distributed reference generation module is used to generate a reference signal based on the local estimation vector, the estimation information sent by neighboring agents, and the local gradient feedback signal. The integration error module is used to construct an integration error state based on the actual output and the reference signal; A dynamic output feedback module is used to construct a dynamic output feedback auxiliary state based on the actual output. The control input generation module is used to generate control inputs based on the integral error state, output tracking error, and dynamic output feedback auxiliary state. The execution control module is used to apply the control input to the corresponding intelligent agent, so that the actual output of the intelligent agent tracks the reference signal in real time.
[0025] The distributed online output feedback control method and system for uncertain multi-agent systems of this invention solves the problem that existing Nash equilibrium search methods, which typically directly update abstract decision variables, are difficult to apply when decision variables are generated by the output of a physical dynamic system; it solves the problem that existing static equilibrium control methods struggle to track dynamic equilibrium trajectories generated by time-varying cost functions; it solves the output tracking control problem under conditions of unknown agent system parameters, high-order dynamics, and incomplete state measurement; it solves the problem of distributed online reference generation in multi-agent systems where each agent can only utilize local information and neighbor communication information; and it solves the problem of online game control using real-time gradients at the actual output when the analytical gradient of the cost function is unavailable.
[0026] To achieve the above objectives, a third aspect of this application provides an electronic device having a computer program stored thereon, which, when executed by a processor, implements the distributed online output feedback control method for uncertain multi-agent systems as described in the first aspect embodiment.
[0027] To achieve the above objectives, a fourth aspect of this application provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the distributed online output feedback control method for uncertain multi-agent systems as described in the first aspect embodiment.
[0028] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0029] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 A flowchart of a distributed online output feedback control method for an uncertain multi-agent system provided in an embodiment of the present invention; Figure 2 A flowchart of another distributed online output feedback control method for an uncertain multi-agent system provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of a two-layer control architecture for a multi-agent system provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the neighbor communication topology provided in an embodiment of the present invention; Figure 5 A schematic diagram of the real-time gradient feedback mechanism provided in an embodiment of the present invention; Figure 6 This is a schematic diagram of the online scheduling and control structure for distributed power sources in a microgrid provided in an embodiment of the present invention; Figure 7 This invention provides a structural diagram of a distributed online output feedback control system for an uncertain multi-agent system. Figure 8 The computer device provided in the embodiments of the present invention. Detailed Implementation
[0030] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0031] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0032] The following description, with reference to the accompanying drawings, describes a distributed online output feedback control method, system, device, and medium for an uncertain multi-agent system according to embodiments of the present invention.
[0033] Example 1 In one embodiment of the present invention, a distributed online output feedback control method for an uncertain multi-agent system is provided. The multi-agent system includes multiple agents, and the output variable of each agent serves as its actual decision variable in a time-varying non-cooperative game, such as... Figure 1 As shown, the method includes: S101, obtain the actual output of each agent, local cost function information, and estimation information sent by neighboring agents; S102, construct a local estimation vector for each agent, which contains the agent's estimates of the decision variables of other agents and the agent's reference signal; S103, Based on the local estimation vector, the estimation information sent by the neighboring agents, and the local gradient feedback signal, construct a distributed online reference generation dynamic to generate the reference signal; S104, construct the output tracking error based on the actual output and the reference signal, and construct the integral error state; S105, construct a dynamic output feedback auxiliary state based on the actual output; S106, generate control input based on the integral error state, the output tracking error, and the dynamic output feedback auxiliary state; S107, the control input is applied to the corresponding intelligent agent, so that the actual output of the intelligent agent tracks the reference signal in real time.
[0034] It is understood that the multi-agent system includes There are several agents, and the output variable of each agent is used as its actual decision variable in a time-varying non-cooperative game. The method includes: Get the The actual output of each agent The local cost function information and the estimation information sent by neighboring agents, among which ; For the first Each agent constructs a local estimation vector. ,in Indicates the first The first agent on the first Estimation of decision variables for an agent; Indicates the first Reference signals for each agent; Based on the local estimation vector, the estimation information sent by neighboring agents, and the local gradient feedback signal, a distributed online reference generation dynamic is constructed to generate the reference signal. ; Based on the actual output With the reference signal Construct the output tracking error and the integral error state; Based on the actual output Construct a dynamic output feedback auxiliary state; The control input is generated based on the integral error state, the output tracking error, and the dynamic output feedback auxiliary state. ; The control input Apply to the The first intelligent agent, making the first The actual output of each agent Tracking the reference signal .
[0035] Furthermore, the first An agent satisfies the following uncertain output dynamics model:
[0036] in, For the first The state variables of an agent, To control the input, For output variables, For unknown or uncertain parameters, , , This refers to the system matrix or vector associated with the unknown or uncertain parameters.
[0037] Furthermore, in the time-varying non-cooperative game, the first... An agent has the following time-varying cost function:
[0038] in, For the first The output decision variables of an agent To exclude the first The set of decision variables output by agents other than the one agent. It is a time variable.
[0039] Furthermore, the neighboring agents interact with each other through a communication graph, which is represented as follows:
[0040] in, A set of intelligent agent nodes. For the set of communication edges, This is the adjacency weight matrix. Indicates the first The agent directs the first The communication weight of each agent sending information.
[0041] Furthermore, the distributed online reference generation dynamics are as follows:
[0042] in, For consistency adjustment parameters, For communication weights, To select a vector, For the first Local gradient feedback signals for each agent.
[0043] Furthermore, the local gradient feedback signal is the local gradient at the reference variable:
[0044] in, Indicates the first The partial derivative of the cost function of an agent with respect to its own decision variables. Indicates the first A set of estimates of decision variables of each agent for other agents besides itself.
[0045] Furthermore, when the first When an agent cannot obtain the local gradient at the reference variable, the local gradient feedback signal is the real-time gradient at the actual output:
[0046] in, It is obtained through online measurement, numerical difference, sensor feedback, operational cost assessment module or learning model estimation.
[0047] Furthermore, the output tracking error is:
[0048] The integral error state satisfies:
[0049] in, For the first The integral error state of each agent.
[0050] Furthermore, the dynamic output feedback auxiliary state includes And satisfy:
[0051] in, For the first The relative order of each agent This is the output feedback gain.
[0052] Furthermore, the control input for:
[0053] in, , This refers to the controller gain parameter.
[0054] Furthermore, the method also includes: The difference between the local gradient at the reference variable and the real-time gradient at the actual output is used as the gradient mismatch. The closed-loop relationship between the distributed online reference generation dynamics and the control input is adjusted or analyzed based on the gradient mismatch. The gradient mismatch is expressed as:
[0055] Furthermore, the multi-agent system can be a microgrid distributed power system, a robot swarm system, a communication network resource allocation system, an intelligent transportation system, or an industrial multi-device collaborative control system.
[0056] Furthermore, when the multi-agent system is a microgrid distributed power system, the first... The agent is the first A distributed generation unit, the actual output For the first The output power of each distributed generation unit, wherein the time-varying cost function is:
[0057] in, This represents the local power generation cost coefficient. This indicates a time-varying electricity price or incentive signal. This represents the power imbalance penalty coefficient. This indicates time-varying load demand or capacity demand.
[0058] It is understood that the multi-agent system of the present invention is applicable to microgrid distributed power systems, robot swarm systems, communication network resource allocation systems, intelligent transportation systems, and industrial multi-device collaborative control systems, and all of the above applications fall within the protection scope of the present invention.
[0059] Example 2 In another embodiment of the present invention, to solve the above-mentioned technical problem, the present invention provides a distributed online output feedback control method for uncertain multi-agent systems oriented towards time-varying non-cooperative game theory, wherein the multi-agent system includes An intelligent agent, such as Figure 3The diagram illustrates the two-layer control architecture of the multi-agent system of this invention, where the output variable of each agent serves as its actual decision variable in a time-varying non-cooperative game. Figure 2 As shown, the method includes the following steps.
[0060] S1: Establish a dynamic model of uncertain output of the agent: For the first An agent establishes the following uncertain output dynamics model:
[0061] in, , For the first The state variables of an agent, To control the input, For output variables, For unknown or uncertain parameters, , , This refers to the system matrix or vector associated with the unknown parameters.
[0062] The output variable As the first The actual decision variables of an agent in a time-varying non-cooperative game.
[0063] S2: Constructing a time-varying non-cooperative game model Set a time-varying cost function for each agent:
[0064] in, For the first The output decision variables of an agent To exclude the first The set of output variables of other agents besides the one agent. The variable is time. Each agent aims to minimize its own time-varying cost function and, under conditions of mutual coupling, forms a dynamic equilibrium trajectory that varies with time.
[0065] S3: Construct neighbor communication topology; such as Figure 4 The diagram shown is a neighbor communication topology diagram of the present invention; Construct a communication graph between agents:
[0066] in, A set of intelligent agent nodes. For the set of communication edges, This is the adjacency weight matrix. For the first The agent directs the first The communication weight of each agent sending information. Each agent only receives... It receives estimation information or reference state information sent by its neighboring intelligent agents.
[0067] Understandably, the communication topology can be a directed or undirected graph that meets connectivity requirements, and the adjacency weights can be adjusted according to network reliability, communication bandwidth, or engineering deployment constraints.
[0068] S4: Constructing an upper-level distributed online reference generator For the first Each agent constructs a local estimation vector:
[0069] in, Indicates the first The first agent on the first Estimation of decision variables for an agent; express No. The reference signal corresponding to each agent.
[0070] Based on neighbor estimation errors and local gradient information, the following distributed online reference generation dynamics are constructed:
[0071] in, For consistency adjustment parameters, To select a vector, For the first Local gradient feedback signals for each agent.
[0072] In one alternative implementation, the local gradient feedback signal is the gradient at the reference variable:
[0073] in, Indicates the first A set of estimates of decision variables of other agents by one agent.
[0074] S5: Constructing the integral error state According to the The actual output of each agent With reference signal Construct the output tracking error:
[0075] And construct the integral error state:
[0076] in, Used to compensate for long-term deviations between the actual output and the reference signal.
[0077] S6: Construct a dynamic output feedback auxiliary state Regarding the first Each agent constructs an auxiliary output feedback state. Its dynamics are as follows:
[0078] in, For the relative order of the system, This is the output feedback gain. The auxiliary output feedback state is used to construct dynamic feedback information based on the measurable output signal.
[0079] S7: Generate lower-level output feedback control input Based on the integral error state, output tracking error, and auxiliary output feedback state, construct the first... Control inputs for each agent:
[0080] in, , This refers to the controller gain parameter. The control input... Apply to the
[0081] An intelligent agent, causing the intelligent agent to actually output Tracking reference signal .
[0082] It is understandable that the integral error state, auxiliary output feedback state, and controller gain in the lower-level output feedback controller can be equivalently adjusted according to the relative order of the agent, output measurement noise, and actuator characteristics.
[0083] S8: Construct a real-time gradient feedback mechanism; such as... Figure 5 The diagram shown is a schematic of the real-time gradient feedback mechanism of the present invention.
[0084] In the An agent cannot obtain the analytical form of the cost function, or cannot calculate the gradient at the reference variable. In this case, the local gradient feedback signal in step S4 is set to the real-time gradient at the actual output:
[0085] This yields the real-time gradient reference generation dynamics:
[0086] This real-time gradient feedback method is used to achieve online reference generation when the analytical gradient of the cost function is unavailable or can only be measured based on the actual output.
[0087] It is understandable that the local gradient feedback signal can be either the local gradient at the reference variable or the real-time gradient at the actual output; the real-time gradient can be estimated by online measurement, numerical difference, sensor feedback, operating cost assessment module or learning model.
[0088] Example 3 In another embodiment of the present invention, a distributed online output feedback control method for uncertain multi-agent systems oriented towards time-varying non-cooperative game theory is proposed. The multi-agent system includes... There are 3 intelligent agents, of which the 1st is the 2nd agent. An intelligent agent satisfies the following uncertain dynamics:
[0089] in, For the agent's state variables, To control the input, For output variables, These are unknown or uncertain parameters. The output variable... This represents the actual decision variables of the agent in a non-cooperative game. Because... The physical output is generated by system dynamics and cannot be directly assigned a value through optimization algorithms; therefore, it requires control input. Indirect regulation is carried out.
[0090] Set a time-varying cost function for each agent:
[0091] in, This represents the set of output variables of other agents. Each agent participates in a time-varying non-cooperative game by minimizing its own cost function.
[0092] Multiple agents communicate through a communication graph To exchange information. Each agent only receives estimation information sent by its neighboring agents.
[0093] For the Each agent constructs a local estimation vector:
[0094] in, Indicates the first The first agent on the first Each agent outputs an estimate of the decision variables, and For the first The reference signal for each agent.
[0095] Constructing a distributed online reference generator:
[0096] The first term in the above formula is used to coordinate neighbor estimation information, and the second term is used for online game updates based on the gradient of the local cost function. Thus, each agent can generate its own reference signal without the need for a central controller.
[0097] Furthermore, construct the output tracking error:
[0098] And introduce the integral error state:
[0099] Then, construct the auxiliary output feedback state:
[0100] Based on the above state, generate control input:
[0101] After the control input is applied to the i-th agent, the actual output is... Tracking reference signal This embodiment achieves distributed online output control of an uncertain multi-agent system under time-varying non-cooperative game conditions through the cooperation of an upper-layer reference generator and a lower-layer output feedback controller.
[0102] Example 4 This embodiment applies the distributed online output feedback control method for uncertain multi-agent systems to the online scheduling of distributed power sources in microgrids. For example... Figure 6 The diagram shown is a schematic of the online scheduling and control structure for distributed power sources in a microgrid according to the present invention. Multiple intelligent agents represent multiple distributed generation units. The output of each agent Indicates the first The output power of each distributed generation unit.
[0103] Set the first The time-varying cost function of a distributed generation unit is:
[0104] in, Indicates the cost of local power generation. This indicates an incentive or benefit item related to electricity pricing. Indicates total power output and demand Penalties for deviations between them For time-varying electricity prices or incentive signals, This refers to time-varying load demand or capacity demand.
[0105] Each distributed generation unit exchanges estimation information through neighbor communication and uses the distributed online reference generator and output feedback controller described in this invention to adjust its own power output, enabling multiple distributed generation units to achieve online coordinated control under changing electricity prices and load demand conditions.
[0106] In summary, compared with the prior art, the present invention has at least the following beneficial effects: This invention is applicable to intelligent agent systems with physical and dynamic constraints. The invention outputs the intelligent agent... As a decision variable in actual game theory, and through controlling the input... Indirectly adjusting the output, compared to distributed optimization methods that directly update abstract decision variables, is more suitable for real-world engineering systems.
[0107] It can track time-varying equilibrium trajectories. This invention generates reference signals in real time through a distributed online reference generator, enabling the agent to track dynamic equilibrium trajectories induced by time-varying cost functions, thus overcoming the problem that traditional static equilibrium search methods are difficult to adapt to time-varying environments.
[0108] This approach achieves hierarchical decoupling between online game reference generation and physical output control. The upper layer is responsible for generating reference signals based on neighbor communication and local gradients, while the lower layer drives the physical output to track the reference signals. This reduces the design complexity when directly applying online game algorithms to complex physical systems.
[0109] Suitable for distributed deployment. Each agent only needs to use local cost function information, actual output information, and neighbor communication information, without relying on a central controller or global information, which can reduce communication burden and the risk of single point of failure in centralized control.
[0110] Enhanced robustness to unknown parameters and unmeasurable states. This invention generates control inputs through integral error states and dynamic output feedback auxiliary states, enabling output tracking even when system parameters are unknown and the complete state is unmeasurable.
[0111] Adapted to real-time gradient feedback scenarios. This invention allows the gradient at the actual output to be used instead of the gradient at the ideal reference variable, making the method applicable to engineering systems where the analytical form of the cost function is unavailable and the gradient can only be obtained through measurement or estimation.
[0112] Reduce cumulative operational losses. Through the synergistic effect of upper-layer dynamic reference generation and lower-layer output tracking control, the actual output of each agent can continuously approach the dynamic equilibrium trajectory, thereby reducing the cumulative performance loss caused by the output deviating from the equilibrium trajectory.
[0113] Example 5 To implement the methods of the above embodiments, the present invention also provides a distributed online output feedback control system 10 for an uncertain multi-agent system, such as... Figure 7 As shown, it includes: The output acquisition module 11 is used to acquire the actual output of each intelligent agent; The local cost information acquisition module 12 is used to acquire the local cost function information or local gradient feedback signal of each agent. The neighbor communication module 13 is used to receive estimation information sent by neighboring intelligent agents; The distributed reference generation module 14 is used to generate a reference signal based on the local estimation vector, the estimation information sent by the neighboring agents, and the local gradient feedback signal. Integration error module 15 is used to construct an integration error state based on the actual output and the reference signal; The dynamic output feedback module 16 is used to construct a dynamic output feedback auxiliary state based on the actual output. The control input generation module 17 is used to generate control input based on the integral error state, output tracking error, and dynamic output feedback auxiliary state. The execution control module 18 is used to apply the control input to the corresponding intelligent agent, so that the actual output of the intelligent agent tracks the reference signal in real time.
[0114] Specifically, the output acquisition module is used to acquire the first... The actual output of each agent ; The local cost information acquisition module is used to acquire the first... Local cost function information or local gradient feedback signal of an agent; The neighbor communication module is used to receive estimation information sent by neighboring intelligent agents; The distributed reference generation module is used to generate a reference signal based on the local estimation vector, the estimation information sent by neighboring agents, and the local gradient feedback signal. ; The integration error module is used to calculate the actual output. With the reference signal Construct the integral error state; The dynamic output feedback module is used to adjust the output based on the actual output. Construct a dynamic output feedback auxiliary state; The control input generation module is used to generate control inputs based on the integral error state, output tracking error, and dynamic output feedback auxiliary state. ; The execution control module is used to transmit the control input. Apply to the The first intelligent agent, making the first The actual output of each agent Tracking the reference signal .
[0115] Furthermore, the distributed reference generation module is configured to perform the following reference generation dynamics:
[0116] in, Local gradient at the reference variable Or the real-time gradient at the actual output. .
[0117] The distributed online output feedback control system for uncertain multi-agent systems in this invention addresses the problem that existing distributed optimization, Nash equilibrium search, and multi-agent control methods struggle to simultaneously adapt to time-varying cost functions, physical output constraints, unknown system parameters, incomplete state measurement, neighbor communication limitations, and real-time gradient feedback.
[0118] Example 6 To implement the above embodiments and the methods described above, the present invention also provides a computer device, such as... Figure 8 As shown, the computer device 600 includes a memory 601 and a processor 602; wherein, the processor 602 reads executable program code stored in the memory 601 to run a program corresponding to the executable program code, so as to implement the various steps of the method described above.
[0119] Example 7 To implement the above embodiments, this application also proposes a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the method described in the foregoing embodiments.
[0120] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
[0121] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0122] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
Claims
1. A distributed online output feedback control method for an uncertain multi-agent system, characterized in that, A multi-agent system comprises multiple agents, each of which uses its output variable as its actual decision variable in a time-varying non-cooperative game. The method includes: Obtain the actual output of each agent, local cost function information, and estimation information sent by neighboring agents; For each agent, a local estimation vector is constructed, which contains the agent's estimates of the decision variables of other agents and the agent's reference signal; Based on the local estimation vector, the estimation information sent by neighboring agents, and the local gradient feedback signal, a distributed online reference generation dynamic is constructed to generate the reference signal; The output tracking error is constructed based on the actual output and the reference signal, and the integral error state is constructed accordingly. Construct a dynamic output feedback auxiliary state based on the actual output; The control input is generated based on the integral error state, the output tracking error, and the dynamic output feedback auxiliary state. The control input is applied to the corresponding intelligent agent, so that the actual output of the intelligent agent tracks the reference signal in real time.
2. The distributed online output feedback control method for an uncertain multi-agent system according to claim 1, characterized in that, Each agent satisfies an uncertain output dynamics model, which includes the agent's state variables, control inputs, output variables, unknown or uncertain parameters, and system matrices or vectors related to the unknown or uncertain parameters.
3. The distributed online output feedback control method for an uncertain multi-agent system according to claim 1, characterized in that, In the time-varying non-cooperative game, each agent has a time-varying cost function, which includes the agent's output decision variables, the set of output decision variables of other agents besides the agent, and the time variable.
4. The distributed online output feedback control method for an uncertain multi-agent system according to claim 1, characterized in that, The neighboring agents interact with each other through a communication graph, which includes a set of agent nodes, a set of communication edges, and an adjacency weight matrix. The elements in the adjacency weight matrix represent the communication weights between the corresponding agents.
5. The distributed online output feedback control method for an uncertain multi-agent system according to claim 1, characterized in that, The construction of the distributed online reference generation dynamic includes: Based on the consistency adjustment parameters, communication weights, selection vectors, and the local gradient feedback signals of the corresponding agents, dynamic equations are established to generate reference signals.
6. The distributed online output feedback control method for an uncertain multi-agent system according to claim 5, characterized in that, The acquisition of the local gradient feedback signal includes: Calculate the partial derivative of the agent's cost function with respect to its own decision variables. The calculation of this partial derivative is based on the agent's estimation set of decision variables for other agents besides itself. The resulting partial derivative is the local gradient at the reference variable.
7. The distributed online output feedback control method for an uncertain multi-agent system according to claim 5, characterized in that, When the corresponding agent cannot obtain the local gradient at the reference variable, the acquisition of the local gradient feedback signal includes: The real-time gradient at the actual output of the agent is obtained through online measurement, numerical difference, sensor feedback, operation cost evaluation module or learning model estimation, and this real-time gradient is used as the local gradient feedback signal.
8. The distributed online output feedback control method for an uncertain multi-agent system according to claim 1, characterized in that, Constructing the output tracking error and the integral error states includes: The output tracking error is constructed based on the difference between the agent's actual output and the reference signal; Establish a dynamic relationship between the integral error states, so that the integral error states change with time integration, and these integral error states are error state variables specific to the corresponding agent.
9. The distributed online output feedback control method for an uncertain multi-agent system according to claim 1, characterized in that, Constructing the dynamic output feedback auxiliary state includes: Determine the relative order and output feedback gain of the corresponding agent; Based on the relative order, output feedback gain, and the actual output of the agent, a dynamic equation is established to generate a dynamic output feedback auxiliary state containing multiple components.
10. The distributed online output feedback control method for an uncertain multi-agent system according to claim 1, characterized in that, Generating the control input includes: Based on preset controller gain parameters, combined with integral error state, output tracking error and dynamic output feedback auxiliary state, control input is generated through linear combination or preset operation relationship. The controller gain parameters include multiple preset adjustment coefficients.
11. The distributed online output feedback control method for an uncertain multi-agent system according to claim 1, characterized in that, The method further includes: Calculate the difference between the local gradient at the reference variable and the real-time gradient at the actual output, and define this difference as the gradient mismatch. Based on the gradient mismatch, adjust the parameter settings of the distributed online reference generation dynamics, or analyze the closed-loop interaction between the distributed online reference generation dynamics and the control input. The gradient mismatch is calculated by subtracting the real-time gradient at the actual output from the local gradient at the reference variable.
12. The distributed online output feedback control method for an uncertain multi-agent system according to claim 1, characterized in that, The multi-agent system is selected from microgrid distributed power systems, robot swarm systems, communication network resource allocation systems, intelligent transportation systems, or industrial multi-device collaborative control systems.
13. The distributed online output feedback control method for an uncertain multi-agent system according to claim 12, characterized in that, When the multi-agent system is a microgrid distributed power system, the specific implementation steps include: Each distributed generation unit is treated as a corresponding intelligent agent; The output power of each distributed generation unit is taken as the actual output of the intelligent agent; Construct a time-varying cost function that includes the local generation cost coefficient, the time-varying electricity price or incentive signal, the power imbalance penalty coefficient, and the time-varying load demand or capacity demand.
14. A distributed online output feedback control system for an uncertain multi-agent system, characterized in that, include: The output acquisition module is used to acquire the actual output of each agent; The local cost information acquisition module is used to acquire the local cost function information or local gradient feedback signal of each agent; The neighbor communication module is used to receive estimation information sent by neighboring intelligent agents; The distributed reference generation module is used to generate a reference signal based on the local estimation vector, the estimation information sent by neighboring agents, and the local gradient feedback signal. The integration error module is used to construct an integration error state based on the actual output and the reference signal; A dynamic output feedback module is used to construct a dynamic output feedback auxiliary state based on the actual output. The control input generation module is used to generate control inputs based on the integral error status, output tracking error, and dynamic output feedback auxiliary status. The execution control module is used to apply the control input to the corresponding intelligent agent, so that the actual output of the intelligent agent tracks the reference signal in real time.
15. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program that, when executed by the processor, implements the method according to any one of claims 1 to 13.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 13.