Layered self-adaptive weighted distributed learning control method and device for multi-agent system, equipment and medium
By adopting a hierarchical adaptive weighted distributed learning control method, the convergence rate of multi-agent systems is improved, and the problem of slow convergence in distributed learning is solved. This method is applicable to large-scale networked systems such as collaborative operation of industrial robots, smart grid regulation and control, and logistics transportation scheduling.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- RENMIN UNIVERSITY OF CHINA
- Filing Date
- 2026-01-09
- Publication Date
- 2026-04-21
AI Technical Summary
Distributed learning convergence rates in multi-agent systems are slow, affecting system deployment efficiency and engineering practicality, while centralized tuning methods compromise system scalability.
A hierarchical adaptive weighted distributed learning control method is adopted. By determining the adaptive weights of the agents' neighbors, cooperative errors, and component adaptive weights, the control signal is updated until the learning stopping condition is met, thereby improving the convergence rate.
It improves the distributed learning convergence rate of multi-agent systems, reduces the time required to reach the global goal, saves resources, avoids the risk of single point of failure, and is suitable for the optimization control of large-scale networked systems.
Smart Images

Figure CN121900169A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent control technology, and in particular to a hierarchical adaptive weighted distributed learning control method, apparatus, device, and medium for multi-agent systems. Background Technology
[0002] Multi-agent systems consist of multiple autonomous agents, which can be software entities, robots, or other units with perception, decision-making, and execution capabilities. Through information sharing, resource coordination, and cooperation or competition mechanisms, multi-agent systems can efficiently handle complex, large-scale tasks that a single agent cannot accomplish, demonstrating excellent flexibility and scalability.
[0003] To overcome the limitations of time-domain feedback in tracking time-varying reference trajectories, related technologies introduce a distributed learning mechanism to learn multi-agent systems. This mechanism is a lightweight learning strategy suitable for multi-agent systems that run repeatedly in batches. Each agent does not need to rely on global information but learns from the historical running data of its neighbors. Ultimately, all agents achieve zero-error trajectory tracking over a finite time interval.
[0004] However, the convergence rate of distributed learning in multi-agent systems is relatively slow in related technologies. Summary of the Invention
[0005] This invention provides a hierarchical adaptive weighted distributed learning control method, apparatus, device, and medium for multi-agent systems, which addresses the slow convergence rate of distributed learning in multi-agent systems in related technologies and improves the convergence rate of distributed learning in multi-agent systems.
[0006] In a first aspect, the present invention provides a hierarchical adaptive weighted distributed learning control method for a multi-agent system, comprising: Based on the dynamic model of the multi-agent system, the desired trajectory, the communication network, and the initial control signal of each agent in the multi-agent system, determine the neighbor adaptive weight of each agent; The cooperation error of each agent is determined based on the neighbor adaptive weights of each agent. Based on the cooperation error of each agent, determine the adaptive weights of each agent's components; Based on the adaptive weights of each agent's components and the cooperation error, the control signal of each agent is updated to obtain a new control signal for each agent. The new control signal of each agent is used as the current control signal. The process of determining the neighbor adaptive weight of each agent based on the dynamic model of the multi-agent system, the desired trajectory, the communication network, and the initial control signal of each agent in the multi-agent system is repeated until the set learning stopping condition is met, and the target multi-agent system that has completed learning is obtained.
[0007] Optionally, the dynamic model of the multi-agent system is as follows: ; in, Indicates the agent's serial number. For runtime, For the running length, This represents the number of iterations. , and The first The agent in the th... k The runtime of the next iteration t The state, control signals and outputs ; , and They represent the first An intelligent agent in The state matrix, input matrix, and output matrix at each time step; This refers to resetting the agent's initial state to a fixed state in each iteration. ; The desired trajectory The multi-agent system includes at least one agent capable of obtaining the desired trajectory and at least one follower agent that fails to obtain the desired trajectory. The communication network consists of a directed graph. express, These represent the set of nodes, the set of directed edges, and the adjacency matrix, respectively. ,express Each node represents one of the aforementioned intelligent agents; This represents the communication connection between the intelligent agents, and the edge In a communication connection, the agent Receiving intelligent agents Information; , express Middle intelligence agent i With intelligent agents j The element corresponding to the connecting edge; node The neighbor set is represented as ; where, if and only if hour, If and only if hour, .
[0008] Optionally, determining the neighbor adaptive weights for each agent based on the dynamic model of the multi-agent system, the desired trajectory, the communication network, and the initial control signals of each agent in the multi-agent system includes: In the directed graph Add a virtual node As a virtual leader; The virtual leader is controlled to operate according to a fixed, desired trajectory; wherein, the intelligent agent... If the desired trajectory information can be obtained, then the intelligent agent is determined. To receive information from the virtual leader, record The intelligent agent If the desired trajectory cannot be obtained, then the intelligent agent is determined to be invalid. To prevent receiving information from the virtual leader, record ; The directed graph is... Adjusted to a new directed graph ;in, ; ; For intelligent agents In the figure Neighbor set in; virtual node It is a picture It is a root node, meaning it is relative to any other node. There exists a path that makes Information can be transmitted along the path to ,Right now .
[0009] Based on the dynamic model of the multi-agent system, the desired trajectory, the new directed graph, and the initial control signal of each agent, the neighbor adaptive weights of each agent are determined.
[0010] Optionally, determining the neighbor adaptive weights of each agent based on the dynamic model of the multi-agent system, the desired trajectory, the new directed graph, and the initial control signals of each agent includes: For any of the aforementioned intelligent agents Utilizing the intelligent agent Calculate the relative error from the output of the neighboring nodes. If the intelligent agent If the desired trajectory information can be obtained, then the tracking error can be calculated. If the intelligent agent If the desired trajectory information is not available, the tracking error is not calculated. The intelligent agent The error data is stacked to obtain error stacked data, which is... , , , T It is the transpose symbol; Using the intelligent agent The error accumulation data is used to calculate the intelligent agent. Adaptive neighbor weights.
[0011] Optionally, the intelligent agent The neighbor adaptive weights are: ; ; in, , and It is a parameter; The intelligent agent The cooperation error is: ; in, It comes from the intelligent agent. The output of the neighboring nodes, It is the expected trajectory represented by a stacked representation.
[0012] Optionally, the intelligent agent The adaptive weights of the components are: ; ; in, Representing vectors The One portion, , and It is a parameter; The intelligent agent The new control signal is: ; in, For step size gain, ; It is a matrix gain; It is the input signal represented by a stack. This is the new control signal, which is the input signal used in the next iteration.
[0013] Optionally, the parameter constraints satisfy set conditions, wherein the set conditions are: It is a diagonal matrix; ; in, For the intelligent agent The system matrix, stacked as a diagonal block matrix. ; The stacked representation of matrix gain satisfies: ; in, express eigenvalues, Indicates the real part; Step gain satisfies: ; This allows the agent network to dynamically compress. The spectral radius is less than 1; Get parameters for and Proportional adjustment for , and Proportional adjustment Starting from zero and gradually increasing the values, ensure the compressibility of the agent network. The spectral radius is less than 1, which makes the error... Approaching zero; .
[0014] In a second aspect, the present invention provides a hierarchical adaptive weighted distributed learning control device for a multi-agent system, comprising: The first determining unit is used to determine the neighbor adaptive weight of each agent based on the dynamic model of the multi-agent system, the desired trajectory, the communication network, and the initial control signal of each agent in the multi-agent system. The second determining unit is used to determine the cooperation error of each agent based on the neighbor adaptive weights of each agent. The third determining unit is used to determine the adaptive weights of each agent's components based on the cooperation error of each agent. The signal update unit is used to update the control signal of each agent according to the adaptive weights and cooperation errors of each agent, so as to obtain a new control signal for each agent. An iterative learning unit is used to take the new control signal of each agent as the current control signal, trigger the first determining unit, until the set learning stop condition is met, and obtain the target multi-agent system that has completed learning.
[0015] Thirdly, the present invention provides a computer device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the hierarchical adaptive weighted distributed learning control method for a multi-agent system described in the first aspect or any corresponding embodiment thereof.
[0016] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the hierarchical adaptive weighted distributed learning control method for a multi-agent system described in the first aspect or any of its corresponding embodiments.
[0017] This invention addresses the trajectory tracking control problem of multi-agent systems with repetitive motion characteristics. It proposes a hierarchical adaptive weighted distributed learning control method for multi-agent systems, including establishing the multi-agent system and the desired trajectory, building a communication network for the multi-agent system, designing a distributed learning control algorithm to improve the input signal, and verifying the convergence conditions of the distributed learning control algorithm to achieve the tracking task. This method can improve the convergence speed of multi-agent systems in iterative environments, reduce the time required to reach the global goal, save resources, and efficiently realize cooperative tasks in networked industrial systems with repetitive motion characteristics.
[0018] This invention constructs lightweight, hierarchical adaptive weights that can be easily embedded into existing distributed learning control frameworks. Employing a distributed architecture, it eliminates the need for centralized information processing, ensuring efficient trajectory tracking of multi-agent systems while effectively avoiding single-point-of-failure risks and weak system robustness. Thanks to its excellent scalability, it is widely applicable to optimization control scenarios in large-scale networked systems such as collaborative industrial robot operations, smart grid control, and logistics transportation scheduling, demonstrating outstanding engineering practical value.
[0019] This invention proposes a hierarchical adaptive weighted distributed learning control method, apparatus, device, and medium for multi-agent systems. Based on the dynamic model of the multi-agent system, the desired trajectory, the communication network, and the initial control signals of each agent in the system, the method determines the neighbor adaptive weights for each agent. Based on these neighbor adaptive weights, the cooperative error of each agent is determined. Based on the cooperative error, the component adaptive weights of each agent are determined. Based on the component adaptive weights and cooperative error, the control signals of each agent are updated, resulting in new control signals for each agent. These new control signals are then used as the current control signals, and the process returns to the previous steps of determining the neighbor adaptive weights based on the dynamic model, desired trajectory, communication network, and initial control signals of each agent in the system, until a predetermined learning termination condition is met, resulting in a target multi-agent system that has completed learning. This invention can effectively improve the convergence rate of distributed learning in multi-agent systems. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in this invention or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 A flowchart of a hierarchical adaptive weighted distributed learning control method for a multi-agent system provided in an embodiment of the present invention; Figure 2 A flowchart of another hierarchical adaptive weighted distributed learning control method for a multi-agent system provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an intelligent agent communication network provided in an embodiment of the present invention; Figure 4 This is a schematic diagram showing the comparison of the number of iterations of a multi-agent system under different learning control methods, provided by an embodiment of the present invention. Figure 5 This is a schematic diagram of the structure of a hierarchical adaptive weighted distributed learning control device for a multi-agent system provided in an embodiment of the present invention; Figure 6 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0023] In related technologies, distributed learning control asymptotically converges with increasing iteration count, but practical systems cannot support an infinite number of runs. Slow convergence significantly reduces deployment efficiency and hinders engineering practicality. On the other hand, using optimization methods that rely on global information or centralized parameter adjustments compromises system scalability and makes it unsuitable for large-scale agent networks. Therefore, this embodiment proposes a hierarchical adaptive weighted distributed learning control method for multi-agent systems to improve the convergence rate of distributed learning in multi-agent systems.
[0024] The following is combined with Figures 1-4 This invention describes a hierarchical adaptive weighted distributed learning control method for multi-agent systems.
[0025] like Figure 1 As shown, this embodiment proposes a first hierarchical adaptive weighted distributed learning control method for multi-agent systems, which may include the following steps: S101. Based on the dynamic model of the multi-agent system, the desired trajectory, the communication network, and the initial control signals of each agent in the multi-agent system, determine the neighbor adaptive weights of each agent.
[0026] Specifically, this embodiment can establish a dynamic model of a multi-agent system, determine the desired trajectory, and the multi-agent system includes several agents capable of obtaining the desired trajectory and other follower agents. A directed graph is constructed. The communication network of a multi-agent system.
[0027] Alternatively, the dynamic model of a multi-agent system is as follows: ; in, Indicates the agent's serial number. For runtime, For the running length, This represents the number of iterations. , and The first The agent in the th... k The runtime of the next iteration t The state, control signals and outputs ; , and They represent the first An intelligent agent in The state matrix, input matrix, and output matrix at each time step; This refers to resetting the agent's initial state to a fixed state in each iteration. ; Expected trajectory A multi-agent system includes at least one agent capable of acquiring the desired trajectory and at least one follower agent that fails to acquire the desired trajectory; only a subset of agents can acquire the desired trajectory. Signal.
[0028] Communication networks consist of directed graphs express, These represent the set of nodes, the set of directed edges, and the adjacency matrix, respectively. ,express There are 1 intelligent agents, and each node represents an intelligent agent; Represents the communication connection between intelligent agents, edge In a communication connection, the agent Receiving intelligent agents Information; , express Middle intelligence agent i With intelligent agents j The element corresponding to the connecting edge; node The neighbor set is represented as ; where, if and only if hour, If and only if hour, .
[0029] Optionally, step S101 includes; In directed graphs Add a virtual node As a virtual leader; The virtual leader is controlled to operate according to a fixed, desired trajectory; among which, the intelligent agent... If the desired trajectory information can be obtained, then the intelligent agent is determined. To receive information from the virtual leader, record intelligent agent If the desired trajectory cannot be obtained, then the agent is determined. To prevent receiving information from the virtual leader, record ; The directed graph is generated by Adjusted to a new directed graph ;in, ; ; For intelligent agents In the figure Neighbor set in; virtual node It is a picture It is a root node, meaning it is relative to any other node. There exists a path that makes Information can be transmitted along the path to ,Right now .
[0030] Based on the dynamic model of the multi-agent system, the desired trajectory, the new directed graph, and the initial control signals of each agent, the adaptive neighbor weights of each agent are determined.
[0031] Specifically, this embodiment can design a distributed learning control algorithm to improve the input signal. In each iteration of the distributed learning control, any agent... Perform locally distributed computation.
[0032] Optional, such as Figure 2 As shown, the above method determines the neighbor adaptive weights for each agent based on the dynamic model of the multi-agent system, the desired trajectory, the new directed graph, and the initial control signals of each agent, including: For any intelligent agent Utilizing from intelligent agents Calculate the relative error from the output of the neighboring nodes. If the intelligent agent If the desired trajectory information can be obtained, then the tracking error can be calculated. If the intelligent agent If the desired trajectory information is not available, the tracking error is not calculated. intelligent agents The error data is stacked to obtain stacked error data, which is: , , , T It is the transpose symbol; Using intelligent agents Error accumulation data, computing intelligent agents Adaptive neighbor weights.
[0033] Optional, intelligent agent The neighbor adaptive weights are: ; ; in, , and It is a parameter.
[0034] S102. Determine the cooperation error of each agent based on the adaptive weights of each agent's neighbors.
[0035] Optional, intelligent agent The cooperation error is: ; in, It comes from the intelligent agent The output of the neighboring nodes, It is the expected trajectory represented by a stacked representation.
[0036] S103. Based on the cooperative error of each agent, determine the adaptive weights of each agent's components.
[0037] Optional, intelligent agent The adaptive weights of the components are: ; ; in, Representing vectors The One portion, , and It is a parameter.
[0038] S104. Based on the adaptive weights of each agent's components and the cooperation error, update the control signal of each agent to obtain a new control signal for each agent.
[0039] Optional, intelligent agent The new control signal is: ; in, For step size gain, ; It is a matrix gain; It is the input signal represented by a stack. This is the new control signal, which is the input signal used in the next iteration.
[0040] S105. Take the new control signal of each agent as the current control signal, return to execute the step of determining the neighbor adaptive weight of each agent based on the dynamic model of the multi-agent system, the desired trajectory, the communication network, and the initial control signal of each agent in the multi-agent system, until the set learning stopping condition is met, and obtain the target multi-agent system that has completed learning.
[0041] The learning termination condition can be a requirement that the iteration termination time, number of iterations, or error magnitude are met, and this embodiment does not impose any limitations.
[0042] Specifically, this embodiment can design parameter constraints to verify the distributed learning control algorithm.
[0043] Optionally, the parameter constraints must satisfy the set conditions, in which: It is a diagonal matrix; ; in, For intelligent agents The system matrix, stacked as a diagonal block matrix. ; The stacked representation of matrix gain satisfies: ; in, express eigenvalues, Indicates the real part; Step gain satisfies: ; This allows the agent network to dynamically compress. The spectral radius is less than 1; Get parameters for and Proportional adjustment for , and Proportional adjustment Starting from zero and gradually increasing the values, ensure the compressibility of the agent network. The spectral radius is less than 1, which makes the error... Approaching zero; in, .
[0044] It should be noted that this embodiment addresses the trajectory tracking control problem of multi-agent systems with repetitive motion characteristics by proposing a hierarchical adaptive weighted distributed learning control method for multi-agent systems. This method includes establishing the multi-agent system and the desired trajectory, building a communication network for the multi-agent system, designing a distributed learning control algorithm to improve the input signal, and verifying the convergence conditions of the distributed learning control algorithm to achieve the tracking task. This method can improve the convergence speed of multi-agent systems operating in an iterative environment, reduce the time required to reach the global goal, save resources, and efficiently realize cooperative tasks in networked industrial systems with repetitive motion characteristics.
[0045] This embodiment constructs a lightweight, hierarchical adaptive weight, which can be easily embedded into existing distributed learning control frameworks. This embodiment can adopt a distributed architecture, eliminating the need for centralized information processing. While ensuring the multi-agent system efficiently completes trajectory tracking tasks, it effectively avoids single-point failure risks and weak system robustness. Thanks to its excellent scalability, it is widely applicable to optimization control scenarios in large-scale networked systems such as collaborative operation of industrial robots, smart grid control, and logistics transportation scheduling, demonstrating outstanding engineering practical value.
[0046] This embodiment proposes a hierarchical adaptive weighted distributed learning control method for multi-agent systems. Based on the dynamic model of the multi-agent system, the desired trajectory, the communication network, and the initial control signals of each agent, it determines the neighbor adaptive weights for each agent. Based on these neighbor adaptive weights, it determines the cooperation error of each agent. Based on the cooperation error, it determines the component adaptive weights for each agent. Based on the component adaptive weights and cooperation error, it updates the control signals of each agent, resulting in new control signals for each agent. Using these new control signals as the current control signals, it returns to the previous step of determining the neighbor adaptive weights based on the dynamic model, desired trajectory, communication network, and initial control signals of each agent, until a predetermined learning termination condition is met, resulting in a target multi-agent system that has completed learning. This embodiment can effectively improve the convergence rate of distributed learning in multi-agent systems.
[0047] To better explain Figure 1 The various steps in the process are described in Example 1 of this embodiment. Example 1 includes the following steps: Step 1: Establish a dynamic model of the multi-agent system and determine the desired trajectory. The multi-agent system includes 48 agents, and the expression for the dynamic model is as follows:
[0048] In the above formula, Indicates the first An intelligent agent in The state matrix, input matrix, and output matrix at time step are as follows: ; ; ; ; ; ; ; ; ; in, This refers to resetting the agent's initial state to a fixed state in each iteration. The desired trajectory can be randomly generated using a Gaussian distribution with a mean of 0 and a variance of 2; determining the desired trajectory includes: given the desired trajectory Generated by dynamic equations, the dynamic equations are as follows: ; ; ; Only agents at nodes 2, 4, 7, 17, 23, and 38 can obtain [the necessary information]. The signal establishes virtual node 49 as the virtual leader. Information from virtual leader 49 can flow to nodes 2, 4, 7, 17, 23, and 38, thus expanding the graph. Contains edges .
[0049] Step 2, as follows Figure 3 As shown, construct a directed graph The communication network of a multi-agent system, wherein the set of nodes express There are *n* agents, each node representing an agent, and a set of directed edges. This represents the communication connections between agents. Agents at nodes 2, 4, 7, 17, 23, and 38 can obtain... Signal, i.e. .
[0050] Step 3: Based on hierarchical adaptive weighted distributed learning control, design a distributed learning control algorithm to improve the input signal; the hierarchical adaptive weighted distributed learning control method includes the following steps: In each iteration, any agent... The computation is performed in a locally distributed manner as follows: (1) Calculate the relative error using the output from the neighbor. If the desired trajectory information can be obtained, then the tracking error is calculated. Otherwise, tracking error is not calculated; (2) Calculate the adaptive neighbor weights using the error: ; ; Among them, parameters ; (3) Calculate the cooperation error: ; (4) Calculate the adaptive weights of the components: ; ; in, Representing vectors The Each component, parameter ; (5) Update input:
[0051] Among them, step size gain Matrix gain .
[0052] Step 4: Set parameter constraints for the distributed learning control algorithm to enable the multi-agent system to quickly achieve consensus tracking as the number of iterations increases; the parameter constraints specifically include: (1) Construction It is a diagonal matrix. ; For intelligent agents The system matrix, stacked as a diagonal block matrix. ; The stacked representation of matrix gain satisfies ; in, express eigenvalues, Indicates the real part; (2) The step size gain satisfies ; This enables the agent network to dynamically compress, i.e. The spectral radius is less than 1; (3) Determine parameters and Proportional adjustment , and Proportional adjustment Starting from zero, gradually increase the value to a smaller value to ensure the compressibility of the agent network, thus guaranteeing... The spectral radius is less than 1, which makes the error... It approaches zero.
[0053] in, .
[0054] like Figure 4 As shown, the horizontal axis represents the given precision. The vertical axis represents the number of iterations required for all agents to achieve the desired accuracy under a given method. The two curves respectively represent the distributed learning control method in related technologies and the hierarchical adaptive weighted distributed learning control method proposed in this embodiment. In this figure, under the same accuracy requirement, the hierarchical adaptive weighted distributed learning control method of this embodiment requires approximately 2000 fewer iterations than the distributed learning control method in related technologies. Clearly, the hierarchical adaptive weighted distributed learning control method proposed in this embodiment can significantly improve the convergence speed of distributed learning in multi-agent systems and achieve high-precision tracking results more quickly.
[0055] like Figure 5 As shown, this embodiment proposes a hierarchical adaptive weighted distributed learning control device for a multi-agent system, the device comprising: The first determining unit 101 is used to determine the neighbor adaptive weight of each agent based on the dynamic model of the multi-agent system, the desired trajectory, the communication network, and the initial control signal of each agent in the multi-agent system. The second determining unit 102 is used to determine the cooperation error of each agent based on the neighbor adaptive weights of each agent; The third determining unit 103 is used to determine the adaptive weights of each agent's components based on the cooperative error of each agent. The signal update unit 104 is used to update the control signal of each agent according to the adaptive weight and cooperation error of each agent's components, so as to obtain the new control signal of each agent. The iterative learning unit 105 is used to take the new control signal of each agent as the current control signal, trigger the first determining unit 101, until the set learning stop condition is met, and obtain the target multi-agent system that has completed learning.
[0056] It should be noted that the processing procedures and beneficial effects of the first determining unit 101, the second determining unit 102, the third determining unit 103, the signal updating unit 104, and the iterative learning unit 105 can be referred to respectively. Figure 1 Steps S101 to S105 are not described in detail here.
[0057] Alternatively, the dynamic model of a multi-agent system is as follows: ; in, Indicates the agent's serial number. For runtime, For the running length, This represents the number of iterations. , and The first The agent in the th... k The runtime of the next iteration t The state, control signals and outputs ; , and They represent the first An intelligent agent in The state matrix, input matrix, and output matrix at each time step; This refers to resetting the agent's initial state to a fixed state in each iteration. ; Expected trajectory A multi-agent system includes at least one agent capable of acquiring the desired trajectory and at least one follower agent that fails to acquire the desired trajectory. Communication networks consist of directed graphs express, These represent the set of nodes, the set of directed edges, and the adjacency matrix, respectively. ,express There are 1 intelligent agents, and each node represents an intelligent agent; Represents the communication connection between intelligent agents, edge In a communication connection, the agent Receiving intelligent agents Information; , express Middle intelligence agent i With intelligent agents j The element corresponding to the connecting edge; node The neighbor set is represented as ; where, if and only if hour, If and only if hour, .
[0058] Optionally, the first determining unit 101 is also used for: In directed graphs Add a virtual node As a virtual leader; The virtual leader is controlled to operate according to a fixed, desired trajectory; among which, the intelligent agent... If the desired trajectory information can be obtained, then the intelligent agent is determined. To receive information from the virtual leader, record intelligent agent If the desired trajectory cannot be obtained, then the agent is determined. To prevent receiving information from the virtual leader, record ; The directed graph is generated by Adjusted to a new directed graph ;in, ; ; For intelligent agents In the figure Neighbor set in; virtual node It is a picture It is a root node, meaning it is relative to any other node. There exists a path that makes Information can be transmitted along the path to ,Right now .
[0059] Based on the dynamic model of the multi-agent system, the desired trajectory, the new directed graph, and the initial control signals of each agent, the adaptive neighbor weights of each agent are determined.
[0060] Optionally, the first determining unit 101 is also used for: For any intelligent agent Utilizing from intelligent agents Calculate the relative error from the output of the neighboring nodes. If the intelligent agent If the desired trajectory information can be obtained, then the tracking error can be calculated. If the intelligent agent If the desired trajectory information is not available, the tracking error is not calculated. intelligent agents The error data is stacked to obtain stacked error data, which is: , , , T It is the transpose symbol; Using intelligent agents Error accumulation data, computing intelligent agents Adaptive neighbor weights.
[0061] Optional, intelligent agent The neighbor adaptive weights are: ; ; in, , and It is a parameter; intelligent agent The cooperation error is: ; in, It comes from the intelligent agent The output of the neighboring nodes, It is the expected trajectory represented by a stacked representation.
[0062] Optional, intelligent agent The adaptive weights of the components are: ; ; in, Representing vectors The One portion, , and It is a parameter; intelligent agent The new control signal is: ; in, For step size gain, ; It is a matrix gain; It is the input signal represented by a stack. This is the new control signal, which is the input signal used in the next iteration.
[0063] Optionally, the parameter constraints must satisfy the set conditions, in which: It is a diagonal matrix; ; in, For intelligent agents The system matrix, stacked as a diagonal block matrix. ; The stacked representation of matrix gain satisfies: ; in, express eigenvalues, Indicates the real part; Step gain satisfies: ; This allows the agent network to dynamically compress. The spectral radius is less than 1; Get parameters for and Proportional adjustment for , and Proportional adjustment Starting from zero and gradually increasing the values, ensure the compressibility of the agent network. The spectral radius is less than 1, which makes the error... Approaching zero; .
[0064] The hierarchical adaptive weighted distributed learning control device for multi-agent systems proposed in this embodiment determines the neighbor adaptive weights of each agent based on the dynamic model of the multi-agent system, the desired trajectory, the communication network, and the initial control signals of each agent in the multi-agent system. Based on these neighbor adaptive weights, the cooperative error of each agent is determined. Based on the cooperative error, the component adaptive weights of each agent are determined. The control signals of each agent are updated according to their component adaptive weights and cooperative errors, resulting in new control signals for each agent. These new control signals are then used as the current control signals, and the process returns to the step of determining the neighbor adaptive weights of each agent based on the dynamic model of the multi-agent system, the desired trajectory, the communication network, and the initial control signals of each agent in the multi-agent system, until a set learning termination condition is met, resulting in a target multi-agent system that has completed learning. This embodiment can effectively improve the convergence rate of distributed learning in multi-agent systems.
[0065] In this embodiment, the hierarchical adaptive weighted distributed learning control device of the multi-agent system is presented in the form of a functional unit. Here, a unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0066] This invention also provides a computer device having the above-described features. Figure 5 The diagram shows a hierarchical adaptive weighted distributed learning control device for a multi-agent system.
[0067] Please see Figure 6The present invention provides a schematic diagram of the structure of a computer device according to an optional embodiment. The computer device includes one or more processors 10, a memory 20, and interfaces for connecting the various components, including high-speed interfaces and low-speed interfaces. The various components are interconnected via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, multiple processors and / or multiple buses can be used with multiple memories, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 6 Take a processor 10 as an example.
[0068] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.
[0069] The memory 20 stores instructions executable by at least one processor 10 to cause at least one processor 10 to perform the method shown in the above embodiments.
[0070] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function. The data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, which can be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0071] Memory 20 may include volatile memory, such as random access memory. Memory may also include non-volatile memory, such as flash memory, hard disk, or solid-state drive. Memory 20 may also include combinations of the above types of memory.
[0072] The computer device also includes a communication interface 30 for communicating with other devices or communication networks.
[0073] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.
[0074] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A hierarchical adaptive weighted distributed learning control method for multi-agent systems, characterized in that, include: Based on the dynamic model of the multi-agent system, the desired trajectory, the communication network, and the initial control signal of each agent in the multi-agent system, determine the neighbor adaptive weight of each agent; The cooperation error of each agent is determined based on the neighbor adaptive weights of each agent. Based on the cooperation error of each agent, determine the adaptive weights of each agent's components; Based on the adaptive weights of each agent's components and the cooperation error, the control signal of each agent is updated to obtain a new control signal for each agent. The new control signal of each agent is used as the current control signal. The process of determining the neighbor adaptive weight of each agent based on the dynamic model of the multi-agent system, the desired trajectory, the communication network, and the initial control signal of each agent in the multi-agent system is repeated until the set learning stopping condition is met, and the target multi-agent system that has completed learning is obtained.
2. The method according to claim 1, characterized in that, The dynamic model of the multi-agent system is as follows: ; in, Indicates the agent's serial number. For runtime, For the running length, This represents the number of iterations. , and The first The agent in the th... k The runtime of the next iteration t The state, control signals and outputs ; , and They represent the first An intelligent agent in The state matrix, input matrix, and output matrix at each time step; This refers to resetting the agent's initial state to a fixed state in each iteration. ; The desired trajectory The multi-agent system includes at least one agent capable of obtaining the desired trajectory and at least one follower agent that fails to obtain the desired trajectory. The communication network consists of a directed graph. express, These represent the set of nodes, the set of directed edges, and the adjacency matrix, respectively. ,express Each node represents one of the aforementioned intelligent agents; This represents the communication connection between the intelligent agents, and the edge In a communication connection, the agent Receiving intelligent agents Information; , express Middle intelligence agent i With intelligent agents j The element corresponding to the connecting edge; node The neighbor set is represented as ; where, if and only if hour, If and only if hour, .
3. The method according to claim 2, characterized in that, The step of determining the neighbor adaptive weights for each agent based on the dynamic model, desired trajectory, communication network, and initial control signals of each agent in the multi-agent system includes: In the directed graph Add a virtual node As a virtual leader; The virtual leader is controlled to operate according to a fixed, desired trajectory; wherein, the intelligent agent... If the desired trajectory information can be obtained, then the intelligent agent is determined. To receive information from the virtual leader, record The intelligent agent If the desired trajectory cannot be obtained, then the intelligent agent is determined to be invalid. To prevent receiving information from the virtual leader, record ; The directed graph is... Adjusted to a new directed graph ;in, ; ; For intelligent agents In the figure Neighbor set in; virtual node It is a picture It is a root node, meaning it is relative to any other node. There exists a path that makes Information can be transmitted along the path to ,Right now ; Based on the dynamic model of the multi-agent system, the desired trajectory, the new directed graph, and the initial control signal of each agent, the neighbor adaptive weights of each agent are determined.
4. The method according to claim 3, characterized in that, The step of determining the neighbor adaptive weights of each agent based on the dynamic model of the multi-agent system, the desired trajectory, the new directed graph, and the initial control signals of each agent includes: For any of the aforementioned intelligent agents Utilizing the intelligent agent Calculate the relative error from the output of the neighboring nodes. If the intelligent agent If the desired trajectory information can be obtained, then the tracking error can be calculated. If the intelligent agent If the desired trajectory information is not available, the tracking error is not calculated. The intelligent agent The error data is stacked to obtain error stacked data, which is... , , , T It is the transpose symbol; Using the intelligent agent The error accumulation data is used to calculate the intelligent agent. Adaptive neighbor weights.
5. The method according to claim 4, characterized in that, The intelligent agent The neighbor adaptive weights are: ; ; in, , and It is a parameter; The intelligent agent The cooperation error is: ; in, It comes from the intelligent agent. The output of the neighboring nodes, It is the expected trajectory represented by a stacked representation.
6. The method according to claim 5, characterized in that, The intelligent agent The adaptive weights of the components are: ; ; in, Representing vectors The One portion, , and It is a parameter; The intelligent agent The new control signal is: ; in, For step size gain, ; It is a matrix gain; It is the input signal represented by a stack. This is the new control signal, which is the input signal used in the next iteration.
7. The method according to any one of claims 1 to 6, characterized in that, The parameter constraints satisfy the set conditions, wherein: It is a diagonal matrix; ; in, For the intelligent agent The system matrix, stacked as a diagonal block matrix. ; The stacked representation of matrix gain satisfies: ; in, express eigenvalues, Indicates the real part; Step gain satisfies: ; This allows the agent network to dynamically compress. The spectral radius is less than 1; Get parameters for and Proportional adjustment for , and Proportional adjustment Starting from zero and gradually increasing the values, ensure the compressibility of the agent network. The spectral radius is less than 1, which makes the error... Approaching zero; 。 8. A hierarchical adaptive weighted distributed learning control device for a multi-agent system, characterized in that, include: The first determining unit is used to determine the neighbor adaptive weight of each agent based on the dynamic model of the multi-agent system, the desired trajectory, the communication network, and the initial control signal of each agent in the multi-agent system. The second determining unit is used to determine the cooperation error of each agent based on the neighbor adaptive weights of each agent. The third determining unit is used to determine the adaptive weights of each agent's components based on the cooperation error of each agent. The signal update unit is used to update the control signal of each agent according to the adaptive weights and cooperation errors of each agent, so as to obtain a new control signal for each agent. An iterative learning unit is used to take the new control signal of each agent as the current control signal, trigger the first determining unit, until the set learning stop condition is met, and obtain the target multi-agent system that has completed learning.
9. A computer device, characterized in that, include: A memory and a processor are interconnected, the memory stores computer instructions, and the processor executes the computer instructions to perform the hierarchical adaptive weighted distributed learning control method for a multi-agent system as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute the hierarchical adaptive weighted distributed learning control method for the multi-agent system according to any one of claims 1 to 7.
Citation Information
Patent Citations
Traffic signal control method for multi-agent reinforcement learning based on neighbor awareness
CN113435112A
Universal nonlinear multi-agent layered adaptive fault-tolerant cooperative control method
CN116661300A
Iterative learning control method based on error adaptive adjustment and medium
CN117193013A
Construction method and system of multi-agent adaptive synchronous iterative learning coordination controller
CN118409507A
Consistency tracking control method and apparatus for multi-agent system, device, and medium
WO2024183286A1