Three-pillar-driven cloud collaborative multi-agent deep reinforcement learning decision-making system

By leveraging a three-pillar cloud-based collaborative multi-agent deep reinforcement learning decision-making system, and utilizing self-attention mechanisms and width-based learning optimization strategy networks, this system addresses the convergence problem of multi-agent deep learning algorithms, enabling efficient collaborative decision-making in resource-constrained environments.

CN121436084AActive Publication Date: 2026-01-30BEIJING INST OF COMP TECH & APPL
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511519579.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2026-01-30
Estimated Expiration
2045-10-23

AI Technical Summary

Technical Problem

Multi-agent deep reinforcement learning algorithms are difficult to converge in real-world environments, especially under non-static and partially observable characteristics, and the complexity of cooperation among multiple agents increases, along with resource constraints.

Method used

A cloud-based collaborative multi-agent deep reinforcement learning decision-making system is adopted, which consists of a three-pillar drive, including an edge execution module, a cloud-based teacher guidance module, a cloud-based action collaboration module, a cloud-based parameter sharing module, and a cloud-based global optimization module. A dynamic collaborative graph is established using a self-attention mechanism, network parameter sharing is achieved through a federated collaboration mechanism, and a width learning optimization strategy is adopted for the network.

Benefits of technology

It accelerates the convergence speed of multi-agent deep learning, enhances the coordination among agents, improves system stability, and solves the problem of limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121436084A_ABST
    Figure CN121436084A_ABST
Patent Text Reader

Abstract

The invention relates to a three-pillar-driven cloud collaborative multi-agent deep reinforcement learning decision-making system, and belongs to the field of agent decision-making. The system comprises an end-side execution module, a cloud-side teacher guidance module, a cloud-side action cooperation module, a cloud-side parameter sharing module and a cloud-side global optimization module. According to the method, a cloud collaboration mode is utilized, a dynamic collaboration diagram is established based on a self-attention mechanism, network parameter sharing of the same task area is realized by utilizing a federated collaboration mechanism, a wide learning optimization strategy network is adopted, and a three-pillar architecture which is used for dynamic collaboration, parameter sharing and global optimization and ensures multi-agent deep reinforcement learning convergence is established; according to the method, the convergence speed of multi-agent deep reinforcement learning can be increased, the coordination among the agents is enhanced, and the stability of the system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of agent decision-making, and particularly relates to a three-pillar driven cloud collaborative multi-agent deep reinforcement learning decision-making system. BACKGROUND

[0002] With the increase of agents, the state dimension of the environment increases exponentially, leading to difficulties in algorithm convergence. In addition, considering the interaction between multiple agents and the time-varying of the global state, the change of an agent strategy will lead to the adjustment of all other agent strategies, so the process of multi-agent training, cooperation and task execution will be more complex, and there are many problems and challenges in the comprehensive application scenarios and requirements. Especially in the case of a large number of multi-agents and a long round of task execution, the current mainstream multi-agent deep reinforcement learning technology of "centralized training-distributed execution" architecture is prone to problems such as "curse of dimensionality" of state space and "history disaster" of iterative update, which needs to be optimized and improved in practical applications. In addition, in most real scenarios, the field of view and perception of the agent is limited, and the agent cannot obtain the state and behavior information of all other agents, and the agent cannot carry high-performance computing and large energy storage devices due to factors such as volume and load, which increases the instability of the agent modeling the environment and increases the difficulty of converging the joint strategy of multiple agents to the global optimum. SUMMARY

[0003] (I) Technical problems to be solved

[0004] The technical problem to be solved by the present application is how to provide a three-pillar driven cloud collaborative multi-agent deep reinforcement learning decision-making system to solve the problem of difficult convergence of multi-agent deep reinforcement learning algorithm caused by non-static and partially observable characteristics of actual environment and the problem of limited resources at the end of multi-agent.

[0005] (II) Technical solutions

[0006] In order to solve the above technical problems, the present application provides a three-pillar driven cloud collaborative multi-agent deep reinforcement learning decision-making system, which comprises: an end-side execution module, a cloud-side teacher guidance module, a cloud-side action cooperation module, a cloud-side parameter sharing module and a cloud-side global optimization module.

[0007] The end-side execution module comprises a graph network and a student action network, the graph network obtains spatio-temporal features based on the information observed by the sensors of the multi-agent platform, and the student action network realizes intelligent decision-making of the multi-agent based on the spatio-temporal features.

[0008] The teacher action network is used for guiding the student action network to complete the update of parameters, and outputting an agent reference action to the cloud side dynamic cooperation module.

[0009] The cloud side dynamic cooperation module uses an attention mechanism to strengthen the differentiated representation of each agent on the attention of different agents, establishes a dynamic cooperation graph, and feeds back to the end side as a reference basis for graph network optimization, and outputs to the cloud side parameter sharing module.

[0010] The cloud side parameter sharing module includes a cloud side strategy network and a parameter aggregation cooperation sub-module, the cloud side strategy network is used for constructing a state-action value function of state and action based on the dynamic cooperation graph, and the parameter aggregation cooperation sub-module aggregates parameters of all cloud side strategy networks into new shared parameters as new parameters of a target strategy network.

[0011] The cloud side global optimization module realizes global optimization of the cloud side strategy network parameters based on width learning, and guides the teacher action network to update network parameters.

[0012] (Three) beneficial effects

[0013] The present application provides a three-pillar driven cloud collaborative multi-agent deep reinforcement learning decision system, which is aimed at the problem that the non-static and partial observability characteristics of the actual environment lead to the difficulty of convergence of multi-agent deep reinforcement learning algorithm, uses cloud collaborative mode, establishes a dynamic cooperation graph based on self-attention mechanism, realizes network parameter sharing in the same task area by using federal cooperation mechanism, adopts width learning optimization strategy network, establishes a three-pillar architecture for guaranteeing the convergence of multi-agent deep reinforcement learning of dynamic cooperation, parameter sharing and global optimization, and proposes a three-pillar driven cloud collaborative multi-agent deep reinforcement learning decision system. Compared with the traditional method, the method can accelerate the convergence speed of multi-agent deep reinforcement learning, enhance the coordination between agents, and improve the stability of the system. The present application has high reference and practical value in the aspect of intelligent decision of resource-limited multi-agent platform. BRIEF DESCRIPTION OF DRAWINGS

[0014] Figure 1 The present application provides a three-pillar driven cloud collaborative multi-agent deep reinforcement learning decision system, which is aimed at the problem that the non-static and partial observability characteristics of the actual environment lead to the difficulty of convergence of multi-agent deep reinforcement learning algorithm, uses cloud collaborative mode, establishes a dynamic cooperation graph based on self-attention mechanism, realizes network parameter sharing in the same task area by using federal cooperation mechanism, adopts width learning optimization strategy network, establishes a three-pillar architecture for guaranteeing the convergence of multi-agent deep reinforcement learning of dynamic cooperation, parameter sharing and global optimization, and proposes a three-pillar driven cloud collaborative multi-agent deep reinforcement learning decision system. Compared with the traditional method, the method can accelerate the convergence speed of multi-agent deep reinforcement learning, enhance the coordination between agents, and improve the stability of the system. The present application has high reference and practical value in the aspect of intelligent decision of resource-limited multi-agent platform. DETAILED DESCRIPTION

[0015] In order to make the purpose, content and advantages of the present application more clear, the specific embodiments of the present application are further described in detail below in combination with the drawings and examples.

[0016] This invention is a three-pillar driven cloud-based collaborative multi-agent deep reinforcement learning decision-making system. The method includes: deploying a lightweight student action network on the edge; outputting decision actions based on sensor observations from a multi-agent platform; and performing dynamic collaboration, parameter sharing, and global optimization on the cloud side, while guiding the edge-side student action network to update its network parameters. It mainly comprises five parts: an edge-side execution module, a cloud-side teacher guidance module, a cloud-side action collaboration module, a cloud-side parameter sharing module, and a cloud-side global optimization module. The overall framework of this invention is as follows: Figure 1 As shown. Among them,

[0017] The edge execution module includes a graph network and a student action network. The graph network obtains spatiotemporal features based on information observed by the sensors of the multi-agent platform, and the student action network realizes intelligent decision-making for the multi-agent based on the spatiotemporal features.

[0018] The cloud-based teacher guidance module includes: a teacher action network, which serves as the optimization target for the student action network, guiding the student action network to update its parameters, and simultaneously outputting agent reference actions to the cloud-based dynamic collaboration module;

[0019] The cloud-side dynamic collaboration module uses an attention mechanism to enhance the differentiated representation of each agent's attention to other agents, establishes a dynamic collaboration graph, feeds it back to the edge as a reference for graph network optimization, and outputs it to the cloud-side parameter sharing module.

[0020] The cloud-side parameter sharing module includes: a cloud-side policy network and a parameter aggregation and collaboration submodule. The cloud-side policy network is used to construct state-action value functions of states and actions based on a dynamic collaboration graph. The parameter aggregation and collaboration submodule aggregates the parameters of all cloud-side policy networks into new shared parameters, which are then used as new parameters for the target policy network.

[0021] The cloud-side global optimization module uses breadth learning to achieve global optimization of cloud-side policy network parameters and guides the teacher's action network in updating network parameters.

[0022] Example 1:

[0023] like Figure 1 The diagram shows the structural block of a cloud-based collaborative multi-agent deep reinforcement learning decision-making system driven by three pillars: dynamic collaboration, parameter sharing, and parameter optimization. The invented method comprises five parts: an edge execution module, a cloud-based teacher guidance module, a cloud-based action collaboration module, a cloud-based parameter sharing module, and a cloud-based global optimization module. The cloud-based action collaboration module, cloud-based parameter sharing module, and cloud-based global optimization module address the convergence challenge in multi-agent deep reinforcement learning. Based on the cloud-based teacher guidance module, the edge execution module is guided to output decision actions to complete the specified task.

[0024] The specific implementation methods of each module are described below.

[0025] (1) End-side execution module

[0026] Each edge execution module includes three sub-modules: a signal acquisition sub-module, a spatiotemporal feature extraction sub-module, and a decision action output sub-module. The spatiotemporal feature extraction sub-module includes a graph network for feature extraction, and the decision action output sub-module includes a student action network.

[0027] First, the signal acquisition submodule of the multi-agent (the agent is a device such as a drone or unmanned vehicle) acquires the state information of itself and the surrounding agents based on its own sensors, such as the position, velocity, acceleration and other information of the agent itself and the neighboring agents, and transmits it to the spatiotemporal feature extraction submodule.

[0028] Secondly, the graph network of the spatiotemporal feature extraction submodule is shared by the end side. Based on the information obtained by all agents and the dynamic collaboration graph output by the cloud-side dynamic collaboration module, the graph network extracts the spatiotemporal information of interest to each agent according to principles such as distance, forms spatiotemporal features, and transmits them to the decision action output submodule.

[0029] Finally, the decision action output submodule inputs the extracted specific spatiotemporal features into each individual student action network. Each agent corresponds to one student action network. The student action network processes the spatiotemporal features, obtains the decision action of each agent through a fully connected layer, and outputs it.

[0030] (2) Cloud-based teacher guidance module

[0031] This module includes: a teacher action network, with each student action network corresponding to a teacher action network. The cloud-side teacher guidance module performs updates to the teacher action network parameters themselves and updates to the student action network parameters on the client side.

[0032] In the first process, the teacher action network updates the parameters of all teacher action networks based on the state-action value function output by the corresponding cloud-side policy network, combined with the spatiotemporal characteristics of each agent, using the backpropagation mechanism and gradient descent algorithm.

[0033] In the second process, the output data of the fully connected layer of the student action network is augmented to the same dimension as the output data of the fully connected layer of the teacher action network using the adaptation layer. The L2 loss of the two vectors is calculated, and the corresponding student action network parameters are updated based on this loss using the backpropagation mechanism and gradient descent algorithm.

[0034] (3) Cloud-based motion collaboration module

[0035] The cloud-based action collaboration module is used to construct dynamic collaboration graphs and update the graph network. In this module, it receives agent reference actions from the teacher's action network, uses a self-attention mechanism to obtain the agent's attention weights towards other agents based on the agent's state-action pair information, constructs a dynamic collaboration graph representing important collaborative relationships, and feeds this dynamic collaboration graph back to the edge-side graph network to improve agent collaboration efficiency. The dynamic collaboration graph is essentially a vector number, which serves as input data to the graph network. Simultaneously, the cloud-based action collaboration module inputs the dynamic collaboration graph into the policy network.

[0036] The agent state-action pair information consists of the state of the environment and other agents observed by the agent's sensors, such as the speed and position of other drones, obstacle and road information in the environment, and the agent reference actions of the teacher's action network.

[0037] When constructing a dynamic collaboration graph, a sorting method is used to select the agents corresponding to the top N weight values ​​as agents with important collaborative relationships with this agent.

[0038] (4) Cloud-side parameter sharing module

[0039] The cloud-side parameter sharing module includes a cloud-side policy network and a parameter aggregation and collaboration submodule. The cloud-side policy network is used to construct state-action value functions based on a dynamic collaboration graph, while the parameter aggregation and collaboration submodule is used to aggregate the parameters of the cloud-side policy network to obtain the target policy network. Each agent corresponds to one cloud-side policy network.

[0040] In this process, the parameter aggregation and collaboration submodule uses a federated averaging method to aggregate the parameters of all cloud-side policy networks. Network parameters are aggregated into a new, shareable target policy network.

[0041] (1)

[0042] Where N represents the number of cloud-side policy networks. New shared parameters will then be... The parameters are distributed to each policy network on the cloud side, replacing the previous policy network's parameters, and are used as the new parameters for the cloud-side policy network at this moment.

[0043] (5) Cloud-side global optimization module

[0044] The cloud-side global optimization module is used for communication and optimization between the cloud-side policy network and the teacher action network. Specifically, it includes: the cloud-side policy network employs a Deep Deterministic Policy Gradient (DDPG) algorithm architecture, and the parameters of the target policy network are obtained through parameter optimization collaboration based on multiple policy networks. For each agent's data transmitted to the cloud based on its spatiotemporal features and dynamic collaboration graph, the module calculates the outputs of the corresponding cloud-side policy network and the target policy network, and then calculates the difference between the output values ​​of the cloud-side policy network and the target policy network. e(k) is used to calculate the link weights from newly added network nodes to the output nodes in the cloud-side policy network using the matrix pseudoinverse. (k+1), ultimately achieving the goal of optimizing the cloud-side policy network using breadth learning.

[0045] b e(k) (2)

[0046] The value of 'b' is obtained by calculating the pseudo-inverse of the matrix using the outputs of newly added network nodes and the outputs of old historical nodes. A width-learning optimization strategy network is employed here; when discrepancies exist, they are eliminated by adding network nodes.

[0047] After optimizing the cloud-side policy network, the obtained state-action value function is output to the corresponding teacher action network to guide the teacher action network in updating its network parameters.

[0048] This invention relates to a multi-agent collaborative decision-making method, specifically a three-pillar driven cloud-based collaborative multi-agent deep reinforcement learning decision-making system. It mainly comprises five parts: an edge execution module, a cloud-based teacher guidance module, a cloud-based dynamic collaboration module, a cloud-based parameter sharing module, and a cloud-based global optimization module. Through dynamic collaboration, parameter sharing, and global optimization, it addresses the convergence difficulties of multi-agent deep reinforcement learning algorithms caused by the non-static and partially observable characteristics of real-world environments. It utilizes cloud collaboration technology to transfer computationally intensive training tasks to the cloud and employs a complex teacher action network to guide a lightweight student action network, thus solving the problem of limited edge resources for multi-agent systems and achieving intelligent collaborative decision-making among resource-constrained multi-agent systems.

[0049] This invention addresses the convergence challenge of multi-agent deep reinforcement learning algorithms due to the non-static and partially observable characteristics of real-world environments. It utilizes a cloud-based collaborative approach, establishing a dynamic collaborative graph based on a self-attention mechanism. A federated collaboration mechanism enables network parameter sharing within the same task region. A width-learning optimization strategy network is employed to establish a three-pillar architecture that guarantees convergence in multi-agent deep reinforcement learning, encompassing dynamic collaboration, parameter sharing, and global optimization. This three-pillar-driven cloud-based collaborative multi-agent deep reinforcement learning decision-making system is proposed. Compared to traditional methods, this approach accelerates the convergence speed of multi-agent deep reinforcement learning, enhances coordination among agents, and improves system stability. This invention has significant reference and practical value for intelligent decision-making in resource-constrained multi-agent platforms.

[0050] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A three-pillar driven cloud collaborative multi-agent deep reinforcement learning decision system, characterized in that, The system comprises: an end-side execution module, a cloud-side teacher guidance module, a cloud-side action cooperation module, a cloud-side parameter sharing module, and a cloud-side global optimization module; The end-side execution module comprises: a graph network and a student action network, the graph network obtains spatiotemporal features based on information observed by sensors of a multi-agent platform, and the student action network realizes intelligent decision-making of the multi-agent based on the spatiotemporal features; The cloud-side teacher guidance module comprises: a teacher action network, which guides the student action network to complete updating of parameters as an optimization target of the student action network, and outputs an agent reference action to the cloud-side dynamic cooperation module; The cloud-side dynamic cooperation module uses an attention mechanism to strengthen differentiated representation of each agent on the rest of different agents, establishes a dynamic cooperation graph, and feeds back to the end side as a reference basis for optimization of the graph network, and outputs to the cloud-side parameter sharing module; The cloud-side parameter sharing module comprises: a cloud-side strategy network and a parameter aggregation cooperation sub-module, the cloud-side strategy network is used to construct a state-action value function of states and actions based on the dynamic cooperation graph, and the parameter aggregation cooperation sub-module aggregates parameters of all cloud-side strategy networks into new shared parameters as new parameters of a target strategy network; The cloud-side global optimization module realizes global optimization of parameters of the cloud-side strategy network based on width learning, and guides the teacher action network to update network parameters.

2. The three-pillar drive cloud synergy multi-agent deep reinforcement learning decision system of claim 1, wherein, The end-side execution module comprises: a signal acquisition sub-module, a spatiotemporal feature extraction sub-module, and a decision action output sub-module, the spatiotemporal feature extraction sub-module comprises: a graph network for extracting features, and the decision action output sub-module comprises: a student action network; Firstly, the signal acquisition sub-module of the multi-agent acquires state information of itself and surrounding agents based on its own sensors, and transmits the state information to the spatiotemporal feature extraction sub-module; Secondly, the graph network of the spatiotemporal feature extraction sub-module is shared by the end side, the graph network extracts spatiotemporal information of each agent according to information acquired by all agents and a dynamic cooperation graph output by the cloud-side dynamic cooperation module, forms spatiotemporal features, and transmits the spatiotemporal features to the decision action output sub-module; Finally, the decision action output sub-module inputs the extracted specific spatiotemporal features into each individual student action network, each agent corresponds to a student action network, the student action network processes the spatiotemporal features, and obtains and outputs a decision action of each agent through a full-link layer.

3. The three-pillar drive cloud synergy multi-agent deep reinforcement learning decision system of claim 2, wherein, The agent is a drone or an unmanned vehicle, and the state information of the agent comprises: position, speed, and acceleration information of the agent itself and neighboring agents.

4. The three-pillar drive cloud synergy multi-agent deep reinforcement learning decision system of claim 2, wherein, Each student action network corresponds to a teacher action network, and the cloud-side teacher guidance module performs updating of parameters of the teacher action network itself and updating of parameters of the end-side student action network; The teacher action network updates parameters of all teacher action networks by using a back propagation mechanism and a gradient descent algorithm, according to values of a state-action value function output by the corresponding cloud-side strategy network, in combination with spatiotemporal features of each agent; The output data of the full link layer of the student action network is expanded to the same dimension as the output data of the full link layer of the teacher action network by using an adaptation layer, an L2 loss of the two vectors is calculated, and the corresponding student action network parameters are updated based on the loss by using a back propagation mechanism and a gradient descent algorithm.

5. The three-pillar drive cloud collaborative multi-agent deep reinforcement learning decision system of claim 4, wherein, The cloud-side action cooperation module receives the agent reference action from the teacher action network, obtains the attention weight of the agent to other agents based on the agent state action pair information by using a self-attention mechanism, constructs a dynamic cooperation graph based on the attention weight, represents the important cooperation relationship in the cooperation graph, and feeds back the dynamic cooperation graph to the graph network on the terminal side to improve the cooperation efficiency of the agent.

6. The three-pillar drive cloud collaborative multi-agent deep reinforcement learning decision system of claim 5, wherein, The agent state action pair information is the state observed by the agent sensor in the environment and other agents, including the speed, position, obstacle road information, road information of other unmanned aerial vehicles, and the agent reference action of the teacher action network.

7. The three-pillar drive cloud collaborative multi-agent deep reinforcement learning decision system of claim 5, wherein, When constructing the dynamic cooperation graph, the sorting method is used to select the top N weight values corresponding to the agents as the agents having important cooperation relationship with the agent.

8. The three-pillar drive cloud collaborative multi-agent deep reinforcement learning decision system of claim 5, wherein, In the cloud-side parameter sharing module, each agent corresponds to a cloud-side policy network; the parameter aggregation and collaboration submodule uses a federated averaging method to aggregate the parameters of all cloud-side policy networks. Network parameters aggregated into a new shared target policy network: (1) Wherein, N is the number of cloud side strategy networks, and the new parameter that can be shared is issued to each strategy network on the cloud side, replaces the parameters of the previous strategy network, and serves as the new parameter of the strategy network on the cloud side at this moment.

9. The three-pillar drive cloud collaborative multi-agent deep reinforcement learning decision system of claim 8, wherein, The cloud side global optimization module comprises: a cloud side strategy network adopts a deep deterministic policy gradient (DDPG) algorithm, and parameters of a target strategy network are obtained by parameter aggregation optimization cooperation based on multiple strategy networks; for each agent, based on respective space-time features and dynamic cooperation graphs transmitted to the cloud side, outputs of the corresponding cloud side strategy network and the target strategy network are respectively obtained, and a difference between the cloud side strategy network output value and the target strategy network output value is calculated e(k), the link weight of the newly added network node of the cloud side strategy network to the output node is calculated by using a matrix pseudo-inverse k+1), and finally the purpose of optimizing the cloud side strategy network by using width learning is achieved; After optimizing the cloud-side strategy network, the obtained state action function is output to the corresponding teacher action network to guide the teacher action network to update the network parameters.

10. The three-pillar drive cloud collaborative multi-agent deep reinforcement learning decision system of claim 9, wherein, adding a link weight to an output node k+1) is: b e(k) (2) The value of b is obtained by using matrix pseudo-inverse calculation based on the output of the new network node and the output of the old historical node.

Citation Information

Patent Citations

  • Multi-agent collaborative decision reinforcement learning method, system and device

    CN117151206A

  • Intelligent agent behavior generation system and method supporting skill drilling

    CN120525024A

  • Multi-agent collaborative decision-making method based on federal reinforcement learning

    CN120745684A

  • Platform for orchestrating fault-tolerant, security-enhanced networks of collaborative and negotiating agents with dynamic resource management

    US20250259043A1