Collaborative design method for multi-agent complete cooperation type task

By combining edge processors with LSTM, COMA, and DRQN ​​networks for joint optimization of state estimation, control strategy, and resource scheduling, the problem of insufficient flexibility of multi-agent systems in complex environments is solved, and highly robust control is achieved.

CN121300152APending Publication Date: 2026-01-09SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511340453.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2026-01-09

Smart Images

  • Figure CN121300152A_ABST
    Figure CN121300152A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of wireless network control systems (WNCSs) and multi-agent cooperative control, in particular to an estimation-control-scheduling cooperative design method based on deep reinforcement learning (DRL), and discloses an estimation-control-scheduling cooperative design method based on the deep reinforcement learning (DRL). The method is suitable for real-time control and resource optimization of a multi-agent complete cooperation type task in an industrial environment, and an innovative collaborative design framework is provided for solving the problems that in the prior art, system modeling dependency is high, dynamic environment adaptability is poor, and resource allocation efficiency is low. The sequential dependency relationship between the observed quantity and the state quantity in the WNCSs is learned through the recurrent neural network, and the adaptability of the method in the complex industrial environment is effectively enhanced. According to the method, joint optimization of state estimation, a control strategy and resource scheduling can be realized, and high-robustness control is realized under the condition of no precise system dynamics modeling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial wireless network control systems, and in particular to a multi-agent estimation-control-scheduling collaborative design method based on deep reinforcement learning. Background Technology

[0002] Multi-agent systems (MAS) are systems in which multiple agents accomplish complex tasks through full cooperation, full competition, or a hybrid of both. Each agent possesses a degree of autonomy, interaction capabilities, and decision-making abilities, enabling it to make independent action decisions based on its own perceptions and the state of the environment. Fully cooperative tasks have wide applications in many industrial fields, such as collaborative material handling and assembly. In these tasks, MAS must coordinate the actions of each agent to achieve a common goal. However, traditional MAS suffers from insufficient flexibility and high installation and deployment costs. To address these issues, wireless networked control systems (WNCSs) utilize edge processors to achieve agent transmission scheduling, significantly improving the flexibility and scalability of MAS collaboration. Especially in complex and dynamic environments, WNCSs exhibit strong robustness and show broad application prospects in fully cooperative tasks.

[0003] Traditional collaborative design methods in WNCSs rely on accurate system dynamics models. However, in complex industrial control scenarios such as robotics and automated guided vehicles, collaborative tasks may fail, limiting the applicability of traditional methods in nonlinear control. Compared to traditional methods that require precise system dynamics modeling, reinforcement learning-based methods can guide agents to explore optimal strategies through reward mechanisms even without precise modeling. In simple scenarios with limited equipment, traditional reinforcement learning methods have shown good control performance. However, in fully cooperative tasks with high environmental complexity and dynamic nature, which typically involve large state and action spaces, reinforcement learning methods struggle to converge effectively. In recent years, deep reinforcement learning-based methods have leveraged deep neural networks for function approximation, automatically extracting key features from high-dimensional data and significantly improving learning efficiency.

[0004] Existing deep reinforcement learning-based communication and control co-design methods mostly employ decoupled designs. For example, the independent design of the scheduler or the omission of the state estimator leads to suboptimal results, thereby reducing the overall control performance of the system. Co-design methods that comprehensively consider the estimator, controller, and scheduler, however, only consider one agent in the network and are not applicable to WNCSs with multiple agents. To address these challenges... Summary of the Invention

[0005] This invention proposes a cooperative design method for fully cooperative multi-agent tasks. Specifically, it utilizes historical observation data to compensate for information loss caused by packet loss in the wireless channel, generating an estimate of the observed state. This invention enables joint optimization of state estimation, control strategy, and resource scheduling, achieving highly robust control even without precise system dynamics modeling. By comprehensively considering state estimation, control strategy, and resource scheduling, this method ensures that the system maintains robust control performance even with lost observation data.

[0006] The technical solution adopted by the present invention to achieve the above objectives is as follows:

[0007] A collaborative design method for multi-agent fully cooperative tasks includes the following steps:

[0008] 1) Collect environmental and agent data as observation data through industrial sensors and transmit it to the edge processor through the shared wireless channel of WNCSs;

[0009] 2) The state estimator in the edge processor analyzes historical observation data and reconstructs the observation information as a state estimate when uplink packet loss occurs;

[0010] 3) The controller in the edge processor generates control actions for multiple agents based on state estimates;

[0011] 4) The state estimator predicts future states using historical observation data, control actions, and observation data;

[0012] 5) The scheduler in the edge processor generates scheduling actions based on historical observation data and future states, and dynamically allocates network resources;

[0013] 6) The edge processor synchronously distributes control and scheduling actions to each agent through a shared wireless network. The agents upload new observation data in the next time step according to the scheduling arrangement, thus completing the closed-loop control cycle.

[0014] The observation data in step 1) includes: agent state, environmental dynamics, number of wireless network resource blocks, and information age used to quantify data freshness.

[0015] Step 2) includes the following steps:

[0016] 2.1) At time step t, the edge processor sets the uplink channel packet loss flag β. i,t Check ∈{0,1}, if β i,t =1, then the observation data is confirmed to be lost and the observation information needs to be reconstructed;

[0017] 2.2) For an agent i, i = 1, 2, ..., N, when it is scheduled at time step t, i.e., when its action is... Furthermore, when the observation data is successfully transmitted, the observation data can be used directly. i,t Otherwise, use an LSTM network. Using historical data Obtain reconstructed observation data Historical data of length L Represented as:

[0018]

[0019] in, To control actions;

[0020] By minimizing the difference between the agent's observed data and estimated state, the network parameters are optimized. Training is performed, therefore, a loss function is constructed.

[0021]

[0022] The estimated value of the observed state is obtained through the LSTM network:

[0023]

[0024] Step 3) specifically refers to:

[0025] The COMA-based controller structure consists of a centralized critic network and N actor networks. The actor networks obtain control strategies as control actions based on the estimated values ​​of the observed states. The critic network evaluates the control strategies obtained by the actor networks, and the actor networks optimize their own network parameters based on the evaluations of the critic network.

[0026] The network of critics specifically refers to:

[0027] The commentator network output state-action value, or Q-value, is represented as:

[0028]

[0029] in, This is the global state. For joint control actions, ψ C For the critic's network parameters;

[0030] Critics Network through Improved TD(λ) Objective Update the loss function of the critic network. Defined as:

[0031]

[0032] Critic network parameter ψ Cand target network parameters ψ C′ The update process is as follows:

[0033]

[0034] ψ C′ ←τψ C +(1-τ)ψ C′

[0035] Where 0<τ≤1 is the soft update rate, α C It is the learning rate of the critic network. It's about the critic network parameter ψ C The gradient operator.

[0036] The actor network specifically refers to:

[0037] The actor network consists of fully connected layers and GRU layers, and the resulting control policy is expressed as follows:

[0038]

[0039] in, For actor network parameters, To control historical information;

[0040] Each agent i calculates an advantage function. Compare the current Q value to its counterfactual baseline:

[0041]

[0042] in, The joint action of all agents except agent i;

[0043] The approximate policy gradient with counterfactual baseline is calculated as follows:

[0044]

[0045] Where K is the number of iterations, E is the policy network parameter after K iterations. π This represents the expectation under the current policy π;

[0046] Obtain the policy gradient g K Then, with a learning rate of α A The gradient ascent is used to update the policy parameters of agent i:

[0047]

[0048] All actors share a single set of network parameters.

[0049] Step 4) specifically involves:

[0050] Using the updated state estimator from step 2), and taking the control action, observations, and historical data as input, the future state is obtained.

[0051]

[0052] Step 5) specifically involves:

[0053] The current scheduling action is obtained through an ε-greedy strategy, that is, a random action is selected with probability ε. Otherwise, the scheduling action is selected by maximizing the Q-value:

[0054]

[0055] in, It is the history of the scheduler. It is a matrix composed of future states;

[0056] During the training iterations of the DRQN-based scheduler, the parameter ε was explored to gradually converge to 0.01; the loss function of DRQN ​​was investigated. Defined as:

[0057]

[0058] Among them, TD target y t The DRQN ​​network parameters θ are obtained through calculation using the target network. Tx and its target network parameters θ Tx′ The update process is as follows:

[0059]

[0060] θ Tx′ ←τθ Tx +(1-τ)θ Tx′

[0061] Where, α Tx It is the learning rate used to control the step size of each parameter update. It's about the network parameter θ Tx The gradient operator;

[0062] Each time the network is updated, the hidden state of the LSTM layer of DRQN ​​is initialized to zero.

[0063] Step 6) includes the following steps:

[0064] 6.1) The edge processor transmits control actions via the downlink wireless channel. and scheduling actions It is encapsulated into a unified data packet and synchronously distributed to each intelligent agent via a shared wireless network;

[0065] 6.2) The intelligent agent actuator immediately executes the corresponding operation based on the received control action, and at the same time writes the scheduling action into the local scheduling queue;

[0066] 6.3) Unscheduled agents enter low-power listening mode, while scheduled agents prioritize using channel resources to upload observation data. i,t This completes the closed-loop control cycle.

[0067] A collaborative design system for multi-agent fully cooperative tasks includes a memory and a processor; the memory is used to store a computer program; the processor is used to implement the collaborative design method for multi-agent fully cooperative tasks when the computer program is executed.

[0068] The present invention has the following beneficial effects and advantages:

[0069] 1. This invention proposes for the first time an estimation-control-scheduling cooperative design framework for multi-agent fully cooperative tasks. It uses recurrent neural networks to learn the temporal dependencies of observables and state variables in WNCSs, which effectively enhances the adaptability of this invention in complex industrial environments.

[0070] 2. This invention can achieve joint optimization of state estimation, control strategy and resource scheduling, and achieve highly robust control without precise system dynamics modeling. Attached Figure Description

[0071] Figure 1 This is a schematic diagram of an LSTM-based estimator structure.

[0072] Figure 2 This is a schematic diagram of a COMA-based controller architecture;

[0073] Figure 3 This is a schematic diagram of a scheduler structure based on DRQN;

[0074] Figure 4 Flowchart of DRL-based collaborative design method Detailed Implementation

[0075] To make the above-mentioned objectives, technical solutions and advantages of the present invention clearer, the specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0076] A collaborative design method for multi-agent fully cooperative tasks includes the following steps:

[0077] Step 1: Collect environmental and agent data using industrial sensors and transmit the observed data to the edge processor via the shared wireless channel of WNCSs;

[0078] Step 2: The state estimator analyzes historical observation data and reconstructs the observation information when uplink packet loss occurs.

[0079] Step 3: The controller generates control actions for the multi-agent system based on the state estimates;

[0080] Step 4: The state estimator predicts the future state using historical data, control actions, and observations;

[0081] Step 5: The scheduler generates scheduling actions based on historical data and future states, and dynamically allocates network resources;

[0082] Step 6: The edge processor synchronously distributes control actions and scheduling actions to each agent through the shared wireless network. The agents upload new observation data in the next time step according to the scheduling arrangement, thus completing the closed-loop control cycle.

[0083] The WNCSs model in step 1 is as follows:

[0084] This invention considers a WNCSs comprising N agents, each observed by a sensor, connected to a remote edge processor via a shared wireless network (e.g., 5G or Wi-Fi 6) based on OFDMA. At time step t, agents i = 1, 2, ..., N, after being scheduled, will transmit the observed data o i,t The data is sent to the wireless controller via a shared wireless network. The edge processor generates control actions based on the observed data and then sends them to the agent i to execute the actions.

[0085] Due to limited wireless network resources and wireless channel fading, WNCSs suffer from packet loss and latency issues. Packet loss is mainly caused by channel fading, interference, and network congestion, leading to the inability to receive signals normally, while latency is caused by data transmission, queuing, etc. Therefore, this invention describes the packet loss model of wireless networks as follows: the packet loss event of uplink channel i at time step t is determined by Bernoulli variables. express.

[0086] This invention uses the Age of Information (AoI) to characterize the freshness of information in WNCSs. The AoI of the agent's observed data is represented as the difference between the current time step t and the generation time step of the data most recently successfully transmitted by each agent. Since the transmission overhead is negligible, this invention assumes that the agent sends a 1-bit acknowledgment message to the edge processor through an ideal wireless channel.

[0087] Furthermore, the transmission of the observation data to the edge processor via the shared wireless channel of the WNCS in step 1 is specifically as follows:

[0088] At each discrete time step t, the sensor network deployed in the industrial field collects multi-dimensional observation data in real time. This data includes: the agent's status (position, velocity, attitude); environmental dynamics (obstacle distribution, target location); the real-time status of the wireless network (number of available resource blocks, channel quality indicators); and information timeliness parameters (AoI). All observation data will be transmitted to the edge processor via an OFDMA-based shared wireless channel uplink according to the transmission priority determined by the current scheduling strategy.

[0089] Furthermore, the reconstruction of observation information by the state estimator in step 2 specifically involves:

[0090] like Figure 1 As shown, the state estimator consists of an LSTM layer and a fully connected layer. For an agent i, i = 1, 2, ..., N, when it is scheduled at time step t... Furthermore, once the observation data is successfully transmitted, it can be used directly. i,t Otherwise, an LSTM network needs to be introduced. Using historical data Obtain reconstructed observation data The parameters are trained using MSE loss. Historical data of length L Represented as:

[0091]

[0092] in, To control actions, the estimation network is designed to minimize the difference between the agent's observed data and estimated state; therefore, the loss function is expressed as:

[0093]

[0094] The estimated value of the observed state can be obtained through an LSTM network:

[0095]

[0096] The obtained observed state estimates This will be used as input to the controller to generate control actions. Ensure that the agent can still operate reliably when observation data is missing.

[0097] Furthermore, in step 3, the controller generates control actions for the multi-agent system based on the state estimate as follows:

[0098] like Figure 2 As shown, the COMA-based controller structure includes a centralized critic network and N actor networks. The controller will generate control actions for each agent based on the controller's history and current observations.

[0099] 1) The commentator network output state-action value (Q-value) is represented as:

[0100]

[0101] in, This is the global state. For joint control actions, ψ C Here are the parameters of the critic network; the critic network is updated using an improved TD(λ) objective, and the loss function of the critic network is defined as:

[0102]

[0103] Critic network parameter ψ C and target network parameters ψ C′ The update process is as follows:

[0104]

[0105] ψ C′ ←τψ C +(1-τ)ψ C′ (2.4)

[0106] Where 0<τ≤1 is the soft update rate, α C It is the learning rate of the critic network;

[0107] 2) The actor network consists of a fully connected layer and a GRU layer, and the control strategy is expressed as:

[0108]

[0109] in, For actor network parameters, To control for historical information, each agent i compares its current Q-value with its counterfactual baseline by calculating an advantage function:

[0110]

[0111] in, The joint action of all agents except agent i; the calculation process of the approximate policy gradient with counterfactual baseline is as follows:

[0112]

[0113] Where K is the number of iterations, These are the policy network parameters after K iterations; the policy gradient g is obtained. K Then, with a learning rate of α A The gradient ascent is used to update the policy parameters of agent i:

[0114]

[0115] All actors share a single set of network parameters.

[0116] Furthermore, in step 4, the state estimator predicts the future state specifically as follows:

[0117] Using the updated state estimator, and taking control actions, observations, and historical data as input, the future state is given by the following equation:

[0118]

[0119] The input structure of the state estimator during prediction is the same as that during estimation, except that the control actions and observations are used as part of the historical data.

[0120] Furthermore, in step 5, the scheduler generates scheduling actions and allocates network resources as follows:

[0121] like Figure 3 As shown, the DRQN-based scheduler consists of an LSTM layer and multiple fully connected layers. The current scheduling action is obtained through an ε-greedy strategy, that is, a random action is selected with probability ε. Otherwise, the scheduling action is selected by maximizing the Q-value:

[0122]

[0123] in, It is the history of the scheduler. It is a matrix composed of future states. During the training iterations, to ensure policy stability, the exploration parameter ε gradually converges to 0.01; the loss function of DRQN ​​is defined as:

[0124]

[0125] Among them, TD target y t The DRQN ​​network parameters θ are calculated through the target network. Tx and its target network parameters θ Tx′ The update process is as follows:

[0126]

[0127] θ Tx′ ←τθ Tx +(1-τ)θ Tx′ (5.4)

[0128] Each time the network is updated, the hidden state of the LSTM layer of DRQN ​​is initialized to zero.

[0129] Furthermore, the reconstruction of observation information by the state estimator in step 6 specifically involves:

[0130] Step 6.1: The edge processor transmits control actions via the downlink wireless channel. and scheduling actions It is encapsulated into a unified data packet and synchronously distributed to each intelligent agent via a shared wireless network;

[0131] Step 6.2: The intelligent agent actuator immediately executes the corresponding operation based on the received control action, and at the same time writes the scheduling action into the local scheduling queue;

[0132] Step 6.3: Unscheduled agents enter low-power listening mode, while scheduled agents prioritize using channel resources to upload observation data. i,t This completes the closed-loop control cycle.

[0133] like Figure 4 As shown, the DRL-based estimation-control-scheduling collaborative design framework designed in this invention comprises N agents, an edge processor, and a shared wireless network. The edge processor consists of a state estimator, a controller, and a scheduler. The estimator reconstructs the current system state using historical observation data, addressing the information loss problem caused by packet loss and providing reliable input for control decisions. The controller generates control actions for multiple agents based on the estimated state. The scheduler optimizes overall system performance under network resource constraints by dynamically allocating network resources and prioritizing the uploading of observation data from key agents. At time step t, the edge processor collects observation data o obtained from the interaction between agents and the environment via the shared wireless network. t Subsequently, the estimator outputs a system state estimate in the event of packet loss. The controller is based on Generate control input Next, the estimator will... and o t Predict the state at the next moment This information is then passed to the scheduler. The scheduler uses this information to generate a schedule table for time step (t+1). final, and This will be simultaneously transmitted to each intelligent agent, which will immediately execute the corresponding control actions and, in accordance with... The scheduling is to send the observation data at time step (t+1). t+1 .

[0134] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.

Claims

1. A collaborative design method for multi-agent fully cooperative tasks, characterized in that, Includes the following steps: 1) Collect environmental and agent data as observation data through industrial sensors and transmit it to the edge processor through the shared wireless channel of WNCSs; 2) The state estimator in the edge processor analyzes historical observation data and reconstructs the observation information as a state estimate when uplink packet loss occurs; 3) The controller in the edge processor generates control actions for multiple agents based on state estimates; 4) The state estimator predicts future states using historical observation data, control actions, and observation data; 5) The scheduler in the edge processor generates scheduling actions based on historical observation data and future states, and dynamically allocates network resources; 6) The edge processor synchronously distributes control and scheduling actions to each agent through a shared wireless network. The agents upload new observation data in the next time step according to the scheduling arrangement, thus completing the closed-loop control cycle.

2. The collaborative design method for multi-agent fully cooperative tasks according to claim 1, characterized in that, The observation data in step 1) includes: agent state, environmental dynamics, number of wireless network resource blocks, and information age used to quantify data freshness.

3. The collaborative design method for multi-agent fully cooperative tasks according to claim 1, characterized in that, Step 2) includes the following steps: 2.1) At time step t, the edge processor sets the uplink channel packet loss flag β. i,t Check ∈{0,1}, if β i,t =1, then the observation data is confirmed to be lost and the observation information needs to be reconstructed; 2.2) For an agent i, i = 1, 2, ..., N, when it is scheduled at time step t, i.e., when its action is... Furthermore, when the observation data is successfully transmitted, the observation data can be used directly. i,t Otherwise, use an LSTM network. Using historical data Obtain reconstructed observation data Historical data of length L Represented as: in, To control actions; By minimizing the difference between the agent's observed data and estimated state, the network parameters are optimized. Training is performed, therefore, a loss function is constructed. The estimated value of the observed state is obtained through the LSTM network:

4. The collaborative design method for multi-agent fully cooperative tasks according to claim 1, characterized in that, Step 3) specifically refers to: The COMA-based controller structure consists of a centralized critic network and N actor networks. The actor networks obtain control strategies as control actions based on the estimated values ​​of the observed states. The critic network evaluates the control strategies obtained by the actor networks, and the actor networks optimize their own network parameters based on the evaluations of the critic network.

5. A collaborative design method for multi-agent fully cooperative tasks according to claim 4, characterized in that, The network of critics specifically refers to: The commentator network output state-action value, or Q-value, is represented as: in, This is the global state. For joint control actions, ψ C For the critic's network parameters; Critics Network through Improved TD(λ) Objective Update the loss function of the critic network. Defined as: Critic network parameter ψ C and target network parameters ψ C′ The update process is as follows: ψ C′ ←ts C +(1-t)ψ C′ Where 0<τ≤1 is the soft update rate, α C It is the learning rate of the critic network. It's about the critic network parameter ψ C The gradient operator.

6. The collaborative design method for multi-agent fully cooperative tasks according to claim 4, characterized in that, The actor network specifically refers to: The actor network consists of fully connected layers and GRU layers, and the resulting control policy is expressed as follows: in, For actor network parameters, To control historical information; Each agent i calculates an advantage function. Compare the current Q value to its counterfactual baseline: in, The joint action of all agents except agent i; The approximate policy gradient with counterfactual baseline is calculated as follows: Where K is the number of iterations, E is the policy network parameter after K iterations. π This represents the expectation under the current policy π; Obtain the policy gradient g K Then, with a learning rate of α A The gradient ascent is used to update the policy parameters of agent i: All actors share a single set of network parameters.

7. The collaborative design method for multi-agent fully cooperative tasks according to claim 1, characterized in that, Step 4) specifically involves: Using the updated state estimator from step 2), and taking the control action, observations, and historical data as input, the future state is obtained.

8. The collaborative design method for multi-agent fully cooperative tasks according to claim 1, characterized in that, Step 5) specifically involves: The current scheduling action is obtained through an ε-greedy strategy, that is, a random action is selected with probability ε. Otherwise, the scheduling action is selected by maximizing the Q-value: in, It is the history of the scheduler. It is a matrix composed of future states; During the training iterations of the DRQN-based scheduler, the parameter ε was explored to gradually converge to 0.01; the loss function of DRQN ​​was investigated. Defined as: Among them, TD target y t The DRQN ​​network parameters θ are obtained through calculation using the target network. Tx and its target network parameters θ Tx′ The update process is as follows: i Tx′ ←tth Tx +(1-τ)θ Tx′ Where, α Tx It is the learning rate used to control the step size of each parameter update. It's about the network parameter θ Tx The gradient operator; Each time the network is updated, the hidden state of the LSTM layer of DRQN ​​is initialized to zero.

9. A collaborative design method for multi-agent fully cooperative tasks according to claim 1, characterized in that, Step 6) includes the following steps: 6.1) The edge processor transmits control actions via the downlink wireless channel. and scheduling actions It is encapsulated into a unified data packet and synchronously distributed to each intelligent agent via a shared wireless network; 6.2) The intelligent agent actuator immediately executes the corresponding operation based on the received control action, and at the same time writes the scheduling action into the local scheduling queue; 6.3) Unscheduled agents enter low-power listening mode, while scheduled agents prioritize using channel resources to upload observation data. i,t This completes the closed-loop control cycle.

10. A collaborative design system for multi-agent fully cooperative tasks, characterized in that, It includes a memory and a processor; the memory is used to store a computer program; the processor is used to implement, when executing the computer program, a cooperative design method for multi-agent fully cooperative tasks as described in any one of claims 1-9.