A Cooperative Control Method for Complex Moving Object Systems Based on Decentralized Networked Multi-Agent Reinforcement Learning

Through the method based on decentralized networked multi-agent reinforcement learning, the graph neural network aggregates the neighbor characteristics of the moving body and reduces the influence of redundant information, the shortcomings of the unmanned moving body system in the traditional method are solved, and efficient collaborative control is achieved in the decentralized environment.

CN119376241BActive Publication Date: 2025-07-11BEIJING UNIV OF TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202411245871.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-06
Publication Date
2025-07-11
Estimated Expiration
2044-09-06

AI Technical Summary

Technical Problem

In a decentralized environment, traditional multi-agent reinforcement learning methods are difficult to promote coordinated control between unmanned moving bodies because they cannot obtain the global state, resulting in insufficient coordination performance.

Method used

Using a method based on decentralized networked multi-agent reinforcement learning, the graph neural network is used to aggregate the node characteristics of the moving body neighbors, and the redundant information impact is reduced through the graph information bottleneck, the task environment and reward functions are designed, and the algorithm is trained to achieve collaborative control.

Benefits of technology

In the decentralized scenario, the coordination performance of the unmanned moving body system is improved, effective collaborative control is achieved, and it is suitable for collaborative tasks of complex moving body systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119376241B_ABST
    Figure CN119376241B_ABST
Patent Text Reader

Abstract

The present invention designs a cooperative control method for a complex moving body system based on decentralized networked multi-agent reinforcement learning, which realizes promoting the cooperation of moving bodies in a decentralized scenario. First, an entity graph is used to model the complex moving body system to represent the spatial correlation between entities. Then, an information aggregation strategy is designed to aggregate the messages of neighbors and focus on important neighbor information by using graph neural networks and attention mechanisms. In addition, graph information bottlenecks are used to reduce the impact of redundant information on the selection of optimal actions. This method can solve the problem that in scenarios where the global state cannot be obtained, traditional multi-agent reinforcement learning cannot promote the cooperation between moving bodies due to limited observation information, improve the cooperative performance of the system, and provide an effective method for the field of cooperative control of various unmanned moving body systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Based on the multi-agent particle environment, the present invention establishes a collaborative control method for complex moving bodies based on decentralized networked multi-agent reinforcement learning. It uses a graph neural network to aggregate the node features with spatial correlation of agent neighbors, and adopts a graph information bottleneck to reduce the impact of redundant information on the optimal action selection, and finally completes the collaborative control of the complex moving body system. This collaborative control method for complex moving body systems based on decentralized networked multi-agent reinforcement learning can solve the problem that in scenarios where the global state cannot be obtained, traditional multi-agent reinforcement learning cannot promote the collaboration between moving bodies due to limited observation information, improve the collaborative performance of the system, and provide technical support for various fields of collaborative control of unmanned moving body systems. Background Art

[0002] With the rapid development of artificial intelligence and machine learning technologies, the application scope of unmanned moving body systems has been increasingly expanded. These systems not only play a key role in the military and aviation fields, but also are widely used in civilian fields such as autonomous driving, intelligent warehousing, and medical services. Traditional collaborative control methods usually rely on accurate dynamic models, which limits their application in complex unmanned moving body systems. Therefore, developing model-free adaptive methods has become the key to realizing the collaborative control of complex moving body systems. The collaborative control method for moving bodies based on multi-agent reinforcement learning effectively promotes the collaboration between moving bodies and significantly improves the collaborative control performance. However, in many actual moving body scenarios, the goal of each moving body is usually to complete its own tasks with as little information sharing as possible. In a decentralized environment, problems such as communication limitations or the absence of a central controller make it difficult for moving bodies to obtain global information even if there is a global goal, thus making it difficult to promote cooperation between agents. Therefore, in practical applications, achieving efficient collaborative control of complex unmanned moving body systems in a decentralized scenario has significant economic and social benefits and shows broad application prospects.

[0003] It is difficult to obtain a dynamic model based on traditional methods, and the accuracy is low. With the development of artificial intelligence technology in China, the application of unmanned moving objects is becoming more and more widespread. Traditional methods have difficulty in achieving precise cooperative control in complex environments. To solve the problem of cooperative control in complex scenarios, some moving object control methods based on reinforcement learning have been proposed. However, when applied to solve practical problems, especially in complex moving object systems, reinforcement learning still faces many challenges. The basic idea of cooperation is that moving objects coordinate and communicate with each other to effectively complete tasks. There are non-stationarity problems in the actual environment, that is, the actions of each moving object affect the decisions of adjacent moving objects, which limits the direct applicability of single-agent reinforcement learning technology in the field of complex multi-unmanned moving objects. In contrast, multi-agent reinforcement learning enables moving objects to interact with each other and effectively solves the non-stationarity of the environment. The effectiveness of traditional multi-agent reinforcement learning algorithms depends on the amount of information contained in the input state of the neural network. For most multi-agent reinforcement learning algorithms, accessing the global state of the system is usually crucial, but in many actual complex moving object systems, there are problems such as the absence of a central controller or network communication, resulting in the infeasibility of sharing the global state. Therefore, how to promote cooperation among moving objects and effectively execute collision-free cooperative control by using limited local environment information in a decentralized scenario has become an important research topic in the field of unmanned moving object systems and has important practical significance.

[0004] The present invention designs a cooperative control method for complex moving objects based on decentralized networked multi-agent reinforcement learning to promote cooperation among moving objects in a decentralized scenario. First, an entity graph is used to model the complex moving object system to represent the spatial correlation between entities, and then an information aggregation strategy is designed to aggregate the messages of neighbors and focus on important neighbor information by using graph neural networks and attention mechanisms. In addition, graph information bottlenecks are used to reduce the impact of redundant information on the selection of optimal actions. This method can solve the problem that traditional multi-agent reinforcement learning cannot promote cooperation among moving objects due to limited observation information in scenarios where the global state cannot be obtained, improve the cooperative performance of the system, and provide an effective method for the field of cooperative control of various unmanned moving object systems. Summary of the Invention

[0005] The present invention proposes a collaborative control method for complex moving bodies based on decentralized networked multi-agent reinforcement learning. First, a collaborative control task for moving bodies is designed based on a multi-agent particle environment, and observation information data is obtained by interacting with the task environment. Secondly, this data is transmitted to the server and converted into a graph structure according to the position of each entity in the global coordinate system to characterize the data features. Then, a graph neural network is used to aggregate the node features of each moving body and its neighboring entities, thereby training a decentralized networked multi-agent reinforcement learning algorithm. Finally, each moving body uses the trained model to independently make decisions based on the data collected by odometry positioning and inertial measurement units, quickly reach the desired position and avoid collisions.

[0006] Odometry is a method that uses mobile sensor data to estimate the change in the position of an object over time and is widely used in many moving body systems to calculate the moving distance of a moving body relative to its initial position. The inertial measurement unit is mainly used to detect the acceleration of a moving body and is usually equipped with an accelerometer and a gyroscope to measure the attitude angle and acceleration of an object on a single axis, double axis, or triple axis.

[0007] The present invention adopts the following technical solutions and implementation steps:

[0008] A collaborative control method for a complex moving body system based on decentralized networked multi-agent reinforcement learning, characterized in that a task environment and a reward function are designed, node features are characterized, information aggregation is performed, an algorithm is trained, and decentralized networked multi-agent reinforcement learning is used to achieve collaborative motion control, including the following steps:

[0009] (1) Design a task environment and a reward function

[0010] A 2D multi-agent particle environment is adopted as the basis. In this environment, the maximum number of time steps T per round is 25, the area of the environment adopts the default value of the multi-agent particle environment, and its area is 4×4. The number of moving bodies is N, the number of landmarks is D, and the number of obstacles is O; i represents the number of any moving body, i ∈ {1,..., N}; t represents the current moment of the environment, t ∈ {1,..., T}; the i-th moving body A i The reward function r at time t during the movement process i,t is designed as:

[0011]

[0012] where pos i,t and respectively represent the position of A i at time t and the position of the target;

[0013] (2) Characterize node features

[0014] The moving object, obstacle, and landmark are collectively referred to as entity e m , where m is the number of any entity, m ∈ {1,..., N + D + O}; the set {1,..., i - 1, i + 1,..., N + D + O} is used to represent the numbers of any entities except A i ; d mn,t ∈ R is used to represent the Euclidean distance between any two entities e m and e n , where m ≠ n and n ∈ {1,..., N + D + O} is the number of any entity other than m; the set of entities within the perception radius ρ of A i is defined as N i,t , and these entities are called the neighbors of A i ; e j is used to represent any neighbor, where j is the number of the neighbor; V i,t = {N i,t ∪ A i} represents the set of A i and its neighbors; the complex moving object system is represented by the graph G t (V, E t ), where V = {e1,..., e N+D+O} represents the set of all entities, represents the set of edges at time t, and the Euclidean distance d i between A j and e ij,t is used as the feature of the edge; A i forms a graph network g i,t ∈ G t at time t. The position, velocity of A i in the two-dimensional coordinate system at time t and the position of the target relative to itself are defined as pos i,t ∈ R 2 , vel i,t ∈ R 2 respectively, and where R represents real numbers;

[0015] pos i,t , vel i,t , and are used to represent the observation information Define and as the position, velocity of e j relative to A i respectively, and the position of e j target relative to A i ; if e j is an obstacle or a target, set Use to represent gi,t The node features of each neighbor entity above, where et(j) is the entity type, defined as et(j) ∈ {"agent": 0, "landmark": 1, "obstacle": 2}, the strings within {·} represent text, the dictionaries agent, landmark, and obstacle represent moving objects, landmarks, and obstacles respectively, and the numbers 0, 1, and 2 represent the indices of the dictionaries;

[0016] (3) Perform information aggregation

[0017] Use an embedding layer with a dictionary size of 3 and an output dimension of 2 to encode et(j) to obtain the encoded entity type et ij,t , concatenate et ij,t and d ij,t to obtain a vector where represents the vector concatenation operation; input into a linear layer with an input dimension of 9 and an output dimension of 16 for feature extraction, thereby obtaining the node features input to the graph neural network Adopt a 3-layer graph neural network as the information aggregation module, with the layer number l ∈ {1,..., L}, the maximum number of layers L = 3, and the input layer dimension of the graph neural network set to 16;

[0018] ① On the first layer of the graph neural network, use

[0019]

[0020] to calculate the attention map where u and v are the numbers of any moving object on g i,t and its neighbor respectively, u ∈ {1,..., N}, v ∈ N u,t , N u,t represents the set of u neighbors; is the node feature of any moving object numbered u on the (l - 1)-th layer, is the node feature of any neighbor numbered v on the (l - 1)-th layer; c is the dimension of the output of each layer of the graph neural network. When l = 1 or 3, c = 16, and when l = 2, c = 32; W Q ∈ R 16×c·h is the learnable weight matrix for the query, is the learnable weight matrix for the key; use

[0021]

[0022] to calculate any moving object A numbered u u and its neighbor entity e numbered vv Attention coefficient between The number of attention heads is set to 3; using

[0023]

[0024] Calculate e at layer l = 1 v Transmit to A u The message of Among them, and Are the node features of A at layer l = 1 u and e v respectively, and W V ∈R 16×c·h is a learnable weight matrix of values; using

[0025]

[0026] Calculate the aggregated information of A u where W x ∈R 16×c·h ;

[0027] ② On the second layer of the graph neural network, use equations (2) and (3) to calculate the attention coefficients between entities, and then use

[0028]

[0029] Calculate the structural sampling term in the graph information bottleneck

[0030] The joint representation of the attention coefficients of A u is Among them, |N u,t | represents the number of neighbors included in N u,t ; for each head of attention, generate |N u,t | independent samples obeying the uniform distribution of (0, 1) Using

[0031]

[0032] Calculate the Gumbel distribution Using

[0033]

[0034] Calculate the reparameterized attention coefficients Among them, tem is the temperature parameter, and its value is 0.1; finally, use equation (4) to calculate

[0035] Calculate the mean of the Gaussian distribution and variance

[0036]

[0037] wherein, [·,·] represents extracting a specific segment within the brackets from the message vector; generating a Gaussian distribution using the mean and variance, and obtaining the actually transmitted message by sampling from the Gaussian distribution; adding one dimension at the end of ; and calculating the log probability under the Gaussian distribution using ; adding one dimension at the end of

[0038]

[0039] ; and calculating the log probability under the Gaussian distribution using using ; calculating the list of scaling parameters after passing through the Softplus activation function using

[0040]

[0041] wherein, σ wherein, σ u,z ∈ R c×z is the list of scaling parameters of the z-th Gaussian distribution in the Gaussian mixture model, which is initialized at the beginning of training, is the learnable parameter after σ u,z passes through the activation function, z is the number of any Gaussian distribution in the Gaussian mixture model, z ∈ {1,..., Z}, Z is the maximum number of Gaussian distributions, which is set to 100; calculating the log probability under any Gaussian distribution z using

[0042]

[0043] using ; calculating the log probability under the Gaussian mixture model using wherein, the weighted log probability μ u,z ∈ R c×Z is the learnable parameter matrix, which is also initialized at the beginning of training; calculating the log probability under the Gaussian mixture model using

[0044]

[0045] using ; calculating the node feature sampling in the graph information bottleneck using using

[0046]

[0047] ; finally, calculating the output of the second layer using Equation (5); Finally, calculate the output of the second layer using Equation (5);

[0048] ③On the 3rd layer of the graph neural network, first, calculate the attention coefficients between entities using Equations (2) and (3), and obtain and Secondly, obtain the output of the graph neural network module using Equation (5); then, select the node data of A from the node data i Finally, obtain as the input of the actor network, and calculate the input of the critic network:

[0049]

[0050] ④Perform the operations in ① to ③ on all moving objects A1,..., A N to obtain the inputs of the actor and critic networks for each moving object, and save the input and output data of the graph neural network and the actor-critic network to the training batch to prepare for the training of the algorithm;

[0051] (4) Train the decentralized multi-agent reinforcement learning algorithm

[0052] ①Set the algorithm training parameters: the maximum number of training steps is 2×10 6 , the number of training times of the proximal policy optimization algorithm is 10, the number of parallel environments is 128, the size B of the training batch is 25×128 = 3200, ρ = 1, N = D = O = 3;

[0053] ②Set the hyperparameters of the actor-critic network: 1-layer multi-layer perceptron, with the input layer dimension of 22, the hidden layer dimension of 64, and the output layer dimension of 64; 1-layer gated recurrent unit, with the input layer, hidden layer, and output layer dimensions all being 64; the fully connected layer of the actor network has an input dimension of 64 and an output dimension of 5; the fully connected layer of the critic network has an input dimension of 64 and an output dimension of 1; the activation function is the ReLU function, and the learning rate is 7×10 -4 ;

[0054] ③Use

[0055]

[0056] to train the actor network of A i ; among them, the clip(·,·,·) function restricts the data of the first item within the range of the latter two items, τ is the number of each group of data in the batch, θ i is the parameter of the actor network, is the policy update range limit parameter, κ = 0.01 is the policy entropy coefficient, β1 = 0.001 is the graph information bottleneck optimization term coefficient of the actor network, ​is the action conditional probability distribution output by the actor network, and is θ i The actor network output P before update i,τ is the probability of each item in and are the structure sampling and node feature sampling data stored in the τ-th group respectively, A i,τ is the advantage function, and its calculation formula is

[0057]

[0058] where λ = 0.95 is the generalized advantage estimator coefficient, γ = 0.99 is the discount factor, are the parameters of the critic network, is the output of the critic network stored in the τ-th group, is the output of the critic network in the (τ + 1)-th group, r i,τ is the reward stored in the τ-th group data in the training batch; adopt

[0059]

[0060] to train the critic network of A i ; where

[0061]

[0062] is the expectation of the cumulative reward starting from time t for A i β2 = 0.001 is the graph information bottleneck optimization term coefficient of the critic network, is the output of the critic network before update, ξ = 0.2 is the value function update range limit parameter;

[0063] ④ Use the training parameters of ① and ②, and the optimization objective of ③ to train the algorithm. After training reaches the maximum number of steps, terminate the training, save the weight parameters of the last training, and complete the model training;

[0064] (5) Use decentralized multi-agent reinforcement learning to achieve cooperative control

[0065] ① Use the hyperparameters and loss function of (4) to train the decentralized multi-agent reinforcement learning algorithm, save the episode reward, and output the reward function curve;

[0066] ② After completing the algorithm training, save the trained model and apply the model to the cooperative control task environment to complete the cooperative control of the complex moving body system. Description of the Drawings

[0067] Figure 1This is a reward curve diagram for the coordinated control of complex motion body systems according to the present invention.

[0068] Figure 2 This is a control effect diagram of the present invention for the coordinated control of complex motion body systems DETAILED DESCRIPTION

[0069] A method for coordinated control of a complex motion system based on decentralized networked multi-agent reinforcement learning, characterized by designing a task environment and a reward function, characterizing node features, performing information aggregation, training algorithms, and using decentralized networked multi-agent reinforcement learning to achieve coordinated motion control, including the following steps:

[0070] (1) Designing the task environment and reward function

[0071] A 2D multi-agent particle environment is used as the basis. The maximum number of time steps per round in this environment is T = 25. The area of ​​the environment adopts the default value of the multi-agent particle environment, which is 4×4. The number of moving bodies is N, the number of landmarks is D, and the number of obstacles is O. i represents the number of any moving body, i∈{1,...,N}; t represents the current time of the environment, t∈{1,...,T}; the i-th moving body A i The reward function r at time t during the movement i,t Designed for:

[0072]

[0073] Among them, pos i,t and Respectively represent A i The position at time t and the position of the target;

[0074] (2) Characterizing node features

[0075] The moving body, obstacles and landmarks are collectively referred to as entities e m , where m is the number of any entity, m∈{1,...,N+D+O}; the set {1,...,i-1,i+1,...,N+D+O} is used to represent all entities except A i The ID of any entity other than mn,t ∈R represents any two entities e m and e n The Euclidean distance between them, where m≠n, n∈{1,...,N+D+O} is the number of any entity except m; the definition is located in A i The set of entities within the perception radius ρ is N i,t , call these entities A i Neighbors, use e j Represents any neighbor, where j is the neighbor's number, and Vi,t = {N i,t ∪ A i} represents the set of A i and its neighbors; using the graph G t (V, E t ) represents a complex moving object system, where V = {e1,..., e N+D+O} represents the set of all entities, represents the set of edges at time t, using A i and e j the Euclidean distance d ij,t between them as the feature of the edge; A i forms a graph network g i,t ∈ G t , define the position, velocity of A i at time t in the two-dimensional coordinate system and the position of the target relative to itself as pos i,t ∈ R 2 , vel i,t ∈ R 2 and where R represents the real numbers;

[0076] Use pos i,t , vel i,t , and to represent the observation information Define and as the position, velocity of e j relative to A i and the position of e j the target relative to A i ; if e j is an obstacle or a target, set Use to represent the node features of each neighbor entity on g i,t , where et(j) is the entity type, defined as et(j) ∈ {"agent": 0, "landmark": 1, "obstacle": 2}, the strings within {·} represent text, the dictionaries agent, landmark, and obstacle represent moving objects, landmarks, and obstacles respectively, and the numbers 0, 1, and 2 represent the indices of the dictionaries;

[0077] (3) Perform information aggregation

[0078] Use an embedding layer with a dictionary size of 3 and an output dimension of 2 to encode et(j) to obtain the encoded entity type et ij,t , concatenate et ij,t and d ij,t to get a vector where represents the vector concatenation operation; input into a linear layer with an input dimension of 9 and an output dimension of 16 for feature extraction, so as to obtain the node features input into the graph neural network Adopt a 3-layer graph neural network as the information aggregation module. The number of each layer is l ∈ {1,..., L}, the maximum number of layers L = 3, and the input layer dimension of the graph neural network is set to 16;

[0079] ① On the first layer of the graph neural network, use

[0080]

[0081] to calculate the attention map where u and v are the numbers of any moving object and its neighbor on g i,t , u ∈ {1,..., N}, v ∈ N u,t , N u,t represents the set of neighbors of A u ; is the node feature of the moving object with any number u on the (l - 1)-th layer, is the node feature of the neighbor with any number v on the (l - 1)-th layer; c is the output dimension of each layer of the graph neural network. When l = 1 or 3, c = 16, and when l = 2, c = 32; W Q ∈ R 16×c·h is the learnable weight matrix for queries, is the learnable weight matrix for keys; use

[0082]

[0083] to calculate the attention coefficient between the moving object A u with any number u and its neighbor entity e v with any number v The number of attention heads is set to 3; use

[0084]

[0085] to calculate the message v transmitted from e u to A where, and are the node features of A u and e v on the first layer respectively, W V ∈ R 16×c·h is the learnable weight matrix for values; adopt

[0086]

[0087] Calculate A u The aggregated information, where W x ∈R 16×c·h ;

[0088] ② On the second layer of the graph neural network, use equations (2) and (3) to calculate the attention coefficients between entities, and then use

[0089]

[0090] Calculate the structural sampling term in the graph information bottleneck

[0091] Let the joint representation of the attention coefficients of A u be where |N u,t | represents the number of neighbors included in N u,t ; for each head of attention, generate |N u,t | independent samples uniformly distributed over (0, 1) Use

[0092]

[0093] to calculate the Gumbel distribution Use

[0094]

[0095] to calculate the reparameterized attention coefficients where tem is the temperature parameter with a value of 0.1; finally, use equation (4) to calculate

[0096] to calculate the mean of the Gaussian distribution and the variance

[0097]

[0098] where [·, ·] represents intercepting a specific segment within the brackets from the message vector ; use the mean and variance to generate a Gaussian distribution, and use sampling from the Gaussian distribution to obtain the actually transmitted message; at the end of add 1 dimension, and use

[0099]

[0100] to calculate the log probability under the Gaussian distribution Use

[0101]

[0102] Calculate the list of scaling parameters after passing through the Softplus activation function where σ u,z ∈ R c×z is the list of scaling parameters of the z-th Gaussian distribution in the Gaussian mixture model, initialized at the start of training, is σ u,z the learnable parameter after passing through the activation function, z is the number of any Gaussian distribution in the Gaussian mixture model, z ∈ {1,..., Z}, Z is the maximum number of Gaussian distributions, set to 100; use

[0103]

[0104] Calculate the log probability under any Gaussian distribution z where the weighted log probability μ u,z ∈ R c×Z is the learnable parameter matrix, which is also initialized at the start of training; use

[0105]

[0106] Calculate the log probability under the Gaussian mixture model Use

[0107]

[0108] Calculate the node feature sampling in the graph information bottleneck Finally, use Equation (5) to calculate the output of the second layer;

[0109] ③ On the third layer of the graph neural network, first, use Equations (2) and (3) to calculate the attention coefficients between entities, and use Equations (4) and (6) to obtain and Secondly, use Equation (5) to obtain the output of the graph neural network module; then, select from the node data the node data of A i Finally, obtain as the input of the actor network and calculate the input of the critic network:

[0110]

[0111] ④ Perform the operations of ① to ③ on all moving bodies A1,..., A N to obtain the actor and critic network inputs for each moving body, and save the input and output data of the graph neural network and the actor-critic network to the training batch, preparing for the training of the algorithm;​

[0112] (4) Train the decentralized networked multi-agent reinforcement learning algorithm

[0113] ① Set the algorithm training parameters: The maximum number of training steps is 2×10 6 , the number of training times of the proximal policy optimization algorithm is 10, the number of parallel environments is 128, the size B of the training batch is 25×128 = 3200, ρ = 1, N = D = O = 3;

[0114] ② Set the hyperparameters of the actor-critic network: 1-layer multi-layer perceptron, with an input layer dimension of 22, a hidden layer dimension of 64, and an output layer dimension of 64; 1-layer gated recurrent unit, with input layer, hidden layer, and output layer dimensions all of 64; the input dimension of the fully connected layer of the actor network is 64, and the output dimension is 5; the input dimension of the fully connected layer of the critic network is 64, and the output dimension is 1; the activation function is the ReLU function, and the learning rate is 7×10 -4 ;

[0115] ③ Use

[0116]

[0117] to train the actor network of A i ; among them, the clip(·,·,·) function limits the data of the first item within the range of the last two items, τ is the number of each group of data in the batch, θ i is the parameter of the actor network, is the policy update range limit parameter, κ = 0.01 is the policy entropy coefficient, β1 = 0.001 is the graph information bottleneck optimization term coefficient of the actor network, is the action conditional probability distribution output by the actor network, and is θ i the output P of the actor network before the update of i,τ is the probability of each item in and are the structure sampling and node feature sampling data stored in the τ-th group respectively, A i,τ is the advantage function, and the calculation formula is

[0118]

[0119] where, λ = 0.95 is the generalized advantage estimator coefficient, γ = 0.99 is the discount factor, is the parameter of the critic network, is the output of the critic network stored in the τ-th group, is the output of the critic network in the τ+1-th group, r i,τ is the reward stored in the τ-th group of data in the training batch; use

[0120]

[0121] Train the critic network for A; where i Among them

[0122]

[0123] is the expected value of the episode reward starting from time t, β2 = 0.001 is the coefficient of the graph information bottleneck optimization term of the critic network, i is the output of the critic network before update, ξ = 0.2 is the value function update range limit parameter; is ④ Use the training parameters of ① and ②, and the optimization objective of ③ to train the algorithm. After the training reaches the maximum number of steps, terminate the training, save the weight parameters of the last training, and complete the model training;

[0124] ④ Use the training parameters of ① and ②, and the optimization objective of ③ to train the algorithm. After the training reaches the maximum number of steps, terminate the training, save the weight parameters of the last training, and complete the model training;

[0125] (5) Use decentralized multi-agent reinforcement learning to achieve cooperative control

[0126] ① Use the hyperparameters and loss function of (4) to train the decentralized multi-agent reinforcement learning algorithm, save the episode reward, and output the reward function curve;

[0127] ② After completing the algorithm training, save the trained model and apply the model to the cooperative control task environment to complete the cooperative control of the complex moving body system.

Claims

1. A cooperative control method for a complex moving body system based on decentralized networked multi-agent reinforcement learning, characterized in that, Design the task environment and reward function, characterize the node features, perform information aggregation, train the algorithm, and use decentralized networked multi-agent reinforcement learning to achieve cooperative motion control, including the following steps: (1) Design the task environment and reward function Based on a 2D multi-agent particle environment, the maximum number of time steps per episode in this environment is \(T = 25\). The area of the environment adopts the default value of the multi-agent particle environment, with an area of \(4\times4\). The number of moving bodies is \(N\), the number of landmarks is \(D\), and the number of obstacles is \(O\). \(i\) represents the number of any moving body, \(i\in\{1,\ldots,N\}\); \(t\) represents the current time of the environment, \(t\in\{1,\ldots,T\}\); the \(i\)-th moving body \(A\) i The reward function \(r\) at time \(t\) during the movement process i,t is designed as: where pos i,t and represent the position of A i at time t and the position of the target, respectively; (2) Characterize the node features The moving object, obstacle, and landmark are collectively referred to as entity e m , where m is the number of any entity, m ∈ {1,..., N + D + O}; the set {1,..., i - 1, i + 1,..., N + D + O} is used to represent the numbers of any entities except A i . The number d mn , t ∈ R is used to represent the Euclidean distance between any two entities e m and e n , where m ≠ n and n ∈ {1,..., N + D + O} is the number of any entity other than m; the set of entities within the perception radius ρ of A i is defined as N i,t . These entities are called the neighbors of A i . The arbitrary neighbor is represented by e j , where j is the number of the neighbor. Let V i,t = {N i,t ∪ A i} represent the set of A i and its neighbors; the complex moving object system is represented by the graph G t (V, E t ), where V = {e1,..., e N+D+O} represents the set of all entities, represents the set of edges at time t. The Euclidean distance d i between A j and e ij,t is used as the feature of the edge; A i forms a graph network g i,t ∈ G t at time t. Define the position, velocity of A i in the two-dimensional coordinate system at time t and the position of the target relative to itself as pos i,t ∈ R 2 , vel i,t ∈ R 2 and where R represents the set of real numbers; Use pos i,t 、vel i,t and to represent the observation information Define and as the position, velocity of e j relative to A i and the position of the e j target relative to A i ; if e j is an obstacle or a target, set Use to represent the node features of each neighbor entity on g i,t where et(j) is the entity type, defined as et(j) ∈ {"agent": 0, "landmark": 1, "obstacle": 2}, the strings within {·} represent text, the dictionaries agent, landmark, and obstacle represent moving objects, landmarks, and obstacles respectively, and the numbers 0, 1, and 2 represent the indices of the dictionaries; (3) Perform information aggregation Encode the entity type et(j) using an embedding layer with a dictionary size of 3 and an output dimension of 2 to obtain the encoded entity type et ij,t , concatenate et ij,t and d ij,t to obtain the vector where represents the vector concatenation operation; input into a linear layer with an input dimension of 9 and an output dimension of 16 for feature extraction to obtain the node features input to the graph neural network Use a 3-layer graph neural network as the information aggregation module, with each layer numbered l ∈ {1,..., L}, the maximum number of layers L = 3, and the input layer dimension of the graph neural network set to 16; ① On the first layer of the graph neural network, use Calculate the attention map where u and v are the numbers of any moving object and its neighbor on g i,t respectively, u ∈ {1,..., N}, v ∈ N u,t , N u,t represents the set of u neighbors; is the node feature of the moving object numbered u in the (l - 1)-th layer, is the node feature of the neighbor numbered v in the (l - 1)-th layer; c is the dimension of the output of each layer of the graph neural network. When l = 1 or 3, c = 16, and when l = 2, c = 32; W Q ∈ R 16×c·h is the learnable weight matrix for the query, is the learnable weight matrix for the key; Using Calculate the moving object A with any number u u and its neighbor entity e with any number v v The attention coefficient between them The number of attention heads is set to 3; Use Calculate e at layer l = 1 v Pass it to A u The message Where And Are the node features of A at layer l = 1 u And e v Respectively, and W V ∈ R 16×c·h Is a learnable weight matrix of values; Use Calculate A u The aggregated information, where W x ∈R 16×c·h ; ② On the second layer of the graph neural network, calculate the attention coefficients between entities using equations (2) and (3), and then use Computing the Structural Sampling Term in the Information Bottleneck of Graphs Combine A u The joint representation of the attention coefficients is where |N u,t | represents the number of neighbors included in N u,t ; for each head of attention, generate |N u,t | independent samples uniformly distributed over (0, 1) Utilize Calculating the Gumbel distribution Using Calculate the attention coefficients after reparameterization where tem is the temperature parameter with a value of 0.1; finally, calculate using Equation (4) Calculate the mean of the Gaussian distribution and variance Among them, [·, ·] represents intercepting specific segments within the brackets for the message vector ; generating a Gaussian distribution using the mean and variance, and obtaining the actually transmitted message by sampling from the Gaussian distribution; at the end, adding one dimension, and using Calculate the log probability under a Gaussian distribution of using Calculate the list of scaling parameters after passing through the Softplus activation function where, σ u,z ∈R c×z is the list of scaling parameters of the z-th Gaussian distribution in the Gaussian mixture model, which is initialized at the beginning of training, is the learnable parameter after σ u,z passes through the activation function, z is the number of any Gaussian distribution in the Gaussian mixture model, z ∈ {1,..., Z}, Z is the maximum value of the number of Gaussian distributions, set to 100; using Compute the log probability under any Gaussian distribution z where the weight log probability μ ∈R u,z is a learnable parameter matrix, which is also initialized at the start of training; using c×Z ​ Calculate the log probability under the Gaussian mixture model Use​ Node Feature Sampling in Computational Graph Information Bottleneck Finally, the output of the second layer is calculated using Equation (5); ③On the third layer of the graph neural network, first, calculate the attention coefficients between entities using equations (2) and (3), and obtain and Second, obtain the output of the graph neural network module using equation (5); then, select the node data of A from the node data i Finally, obtain as the input of the actor network and calculate the input of the critic network: as the input of the actor network and calculate the input of the critic network: ④For all moving bodies A1,..., A N Perform the operations ① to ③ to obtain the actor and critic network inputs for each moving body, save the input and output data of the graph neural network and the actor-critic network to the training batch, and prepare for the training of the algorithm; (4) Train the decentralized networked multi-agent reinforcement learning algorithm ① Set the algorithm training parameters: the maximum number of training steps is 2×10 6 , the number of training times of the proximal policy optimization algorithm is 10, the number of parallel environments is 128, the size B of the training batch is 25×128 = 3200, ρ = 1, N = D = O = 3; ②Set the hyperparameters of the actor-critic network: a 1-layer multi-layer perceptron with an input layer dimension of 22, a hidden layer dimension of 64, and an output layer dimension of 64; a 1-layer gated recurrent unit with input, hidden, and output layer dimensions all of 64; the fully connected layer of the actor network has an input dimension of 64 and an output dimension of 5; the fully connected layer of the critic network has an input dimension of 64 and an output dimension of 1; the activation function is the ReLU function, and the learning rate is 7×10 -4 ; ③ Use Train the actor network of A i ; where the clip(·,·,·) function restricts the data of the first item within the range of the last two items, τ is the number of each group of data in the batch, and θ i are the parameters of the actor network, is the policy update range limit parameter, κ = 0.01 is the policy entropy coefficient, and β1 = 0.001 is the graph information bottleneck optimization term coefficient of the actor network, is the action conditional probability distribution output by the actor network, and is θ i The output P of the actor network before the update i,τ is the probability of each item in and are the structure sampling and node feature sampling data stored in the τ-th group respectively, and A i,τ is the advantage function, and the calculation formula is where λ = 0.95 is the coefficient of the generalized advantage estimator, γ = 0.99 is the discount factor, are the parameters of the critic network, is the output of the critic network stored in the τ-th group, is the output of the critic network in the (τ + 1)-th group, and r i,τ is the reward stored for the data in the τ-th group in the training batch; The Train the critic network for A i ; wherein It is A i The expected value of the round reward starting from time t, where β2 = 0.001 is the coefficient of the graph information bottleneck optimization term of the critic network, is the output of the critic network before update, and ξ = 0.2 is the value function update range limit parameter; ④ Use the training parameters of ① and ②, and the optimization objective of ③ to train the algorithm. After training reaches the maximum number of steps, terminate the training, save the weight parameters of the last training, and complete the model training; (5) Use decentralized networked multi-agent reinforcement learning to achieve cooperative control ① Use the hyperparameters and loss function of (4) to train the decentralized networked multi-agent reinforcement learning algorithm, save the episode rewards, and output the reward function curve; ② After completing the algorithm training, save the trained model and apply the model to the cooperative control task environment to complete the cooperative control of the complex moving body system.

2. The method according to claim 1, wherein: This method first designs the cooperative control task of the moving body based on the multi-agent particle environment, and obtains the observed information data by interacting with the task environment; secondly, transmits these data to the server and converts them into a graph structure according to the position of each entity in the global coordinate system to characterize the data features; Then, use the graph neural network to aggregate the node features of each moving body and its neighbor entities, thereby training the decentralized networked multi-agent reinforcement learning algorithm; finally, each moving body uses the trained model to independently make decisions through the data collected by odometer positioning and inertial measurement units, quickly reach the desired position and avoid collisions.

Citation Information

Patent Citations

  • Shiftable table of hydraulic forming machine

    CN2440618Y

  • Conveying gearing mechanism without pulley track

    CN2448893Y

  • Couplings (2)

    CN3160129D

  • Packaging bag (Dongguan sausage)

    CN3188332D

  • Multi-agent reinforcement learning adversarial strategy detection method based on graph neural network

    CN118133879A