Mobile multi-agent knowledge migration method based on replacement strategy network

By embedding a hypernetwork framework into a permutation policy network, the problem of inefficient knowledge transfer in multi-agent systems under varying agent numbers and dynamic environmental conditions is solved. This achieves efficient and stable policy transfer and improved learning performance, making it suitable for collaborative decision-making in multi-agent systems under complex dynamic environments.

CN120893518AActive Publication Date: 2025-11-04BEIJING INFORMATION SCI & TECH UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511002106.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-11-04
Estimated Expiration
2045-07-21

AI Technical Summary

Technical Problem

Traditional multi-agent reinforcement learning methods suffer from the curse of dimensionality and inefficient knowledge transfer when faced with an increasing number of agents and dynamic environmental changes. In particular, policy learning efficiency is low in complex scenarios, and existing methods are unable to capture the permutation invariance and permutation isovariance of multi-agent systems, leading to policy oscillations or negative transfer.

Method used

A mobile multi-agent knowledge transfer method based on permutation policy networks is adopted. By embedding a hypernetwork framework with permutation invariance and permutation covariance policy networks, a dynamic adaptation relationship between agent size and environmental changes is established. Combined with centralized training and distributed execution, a knowledge transfer model containing permutation invariance and covariance constraints is constructed.

Benefits of technology

It achieves efficient policy transfer under varying agent numbers and dynamic environments, improving learning speed by over 50%, asymptotic performance by 60%, significantly enhancing training stability, reducing threshold time by 64-78%, increasing final reward by 59.4-73.7%, and improving policy convergence speed by over 50%, while avoiding the uninterpretability of black-box models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120893518A_ABST
    Figure CN120893518A_ABST
Patent Text Reader

Abstract

The invention discloses a mobile multi-agent knowledge migration method based on a replacement strategy network, and relates to the technical field of multi-agent reinforcement learning. Comprising the following steps: embedding a replacement invariance strategy network and a replacement covariance strategy network into a super network framework, dynamically generating weight matrixes of an input layer and an output layer through a super network, and establishing a dynamic adaptation relationship between a joint state-action space and an intelligent agent scale as well as environment change; a permutation matrix characteristic is introduced to realize decoupling of agent sequence independence and task target responsiveness, and strategy network parameters are optimized through a centralized training-distributed execution architecture; constructing a knowledge migration model containing permutation invariance and homodenaturation constraints; and efficient migration of strategies between similar domain tasks is realized. According to the method, the problem of low knowledge migration efficiency caused by intelligent agent scale change and joint state-action space dimension explosion in a dynamic complex environment is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multi-agent reinforcement learning technology, and more specifically to a mobile multi-agent knowledge transfer method based on a permutation policy network. Background Technology

[0002] With the widespread application of mobile multi-agent systems in complex scenarios such as industrial collaboration, intelligent navigation, and emergency rescue, the increasing number of agents and dynamic environmental changes lead to an exponential expansion of the joint state-action space. This poses a dual challenge to traditional multi-agent reinforcement learning (MARL) methods: the curse of dimensionality and inefficient knowledge transfer. In multi-agent collaborative tasks, dynamic interactions between agents and environmental disturbances (such as obstacle movement and task objective reassignment) exacerbate the non-stationarity of the system, resulting in extremely low sample efficiency for policy learning.

[0003] Traditional multi-agent knowledge transfer methods primarily rely on cross-domain feature alignment (such as partial domain adaptation PADA). However, these methods fail to capture the permutation invariance (agent order does not affect the global collaborative goal) and permutation isovariance (actions need to be adjusted according to the goal allocation order) of multi-agent systems. For example, in cooperative navigation tasks, the physical identity independence of agents requires the input layer to be insensitive to order, while in goal point reassignment scenarios, action outputs need to change synchronously with the agent order. Traditional methods, lacking such inductive biases, are prone to policy oscillations or negative transfer, resulting in transfer efficiency improvements of less than 20% in scenarios with varying agent scale.

[0004] In recent years, Physical Information Neural Networks (PINNs) have improved model interpretability by embedding physical constraints, but their application in multi-agent scenarios still has limitations. For example, while Deep Lagrange Networks (DeLaNs) can guarantee the physical consistency of the inertia matrix through mechanical principles, they cannot model non-conservative interactions between agents (such as communication delays and cooperation conflicts), resulting in a policy error rate exceeding 30% in dynamic task allocation. Existing improvement methods (such as introducing dissipative force networks) rely on manually designed physical models, making it difficult to generalize complex interactions; while purely data-driven black-box models (such as FFNNs) lack structural constraints, leading to a policy divergence risk as high as 40% during sample extrapolation.

[0005] Furthermore, the contradiction between the distributed execution requirements of multi-agent systems and the overhead of centralized training is becoming increasingly prominent. Traditional centralized training methods exhibit exponentially increasing computational complexity with the number of agents, while independent learning suffers from policy collapse due to environmental non-stationarity. How to achieve both high efficiency and robustness in distributed execution while ensuring global policy optimization has become a key bottleneck restricting the large-scale application of multi-agent systems.

[0006] Therefore, proposing a mobile multi-agent knowledge transfer method based on permutation policy networks to address the difficulties in existing technologies is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0007] In view of this, the present invention provides a mobile multi-agent knowledge transfer method based on permutation policy networks, which aims to solve the problem of inefficient knowledge transfer caused by changes in agent size and the explosion of joint state-action space dimensions in dynamic and complex environments.

[0008] To achieve the above objectives, the present invention provides the following technical solution:

[0009] A mobile multi-agent knowledge transfer method based on a permutation policy network includes the following steps:

[0010] S1. Embed the permutation-invariant policy network and the permutation-equivariant policy network into the supernetwork framework. The supernetwork dynamically generates the input and output layer weight matrices to establish a dynamic adaptation relationship between the joint state-action space and the agent size and environmental changes.

[0011] S2. The permutation matrix property is introduced to decouple the order independence of the agent from the responsiveness of the task objective, and the policy network parameters are optimized through a centralized training-distributed execution architecture.

[0012] S3. Construct a knowledge transfer model that includes constraints on permutation invariance and isovariance;

[0013] S4. For similar domain tasks with varying numbers of agents or dynamic environmental adjustments, the constructed knowledge transfer model is used to achieve efficient policy transfer between similar domain tasks.

[0014] Optionally, the input layer of the permutation-invariant policy network in S1 satisfies permutation invariance, specifically:

[0015] For any permutation matrix g, the input layer outputs h. in Satisfy h in (g·X)=h in (X), where X is the agent's observed feature vector, and the input layer weight matrix W is generated through a hypernetwork. i The calculation formula is:

[0016]

[0017] Where, x i W represents the observational features of a single agent. i Based on x by the hypernet i generate.

[0018] Optionally, the output layer of the permutation isovariance policy network in S1 satisfies permutation isovariance, specifically:

[0019] For any permutation matrix g, the output layer action a satisfies a(g·X)=g·a(X), and the output layer weight matrix W is generated through the hypernetwork. ji The calculation formula is:

[0020]

[0021] Among them, h hidden W is the output of the hidden layer of the neural network. ji It dynamically adjusts according to the input order.

[0022] Optionally, the hypernetwork framework in S1 is a neural network, with the input being the agent's observed features or task environment parameters, and the output being the weight matrix of the policy network;

[0023] The hypernetwork framework uses two fully connected layers for both the input and output layers, with a hidden layer dimension of 64 and the activation function being ReLU. The output layer generates a weight matrix through a linear transformation.

[0024] Optionally, S2 introduces the permutation matrix property to decouple the agent's order independence from the task objective responsiveness, and optimizes the strategy network parameters through a centralized training-distributed execution architecture. The specific details are as follows:

[0025] Based on the combination of policy network and multi-agent deep reinforcement learning algorithm, a centralized training-distributed execution framework is adopted, with global observation optimizing policy training and local observation enabling independent decision-making and execution.

[0026] The objective function for training the global observation optimization strategy is:

[0027]

[0028] Where r(θ) is the strategy ratio, For generalized advantage estimation, α is the entropy regularization coefficient, S(π) θ ) represents the policy entropy. Let be the mathematical expectation, and ∈ be the clipping parameter of PPO.

[0029] Optionally, in S4, similar domain task migration for changes in the number of agents or dynamic adjustments to the environment includes scenarios where the number of agents changes or the environment is dynamically adjusted. Permutation invariance ensures that the input layer is insensitive to the number of agents, and permutation isovariance adapts to changes in target point allocation or obstacle changes in task target adjustment.

[0030] As can be seen from the above technical solution, compared with the prior art, the present invention provides a mobile multi-agent knowledge transfer method based on a permutation policy network, which has the following beneficial effects:

[0031] (1) This invention effectively solves the problems of joint state-action space dimension explosion and inefficient knowledge transfer in dynamic environments in multi-agent reinforcement learning; it integrates permutation invariance and permutation isovariance policy networks into a super network framework, and establishes a dynamic adaptation relationship between agent size, environmental changes and policy space by dynamically generating input and output layer weight matrices; it proposes an inductive bias based on the characteristics of permutation matrix to decouple agent order independence from task target responsiveness, and constructs a knowledge transfer model with physical consistency constraints;

[0032] (2) The multi-agent cooperative navigation experiment in the Gym environment shows that the present invention can adapt to changes in the number of agents and dynamic obstacle scenarios without retraining, and achieve efficient policy transfer. Compared with existing methods (such as MAPPO and PADA), the learning speed is improved by more than 50%, the asymptotic performance (final reward) is improved by 60%, and the training stability is significantly enhanced.

[0033] (3) In tasks with an increased number of agents, the threshold time of PPN is reduced by 64% compared to the baseline MAPPO (from 10e6 steps to 3.6e6 steps), and the final reward is increased by 59.4% (from 496.11 to 790.60); in dynamic environment tasks, the threshold time is reduced by 78% (from 10e6 steps to 2.2e6 steps), and the final reward is increased by 73.7% (from 410.56 to 713.14);

[0034] (4) In untrained dynamic obstacle scenarios, the reward fluctuation amplitude of PPN is reduced by 42% compared with the PADA method, and the asymptotic performance (R) is significantly improved. 2 The convergence speed of the strategy is significantly better than that of traditional methods (>0.95), and the convergence speed is improved by more than 50% under external interference.

[0035] (5) This invention ensures that the agent's strategy conforms to the physical laws of multi-agent interaction (such as order independence and target responsiveness) by permutation invariance / homogeneity constraints, avoids the uninterpretability of black box models, and improves the reliability of the model.

[0036] (6) Through theoretical modeling and experimental verification, this invention provides an efficient and interpretable technical solution for real-time collaborative decision-making of multi-agent systems in complex dynamic environments, and has broad application prospects in fields such as industrial collaboration and intelligent navigation. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0038] Figure 1 A flowchart of a mobile multi-agent knowledge transfer method based on a permutation policy network provided by the present invention;

[0039] Figure 2 This is a schematic diagram of the network structure for the output layer permutation and assimilation strategy provided by the present invention;

[0040] Figure 3 This is a schematic diagram of the network structure for the input layer permutation invariant strategy provided by the present invention. Detailed Implementation

[0041] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0042] See Figure 1 As shown, this invention discloses a mobile multi-agent knowledge transfer method based on a permutation policy network, comprising the following steps:

[0043] S1. Embed the permutation-invariant policy network and the permutation-equivariant policy network into the supernetwork framework. The supernetwork dynamically generates the input and output layer weight matrices to establish a dynamic adaptation relationship between the joint state-action space and the agent size and environmental changes.

[0044] S2. The permutation matrix property is introduced to decouple the order independence of the agent from the responsiveness of the task objective, and the policy network parameters are optimized through a centralized training-distributed execution architecture.

[0045] S3. Construct a knowledge transfer model that includes constraints on permutation invariance and isovariance;

[0046] S4. For similar domain tasks with varying numbers of agents or dynamic environmental adjustments, the constructed knowledge transfer model is used to achieve efficient policy transfer between similar domain tasks.

[0047] Furthermore, the input layer of the permutation-invariant policy network in S1 satisfies permutation invariance, specifically as follows:

[0048] For any permutation matrix g, the input layer outputs h. in Satisfy h in (g·X)=h in (X), where X is the agent's observed feature vector, and the input layer weight matrix W is generated through a hypernetwork. i The calculation formula is:

[0049]

[0050] Where, x i W represents the observational features of a single agent. i Based on x by the hypernet i generate.

[0051] Furthermore, the output layer of the permutation isovariance policy network in S1 satisfies permutation isovariance, specifically:

[0052] For any permutation matrix g, the output layer action a satisfies a(g·X)=g·a(X), and the output layer weight matrix W is generated through the hypernetwork. ji The calculation formula is:

[0053]

[0054] Among them, h hidden W is the output of the hidden layer of the neural network. ji It dynamically adjusts according to the input order.

[0055] Specifically, the permutation matrix property constraint: utilizing the orthogonality of the permutation matrix ((P T P = E)) and the transformation property ((PTAP)) ensure the mathematical consistency between the permutation-invariant policy network (PI) and the permutation-equivariant policy network (PE), where (P) is the permutation matrix and (A) is the state-action space matrix.

[0056] Furthermore, the hypernetwork framework in S1 is a neural network, with the input being the agent's observed features or task environment parameters, and the output being the weight matrix of the policy network;

[0057] The hypernetwork framework uses two fully connected layers for both the input and output layers, with a hidden layer dimension of 64 and the activation function being ReLU. The output layer generates a weight matrix through a linear transformation.

[0058] Furthermore, S2 introduces the permutation matrix property to decouple the agent's order independence from the task objective responsiveness, and optimizes the strategy network parameters through a centralized training-distributed execution architecture. The specific details are as follows:

[0059] Based on the combination of policy network and multi-agent deep reinforcement learning algorithm, a centralized training-distributed execution framework is adopted, with global observation optimizing policy training and local observation enabling independent decision-making and execution.

[0060] The objective function for training the global observation optimization strategy is:

[0061]

[0062] Where r(θ) is the strategy ratio, For generalized advantage estimation, α is the entropy regularization coefficient, S(π) θ ) represents the policy entropy. Let be the mathematical expectation, and ∈ be the clipping parameter of PPO.

[0063] Furthermore, in S4, similar domain task migration for changes in the number of agents or dynamic adjustments to the environment includes scenarios where the number of agents changes or the environment is dynamically adjusted. Permutation invariance ensures that the input layer is insensitive to the number of agents, and permutation isovariance adapts to changes in target point allocation or obstacle changes in task target adjustment.

[0064] The experimental verification framework of this invention is as follows: a two-dimensional grid environment is constructed based on OpenAI Gym, and two types of tasks are set up: increasing the number of agents (from 3 to 5) and increasing the complexity of the environment (dynamic obstacles). The average round reward and the threshold time (the number of training rounds to reach the convergence reward of the baseline algorithm) are used as evaluation indicators.

[0065] Example 1

[0066] To address the issues of dimensionality explosion in the joint state-action space and inefficiency in knowledge transfer from dynamic environments in multi-agent reinforcement learning, this invention embeds permutation invariance (PI) and permutation isovariance (PE) policy networks into a supernetwork framework, proposing a similar domain knowledge transfer method based on permutation policy networks (PPN).

[0067] To address the input order sensitivity issue caused by changes in agent size, a hypernetwork dynamically generates the input layer weight matrix, mapping local agent observation features to permutation-invariant global features, ensuring the input layer is insensitive to increases or decreases in the number of agents. For example, in scenarios expanding from 3 to 5 agents, the PI network uses summation pooling operations h... in =Σ i W i x i Eliminate sequence interference, where W i The hypernetwork is based on the observations of a single agent x i generate.

[0068] To address the action response requirements arising from adjustments to task objectives, a PE network is designed to generate a dynamic weight matrix for the output layer via a supernetwork, ensuring that the action output maintains a consistent relationship with the agent's sequence. Taking a dynamic target point allocation task as an example, the PE network ensures that when the order of target points is changed, the agent's action 'a' remains consistent. j Synchronous adjustments are made to satisfy a(g·X)=g·a(X) (where g is the permutation matrix), maintaining the consistency of system cooperation.

[0069] Experiments were conducted in the OpenAI Gym multi-agent cooperative navigation environment to validate the PPN, setting up two tasks: expanding the number of agents (from 3 to 5) and dynamic obstacles. Results show that in the agent expansion task, PPN reduces the threshold time by 64% compared to the baseline MAPPO, and improves the final reward by 59.4%; in the dynamic environment task, it reduces the threshold time by 78% and decreases the reward fluctuation by 42%. Compared to the partial domain adaptation method PADA, PPN significantly optimizes learning speed, asymptotic performance, and stability, validating the effectiveness of the permutation policy network and the supernetwork dynamic adaptation mechanism.

[0070] Example 2

[0071] In multi-agent systems, when the number of agents changes (e.g., from 3 to 5) or the input order is randomized, traditional fixed-architecture networks cannot maintain the consistency of feature extraction, leading to knowledge transfer failure. This embodiment is based on... Figure 2 The PI network structure addresses the input order sensitivity issue and ensures permutation invariance of the joint state space, including the following:

[0072] 1. PI Network Core Architecture:

[0073] Hypernetwork design: such as Figure 2 As shown, the supernetwork of the PI network is a two-layer fully connected neural network, and the input is the local observation of a single agent. (e.g., position, velocity, distance to the target point), the hidden layer uses the ReLU activation function, and the output dimension is d. in ×d feat The weight matrix W i , where d in For the input layer dimension.

[0074] Summation pooling operation: via The features of all agents are fused, where m is the number of agents. This operation ensures that the global features h are preserved even when the input order is changed. in Keeps unchanged, i.e., h in (g·X)=h in (X), where g is the permutation matrix.

[0075] 2. Dynamically adapting to changes in agent size:

[0076] Parameter generation mechanism: When the number of agents increases from m to m', the supernetwork does not need to be retrained and directly generates the corresponding W for the newly added agents. i Matrix. For example, when expanding from 3 agents to 5 agents, the hypernetwork generates W4 and W5 for the observations x4 and x5 of the 4th and 5th agents, and then uses summation pooling to convert h... in The dimension remains d in To avoid dimensional explosion.

[0077] Physical meaning: Through permutation invariance, the physical identity (such as number) of an agent does not affect global features, but only depends on its observation content (such as positional relationships), which is consistent with the isomorphism assumption of multi-agent systems.

[0078] 3. Integration with the MAPPO algorithm:

[0079] Training phase: In intensive training, the output h of the PI network is... in The Critic network, which uses MAPPO as the global feature input, optimizes the PPO loss function to ensure that the hypernetwork parameters {W} are consistent. i Co-convergence with the policy network;

[0080]

[0081] Where r(θ) is the strategy ratio, For generalized advantage estimation, α is the entropy regularization coefficient, S(π) θ ) represents the policy entropy. Let be the mathematical expectation, and ∈ be the clipping parameter of PPO.

[0082] Execution phase: The agent only needs to input its own observation x into the supernetwork. i You can then generate your own W. i It eliminates the need to share sequence information with other intelligent agents, thus reducing communication overhead.

[0083] Example 3

[0084] In multi-agent collaborative tasks, dynamic adjustments to the task objective (such as target point reassignment or agent role switching) can lead to a mismatch between the action space and the agent order. Traditional fixed-weight networks cannot synchronously adjust action outputs, resulting in policy confusion. This embodiment is based on... Figure 3 The PE network structure addresses the issue of sequential dependencies in action responses, ensuring consistent system collaboration. This includes the following:

[0085] 1. PE Network Core Architecture:

[0086] Hypernetwork design: such as Figure 3 As shown, the supernetwork input of the PE network is the hidden layer features. (Includes global collaboration information), output dimension is d action ×d hidden The weight matrix W ji , where j is the agent index and i is the action dimension index.

[0087] Generation of substitution-based isomorphic actions: through The calculation of actions ensures that when the input order is affected by the permutation matrix g, the action output satisfies a(g·X)=g·a(X). For example, when the target point order is permuted from [T1,T2,T3] to [T2,T1,T3], the action sequence generated by the PE network is synchronously adjusted to [a2,a1,a3] to maintain the one-to-one correspondence between the agent and the target.

[0088] 2. Dynamic task target adaptation:

[0089] The weight matrix is ​​dynamically updated: when the task objectives are replaced (such as the reassignment of search and rescue areas in emergency rescue), the hypernetwork updates the weight matrix according to the new objective order information (embedded h). hidden Generate a new W ji 'Matrix, no need to retrain the entire network.'

[0090] Action space decoupling: Actions are divided into entity-related actions (such as target point assignment actions) and entity-independent actions (such as movement speed). Entity-related actions achieve permutation isovariance through PE networks, while entity-independent actions maintain permutation invariance through PI networks, ensuring that the action logic conforms to physical constraints.

[0091] 3. Integration with the MAPPO algorithm:

[0092] Training phase: In centralized training, the target point order information is treated as part of the global state, and the hypernetwork parameters are optimized through the PPO loss function to strengthen the covariance between action and target order.

[0093] Execution phase: The agent, based on the target point index observed locally, uses W... ji The matrix is ​​directly mapped to the corresponding action, with a decision latency of less than 10ms, meeting the requirements for real-time task adjustment.

[0094] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0095] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A mobile multi-agent knowledge transfer method based on a permutation policy network, characterized in that, Includes the following steps: S1. Embed the permutation-invariant policy network and the permutation-equivariant policy network into the supernetwork framework. The supernetwork dynamically generates the input and output layer weight matrices to establish a dynamic adaptation relationship between the joint state-action space and the agent size and environmental changes. S2. The permutation matrix property is introduced to decouple the order independence of the agent from the responsiveness of the task objective, and the policy network parameters are optimized through a centralized training-distributed execution architecture. S3. Construct a knowledge transfer model that includes constraints on permutation invariance and isovariance; S4. For similar domain tasks with varying numbers of agents or dynamic environmental adjustments, the constructed knowledge transfer model is used to achieve efficient policy transfer between similar domain tasks.

2. The mobile multi-agent knowledge transfer method based on a permutation policy network according to claim 1, characterized in that, The input layer of the permutation-invariant policy network in S1 satisfies permutation invariance, specifically: For any permutation matrix g, the input layer outputs h. in Satisfy h in (g·X)=h in (X), where X is the agent's observed feature vector, and the input layer weight matrix W is generated through a hypernetwork. i The calculation formula is: Where, x i W represents the observational features of a single agent. i Based on x by the hypernet i generate.

3. The mobile multi-agent knowledge transfer method based on a permutation policy network according to claim 1, characterized in that, The output layer of the permutation-homogeneous policy network in S1 satisfies permutation-homogeneousness, specifically: For any permutation matrix g, the output layer action a satisfies a(g·X)=g·a(X), and the output layer weight matrix W is generated through the hypernetwork. ji The calculation formula is: Among them, h hidden W is the output of the hidden layer of the neural network. ji It dynamically adjusts according to the input order.

4. The mobile multi-agent knowledge transfer method based on a permutation policy network according to claim 1, characterized in that, The S1 super network framework is a neural network. The input is the agent's observed features or task environment parameters, and the output is the weight matrix of the policy network. The hypernetwork framework uses two fully connected layers for both the input and output layers, with a hidden layer dimension of 64 and the activation function being ReLU. The output layer generates a weight matrix through a linear transformation.

5. The mobile multi-agent knowledge transfer method based on a permutation policy network according to claim 1, characterized in that, S2 introduces the permutation matrix property to decouple the agent's order independence from the task objective responsiveness, and optimizes the strategy network parameters through a centralized training-distributed execution architecture. The specific details are as follows: Based on the combination of policy network and multi-agent deep reinforcement learning algorithm, a centralized training-distributed execution framework is adopted, with global observation optimizing policy training and local observation enabling independent decision-making and execution. The objective function for training the global observation optimization strategy is: Where r(θ) is the strategy ratio, For generalized advantage estimation, α is the entropy regularization coefficient, S(π) θ ) represents the policy entropy. Let be the mathematical expectation, and ∈ be the clipping parameter of PPO.

6. The mobile multi-agent knowledge transfer method based on a permutation policy network according to claim 1, characterized in that, S4 addresses similar domain task transfer scenarios involving changes in the number of agents or dynamic environmental adjustments. It ensures that the input layer is insensitive to the number of agents through permutation invariance and adapts to changes in target point allocation or obstacle changes through permutation isovariance.

Citation Information

Patent Citations

  • Link prediction method of few-sample learning heterogeneous information network

    CN114219075A

  • Unmanned aerial vehicle control method and system, electronic equipment and storage medium

    CN119225412A

  • Cross-scene path planning optimization method and device based on domain knowledge enhancement

    CN120181187A

  • Transfer learning model training

    WO2025103130A1