Multi-agent virtual-real migration method and device

CN122472083BActive Publication Date: 2026-10-09INST OF AUTOMATION CHINESE ACAD OF SCI +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610956811.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-10-09
Estimated Expiration
2046-06-30

AI Technical Summary

Technical Problem

[0004]本发明提供一种多智能体虚实迁移方法及装置,用以解决现有技术中多智能体策略迁移时存在的稳定性与泛化能力较差的技术问题

Benefits of technology

[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multi-agent virtual-real migration method as described above.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122472083B_ABST
    Figure CN122472083B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of artificial intelligence, and provides a multi-agent virtual-real migration method and device, which comprises the following steps: inputting local observation information obtained by each agent, including itself and neighbor agents in its neighborhood, into an equivariant map policy model to obtain an action decided by the equivariant map policy model based on the local observation information, wherein the local observation information comprises spatial features that vary with task scene geometry and mode features that do not vary with task scene geometry; training the equivariant map policy model based on the reward, state transition information, spatial features, mode features and action of the agent in a simulation environment; and migrating the trained equivariant map policy model to each agent in a real environment. The present application significantly improves the stability and generalization capability of a multi-agent system in the virtual-real migration process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method and apparatus for multi-agent virtual-real migration. Background Technology

[0002] Multi-agent systems, due to their strong parallel collaboration capabilities, high system robustness, and good scalability, have been widely applied in fields such as UAV swarm collaboration, autonomous vehicle formation control, distributed robotic operations, and intelligent manufacturing. Multi-agent reinforcement learning methods, capable of automatically learning collaborative behavior through interactive data, have become an important direction in current multi-agent decision-making research. In practical applications, due to factors such as high training costs and significant security risks in real-world environments, multi-agent strategies are typically trained in simulation environments before being deployed to real-world environments—a technique known as virtual-to-real transfer learning. However, simulation and real-world environments generally differ in physical dynamics, perceptual noise, and interaction structures, leading to a significant performance degradation of strategies that perform well in simulation environments in real-world environments, severely hindering the practical application of multi-agent systems. Therefore, improving the stability and generalization ability of multi-agent strategies during migration from simulation to real-world environments has become a critical technical problem that urgently needs to be solved in this field.

[0003] Currently, existing methods generally focus on "reducing the difference" and regard the virtual-to-real migration problem as a passive compensation process for environmental differences. Although they can reduce the difference between the simulation environment and the real environment to a certain extent, they are still difficult to fully cover the complex uncertainties in the real environment. When migrating multi-agent policies, they still have problems with poor stability and generalization ability. Summary of the Invention

[0004] This invention provides a method and apparatus for multi-agent virtual-real migration, which solves the technical problems of poor stability and generalization ability in the multi-agent policy migration of the prior art.

[0005] This invention provides a multi-agent virtual-real migration method, comprising the following steps: The local observation information acquired by each agent, including its own and its neighboring agents, is input into the isomorphic graph strategy model to obtain the action decided by the isomorphic graph strategy model based on the local observation information. The local observation information includes: spatial features that change with the geometric transformation of the task scene and pattern features that do not change with the geometric transformation of the task scene. In a simulation environment, the isomorphic graph policy model is trained based on the agent's reward, state transition information, spatial features, pattern features, and actions. The trained isomorphic graph policy model is then transferred to each agent in a real-world environment.

[0006] According to a multi-agent virtual-real migration method provided by the present invention, local observation information acquired by each agent, including its own and its neighboring agents, is input into an isomorphic graph policy model to obtain the actions decided by the isomorphic graph policy model based on the local observation information, including: For the pattern features The spatial features Neighborhood feature aggregation is performed separately to obtain the initial aggregated pattern features of each agent. and initial aggregation space features ; For the initial aggregation mode features and initial aggregation space features Joint iterative updates are performed until the maximum number of iterations is reached, in order to obtain equivariant spatial features that characterize the cooperative relationships and spatial structure information among the agents. ; Based on the aforementioned isovariant space characteristics The action is determined by a learnable linear mapping weight matrix, which is obtained during the training of the isovariant graph policy model.

[0007] According to the multi-agent virtual-real migration method provided by the present invention, the pattern features are... The spatial features Neighborhood feature aggregation is performed separately to obtain the initial aggregated pattern features of each agent. and initial aggregation space features ,include: For pattern features , , Represent any neighboring intelligent agent Compared to intelligent agents The pattern feature vector, Represents intelligent agents The set of neighboring intelligent agents is used to aggregate neighborhood features of the pattern features using the following formula: ; in, This is the first multilayer perceptron, used to generate corresponding weights based on the pattern feature vector; For spatial features , , Represent any neighboring intelligent agent Compared to intelligent agents The spatial feature vector is aggregated into neighborhood features using the following formula: ; in, This is a second multilayer perceptron used to modulate the weight matrix of the spatial feature vector. This is an intermediate representation that is invariant to geometric transformations of the task scene.

[0008] According to a multi-agent virtual-real migration method provided by the present invention, the initial aggregation pattern features are... and initial aggregation space features Joint iterative updates are performed until the maximum number of iterations is reached, in order to obtain equivariant spatial features that characterize the cooperative relationships and spatial structure information among the agents. ,include: For the Layer iteration, Less than the maximum number of iterations, based on any two adjacent agents With intelligent agents Each Layer pattern characteristics and Two-agent system constructed from layer spatial features With intelligent agents Interaction messages between ; Regarding the interactive message Aggregate the messages and combine them into a single message. and The layer pattern features are input into the third multilayer perceptron, and the output of the third multilayer perceptron is obtained. Layer pattern features ; Based on the interaction message , Layer pattern features , Layer space features and Neighbor Intelligent Agent of Layer space features Updated using the following formula Layer space features : ; in, This is the fourth multilayer perceptron, used to encode the pattern features of the agent. This is a fifth multilayer perceptron, used to encode the interaction messages. It is a constant.

[0009] According to the multi-agent virtual-real migration method provided by the present invention, for the first Layer iteration, Less than the maximum number of iterations, based on any two adjacent agents With intelligent agents Each Layer pattern characteristics and Two-agent system constructed from layer spatial features With intelligent agents Interaction messages between ,include: intelligent agents and intelligent agents Each Layer pattern characteristics, and intelligent agents and intelligent agents Both The Euclidean distance of the layer space features is input into the sixth multilayer perceptron to obtain the interaction message output by the sixth multilayer perceptron. .

[0010] According to the present invention, a multi-agent virtual-real migration method is provided, based on the aforementioned isotropic spatial features. Determining the action using a learnable linear mapping weight matrix includes: applying the learnable linear mapping weight matrix to equivariant space features. The action is obtained by weighting and fusing the components according to the following formula: ; in, Indicates the first The actions of an intelligent agent This represents the learnable linear mapping weight matrix.

[0011] According to the multi-agent virtual-real migration method provided by the present invention, when the multi-agent system is applied to a two-dimensional scene, any neighboring agent... Compared to intelligent agents The pattern feature vector and spatial feature vector are defined as follows: ; in, Represents intelligent agents With intelligent agents The relative distance between them This represents the relative orientation angle between two agents. and Representing intelligent agents respectively Position vector and velocity vector and Representing intelligent agents respectively The position vector and velocity vector.

[0012] The present invention also provides a multi-agent virtual-real migration device, comprising the following modules: The model decision module is used to input the local observation information acquired by each agent, including its own and its neighboring agents, into the isomorphic graph strategy model to obtain the action decided by the isomorphic graph strategy model based on the local observation information. The local observation information includes: spatial features that change with the geometric transformation of the task scene and pattern features that do not change with the geometric transformation of the task scene. The model training module is used to train the isomorphic graph policy model in a simulation environment based on the agent's reward, state transition information, spatial features, pattern features, and actions. The model transfer module is used to transfer the trained isomorphic graph policy model to each agent in the real environment.

[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the program to implement the multi-agent virtual-real migration method as described above.

[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multi-agent virtual-real migration method as described above.

[0015] The multi-agent virtual-to-real migration method and apparatus provided by this invention utilizes local observation information, which includes spatial features that change with the geometric transformation of the task scene and pattern features that do not. The equivariant graph policy model outputs actions based on this local observation information and is trained by combining the agent's reward, state transition information, the spatial features, the pattern features, and the actions. This explicitly encodes the geometric equivariance of the multi-agent system within the equivariant graph policy model structure, allowing the policy learning process to focus on the shared structural relationships between the simulation and real environments, thereby reducing dependence on simulation-specific parameters. Therefore, when the equivariant graph policy model migrates from the simulation environment to the real environment, even if the real system differs in dynamic characteristics, perceptual noise, or execution errors, the equivariant graph policy model can maintain stable and consistent decision-making behavior. This effectively weakens or even eliminates the impact of differences between the simulation and real environments on the decision-making process, significantly improving the stability and generalization ability of the multi-agent system during virtual-to-real migration, and enabling reliable deployment and operation of the multi-agent system in a real environment. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0017] Figure 1 This is a schematic diagram of the isovariance of a multi-agent system.

[0018] Figure 2 This is a flowchart illustrating the multi-agent virtual-real migration method provided by the present invention.

[0019] Figure 3 This is a schematic diagram illustrating the simulation training and real deployment of the variable graph strategy model in the multi-agent virtual-real migration method provided by the present invention.

[0020] Figure 4 This is a schematic diagram of the decision-making process of the variable graph strategy model in the multi-agent virtual-real migration method provided by the present invention.

[0021] Figure 5 This is a schematic diagram of the structure of the multi-agent virtual-real migration device provided by the present invention.

[0022] Figure 6 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0024] Currently, the virtual-to-real migration of multi-agent systems in existing technologies is mainly limited by the data distribution differences between the simulation and real environments. Existing methods mitigate this through data augmentation, domain randomization, or the construction of high-precision simulation environments, but the overall effectiveness remains limited. Specifically, existing technologies suffer from the following shortcomings: First, data augmentation-based methods expand the data distribution range by introducing simulated noise or perturbations into virtual observations. These methods typically require a significant increase in the size of training data and computational overhead, making them difficult to apply efficiently in large-scale multi-agent scenarios.

[0025] Secondly, domain randomization-based methods improve the adaptability of policies by randomly perturbing physical parameters or environmental configurations during the training phase. However, excessive randomization often leads to instability in the policy learning process, or even convergence to a suboptimal solution, affecting the final transfer performance.

[0026] Third, while building high-precision simulators to approximate the real environment can reduce the difference between the simulated and real environments to some extent, the simulators are costly to build and maintain, and it is difficult to fully cover the complex and uncertain factors in the real environment, thus failing to fundamentally solve the problem of virtual-real migration.

[0027] It is evident that most existing methods start from the perspective of "reducing differences" and regard the virtual-to-real migration problem as a passive compensation process for environmental differences. They fail to delve into the inherent structural features shared between the simulation environment and the real environment, lack system modeling of the common laws of multi-agent systems, and still have problems with poor stability and generalization ability when migrating multi-agent policies.

[0028] However, despite deviations in state transition functions, observation noise distribution, and actuator accuracy between simulated and real environments, these deviations often do not disrupt the inherent structural equivariance of multi-agent systems. By explicitly introducing equivariance constraints during the policy learning phase, policies can remain invariant or approximately invariant to such unstructured disturbances, thereby achieving a natural cancellation of virtual-real differences and improving the stability and transfer reliability of policies in real environments. Equivariance is a ubiquitous and stable underlying structural characteristic in multi-agent systems, reflecting the common decision-making rules followed by the system under different environmental conditions. Specifically, in Figure 1 In the cooperative navigation task shown in (a), when the environmental state undergoes a rotational transformation, the agent's optimal policy outputs a corresponding rotational action, exhibiting the equivariant characteristics of the optimal policy. Similarly, Figure 1 (b) In the rotating intersection scenario shown, the traffic light phase control strategy also adjusts in tandem with changes in road direction. This equivariance of agent behavior strategy is a common objective law in real-world group tasks such as drone formation control and autonomous driving.

[0029] To address the aforementioned technical problems in existing technologies, based on the equivariance law of agent behavior strategies, embodiments of the present invention provide a multi-agent virtual-real migration method, such as... Figure 2 As shown, it includes the following steps S210 to S230.

[0030] Step S210: Input the local observation information acquired by each agent, including its own and its neighboring agents, into the isomorphic graph policy model to obtain the action decided by the isomorphic graph policy model based on the local observation information. The local observation information includes: spatial features that change with the geometric transformation of the task scene and pattern features that do not change with the geometric transformation of the task scene; that is, spatial features are quantities that change synchronously with the geometric transformation of the task scene, and pattern features are quantities that do not change synchronously with the geometric transformation of the task scene.

[0031] The geometric transformations of the task scene include rotation transformation and flip transformation. The spatial features include the position vectors and velocity vectors of each agent in the task scene, and may also include the position vectors of each marker point (obstacle or target point to be reached by the agent) in the task scene. The pattern features include the relative distance and relative orientation angle between each agent, and may also include the relative distance and relative orientation angle between each agent and each marker point.

[0032] For example: Figure 1 In (a), a rotation transformation is performed on the entire cooperative navigation task scenario. After the rotation transformation, the position vectors and velocity vectors of each agent (e.g., robot or vehicle) change, that is, the spatial characteristics change, while the relative distance and relative orientation angle between each agent do not change, that is, the pattern characteristics do not change.

[0033] It is understandable that a multi-agent system consists of multiple agents, each making collaborative decisions based on local observation information. In order to achieve collaborative control of the multi-agent system, in this step, the isomorphic graph strategy model is used as the unified decision core for each agent. That is, the isomorphic graph strategy model can generate and output action decisions based on the local observation information of each agent.

[0034] In decentralized multi-agent decision-making scenarios, each agent can only obtain local observation information about itself and its neighboring agents within its neighborhood. For any given agent... The set of its neighboring intelligent agents is denoted as intelligent agent With any neighboring intelligent agent The information exchanged includes: spatial features Pattern features Two categories. Among them, spatial features Model features are used to characterize observations affected by geometric transformations, i.e., observations that change with geometric transformations. Used to characterize observations that are invariant to geometric transformations. Based on the above classification, an intelligent agent is defined. Pattern features received in its neighborhood Spatial features They are respectively: (1).

[0035] Therefore, intelligent agents Current local observation information can be uniformly represented as: What can be understood is: pattern features and spatial features Both are multidimensional feature matrices.

[0036] For example, the method is illustrated using a multi-agent cooperative motion task in a two-dimensional space (such as ground mobile robot cooperation, multi-vehicle cooperation in autonomous driving transportation systems, and regional coverage and search tasks). However, the method is not limited to two-dimensional scenarios and can be extended to higher-dimensional spaces, such as a drone swarm formation control task in a three-dimensional scenario, where the positions of each agent (drone) and each marker point are three-dimensional coordinates.

[0037] In the context of a multi-agent system applied to a two-dimensional scene, any neighboring agent... Compared to intelligent agents The pattern feature vector and spatial feature vector are defined as follows: (2).

[0038] in, For pattern feature vectors, Represents intelligent agents With intelligent agents The relative distance between them Represents intelligent agents With intelligent agents The relative orientation angle between them, i.e., the intelligent agent With intelligent agents The relative distance and relative orientation angle between the two do not change with geometric transformation. For spatial feature vectors, and Representing intelligent agents respectively Position vector and velocity vector and Representing intelligent agents respectively The position vector and velocity vector, i.e., the agent's position vector and velocity vector. With intelligent agents Their respective positions and velocities will change with geometric transformations.

[0039] Step S220: In a simulation environment, train the isomorphic graph policy model based on the agent's reward, state transition information, spatial features, pattern features, and actions.

[0040] like Figure 3As shown, the isomorphic graph policy model interacts with the simulation environment as the decision-making module of the agent. Specifically, at each moment, each agent first obtains interaction information from the simulation environment, including its own state and local observation information of neighboring agents, and inputs this local observation information into the isomorphic graph policy model for forward computation (as in step S210). The isomorphic graph policy model outputs the action of each agent based on the input local observation information and neighborhood structure relationship. Subsequently, each agent executes the corresponding behavior in the simulation environment according to the generated action. The simulation environment updates the system state according to the agent's action and returns new local observation information and reward signal.

[0041] like Figure 3 As shown, during the above interaction process, the local observation information (spatial features and pattern features), actions, rewards, and state transition information generated by the simulation environment and the multi-agent system are recorded as training data and stored in the data storage pool. The multi-agent reinforcement learning algorithm samples the interaction data from the data storage pool and updates the parameters of the equivariant graph policy model according to the training objective, thereby completing the training of the equivariant graph policy model. During training, a multi-agent reinforcement learning algorithm based on policy gradients can be used, with the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm being preferred for training the equivariant graph policy model.

[0042] Step S230: Transfer the trained isomorphic graph policy model to each agent in the real environment. For example... Figure 3 As shown, during the execution phase in the real environment, the isovariant graph policy model outputs upper-level action commands (i.e., actions) based on the local observation information obtained by each agent in the real environment. These action commands are then input to the control unit as decision results. The control unit, according to the specific hardware structure and execution constraints of the system, maps the upper-level action commands to corresponding lower-level driving signals and sends them down to the execution unit to control each agent to perform actions. Through this execution chain of "policy decision—control mapping—lower-level driving," the isovariant cooperative policy learned in the simulation environment can be accurately executed in the real hardware system.

[0043] In the multi-agent virtual-to-real migration method of this embodiment, since the local observation information includes spatial features that change with the geometric transformation of the task scene and pattern features that do not change with the geometric transformation of the task scene, the equivariant graph policy model outputs actions based on this local observation information. It is then trained by combining the agent's reward, state transition information, the spatial features, the pattern features, and the actions. This explicitly encodes the geometric equivariance of the multi-agent system within the equivariant graph policy model structure, allowing the policy learning process to focus on the shared structural relationships between the simulation and real environments, thereby reducing dependence on simulation-specific parameters. Therefore, when the equivariant graph policy model migrates from the simulation environment to the real environment, even if the real system differs in dynamic characteristics, perceptual noise, or execution errors, the equivariant graph policy model can still maintain stable and consistent decision-making behavior. This effectively weakens or even eliminates the impact of differences between the simulation and real environments on the decision-making process, significantly improving the stability and generalization ability of the multi-agent system during virtual-to-real migration, and enabling reliable deployment and operation of the multi-agent system in a real environment.

[0044] In some embodiments, such as Figure 4 As shown, in step S210, the local observation information acquired by each agent, including its own and its neighboring agents, is input into the isomorphic graph policy model to obtain the action decided by the isomorphic graph policy model based on the local observation information, specifically including: Step S211: For the pattern features The spatial features Neighborhood feature aggregation is performed separately to obtain the initial aggregated pattern features of each agent. and initial aggregation space features .

[0045] Specifically, the isomorphic graph strategy model in the input pattern features and spatial features Then, the features of agents within neighborhoods of different sizes are further aggregated in a unified manner. Let the agents be... The pattern features and spatial features are respectively and The neighborhood feature aggregation algorithm is denoted as Its output initial aggregation features Represented as: (3).

[0046] in, This represents the adjacency relationship within the neighborhood, i.e., the edge connection relationship in the multi-agent graph structure. The superscript (0) indicates that this aggregated feature is used as the initial input for the subsequent feature update process.

[0047] Step S212: For the initial aggregation mode features and initial aggregation space features Joint iterative updates are performed until the maximum number of iterations is reached, in order to obtain equivariant spatial features that characterize the cooperative relationships and spatial structure information among the agents. .

[0048] After completing the neighborhood feature aggregation, this step involves processing the initial aggregated pattern features. and initial aggregation space features The update process, while maintaining equivariance, jointly updates the pattern features and spatial features of multiple agents, thereby extracting the collaborative relationships and structural information among the agents layer by layer. Feature updating specifically includes two processes: pattern feature updating and spatial feature updating, which are executed collaboratively in each iteration layer.

[0049] Let the first During layer iteration, the agent The pattern features and spatial features are respectively and The adjacency relationship between intelligent agents is Pattern feature update algorithm and spatial feature update algorithm They are represented as follows: (4).

[0050] in, and These represent the nth feature after feature update. Pattern and spatial characteristics during layer iteration.

[0051] It is understandable that the total number of iterations can be preset, for example, 1 to 3 layers. Once the total number of layers is reached, iteration stops, and at this point, the isomorphic pattern features are output. and equivariant space features .

[0052] Step S213: Based on the aforementioned equivariant spatial features The action is determined by a learnable linear mapping weight matrix, which is obtained during the training of the isovariant graph policy model.

[0053] Specifically, after step S212, the intelligent agent is obtained. Final equivariant mode features and Equivariant Space Features For example, equivariant space features Composed of position vectors and velocity vectors, etc., when a multi-agent system is in a two-dimensional space, the equivariant spatial characteristics... The dimensions are: , Represents the real number field. This embodiment introduces a learnable linear mapping weight matrix. Through a learnable linear mapping weight matrix Equivalent spatial features Each component (i.e.) The action is obtained by weighting and fusing the elements in the formula below: (5).

[0054] in, Indicates the first The actions of each agent, a learnable linear mapping weight matrix. The learned linear mapping weight matrix is ​​acquired through centralized training and remains fixed during the execution phase. Acting on the characteristics of isovariant space Above, the output action When the system undergoes geometric transformations such as rotation or translation, a corresponding uniform transformation will be generated, thereby ensuring the equivariance of the action output.

[0055] In some embodiments, step S211 involves processing the pattern features. The spatial features Neighborhood feature aggregation is performed separately to obtain the initial aggregated pattern features of each agent. and initial aggregation space features Specifically, it includes: For pattern features , , Represent any neighboring intelligent agent Compared to intelligent agents The pattern feature vector, Represents intelligent agents The set of neighboring intelligent agents is used to aggregate neighborhood features of the pattern features using the following formula: (6).

[0056] in, This is the first multilayer perceptron, used to generate corresponding weights based on the pattern feature vector. In this step, for the agent... By analyzing the pattern features of each neighborhood The initial aggregation pattern features are obtained by weighted summation with the corresponding weights. .

[0057] For spatial features , , Represent any neighboring intelligent agent Compared to intelligent agents The spatial eigenvectors are first used to construct an intermediate representation invariant to geometric transformations based on the inner product form of the spatial eigenvectors. The intermediate representation is input into the second multilayer perceptron to obtain the spatial feature vector output by the second multilayer perceptron. The weights for the agent By analyzing the spatial feature vectors of each neighborhood The initial aggregate space features are obtained by weighting and summing the features with their corresponding weights. The formula for aggregating neighborhood features of spatial features is expressed as follows: (7).

[0058] in, This is a second multilayer perceptron used to modulate the weight matrix of the spatial feature vector. , Representing spatial dimension, for two-dimensional space, Three-dimensional space , This represents the total number of feature vectors in the spatial features. For quantities that do not change with geometric transformations such as rotation and translation, the second multilayer perceptron... A weight matrix is ​​generated to modulate the spatial feature vector, thereby ensuring that the spatial features remain equivariant to geometric transformations during the neighborhood feature aggregation process.

[0059] In practical applications, spatial features can simultaneously include multiple feature vectors such as position and velocity, parameters This represents the total number of feature vectors in a spatial feature. Therefore, the neighborhood aggregation method can be naturally extended to task scenarios that include spatial features composed of multiple feature vectors.

[0060] It is understandable that the first and second multilayer perceptrons can adopt existing multilayer perceptron structures, and their respective parameters are updated during the training of the equivariant graph policy model.

[0061] In some embodiments, step S212 involves processing the initial aggregation mode features. and initial aggregation space features Joint iterative updates are performed until the maximum number of iterations is reached, in order to obtain equivariant spatial features that characterize the cooperative relationships and spatial structure information among the agents. Specifically, it includes: For the Layer iteration, Less than the maximum number of iterations, based on any two adjacent agents With intelligent agents Each Layer pattern characteristics and Two-agent system constructed from layer spatial features With intelligent agents Interaction messages between .

[0062] Regarding the interactive message Aggregate the messages and combine them into a single message. and The layer pattern features are input into the third multilayer perceptron, and the output of the third multilayer perceptron is obtained. Layer pattern features Third multilayer perceptron The function can be represented as follows: (8).

[0063] This represents the aggregated interactive message.

[0064] Based on the interaction message , Layer pattern features , Layer space features and Neighbor Intelligent Agent of Layer space features Updated using the following formula Layer space features : (9).

[0065] in, This is the fourth multilayer perceptron, used to encode the pattern features of the agent. This is a fifth multilayer perceptron, used to encode the interaction messages. is a constant used to normalize the neighborhood size. Through the above formula (9), the spatial features maintain a consistent isotropic response under geometric transformations such as rotation and translation.

[0066] Specifically, for the first Layer iteration, Less than the maximum number of iterations, based on any two adjacent agents With intelligent agents Each Layer pattern characteristics and Two-agent system constructed from layer spatial features With intelligent agents Interaction messages between Specifically, it includes: intelligent agents and intelligent agents Each Layer pattern characteristics, and intelligent agents and intelligent agents Both The Euclidean distance of the layer space features is input into the sixth multilayer perceptron to obtain the interaction message output by the sixth multilayer perceptron. The sixth multilayer perceptron function is represented as follows: (10).

[0067] The sixth multilayer perceptron is used to extract interaction messages between multiple agents. and Representing intelligent agents respectively and intelligent agents Each Layer pattern features, Represents intelligent agents and intelligent agents Both Euclidean distance of layer space features.

[0068] It is understandable that the third, fourth, fifth, and sixth multilayer perceptrons can adopt existing multilayer perceptron structures, and their respective parameters are updated during the training of the variable graph strategy model.

[0069] In this embodiment, by jointly updating pattern features and spatial features, it is possible to gradually extract high-level collaborative relationships and spatial structure information among multiple agents during multi-layer iteration, and use the updated features as input for feature updates in the next layer of iteration or subsequent action decisions.

[0070] The multi-agent virtual-real migration device provided by the present invention is described below. The multi-agent virtual-real migration device described below can be referred to in correspondence with the multi-agent virtual-real migration method described above.

[0071] The multi-agent virtual-real migration device of this invention, such as Figure 5 As shown, it includes: The model decision module 510 is used to input the local observation information acquired by each agent, including its own and its neighboring agents, into the isomorphic graph strategy model to obtain the action decided by the isomorphic graph strategy model based on the local observation information. The local observation information includes: spatial features that change with the geometric transformation of the task scene and pattern features that do not change with the geometric transformation of the task scene.

[0072] The model training module 520 is used to train the isomorphic graph policy model in a simulation environment based on the agent's reward, state transition information, spatial features, pattern features, and actions.

[0073] The model transfer module 530 is used to transfer the trained isomorphic graph policy model to each agent in the real environment.

[0074] The multi-agent virtual-to-real migration device in this embodiment utilizes local observation information, including spatial features that change with the geometric transformation of the task scene and pattern features that do not. The equivariant graph policy model outputs actions based on this local observation information and is trained by combining the agent's reward, state transition information, the spatial features, the pattern features, and the actions. This explicitly encodes the geometric equivariance of the multi-agent system within the equivariant graph policy model structure, allowing the policy learning process to focus on the shared structural relationships between the simulation and real environments, thus reducing dependence on simulation-specific parameters. Therefore, when the equivariant graph policy model migrates from the simulation environment to the real environment, even if the real system differs in dynamic characteristics, perceptual noise, or execution errors, the equivariant graph policy model can maintain stable and consistent decision-making behavior. This effectively weakens or even eliminates the impact of differences between the simulation and real environments on the decision-making process, significantly improving the stability and generalization ability of the multi-agent system during virtual-to-real migration, and enabling reliable deployment and operation of the multi-agent system in a real environment.

[0075] In some embodiments, the model decision module 510 includes: The neighborhood feature aggregation module is used to aggregate the pattern features. The spatial features Neighborhood feature aggregation is performed separately to obtain the initial aggregated pattern features of each agent. and initial aggregation space features ; The joint iteration module is used to process the initial aggregation pattern features. and initial aggregation space features Joint iterative updates are performed until the maximum number of iterations is reached, in order to obtain equivariant spatial features that characterize the cooperative relationships and spatial structure information among the agents. ; Action determination module, used to determine actions based on the aforementioned isotropic spatial features. The action is determined by a learnable linear mapping weight matrix, which is obtained during the training of the isovariant graph policy model.

[0076] In some embodiments, the neighborhood feature aggregation module includes: The pattern feature aggregation module is used for pattern features. , , Represent any neighboring intelligent agent Compared to intelligent agents The pattern feature vector, Represents intelligent agents The set of neighboring intelligent agents is used to aggregate neighborhood features of the pattern features using the following formula: ; in, This is the first multilayer perceptron, used to generate corresponding weights based on the pattern feature vector; The spatial feature aggregation module is used for spatial features. , , Represent any neighboring intelligent agent Compared to intelligent agents The spatial feature vector is aggregated into neighborhood features using the following formula: ; in, This is a second multilayer perceptron used to modulate the weight matrix of the spatial feature vector. This is an intermediate representation that is invariant to geometric transformations of the task scene.

[0077] In some embodiments, the joint iteration module includes: Interactive message building module, used for the first Layer iteration, Less than the maximum number of iterations, based on any two adjacent agents With intelligent agents Each Layer pattern characteristics and Two-agent system constructed from layer spatial features With intelligent agents Interaction messages between ; The pattern feature update module is used to update the interaction message. Aggregate the messages and combine them into a single message. and The layer pattern features are input into the third multilayer perceptron, and the output of the third multilayer perceptron is obtained. Layer pattern features ; The spatial feature update module is used to update the spatial feature based on the interaction message. , Layer pattern features , Layer space features and Neighbor Intelligent Agent of Layer space features Updated using the following formula Layer space features : ; in, This is the fourth multilayer perceptron, used to encode the pattern features of the agent. This is a fifth multilayer perceptron, used to encode the interaction messages. It is a constant.

[0078] In some embodiments, the interaction message construction module is specifically used to connect the intelligent agent. and intelligent agents Each Layer pattern characteristics, and intelligent agents and intelligent agents Both The Euclidean distance of the layer space features is input into the sixth multilayer perceptron to obtain the interaction message output by the sixth multilayer perceptron. .

[0079] In some embodiments, the action determination module is specifically used to apply the learnable linear mapping weight matrix to the equivariant space features. The action is obtained by weighting and fusing the components according to the following formula: ; in, Indicates the first The actions of an intelligent agent This represents the learnable linear mapping weight matrix.

[0080] In some embodiments, when a multi-agent system is applied to a two-dimensional scene, any neighboring agent Compared to intelligent agents The pattern feature vector and spatial feature vector are defined as follows: ; in, Represents intelligent agents With intelligent agents The relative distance between them This represents the relative orientation angle between two agents. and Representing intelligent agents respectively Position vector and velocity vector and Representing intelligent agents respectively The position vector and velocity vector.

[0081] Figure 6 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 6As shown, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communications interface 620, and the memory 630 communicate with each other via the communication bus 640. The processor 610 can call logical instructions in the memory 630 to execute a multi-agent virtual-real migration method, which includes: The local observation information acquired by each agent, including its own and its neighboring agents, is input into the isomorphic graph strategy model to obtain the action decided by the isomorphic graph strategy model based on the local observation information. The local observation information includes: spatial features that change with the geometric transformation of the task scene and pattern features that do not change with the geometric transformation of the task scene.

[0082] In a simulation environment, the isomorphic graph strategy model is trained based on the agent's reward, state transition information, spatial features, pattern features, and actions.

[0083] The trained isomorphic graph policy model is then transferred to each agent in a real-world environment.

[0084] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0085] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program that can be stored on a non-transitory computer-readable storage medium, wherein when the computer program is executed by a processor, the computer is able to execute the multi-agent virtual-real migration method provided by the above methods, the method comprising: The local observation information acquired by each agent, including its own and its neighboring agents, is input into the isomorphic graph strategy model to obtain the action decided by the isomorphic graph strategy model based on the local observation information. The local observation information includes spatial features that change with the geometric transformation of the task scene and pattern features that do not change with the geometric transformation of the task scene.

[0086] In a simulation environment, the isomorphic graph strategy model is trained based on the agent's reward, state transition information, spatial features, pattern features, and actions.

[0087] The trained isomorphic graph policy model is then transferred to each agent in a real-world environment.

[0088] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the multi-agent virtual-to-real migration method provided by the above methods, the method comprising: The local observation information acquired by each agent, including its own and its neighboring agents, is input into the isomorphic graph strategy model to obtain the action decided by the isomorphic graph strategy model based on the local observation information. The local observation information includes spatial features that change with the geometric transformation of the task scene and pattern features that do not change with the geometric transformation of the task scene.

[0089] In a simulation environment, the isomorphic graph strategy model is trained based on the agent's reward, state transition information, spatial features, pattern features, and actions.

[0090] The trained isomorphic graph policy model is then transferred to each agent in a real-world environment.

[0091] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0092] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0093] All actions involving the acquisition of signal information or data in this invention are carried out in compliance with the relevant data protection laws and policies of the country where the device is located, and with the authorization granted by the owner of the device.

[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-agent virtual-real migration method, characterized in that, include: The local observation information acquired by each agent, including its own and its neighboring agents, is input into the isomorphic graph strategy model to obtain the action decided by the isomorphic graph strategy model based on the local observation information. The local observation information includes: spatial features that change with the geometric transformation of the task scene and pattern features that do not change with the geometric transformation of the task scene. In a simulation environment, the isomorphic graph policy model is trained based on the agent's reward, state transition information, spatial features, pattern features, and actions. The trained isomorphic graph policy model is then transferred to each agent in a real-world environment. Specifically, the local observation information acquired by each agent, including its own and its neighboring agents, is input into the isomorphic graph policy model to obtain the actions decided by the isomorphic graph policy model based on the local observation information, including: For the pattern features The spatial features Neighborhood feature aggregation is performed separately to obtain the initial aggregated pattern features of each agent. and initial aggregation space features ; For the initial aggregation mode features and initial aggregation space features Joint iterative updates are performed until the maximum number of iterations is reached, in order to obtain equivariant spatial features that characterize the cooperative relationships and spatial structure information among the agents. ; Based on the aforementioned isovariant space characteristics The action is determined by a learnable linear mapping weight matrix, which is obtained during the training of the isovariant graph policy model. Among them, the pattern features The spatial features Neighborhood feature aggregation is performed separately to obtain the initial aggregated pattern features of each agent. and initial aggregation space features ,include: For pattern features , , Represent any neighboring intelligent agent Compared to intelligent agents The pattern feature vector, Represents intelligent agents The set of neighboring intelligent agents is used to aggregate neighborhood features of the pattern features using the following formula: ; in, This is the first multilayer perceptron, used to generate corresponding weights based on the pattern feature vector; For spatial features , , Represent any neighboring intelligent agent Compared to intelligent agents The spatial feature vector is aggregated into neighborhood features using the following formula: ; in, This is a second multilayer perceptron used to modulate the weight matrix of the spatial feature vector. This is an intermediate representation that is invariant to geometric transformations of the task scene.

2. The multi-agent virtual-real migration method according to claim 1, characterized in that, For the initial aggregation mode features and initial aggregation space features Joint iterative updates are performed until the maximum number of iterations is reached, in order to obtain equivariant spatial features that characterize the cooperative relationships and spatial structure information among the agents. ,include: For the Layer iteration, Less than the maximum number of iterations, based on any two adjacent agents With intelligent agents Each Layer pattern characteristics and Two-agent system constructed from layer spatial features With intelligent agents Interaction messages between ; Regarding the interactive message Aggregate the messages and combine them into a single message. and The layer pattern features are input into the third multilayer perceptron, and the output of the third multilayer perceptron is obtained. Layer pattern features ; Based on the interaction message , Layer pattern features , Layer space features and Neighbor Intelligent Agent of Layer space features Updated using the following formula Layer space features : ; in, This is the fourth multilayer perceptron, used to encode the pattern features of the agent. This is a fifth multilayer perceptron, used to encode the interaction messages. It is a constant.

3. The multi-agent virtual-real migration method according to claim 2, characterized in that, For the Layer iteration, Less than the maximum number of iterations, based on any two adjacent agents With intelligent agents Each Layer pattern characteristics and Two-agent system constructed from layer spatial features With intelligent agents Interaction messages between ,include: intelligent agents and intelligent agents Each Layer pattern characteristics, and intelligent agents and intelligent agents Both The Euclidean distance of the layer space features is input into the sixth multilayer perceptron to obtain the interaction message output by the sixth multilayer perceptron. .

4. The multi-agent virtual-real migration method according to claim 1, characterized in that, Based on the aforementioned isovariant space characteristics Determining the action using a learnable linear mapping weight matrix includes: applying the learnable linear mapping weight matrix to equivariant space features. The action is obtained by weighting and fusing the components according to the following formula: ; in, Indicates the first The actions of an intelligent agent This represents the learnable linear mapping weight matrix.

5. The multi-agent virtual-real migration method according to any one of claims 1 to 4, characterized in that, In the context of a multi-agent system applied to a two-dimensional scene, any neighboring agent... Compared to intelligent agents The pattern feature vector and spatial feature vector are defined as follows: ; in, Represents intelligent agents With intelligent agents The relative distance between them This represents the relative orientation angle between two agents. and Representing intelligent agents respectively Position vector and velocity vector and Representing intelligent agents respectively The position vector and velocity vector.

6. A multi-agent virtual-real migration device, characterized in that, include: The model decision module is used to input the local observation information acquired by each agent, including its own and its neighboring agents, into the isomorphic graph strategy model to obtain the action decided by the isomorphic graph strategy model based on the local observation information. The local observation information includes: spatial features that change with the geometric transformation of the task scene and pattern features that do not change with the geometric transformation of the task scene. The model training module is used to train the isomorphic graph policy model in a simulation environment based on the agent's reward, state transition information, spatial features, pattern features, and actions. The model transfer module is used to transfer the trained isomorphic graph policy model to each agent in the real environment; The model decision module includes: The neighborhood feature aggregation module is used to aggregate the pattern features. The spatial features Neighborhood feature aggregation is performed separately to obtain the initial aggregated pattern features of each agent. and initial aggregation space features ; The joint iteration module is used to process the initial aggregation pattern features. and initial aggregation space features Joint iterative updates are performed until the maximum number of iterations is reached, in order to obtain equivariant spatial features that characterize the cooperative relationships and spatial structure information among the agents. ; Action determination module, used to determine actions based on the aforementioned isotropic spatial features. The action is determined by a learnable linear mapping weight matrix, which is obtained during the training of the equivariant graph policy model. The neighborhood feature aggregation module includes: The pattern feature aggregation module is used for pattern features. , , Represent any neighboring intelligent agent Compared to intelligent agents The pattern feature vector, Represents intelligent agents The set of neighboring intelligent agents is used to aggregate neighborhood features of the pattern features using the following formula: ; in, This is the first multilayer perceptron, used to generate corresponding weights based on the pattern feature vector; The spatial feature aggregation module is used for spatial features. , , Represent any neighboring intelligent agent Compared to intelligent agents The spatial feature vector is aggregated into neighborhood features using the following formula: ; in, This is a second multilayer perceptron used to modulate the weight matrix of the spatial feature vector. This is an intermediate representation that is invariant to geometric transformations of the task scene.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the multi-agent virtual-real migration method as described in any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the multi-agent virtual-real migration method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Multi-agent reinforcement learning method based on isotropic graph neural network

    CN118428444A

  • Multi-agent geometric graph reinforcement learning method and device and storage medium

    CN118674002A