Hierarchical multi-agent game confrontation and collaborative decision-making method based on federated learning

By adopting a layered multi-agent game confrontation and collaborative decision-making method based on federated learning in multi-agent systems, challenges in privacy protection, strategy generation and collaborative decision-making are solved, efficient privacy protection and collaborative decision-making are achieved, and the scalability and adaptability of the system are improved.

CN119443312BActive Publication Date: 2025-05-02NANJING TONGFANG BEIDOU TECH CO LTD +1

Patent Information

Application Number
CN202510027115.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-02
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

The existing multi-agent system has many challenges in privacy protection, strategy generation, game confrontation and collaborative decision-making, including privacy information leakage, partial observability leading to strategy generation difficulties, restricted policy optimization, and conflict and deadlock problems in complex joint designs.

Method used

The hierarchical multi-agent game adversity and collaborative decision-making method based on federated learning is adopted. By establishing a federated learning framework, the agent conducts model training locally, and performs game adversity and collaborative decision-making through a hierarchical reinforcement learning model. The monitoring data is used to introduce collaborative decision-making algorithms to avoid conflicts and deadlocks, and global model optimization is carried out when preset conditions are met.

Benefits of technology

It realizes privacy protection, efficient coordination and game confrontation between agents, reduces the risk of data leakage, improves the efficiency of strategy generation, and enhances the scalability and adaptability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119443312B_ABST
    Figure CN119443312B_ABST
Patent Text Reader

Abstract

The present invention discloses a hierarchical multi-agent game confrontation and collaborative decision-making method based on federated learning, including: by setting up a federated learning framework containing multiple agents, the agents are located in different areas, and the data of the area is used to train the local model; within the federated learning framework, a hierarchical reinforcement learning model is deployed inside each agent; using the hierarchical reinforcement learning model, multiple agents are trained in a game environment; on the basis of game confrontation, the agents are monitored, and a collaborative decision-making algorithm is introduced based on the monitoring data to coordinate the behavior of multiple agents; when the preset training rounds are reached or specific conditions are met, the local model parameters are uploaded to the central server for global optimization. The method of the present invention can ensure that efficient collaboration and strategy optimization between multiple agents are achieved under the premise of protecting privacy, and maintain a high degree of adaptability and robustness in a complex and changing environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence and machine learning technology, and in particular to a hierarchical multi-agent game confrontation and collaborative decision-making method based on federated learning. Background Art

[0002] With the widespread application of multi-agent systems in various fields, the issue of privacy protection has become increasingly prominent. In the process of multi-agent reinforcement learning, agents need to constantly interact with the environment, collect data and update strategies. However, these data often contain the private information of the agents, such as location, action, reward, etc. If these data are leaked to unauthorized third parties, it will pose a serious threat to the security and interests of the agents.

[0003] In a multi-agent system, agents usually only have access to partial observable information, that is, they can only observe part of the environment or are restricted from obtaining complete information. This partial observability makes it difficult for agents to generate an overall optimal strategy because they cannot accurately understand the states and actions of other agents. To solve this problem, researchers have proposed various methods, such as model-based methods, memory-based methods, etc. However, these methods still have limitations in terms of computational complexity and applicability.

[0004] In multi-agent game confrontation, agents need to constantly adjust their strategies according to the strategies and actions of their opponents in order to maximize their own interests. However, since the strategies of agents influence each other and the strategies of opponents are often unknown or dynamically changing, strategy optimization becomes very difficult. To solve this problem, researchers have proposed various game confrontation algorithms, such as Nash equilibrium-based methods and policy gradient-based methods. However, these methods still have limitations when dealing with complex environments and changing opponents.

[0005] In the complex joint design of multi-agent systems, there is often a relationship of mutual dependence and constraint between agents. This relationship makes it easy for agents to have individual conflicts and local deadlock problems in the collaborative decision-making process. Individual conflicts refer to the contradictions or confrontations between agents due to competition for resources or goals; local deadlock refers to the deadlock of agents because they cannot find a feasible solution. These problems not only affect the stability and convergence efficiency of the system, but also lead to the failure of collaboration between agents.

[0006] In summary, existing multi-agent systems still face many challenges in collaborative decision-making and game confrontation. In order to solve these problems, this paper proposes a hierarchical multi-agent game confrontation and collaborative decision-making method based on federated learning, aiming to achieve privacy protection, efficient collaboration and game confrontation among agents. Summary of the invention

[0007] The purpose of the present invention is to propose a hierarchical multi-agent game confrontation and collaborative decision-making method based on federated learning in order to solve the problems of insufficient privacy protection in the prior art, difficulty in strategy generation under partial observable information conditions, limited strategy optimization in game confrontation, and conflict and deadlock in complex joint design.

[0008] In order to achieve the above object, the present invention adopts the following technical solution: a hierarchical multi-agent game confrontation and collaborative decision-making method based on federated learning, comprising the following steps:

[0009] Step S1, establish a federated learning framework including multiple agents, where the agents are located in different regions and use the data in their respective regions to train local models;

[0010] Step S2: deploy a hierarchical reinforcement learning model inside each agent within the federated learning framework, including a top-level policy control model and an individual policy execution model;

[0011] Step S3, using a hierarchical reinforcement learning model to enable multiple agents to conduct adversarial training in a game environment;

[0012] Step S4, based on the game confrontation, the intelligent agent is monitored, and a collaborative decision-making algorithm is introduced based on the monitoring data to coordinate the behaviors of multiple intelligent agents to avoid conflicts and deadlocks, wherein the monitoring data includes the position, speed, decision-making choice, and interaction history of the intelligent agent;

[0013] Step S5, when the preset training rounds are reached or specific conditions are met, the local model parameters are uploaded to the central server, and the model parameters are shared through the federated learning framework for global optimization, where the preset training rounds refer to the rounds in which the agent uses the data in the area to train the local model, and the local model is the local model of the agent in different areas;

[0014] Wherein step S4 also includes the following sub-steps:

[0015] S4-1, during the game confrontation, monitor and record the behavior, decision-making and interactions of each intelligent agent;

[0016] S4-2, based on monitoring data, starts the collaborative decision-making algorithm and uses the top-level policy control layer of the hierarchical reinforcement learning model to coordinate the behaviors of multiple agents. The specific formula is:

[0017] ;

[0018] in, represents the strategy of agent i, represents the reward function of agent i, represents the global state at time t, represents the action of agent i at time t, represents the cost caused by the action conflict between agents i and j, is the weight coefficient for adjusting conflict costs;

[0019] S4-3, according to the output of the collaborative decision-making algorithm, adjust the strategy selection of each intelligent agent, eliminate conflicts and deadlocks, and optimize the overall performance.

[0020] Furthermore, in step S1, the following sub-steps are also included:

[0021] S1-1, determine the global model update frequency of federated learning, the threshold for agents to participate in federated averaging, and the communication protocol;

[0022] S1-2, establishing multiple intelligent agents, assigning a unique identifier to each intelligent agent, and initializing its local data set, wherein the local data set includes historical data required for the intelligent agent to make decisions in a specific environment;

[0023] S1-3, each agent uses a local data set to train a reinforcement learning model, which is used to generate an action strategy based on the current observation information. The specific formula is:

[0024] ;

[0025] Among them, s is the current state, a is the action taken, and r is the reward. is the discount factor, is the learning rate, is the next state, is the next action to take.

[0026] Furthermore, in step S2, the following sub-steps are also included:

[0027] S2-1, configure a top-level strategy control model and an individual strategy execution model for each agent, and set initial parameters. The top-level strategy control model is responsible for the formulation and adjustment of the global strategy, and the individual strategy execution model is responsible for executing specific actions in the local environment. The initial parameters of the model include the weights and biases of the neural network;

[0028] S2-2, within each agent, the individual strategy execution model provides its current state, observation information, and historical behavior data to the top-level strategy control model;

[0029] S2-3, based on the information provided by the individual policy execution model, the top-level policy control model performs global policy analysis, generates corresponding control information, and passes it to the individual policy execution model. The top-level policy control model generates control information through the following formula:

[0030] ;

[0031] Among them, G represents global information, C represents control information, The mapping function representing the top-level policy control model;

[0032] S2-4, after receiving the control information, the individual agent performs a comprehensive analysis based on its own partial observable information and makes a decision using the individual strategy execution model. The individual strategy execution model makes a decision using the following formula:

[0033] ;

[0034] in, represents the action value of the ith agent, represents the control information of the ith agent, represents the partial observable information of the ith agent, A mapping function representing an individual policy execution model.

[0035] Furthermore, in step S3, the following sub-steps are also included:

[0036] S3-1, build a game environment, set game rules and reward mechanism, the game rules include action sequence, observation and information transmission, legal action set, state transition rules, game termination conditions, the reward mechanism includes instant reward, collaborative reward, global value maximization, and punishment mechanism;

[0037] S3-2, initialize the state of each agent, place it at the starting position of the game environment, and allocate initial resources or conditions;

[0038] S3-3, start the adversarial training process, each agent makes decisions and takes actions in the game environment according to its own hierarchical reinforcement learning model. The specific formula is:

[0039] ;

[0040] in, is the observation information of the agent at time step t, is the action taken by the agent, is the policy function, A is the number of actions, Indicates that at time step t, given the observation information When taking action a, the overall action value is calculated by a neural network with parameter θ;

[0041] S3-4, records each agent’s action sequence, environmental feedback, and game results;

[0042] S3-5, based on the game results and recorded data, evaluate the performance of each agent, including its strategy selection, action efficiency, coordination ability, and confrontation ability in the game;

[0043] S3-6, according to the evaluation results, adjust and optimize the parameters of the hierarchical reinforcement learning model of the agent. The specific formula is:

[0044] ;

[0045] Among them, L represents the loss, m represents the number of samples, represents the true value, Represents the predicted value.

[0046] Furthermore, in step S5, the following sub-steps are also included:

[0047] S5-1, each agent makes decisions in different environments and collects new data to update the local data set. When the preset training rounds are reached or specific conditions are met, the agent uploads the local model parameters to the central server;

[0048] S5-2, the central server receives the model parameters from each agent and performs a federated average operation to generate global model parameters. The specific formula is:

[0049] ;

[0050] in, is a global model parameter, are the local model parameters of the ith agent, and N is the number of agents participating in the federated average;

[0051] S5-3, the central server sends the updated global model parameters to each agent, and each agent updates the local model using the global model parameters;

[0052] S5-4, during parameter upload and download, multiplication homomorphic encryption technology is used for secure transmission. The specific formula is:

[0053] ;

[0054] Among them, CT is the ciphertext, and There are two parts of the ciphertext, e is a random number selected during the encryption process, g is a generator, h is the public key, and m is the corresponding plaintext.

[0055] The beneficial effects brought about by the technical solution provided by the present invention include at least:

[0056] The present invention establishes a federated learning framework including multiple intelligent agents. The intelligent agents are located in different areas and use the data in the area to train local models. In the federated learning framework, a hierarchical reinforcement learning model is deployed inside each intelligent agent, including a top-level policy control model and an individual policy execution model. The hierarchical reinforcement learning model is used to enable multiple intelligent agents to perform adversarial training in a game environment. On the basis of game confrontation, the intelligent agents are monitored, and a collaborative decision-making algorithm is introduced based on the monitoring data to coordinate the behaviors of multiple intelligent agents to avoid conflicts and deadlocks. When the preset training rounds are reached or specific conditions are met, the local model parameters are uploaded to the central server, and the model parameters are shared through the federated learning framework for global optimization.

[0057] By introducing the federated learning mechanism, the present invention realizes distributed storage and encryption processing of data in a multi-agent system, effectively protecting user privacy and data security. During the game confrontation and collaborative decision-making process, each agent does not need to share original data, but only needs to transmit aggregated model update information, thereby reducing the risk of data leakage.

[0058] The present invention decomposes the complex game confrontation problem into multiple sub-problems by constructing a hierarchical reinforcement learning framework, which is handled separately by agents at different levels. This hierarchical structure not only improves the efficiency of strategy generation, but also enhances the scalability and flexibility of the system.

[0059] The method of the present invention realizes a complex interaction of competition and cooperation in a multi-agent system by combining game confrontation with collaborative decision-making. The dual mechanism of game confrontation and collaborative decision-making enables the method of the present invention to maintain a high degree of adaptability and robustness in a complex and changing environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0061] Figure 1 A flow chart of a method provided by an embodiment of the present invention;

[0062] Figure 2 A diagram of a federated reinforcement learning framework provided for an embodiment of the present invention;

[0063] Figure 3 A model framework diagram of hierarchical reinforcement learning provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0064] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following is a detailed description of the specific implementation method, structure, features and effects of the hierarchical multi-agent game confrontation and collaborative decision-making method based on federated learning proposed by the present invention in combination with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures, or characteristics in one or more embodiments may be combined in any suitable form.

[0065] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0066] The following examples are for illustrative purposes only and are not intended to limit the scope of the present invention.

[0067] The specific scheme of the hierarchical multi-agent game confrontation and collaborative decision-making method based on federated learning provided by the present invention is described in detail below with reference to the accompanying drawings. Example

[0068] See also Figure 1 , which shows a method flow chart of a hierarchical multi-agent game confrontation and collaborative decision-making method based on federated learning provided by an embodiment of the present invention, the method comprising the following steps:

[0069] Step S1: Establish a federated learning framework including multiple agents, where the agents are located in different regions and use the data in their respective regions to train local models;

[0070] See also Figure 2 A diagram of a federated reinforcement learning framework provided for an embodiment of the present invention.

[0071] Wherein step S1 also includes the following sub-steps:

[0072] S1-1, determine the global model update frequency of federated learning, the threshold for agents to participate in federated averaging, and the communication protocol;

[0073] S1-2, establishing multiple intelligent agents, assigning a unique identifier to each intelligent agent, and initializing its local data set, wherein the local data set includes historical data required for the intelligent agent to make decisions in a specific environment;

[0074] S1-3, each agent uses a local data set to train a reinforcement learning model, which is used to generate an action strategy based on the current observation information. The specific formula is:

[0075] ;

[0076] Among them, s is the current state, a is the action taken, and r is the reward. is the discount factor, is the learning rate, is the next state, is the next action to take.

[0077] It should be noted that federated learning allows intelligent agents to train models locally and only upload model parameters rather than original data to the central server, thereby protecting data privacy. Through collaborative training, the computing resources and data of multiple intelligent agents can be fully utilized to accelerate the model training process. In addition, federated learning can integrate the learning results of multiple intelligent agents and improve the generalization ability and performance of the global model.

[0078] A unique identifier can ensure that each agent is unique in the system, facilitating management and communication. The quality and diversity of the data set should be ensured during initialization to support the initial learning of the agent. The data set should include the historical data required for the agent to make decisions in a specific environment, including state, action, and reward.

[0079] The reinforcement learning model is a commonly used policy gradient method in reinforcement learning, where γ is a discount factor used to balance the importance of current rewards and future rewards. is the learning rate, which indicates the update speed of the model parameters; the model aims to learn a strategy that enables the agent to choose the best action in a given state to maximize the long-term reward.

[0080] The global model is updated every two hours to ensure timely updates of the global model. The threshold for agents to participate in the federated average is that only local models with an accuracy rate of more than 80% are included in the update of the global model, which can improve the performance of the global model. The communication protocol uses the WiFi protocol, which has fast speed and large data transmission capacity, and can ensure stable and secure data transmission between the agent and the central server.

[0081] Step S2: In the federated learning framework, a hierarchical reinforcement learning model is deployed inside each agent, including a top-level policy control model and an individual policy execution model;

[0082] See also Figure 3 A model framework diagram of hierarchical reinforcement learning provided in an embodiment of the present invention.

[0083] Wherein step S2 also includes the following sub-steps:

[0084] S2-1, configure a top-level strategy control model and an individual strategy execution model for each agent, and set initial parameters. The top-level strategy control model is responsible for the formulation and adjustment of the global strategy, and the individual strategy execution model is responsible for executing specific actions in the local environment. The initial parameters of the model include the weights and biases of the neural network;

[0085] S2-2, within each agent, the individual strategy execution model provides its current state, observation information, and historical behavior data to the top-level strategy control model;

[0086] S2-3, based on the information provided by the individual policy execution model, the top-level policy control model performs global policy analysis, generates corresponding control information, and passes it to the individual policy execution model. The top-level policy control model generates control information through the following formula:

[0087] ;

[0088] Among them, G represents global information, C represents control information, The mapping function representing the top-level policy control model;

[0089] S2-4, after receiving the control information, the individual agent conducts a comprehensive analysis based on its own observable information and makes a decision using the individual strategy execution model. The above individual strategy execution model makes a decision using the following formula:

[0090] ;

[0091] in, represents the action value of the ith agent, represents the control information of the ith agent, represents the partial observable information of the ith agent, A mapping function representing an individual policy execution model.

[0092] It should be noted that deploying a hierarchical reinforcement learning model within a federated learning framework means that each agent trains the model on its own data and only shares the model updates with other agents, rather than the original data, which helps protect data privacy while leveraging the data of multiple agents to improve the model.

[0093] Hierarchical reinforcement learning is a method that divides strategies into different levels. The hierarchical structure helps intelligent agents learn and make decisions more effectively in complex environments. The top-level policy control model is responsible for formulating and adjusting global strategies based on global information, which requires the model to have a global vision and policy analysis capabilities; the individual policy execution model is responsible for making decisions in the local environment based on the control information provided by the top-level policy control model and some of its own observable information, which requires the model to be able to quickly adapt to changes in the local environment and accurately perform actions.

[0094] The initial parameters of the model include the weight matrix W1 (3x4) from the input layer to the hidden layer; the weight matrix W2 (4x1) from the hidden layer to the output layer; the bias vector b1 (4x1) from the input layer to the hidden layer; and the bias vector b2 (1x1) from the hidden layer to the output layer.

[0095] The current state refers to the state of the environment in which the individual agent is located before performing an action, including the agent's position, speed, and the distribution of obstacles in the surrounding environment (in a physical environment), or the character status, number of resources, and enemy position in an online game (in a virtual environment).

[0096] Observation information: It is the information about the environment obtained by individual agents through sensors or other means, including temperature, humidity, light intensity (in physical environments), or scores, enemy status, and teammate positions in games (in virtual environments).

[0097] Historical behavior data is a record of the actions performed and feedback received by an individual agent over a period of time in the past, including the actions taken by the agent in the past, the results of the actions (rewards or penalties), and the impact of these actions on the state of the environment.

[0098] Step S3: using a hierarchical reinforcement learning model to enable multiple agents to conduct adversarial training in a game environment;

[0099] Wherein step S3 also includes the following sub-steps:

[0100] S3-1, build a game environment, set game rules and reward mechanism, the game rules include action sequence, observation and information transmission, legal action set, state transition rules, game termination conditions, the reward mechanism includes instant reward, collaborative reward, global value maximization, and punishment mechanism;

[0101] S3-2, initialize the state of each agent, place it at the starting position of the game environment, and allocate initial resources or conditions;

[0102] S3-3, start the adversarial training process, each agent makes decisions and takes actions in the game environment according to its own hierarchical reinforcement learning model. The specific formula is:

[0103] ;

[0104] in, is the observation information of the agent at time step t, is the action taken by the agent, is the policy function, A is the number of actions, Indicates that at time step t, given the observation information When taking action a, the overall action value is calculated by a neural network with parameter θ;

[0105] S3-4, records each agent’s action sequence, environmental feedback, and game results;

[0106] S3-5, based on the game results and recorded data, evaluate the performance of each agent, including its strategy selection, action efficiency, coordination ability, and confrontation ability in the game;

[0107] S3-6, according to the evaluation results, adjust and optimize the parameters of the hierarchical reinforcement learning model of the agent. The specific formula is:

[0108] ;

[0109] Among them, L represents the loss, m represents the number of samples, represents the true value, Represents the predicted value.

[0110] It should be noted that the game rules are the basis for the agent's actions in the environment. Action order: stipulates the order of the agent's actions in the game, which can be synchronous or asynchronous actions; Observation and information transmission: defines how the agent obtains environmental information and how to transmit information to other agents; Legal action set: lists all legal actions that the agent can take in each state; State transition rules: describe how the environment state changes according to the agent's actions; Game termination conditions: define when the game ends, such as reaching a certain state or after a certain number of time steps.

[0111] The reward mechanism is used to motivate agents to take actions that are conducive to maximizing global value. Instant reward: the reward that the agent receives immediately after taking a certain action; collaborative reward: in order to encourage collaboration between agents, some rewards are designed to reflect the overall performance of the team; global value maximization: the reward mechanism should be able to guide agents to take strategies that are conducive to global optimization; punishment mechanism: for actions that are not conducive to maximizing global value, some penalties are designed to suppress these behaviors.

[0112] The multi-agent game strategy generation training algorithm based on hierarchical reinforcement learning includes:

[0113] Initializing a multi-agent game strategy generation framework with random network parameters θ

[0114] Initialize experience pool R

[0115] for the curtain sequence e=1→E do

[0116] Get environment initialization observation data

[0117] for time step t=1→T do

[0118] According to the model framework of the current network parameter θ, the individual selects action a with a greedy strategy

[0119] Execute action a, get reward r, and the observed data becomes

[0120] The experience ( ) is stored in the experience pool R

[0121] If the amount of data in the experience pool R exceeds K, K data are randomly extracted from the experience pool R.

[0122] Minimize the objective function , and update the network parameters θ

[0123] end for

[0124] If e % testNum == 0 do

[0125] for e'=1→E'do

[0126] Get environment initialization observation data

[0127] for time step t=1→T do

[0128] According to the model framework of the current network parameter θ, the individual selects the action a' with the largest action value.

[0129] Update the observation data to

[0130] end for

[0131] if endFlag do

[0132] Update Win Rate

[0133] end if

[0134] end if

[0135] end if

[0136] end if

[0137] Step S4: On the basis of game confrontation, the intelligent agent is monitored, and a collaborative decision-making algorithm is introduced based on the monitoring data to coordinate the behaviors of multiple intelligent agents to avoid conflicts and deadlocks. The monitoring data includes the position, speed, decision-making choice, and interaction history of the intelligent agent.

[0138] Wherein step S4 also includes the following sub-steps:

[0139] S4-1, during the game confrontation, monitor and record the behavior, decision-making and interactions of each intelligent agent;

[0140] S4-2, based on monitoring data, starts the collaborative decision-making algorithm and uses the top-level policy control layer of the hierarchical reinforcement learning model to coordinate the behaviors of multiple agents. The specific formula is:

[0141] ;

[0142] in, represents the strategy of agent i, represents the reward function of agent i, represents the global state at time t, represents the action of agent i at time t, represents the cost caused by the action conflict between agents i and j, is the weight coefficient for adjusting conflict costs;

[0143] S4-3, according to the output of the collaborative decision-making algorithm, adjust the strategy selection of each intelligent agent, eliminate conflicts and deadlocks, and optimize the overall performance.

[0144] It should be noted that the monitoring data includes the agent's position, speed, decision-making choices, and interaction history. The agent's position can be monitored in real time, which helps the system accurately track the dynamics of each agent and ensure that they are within the expected range. Understanding the location distribution of the agents can help the algorithm avoid spatial conflicts and ensure that the agents do not appear in the same place at the same time, thereby preventing collisions or deadlocks.

[0145] The speed of an agent is an important indicator of its dynamic response capability. By monitoring the speed, the algorithm can predict the future position of the agent, thereby performing path planning or obstacle avoidance operations in advance; the speed of the agent also reflects the efficiency of its task execution. The algorithm can adjust the task allocation of the agent based on the speed data to improve the efficiency of the overall system.

[0146] By monitoring the decision choices of intelligent agents, the algorithm can analyze the strategic preferences and decision-making logic of intelligent agents, thereby discovering potential conflict points and optimization space; based on the understanding of intelligent agent decisions, the algorithm can adjust the strategies of other intelligent agents to achieve better synergy and reduce the occurrence of conflicts and deadlocks.

[0147] By analyzing the interaction history of the agents, the algorithm can identify the behavior patterns and interaction rules of the agents, which helps predict the future behavior of the agents. The interaction history can also be used to evaluate the trust and willingness to collaborate between agents. The algorithm can adjust the collaboration strategy of the agents based on this information to enhance the stability and reliability of the system.

[0148] The symbols in the collaborative decision-making algorithm formula represent the strategy, reward function, global state, action and conflict cost between agent i and other agents. These elements jointly determine the agent's behavior choices. The weight coefficient for adjusting the conflict cost is a key parameter that determines the extent to which the algorithm focuses on reducing conflicts between agents. The weight coefficient can be adjusted according to the specific application scenario.

[0149] Step S5: When the preset training rounds are reached or specific conditions are met, the local model parameters are uploaded to the central server, and the model parameters are shared through the federated learning framework for global optimization. The preset training rounds refer to the rounds in which the agent uses the data in the area to train the local model, and the local model is the local model of the agent in different areas.

[0150] Wherein, in step S5, the following sub-steps are also included:

[0151] S5-1, each agent makes decisions in different environments and collects new data to update the local data set. When the preset training rounds are reached or specific conditions are met, the agent uploads the local model parameters to the central server;

[0152] S5-2, the central server receives the model parameters from each agent and performs a federated average operation to generate global model parameters. The specific formula is:

[0153] ;

[0154] in, is a global model parameter, are the local model parameters of the ith agent, and N is the number of agents participating in the federated average;

[0155] S5-3, the central server sends the updated global model parameters to each agent, and each agent updates the local model using the global model parameters;

[0156] S5-4, during parameter upload and download, multiplication homomorphic encryption technology is used for secure transmission. The specific formula is:

[0157] ;

[0158] Among them, CT is the ciphertext, and There are two parts of the ciphertext, e is a random number selected during the encryption process, g is a generator, h is the public key, and m is the corresponding plaintext.

[0159] It should be noted that the preset training rounds are 10 rounds. Each agent needs to be trained on its local dataset for 10 rounds before uploading the local model parameters to the central server. By setting the training rounds, it can be ensured that each agent has enough time to train the model on its local dataset, thereby improving the accuracy and generalization ability of the model.

[0160] Specific conditions that are met include model performance improvement: when the performance of the agent's local model on the validation set reaches or exceeds a certain preset threshold, it uploads the local model parameters, which helps ensure that only high-quality model parameters are incorporated into the update of the global model; data diversity: if the agent's local dataset is significantly diverse from that of other agents, then its model may contain unique information. In this case, the agent may be allowed to upload its local model parameters to enrich the data representation of the global model even if the preset training rounds have not been reached.

[0161] After receiving the model parameters from each agent, the central server should perform aggregation operations to generate global model parameters. The aggregation method can be weighted averaging. The updated global model parameters should be able to reflect the learning results of all agents participating in the federated averaging to improve the performance of the global model.

[0162] After receiving the updated global model parameters, each agent should use these parameters to update its local model, which helps the agent to utilize the learning results of other agents and accelerate its own learning process. When updating the local model, attention should be paid to synchronization with the global model to avoid model drift or performance degradation.

[0163] Multiplicative homomorphic encryption technology allows calculations on data in an encrypted state without decryption, ensuring the security of communication between the intelligent agent and the central server and preventing the risk of data leakage; CT is the encrypted data (ciphertext), which consists of two parts (one part encrypted with the public key and one part encrypted with the private key), e is a random number selected during the encryption process, g is a generator (the basic element used to construct the encryption system), h is the public key (used to encrypt data), and m is the corresponding plaintext (original data).

[0164] In this way, the hierarchical multi-agent game confrontation and collaborative decision-making method based on federated learning can effectively solve the key issues of privacy protection, strategy generation, and collaborative cooperation in multi-agent systems, and can also provide strong technical support for multiple fields such as smart grids, autonomous driving, and robot collaboration.

[0165] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A hierarchical multi-agent game confrontation and collaborative decision-making method based on federated learning, characterized in that: The method includes: Step S1, establish a federated learning framework including multiple agents, where the agents are located in different regions and use the data in their respective regions to train local models; Step S2: deploy a hierarchical reinforcement learning model inside each agent within the federated learning framework, including a top-level policy control model and an individual policy execution model; Step S3, using a hierarchical reinforcement learning model to enable multiple agents to conduct adversarial training in a game environment; Step S4, based on the game confrontation, the intelligent agent is monitored, and a collaborative decision-making algorithm is introduced based on the monitoring data to coordinate the behaviors of multiple intelligent agents to avoid conflicts and deadlocks, wherein the monitoring data includes the position, speed, decision-making choice, and interaction history of the intelligent agent; Step S5, when the preset training rounds are reached or specific conditions are met, each agent uploads the local model parameters to the central server, and shares the model parameters through the federated learning framework for global optimization, where the preset training rounds refer to the rounds in which the agent uses the data in the area to train the local model, and the local model is the local model of the agent in different areas; Wherein step S4 also includes the following sub-steps: S4-1, during the game confrontation, monitor and record the behavior, decision-making and interactions of each intelligent agent; S4-2, based on monitoring data, starts the collaborative decision-making algorithm and uses the top-level policy control layer of the hierarchical reinforcement learning model to coordinate the behaviors of multiple agents. The specific formula is: ; in, represents the strategy of agent i, represents the reward function of agent i, represents the global state at time t, represents the action of agent i at time t, represents the cost caused by the action conflict between agents i and j, is the weight coefficient for adjusting conflict costs; S4-3, according to the output of the collaborative decision-making algorithm, adjust the strategy selection of each intelligent agent, eliminate conflicts and deadlocks, and optimize the overall performance.

2. The hierarchical multi-agent game confrontation and collaborative decision-making method based on federated learning as claimed in claim 1, characterized in that: Wherein step S1 also includes the following sub-steps: S1-1, determine the global model update frequency of federated learning, the threshold for agents to participate in federated averaging, and the communication protocol; S1-2, establishing multiple intelligent agents, assigning a unique identifier to each intelligent agent, and initializing its local data set, wherein the local data set includes historical data required for the intelligent agent to make decisions in a specific environment; S1-3, each agent uses a local data set to train a reinforcement learning model, which is used to generate an action strategy based on the current observation information. The specific formula is: ; Among them, s is the current state, a is the action taken, and r is the reward. is the discount factor, is the learning rate, is the next state, is the next action to take.

3. The hierarchical multi-agent game confrontation and collaborative decision-making method based on federated learning as claimed in claim 1, characterized in that: The steps also include the following sub-steps: S2-1, configure a top-level strategy control model and an individual strategy execution model for each agent, and set initial parameters. The top-level strategy control model is responsible for the formulation and adjustment of the global strategy, and the individual strategy execution model is responsible for executing specific actions in the local environment. The initial parameters of the model include the weights and biases of the neural network; S2-2, within each agent, the individual strategy execution model provides its current state, observation information, and historical behavior data to the top-level strategy control model; S2-3, based on the information provided by the individual policy execution model, the top-level policy control model performs global policy analysis, generates corresponding control information, and passes it to the individual policy execution model. The top-level policy control model generates control information through the following formula: ; Among them, G represents global information, C represents control information, The mapping function representing the top-level policy control model; S2-4, after receiving the control information, the individual agent performs a comprehensive analysis based on its own partial observable information and makes a decision using the individual strategy execution model. The individual strategy execution model makes a decision using the following formula: ; in, represents the action value of the ith agent, represents the control information of the ith agent, represents the partial observable information of the ith agent, A mapping function representing an individual policy execution model.

4. The hierarchical multi-agent game confrontation and collaborative decision-making method based on federated learning as claimed in claim 1, characterized in that: Wherein step S3 also includes the following sub-steps: S3-1, build a game environment, set game rules and reward mechanism, the game rules include action sequence, observation and information transmission, legal action set, state transition rules, game termination conditions, the reward mechanism includes instant reward, collaborative reward, global value maximization, and punishment mechanism; S3-2, initialize the state of each agent, place it at the starting position of the game environment, and allocate initial resources or conditions; S3-3, start the adversarial training process, each agent makes decisions and takes actions in the game environment according to its own hierarchical reinforcement learning model. The specific formula is: ; in, is the observation information of the agent at time step t, is the action taken by the agent, is the policy function, A is the number of actions, Indicates that at time step t, given the observation information When taking action a, the overall action value is calculated by a neural network with parameter θ; S3-4, records each agent’s action sequence, environmental feedback, and game results; S3-5, based on the game results and recorded data, evaluate the performance of each agent, including its strategy selection, action efficiency, coordination ability, and confrontation ability in the game; S3-6, according to the evaluation results, adjust and optimize the parameters of the hierarchical reinforcement learning model of the agent. The specific formula is: ; Among them, L represents the loss, m represents the number of samples, represents the true value, Represents the predicted value.

5. The hierarchical multi-agent game confrontation and collaborative decision-making method based on federated learning as claimed in claim 1, characterized in that: Wherein, in step S5, the following sub-steps are also included: S5-1, each agent makes decisions in different environments and collects new data to update the local data set. When the preset training rounds are reached or specific conditions are met, the agent uploads the local model parameters to the central server; S5-2, the central server receives the model parameters from each agent and performs a federated average operation to generate global model parameters. The specific formula is: ; in, is a global model parameter, are the local model parameters of the ith agent, and N is the number of agents participating in the federated average; S5-3, the central server sends the updated global model parameters to each agent, and each agent updates the local model using the global model parameters; S5-4, during parameter upload and download, multiplication homomorphic encryption technology is used for secure transmission. The specific formula is: ; Among them, CT is the ciphertext and There are two parts of the ciphertext, e is a random number selected during the encryption process, g is a generator, h is the public key, and m is the corresponding plaintext.

Citation Information

Patent Citations

  • Multi-agent federated cooperation method based on deep reinforcement learning

    CN112465151A

  • Multi-unmanned aerial vehicle air combat decision-making method based on multi-agent layered reinforcement learning

    CN115291625A

Cited By

  • Multi-agent layered collaborative optimization method based on knowledge graph and graph neural network

    CN121901024A