Unmanned ship cluster confrontation learning hybrid network reinforcement method and device
By introducing hybrid networks and attention mechanisms into unmanned ship swarm combat, the problem of low learning efficiency of value decomposition algorithms is solved, enabling rapid decision-making and optimized strategy selection for ship swarms, thereby improving naval combat performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-26
- Publication Date
- 2026-04-07
AI Technical Summary
In unmanned vessel swarm combat scenarios, value decomposition-related algorithms have low learning efficiency and cannot effectively utilize state information, resulting in the algorithm failing to converge and thus failing to fit the true joint action value function.
A hybrid network reinforcement method based on adversarial learning for unmanned ship swarms is adopted. By collecting the state and action information of the ships, the hybrid network and attention mechanism are used to extract important global state information, optimize the action value function of the ships, and introduce reward value as a query to improve the algorithm model.
It improves the convergence speed and decision-making performance of the algorithm, enabling ships to make optimal strategies against the enemy more quickly and enhancing the combat effectiveness in naval warfare.
Smart Images

Figure CN116994128B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the technical field of sea battle confrontation strategy selection, and in particular to reinforcement of a confrontation learning network. BACKGROUND
[0002] In recent years, unmanned ship technology has been widely concerned. Unmanned ships have the advantages of good stealth effect, low cost, small casualties and the like, and can be used for various combat tasks such as sea electronic reconnaissance, electronic countermeasures and underwater sound interference. Intelligence has always been a major trend in the development of ships. Recently, with the breakthroughs in concepts and technologies such as the Internet of Things, cloud computing, big data and artificial intelligence, unmanned ships have been continuously improved. Unmanned ships include fully autonomous unmanned ships with autonomous navigation, autonomous strategy and autonomous environmental perception capabilities.
[0003] Partially observable and confrontational multi-unmanned ship problems are particularly important, which require a single ship to observe and make decisions within a limited field of view, greatly reducing the range of ship selection, and such problems are more in line with the confrontation environment in real sea battles. At present, the mainstream method to solve this problem is the centralized training and distributed execution framework, which requires the system to know all the global information between ships during training, equivalent to observing the entire field of view, and the communication restrictions between ships are also eliminated. During testing, the communication restrictions between ships are opened again, and the ships also return to the partially observable state. The centralized training and distributed execution framework can usually be simulated in a laboratory.
[0004] In the unmanned ship cluster confrontation scene, there are problems such as partial observability, exponential growth of state space and action space, non-stationary environment and credit allocation. In order to solve these problems, the current mainstream algorithms mainly include the policy gradient algorithm, the actor-critic algorithm and the value decomposition algorithm. The value decomposition algorithm includes VDN, QIMX and MAVEN.
[0005] The value decomposition algorithm still has some problems that cannot be solved:
[0006] (1) The value decomposition related algorithm still has low learning efficiency, cannot effectively utilize state information, and cannot converge to find the optimal strategy;
[0007] (2) The joint action value function fitting is not good enough, and the real joint action value function cannot be fitted. SUMMARY
[0008] To address the existing technical problems in unmanned vessel swarm combat scenarios, where value decomposition algorithms still suffer from low learning efficiency, fail to effectively utilize state information, and thus fail to converge, ultimately failing to find the optimal strategy, and also exhibit poor fitting of the joint action value function, failing to accurately reproduce the true joint action value function, the technical solution provided by this invention is as follows:
[0009] A hybrid network reinforcement method for adversarial learning in unmanned vessel swarms, based on at least one friendly vessel and at least one enemy vessel, the method comprising:
[0010] The steps include collecting status and action information of all the aforementioned ships within a preset visual range from other ships;
[0011] The steps to obtain the value function of each ship's action based on the state information and action information;
[0012] The preset time step action with the highest value function value for each ship's action is selected as the action strategy selection step;
[0013] The step of inputting the value functions of all the aforementioned action strategies into a preset hybrid network to obtain the joint action value function;
[0014] The steps are as follows: obtaining an attention distribution based on the preset reward value of each ship and other ships within a preset visual range, and the current state information; updating the ship's own state information based on the attention distribution; and inputting the information into the hybrid network.
[0015] The step of optimizing the motion value function of each of the ships based on the current hybrid network.
[0016] Furthermore, in a preferred embodiment, the data acquisition step further includes a step of batch regularizing the acquired data.
[0017] Furthermore, a preferred embodiment is provided in which the actions of the ship include: moving, attacking, and stopping.
[0018] Furthermore, a preferred implementation is provided in which the selection step is specifically implemented through an individual intelligent agent network.
[0019] Furthermore, a preferred embodiment is provided in which the attention distribution is obtained based on the action strategy and the current state information of all the ships.
[0020] Furthermore, a preferred embodiment is provided, wherein the specific method for obtaining the attention distribution is as follows:
[0021] The preset reward value matrix obtained based on the action strategy is used as the query vector value and fitted with the current state information of all the ships.
[0022] Furthermore, a preferred embodiment is provided, wherein the preset hybrid network is a network in the QMIX algorithm.
[0023] Based on the same inventive concept, this invention also provides an unmanned ship swarm adversarial learning hybrid network enhancement device, based on at least one friendly ship and at least one enemy ship, the device comprising:
[0024] A module that collects status and action information of all the aforementioned ships within a preset visual range from other ships;
[0025] A module that obtains the value function of each ship's action based on the state information and action information;
[0026] The preset time step action with the highest value function value for each ship's action is selected as the action strategy selection module;
[0027] The module that inputs the value functions of all the action strategies into a preset hybrid network to obtain the joint action value function;
[0028] Based on the preset reward value obtained for each ship and other ships within the preset visual range, and the current state information, an attention distribution is obtained. The ship's own state information is updated according to the attention distribution and then input into the module of the hybrid network.
[0029] The module that optimizes the action value function of each of the ships based on the current hybrid network.
[0030] Based on the same inventive concept, the present invention also provides a computer storage medium for storing a computer program, which, when read by a computer, executes the unmanned ship swarm adversarial learning hybrid network reinforcement method.
[0031] Based on the same inventive concept, the present invention also provides a computer, including a processor and a storage medium, wherein when the processor reads the computer program stored in the storage medium, the computer executes the unmanned ship swarm adversarial learning hybrid network reinforcement method.
[0032] Compared with the prior art, the advantages of the technical solution provided by the present invention are as follows:
[0033] The hybrid network reinforcement method for adversarial learning in unmanned vessel swarms provided by this invention differs from the application of traditional reinforcement learning algorithms in adversarial learning of unmanned vessel swarms.
[0034] The hybrid network reinforcement method for adversarial learning of unmanned ship swarms provided by this invention introduces an attention mechanism structure and a reward value as a query. It extracts more important global state information in the current combat mission, allowing the algorithm model to focus on more important state information, thus obtaining new global state information. Finally, it is input into the super network to obtain weights and bias parameters, and finally, a joint action value function is fitted.
[0035] The hybrid network reinforcement method for adversarial learning in unmanned ship swarms provided by this invention avoids the problem of algorithm redundancy and decision-making failure caused by massive state information. At the same time, it is this method of extracting global state information that makes the model converge faster, enabling it to make rapid decisions and deployments against the enemy, and greatly improving performance.
[0036] The hybrid network reinforcement method for adversarial learning of unmanned vessel swarms provided by this invention is suitable for use in naval warfare when unmanned vessel swarms encounter enemy unmanned vessel swarms, and for selecting the optimal strategy to engage in combat with enemy vessels using the optimal strategy to achieve combat objectives. Attached Figure Description
[0037] Figure 1 A flowchart illustrating the hybrid network reinforcement method for adversarial learning in unmanned vessel swarms provided in Implementation Method 1;
[0038] Figure 2 A schematic diagram of the overall structure of the hybrid network reinforcement method for adversarial learning in unmanned ship swarms provided in Implementation Method 1;
[0039] Where Environment represents the environment, Agent 1 represents intelligent agent 1, Mixing Network represents the hybrid network, Attention represents the attention mechanism network, and Target Network represents the target network;
[0040] Figure 3 This is a schematic diagram of the reward query attention mechanism layer structure mentioned in Implementation Method 1;
[0041] Where FC represents encoding, Scores represents the scoring function, Softmax represents normalization, and Weighted sum represents weighted summation;
[0042] Figure 4 This is a schematic diagram of the experimental results of the method mentioned in Implementation Method Eleven on the StarCraft II adversarial simulation platform SMAC;
[0043] Figure 5 This is a schematic diagram of the average win rate and average reward results in the 3s5z_vs_3s6z scenario, as mentioned in Implementation Method Twelve. Detailed Implementation
[0044] To make the advantages and benefits of the technical solution provided by the present invention clearer, the technical solution provided by the present invention will now be described in further detail with reference to the accompanying drawings, specifically:
[0045] Implementation Method 1: Combination Figures 1-3 This embodiment describes a hybrid network reinforcement method for adversarial learning in unmanned ship swarms, based on at least one friendly ship and at least one enemy ship. The method includes:
[0046] The steps include collecting status and action information of all the aforementioned ships within a preset visual range from other ships;
[0047] The steps to obtain the value function of each ship's action based on the state information and action information;
[0048] The preset time step action with the highest value function value for each ship's action is selected as the action strategy selection step;
[0049] The step of inputting the value functions of all the aforementioned action strategies into a preset hybrid network to obtain the joint action value function;
[0050] The steps are as follows: obtaining an attention distribution based on the preset reward value of each ship and other ships within a preset visual range, and the current state information; updating the ship's own state information based on the attention distribution; and inputting the information into the hybrid network.
[0051] The step of optimizing the motion value function of each of the ships based on the current hybrid network.
[0052] Step 1: Unmanned ships use radar and base station-related equipment to obtain status and movement information of both friendly and enemy ships.
[0053] Step 2: Perform batch regularization processing on the status and action information collected from individual ships.
[0054] Step 3: Treat each ship as an intelligent agent. The batch-regularized state and action information is input into the individual agent network, which incorporates a GRU recurrent neural network to obtain the action value function for each agent. Based on a greedy strategy, the action policy to be adopted by the agent in the next time step is selected. The action policies include move, attack, and stop, with movement having four directions: east, south, west, and north.
[0055] Step 4: Using the reward matrix obtained by the agent in the previous time step as the query vector value, calculate the correlation value with the agent's global state information to obtain the attention distribution of the global state. This represents the probability of each piece of information appearing in the i-th agent. The higher the probability, the more attention we need to pay to the information. Finally, a weighted average is performed to obtain the attention value of the global state. This approach allows for better utilization of global state information and prevents the algorithm's learning process from becoming overly complex.
[0056] Step 5: Input the new global state information, i.e., the attention distribution, into the hybrid network. The hybrid network is the network in the original algorithm QMIX. This implementation is mainly an improvement on the QMIX algorithm to obtain the weights and biases of the hybrid network.
[0057] Step 6: Input the agent's action value function into the hybrid network and fit it to obtain the joint action value function;
[0058] This joint action value function is the output of the hybrid network, and the input of the hybrid network is the action value function of each agent.
[0059] Step 7: Calculate the loss value of the joint action value function, specifically by calculating the difference between the action value function value of the target network and the action value function value obtained from the hybrid network, and then update the network parameters of the individual agents.
[0060] In step three, the process of inputting the regularized state and action information into the individual agent network is as follows: the data is input to the fully connected layer. Since the input values are partially observable, a GRU recurrent neural network is added to address this issue. Given the hidden layer state in the previous time step, the output is the action value function for a single agent. Finally, a greedy strategy is used to select the action value function at the current time. and action value
[0061] In step four, the parameters of the fitted joint action value function are generated by the hypernetwork, and the input of the hypernetwork is the global state s. t s t This generally refers to information about all units in the environment, including both friendly and enemy units. However, not all of this information is relevant to our needs. A significant portion of the state information is unnecessary for current task decisions; focusing on it would only complicate computation and often yield poor results. Such a large amount of information can sometimes drastically reduce training efficiency, leading to information overload. Therefore, we propose a reward-based attention mechanism network layer. First, we will... t output f via encoder t i The encoder is a fully connected network layer, f t i This represents the state encoding information of the i-th agent, f ti The score function is used to calculate f from the previous time step. t i The reward matrix r of all agents within the observation range t-1 As the query value and f t i Calculations are performed, and an additive model is used to derive the correlation between the agent and the reward, thus obtaining the scoring function.
[0062]
[0063] Where V1, W1, and W2 are all learnable parameter matrices, r t-1 f represents the reward matrix of the agent within the observation range of the previous time step. t i V1 is the encoded state information of the i-th agent. T This represents the learnable parameters.
[0064] Then the attention distribution is derived:
[0065]
[0066] Where N represents the Nth piece of information from the agent;
[0067] Normalize it, It can be obtained using the softmax function. This represents the probability of each piece of information appearing in the i-th agent. The higher the probability, the more information we need to pay attention to. Finally, the global state f of the i-th agent is... t i The attention value is finally obtained by weighting the attention distribution. Finally, the merged outputs yield a new global state value G. t This state value reflects the state information that has the greatest impact on decision-making at the current time step. Its dimensionality is mainly determined by the number of agents.
[0068] In step five, the input to the hybrid network is the action-value function of the individual agent network. The hybrid network is a feedforward neural network, and its parameters are mainly determined by the supernetwork. The input G of the supernetwork is... t W is obtained after using the softmax function. a and W b Satisfying the non-negativity condition, the softmax function has been verified to perform better than the abs function. t The bias value is obtained through the ReLU function and the activation function. The bias value does not need to satisfy the non-negativity condition.
[0069] In step seven, the joint action value function is fed into the target network to calculate and minimize the loss value, and the network parameters of the individual agent are updated. The loss value is:
[0070]
[0071] Where b represents the number of samples, Q tot Let τ represent the joint action value function, τ represent past historical observations, a represent the joint action, and θ represent the network parameters.
[0072] r all The total reward value in the entire environment, γ represents the discount factor, and max a' τ' represents the largest action, τ' represents the historical observation, a' represents the action at the next time step, and θ' represents the target network parameters.
[0073] Unlike traditional reinforcement learning algorithms applied to unmanned surface vessel swarm warfare, this method introduces an attention mechanism. The reward value serves as a query, extracting more relevant global state information for the current combat mission. This allows the algorithm to focus on more critical state information, resulting in new global state data. This new global state data is then input into a hypernetwork to obtain weights and bias parameters, and finally, a joint action value function is fitted. This method avoids algorithmic redundancy and convergence problems caused by massive amounts of state information. Furthermore, this method of extracting global state information enables faster model convergence, allowing for rapid decision-making and deployment against the enemy, significantly improving performance.
[0074] Implementation Method 2: This implementation method further defines the unmanned vessel swarm adversarial learning hybrid network reinforcement method provided in Implementation Method 1. The data acquisition step also includes a step of batch regularizing the acquired data.
[0075] Implementation Method 3: This implementation method further defines the unmanned ship swarm adversarial learning hybrid network reinforcement method provided in Implementation Method 1. The ship actions include: moving, attacking, and stopping.
[0076] Implementation Method 4: This implementation method further defines the unmanned ship swarm adversarial learning hybrid network reinforcement method provided in Implementation Method 1. The selection step is specifically implemented through an individual agent network.
[0077] Implementation Method 5: This implementation method is a further limitation of the unmanned ship swarm adversarial learning hybrid network reinforcement method provided in Implementation Method 1. The attention distribution is obtained based on the action strategy and the current state information of all the ships.
[0078] Implementation Method Six: This implementation method further defines the unmanned vessel swarm adversarial learning hybrid network reinforcement method provided in Implementation Method Five. The specific method for obtaining the attention distribution is as follows:
[0079] The preset reward value matrix obtained based on the action strategy is used as the query vector value and fitted with the current state information of all the ships.
[0080] Implementation Method Seven: This implementation method further defines the unmanned vessel swarm adversarial learning hybrid network reinforcement method provided in Implementation Method One. The preset hybrid network is the network in the QMIX algorithm.
[0081] Implementation Method Eight: This implementation method provides an unmanned ship swarm adversarial learning hybrid network enhancement device, based on at least one friendly ship and at least one enemy ship. The device includes:
[0082] A module that collects status and action information of all the aforementioned ships within a preset visual range from other ships;
[0083] A module that obtains the value function of each ship's action based on the state information and action information;
[0084] The preset time step action with the highest value function value for each ship's action is selected as the action strategy selection module;
[0085] The module that inputs the value functions of all the action strategies into a preset hybrid network to obtain the joint action value function;
[0086] Based on the preset reward value obtained by each ship and other ships within the preset visual range, and the current state information, an attention distribution is obtained. The ship's own state information is updated according to the attention distribution and then input into the module of the hybrid network.
[0087] The module that optimizes the action value function of each of the ships based on the current hybrid network.
[0088] Implementation Method Nine: This implementation method provides a computer storage medium for storing a computer program. When the computer program is read by the computer, the computer executes the unmanned ship swarm adversarial learning hybrid network reinforcement method provided in any one of Implementation Methods One to Seven.
[0089] Implementation Method 10: This implementation method provides a computer, including a processor and a storage medium. When the processor reads the computer program stored in the storage medium, the computer executes the unmanned ship swarm adversarial learning hybrid network reinforcement method provided in any one of Implementation Methods 1 to 7.
[0090] Implementation Method Eleven: CombinationFigure 4 This embodiment describes a specific implementation of the unmanned ship swarm adversarial learning hybrid network reinforcement method provided in Embodiment 1, applied to StarCraft II. Specifically:
[0091] In this embodiment, the SMAC simulation platform is used to simulate an unmanned vessel swarm confrontation scenario. The experiment is conducted on an 8m_vs_9m scenario with average win rate and average reward as indicators. The experimental results show that the method proposed in Implementation 1 is superior to other contrastive reinforcement learning methods in terms of average win rate and average reward.
[0092] Implementation Method Twelve, Combination Figure 5 This embodiment describes a specific implementation of the adversarial learning hybrid network reinforcement method for unmanned ship swarms provided in Embodiment 1, applied to a 3s5z vs 3s6z scenario. Specifically:
[0093] Figure 4 This is a graph showing the average win rate and average reward results of the method provided in Implementation Method 1 in the 3s5z_vs_3s6z scenario;
[0094] The experimental results show that this method outperforms other contrastive reinforcement learning methods in terms of average win rate and average reward.
[0095] The above description of several specific embodiments further details the technical solution provided by the present invention in order to highlight the advantages and benefits of the technical solution provided by the present invention. However, the above-described specific embodiments are not intended to limit the present invention. Any reasonable modifications and improvements to the present invention, combinations of embodiments, and equivalent substitutions based on the spirit and principles of the present invention should be included within the protection scope of the present invention.
[0096] In the description of this specification, only preferred embodiments of the present invention are described, and should not be construed as limiting the scope of the invention. Furthermore, the use of terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples" indicates that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or N embodiments or examples. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described in this specification and the features of different embodiments or examples without contradiction. Additionally, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of the present invention, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified. Any process or method described in the flowcharts or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more N executable instructions for implementing custom logical functions or processes, and the scope of preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order according to the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain. The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, a “computer-readable medium” can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection having one or N wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic device, and portable optical disc read-only memory (CDROM).Furthermore, the computer-readable medium can even be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory. It should be understood that various parts of the invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0097] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it includes one or a combination of the steps of the method embodiments. Furthermore, the functional units in the various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
Claims
1. A hybrid network reinforcement method for adversarial learning in unmanned ship swarms, based on at least one friendly ship and at least one enemy ship, characterized in that... The method includes: The steps include collecting status and action information of all the aforementioned ships within a preset visual range from other ships; The steps to obtain the value function of each ship's action based on the state information and action information; The preset time step action with the highest value function value for each ship's action is selected as the action strategy selection step; The step of inputting the value functions of all the aforementioned action strategies into a preset hybrid network to obtain the joint action value function; The steps involve obtaining an attention distribution based on the preset reward value of each ship and other ships within a preset visual range, along with the current state information; updating the ship's own state information based on the attention distribution; and inputting the updated state information into the hybrid network. The preset hybrid network is the network in the QMIX algorithm. The step of optimizing the motion value function of each of the ships based on the current hybrid network.
2. The unmanned vessel swarm adversarial learning hybrid network reinforcement method according to claim 1, characterized in that, The data acquisition step also includes a step of batch regularizing the acquired data.
3. The unmanned vessel swarm adversarial learning hybrid network reinforcement method according to claim 1, characterized in that, The ship's actions include: moving, attacking, and stopping.
4. The unmanned vessel swarm adversarial learning hybrid network reinforcement method according to claim 1, characterized in that, The selection step is specifically achieved through an individual intelligent agent network.
5. The unmanned vessel swarm adversarial learning hybrid network reinforcement method according to claim 1, characterized in that, The attention distribution is obtained based on the action strategy and the current state information of all the ships.
6. The unmanned vessel swarm adversarial learning hybrid network reinforcement method according to claim 5, characterized in that, The specific method for obtaining the attention distribution is as follows: The preset reward value matrix obtained from the action strategy is used as the query vector value and fitted with the current state information of all the ships.
7. An unmanned ship swarm adversarial learning hybrid network enhancement device, based on at least one friendly ship and at least one enemy ship, characterized in that, The device includes: A module that collects status and action information of all the aforementioned ships within a preset visual range from other ships; A module that obtains the value function of each ship's action based on the state information and action information; The preset time step action with the highest value function value for each ship's action is selected as the action strategy selection module; The module that inputs the value functions of all the action strategies into a preset hybrid network to obtain the joint action value function; Based on the preset reward value obtained for each ship and other ships within a preset visual range, and the current state information, an attention distribution is obtained. The ship's own state information is updated according to the attention distribution and then input into the module of the hybrid network. The preset hybrid network is the network in the QMIX algorithm. The module that optimizes the action value function of each of the ships based on the current hybrid network.
8. A computer storage medium for storing computer programs, characterized in that, When the computer program is read by the computer, the computer executes the method according to any one of claims 1-6.
9. A computer, comprising a processor and a storage medium, characterized in that, When the processor reads the computer program stored in the storage medium, the computer executes the method according to any one of claims 1-6.
Citation Information
Patent Citations
Multi-agent game AI design method based on attention mechanism and reinforcement learning
CN114130034A
Embedded multi-agent reinforcement learning method using sparse attention aided decision
CN114626499A