Marine ship formation attack and defense deduction method based on multi-agent game and expert knowledge and computer device

By adopting a hybrid drive method of multi-agent game and expert knowledge in maritime ship formations, combined with Bayesian estimation and reinforcement learning, the explanatory and accuracy problems of overall mission decision-making in ship formations are solved, and a more efficient and credible decision-making process is achieved.

CN120217847APending Publication Date: 2025-06-27ZHENGZHOU UNIV

Patent Information

Application Number
CN202510272890.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, the overall task decision of ship formations is poor, and reinforcement learning based on value functions is difficult to deal with continuous action space, resulting in poor accuracy.

Method used

A hybrid drive method based on multi-agent game and expert knowledge is adopted to construct high-level task decisions through Bayesian expert knowledge estimation, and low-level action decisions are made in combination with reinforcement learning to form a hybrid decision framework.

Benefits of technology

It improves the interpretability and autonomous learning ability of ship formation task decisions, enhances the system's adaptability, and improves the credibility of the decision-making process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217847A_ABST
    Figure CN120217847A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of ship formation, and particularly relates to a marine ship formation attack and defense deduction method based on multi-agent game and expert knowledge and a computer device. The method comprises the following steps: S1, acquiring environment data information; the environment data information comprises formation situation information used for representing ship states of the friend and the foe; s2, obtaining a likelihood function value for Bayesian estimation according to the formation situation information in combination with pre-designed expert knowledge, and then obtaining an initial task overall strategy of the marine ship formation through Bayesian estimation; and S3, inputting the environment data information into a pre-trained strategy decision model corresponding to the initial task overall strategy to obtain an action instruction of each agent in the marine ship formation. According to the invention, the technical problem of poor interpretation of formation overall task decision in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of ship formations, and particularly relates to a method and a computer device for offensive and defensive deduction of maritime ship formations based on multi-agent games and expert knowledge. Background Art

[0002] In modern maritime operations and ship formation management, efficiently realizing cooperative task allocation and tactical adjustment of ships in a simulation environment is one of the key factors for improving the effectiveness of a fleet. With the development of technology, multi-agent systems based on reinforcement learning have gradually been applied to the cooperative control of ship formations. The simulation method of a multi-agent ship formation based on reinforcement learning simulates the interaction and cooperation process among different ships, combines environmental information and task requirements, and automatically adjusts parameters such as the formation form, speed, and formation of the fleet to ensure that the fleet can operate efficiently under different sea conditions, allowing each ship to make autonomous decisions based on real-time data and the interaction among them and complete tasks cooperatively.

[0003] The published text of a Chinese invention patent application with the application publication number CN115755593A and the application publication date of March 7, 2023 discloses a cooperative multi-agent control method based on value function supervision. This method first generates formation-level action instructions composed of multiple agents through a formation-level controller, and then each agent makes platform-level decision action instructions of each agent under the supervision of the formation-level action instructions through a platform-level controller. The entire decision-making process consists of two stages: the first stage is for the formation-level controller to make formation-level macroscopic decisions, and the second stage is for the platform-level controller to make individual decisions for each agent based on the macroscopic decisions in the first stage to generate decision action instructions. However, on the one hand, the formation-level controller in the above solution uses an actor-critic network based on reinforcement learning, which requires a large amount of data to be learned in advance, especially a large amount of data is required for training in both stages, and the interpretability of the decision is also poor, which is not conducive to expansion and debugging, and the unclear decision-making process leads to poor credibility; on the other hand, the platform-level controller in the above solution uses a reinforcement learning method based on value function; however, in the simulation deduction of maritime ship formations, the output action instructions are generally continuous actions, and the reinforcement learning method based on value function often has difficulty dealing with continuous action spaces, resulting in poor accuracy. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and a computer device for offensive and defensive deduction of maritime ship formations based on multi-agent games and expert knowledge to solve the technical problem of poor interpretability of the overall task decision of the formation in the prior art.

[0005] To solve the above technical problems, the technical solution of a method for attacking and defending deduction of a maritime ship formation based on multi-agent game and expert knowledge provided by the present invention is as follows: A method for attacking and defending deduction of a maritime ship formation based on multi-agent game and expert knowledge, the method comprising:

[0006] S1. Obtain environmental data information; the environmental data information includes formation situation information and combat environment information; the formation situation information is the state information of the intelligent agents on both sides of the enemy and us, and the combat environment information is the environmental information of the battlefield;

[0007] S2. According to the formation situation information, combine with the pre-designed expert knowledge to obtain the likelihood function value for Bayesian estimation, and then obtain the initial task overall strategy of the maritime ship formation through Bayesian estimation;

[0008] S3. According to the initial task overall strategy, input the environmental data into the pre-trained strategy decision model to obtain the action instructions of each intelligent agent in the maritime ship formation.

[0009] The beneficial effects of the above technical solution are as follows: The technical solution of a method for attacking and defending deduction of a maritime ship formation based on multi-agent game and expert knowledge of the present invention belongs to an improved invention creation. The present invention provides a multi-agent game model architecture with hybrid drive by combining two methods of Bayesian expert knowledge estimation and reinforcement learning. Among them, Bayesian expert knowledge estimation is used to construct high-level task decision-making, and reinforcement learning is mainly used for low-level action decision-making. This hybrid decision-making framework can effectively combine the advantages of both, and can not only provide clear and reasonable guidance through Bayesian expert knowledge estimation, but also use reinforcement learning to improve the system's autonomous learning ability and adaptability. The present invention solves the technical problem of poor interpretability of the overall task decision-making of the formation in the prior art.

[0010] Further, the ship state includes strength, and the strength is obtained according to the ship type and ship category.

[0011] Further, the initial task overall strategy includes an attack strategy, a defense strategy, and a search strategy for obtaining the state information of enemy ships; the likelihood function value satisfies:

[0012] If our strength is higher than the enemy's strength and the strength difference between the two sides is higher than the strength difference threshold, the likelihood function value corresponding to the attack strategy is the highest;

[0013] If the strength difference between the two sides is not higher than the strength difference threshold, the likelihood function value corresponding to the defense strategy is the highest;

[0014] If our strength is lower than the enemy's strength and the strength difference between the two sides is higher than the strength difference threshold, the likelihood function value corresponding to the search strategy is the highest.

[0015] Furthermore, the attack strategy includes a target selection rule for selecting an attack target and a weapon allocation rule for selecting a weapon.

[0016] Furthermore, the target selection rule is to preferentially select a target with a high priority for attack, and the priority is determined as follows:

[0017]

[0018] where P target is the priority; d target is the distance between the agent that launches the attack on our side and the target to be attacked; T target is the threat level of the target to be attacked; C target is the type of the target to be attacked; S target is the ship type of the target to be attacked; w1, w2, w3, and w4 are all weight coefficients.

[0019] Furthermore, the weapon allocation rule is as follows:

[0020] If the target to be attacked is an aircraft carrier or a cruiser and the distance between the agent that launches the attack on our side and the target to be attacked is greater than the first threshold, then the selected weapon is an anti-ship missile;

[0021] If the target to be attacked is a destroyer or a submarine and the distance between the agent that launches the attack on our side and the target to be attacked is less than or equal to the first threshold, then the selected weapon is a torpedo;

[0022] If the target to be attacked is a light frigate or a medium frigate and the threat level of the target to be attacked is less than or equal to the second threshold, then the selected weapon is artillery fire.

[0023] Furthermore, the policy decision model is a reinforcement learning model, and the reward function for training the policy decision model corresponding to the search strategy includes R patrol :

[0024] R patrol = R critical + R detect + R fuel

[0025] R critical = r1 × I critical

[0026] R detect = r2 × I detect

[0027] R fuel = -λ1 × (Fuel total - Fuel remain )

[0028] Among them, r1 is a positive reward value, representing the reward for searching the key area; I critical is an indicator function, which is 1 when the agent enters the key area and 0 otherwise; r2 is a positive reward value, representing the reward for discovering the enemy target; I detect is an indicator function, which is 1 when the agent discovers the enemy target and 0 otherwise; λ1 is a positive penalty coefficient, representing the penalty intensity of fuel consumption; Fuel total is the initial fuel quantity; Fuel remain is the remaining fuel quantity;

[0029] The reward function for training the policy decision model corresponding to the attack strategy includes R attack :

[0030] R attack =R destroy +R fail +R ammo

[0031] R destroy =r4×I destroy

[0032] R fail =-r5×I fail

[0033] R ammo =-λ2×(Ammo total -Ammo remain )

[0034] Among them, r4 is a positive reward value, representing the reward for destroying the enemy target; I destroy is an indicator function, which is 1 when the agent destroys the enemy target and 0 otherwise; r5 is a positive penalty value, representing the penalty for not destroying the enemy target; I fail is an indicator function, which is 1 when the agent does not destroy the enemy target and 0 otherwise; λ2 is a positive penalty coefficient, representing the penalty intensity of ammunition consumption; Ammo total is the initial ammunition quantity; Ammo remain is the remaining ammunition quantity;

[0035] The reward function for training the policy decision model corresponding to the attack strategy includes R defense :

[0036] R defense =R defend +R region-protect +R resource-cost

[0037] R defend =r6×I defend

[0038] R region-protect = r7 × I region-protect

[0039] R resource-cost = -λ3 × (Resource total - Resource remain )

[0040] Wherein, r6 is a positive reward value, representing the reward for successfully defending against the enemy's attack; I defend is an indicator function, which is 1 when the agent successfully defends against the enemy's attack and 0 otherwise; r7 is a positive reward value, representing the reward for protecting the important area; I region-protect is an indicator function, which is 1 when the agent protects the important area and 0 otherwise; λ3 is a positive penalty coefficient, representing the penalty intensity of resource consumption; Resource total is the initial resource quantity; Resource remain is the remaining resource quantity.

[0041] Furthermore, the policy decision-making model is a reinforcement learning model, and the loss function during the training of the policy decision-making model includes at least one of policy loss, value loss, entropy loss, and global policy coordination loss.

[0042] Furthermore, the policy decision-making model is a reinforcement learning model, and the reward function during the training of the policy decision-making model includes a cooperation reward function:

[0043]

[0044] Wherein, R coop is the global cooperation reward; w i is the weight of agent i; R i is the local return of agent i; N is the number of agents.

[0045] The present invention also provides a technical solution for a computer device: a computer device, including a processor, and the processor is used to execute a computer program to implement the steps of the above-mentioned method for maritime ship formation attack and defense deduction based on multi-agent game and expert knowledge. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is a schematic diagram of the overall system architecture of the embodiment of the method for maritime ship formation attack and defense deduction based on multi-agent game and expert knowledge of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The present invention provides a multi-agent game model architecture with hybrid drive combining Bayesian expert knowledge estimation and reinforcement learning. Among them, Bayesian expert knowledge estimation is used to construct high-level task decisions, and reinforcement learning is mainly used for low-level action decisions. This hybrid decision-making framework can effectively combine the advantages of both. It can not only provide clear and reasonable guidance through Bayesian expert knowledge estimation but also utilize reinforcement learning to enhance the system's autonomous learning ability and adaptability. The present invention solves the technical problem of poor interpretability of the overall task decision-making of the formation in the prior art.

[0048] Example of a method for offensive and defensive deduction of a naval ship formation based on multi-agent game and expert knowledge:

[0049] As Figure 1 shown, the method for offensive and defensive deduction of a naval ship formation based on multi-agent game and expert knowledge in this embodiment can be implemented based on the simulation deduction system described below. This system includes a platform layer, a rule layer, and a reinforcement learning layer.

[0050] Among them, the platform layer is responsible for environmental interaction and has two interfaces: an initialization and running system interface and an action list and micro-action list interface. The initialization and running system interface is used to initialize, execute simulation steps, and reset the simulation environment, and send the environmental data information in the simulation environment to the rule layer. The action list and micro-action list interface is used to receive the action instruction list and micro-action set generated by multiple agents, and return feedback information after executing specific actions. The platform layer provides an interface for interacting with the environment and is the basic operation platform for the entire multi-agent game. Agents collect all environmental data information from the platform layer, providing sufficient auxiliary information for key decisions in the subsequent decision-making stage. Among them, the environmental data information includes formation situation information and combat environment information. The formation situation information is used to represent the states of friendly and enemy agents; such as whether friendly operators, opponent operators, enemy agents are visible, their formation distribution and heading, etc., and the current task status and position of friendly agents. The combat environment information is the environmental information of the maritime battlefield, such as sea conditions, wind speed, water flow, obstacle distribution, etc.

[0051] The rule layer includes a situation analysis module, a decision-making module, and an action module. The situation analysis module is mainly responsible for task decisions driven by expert knowledge rules. Based on the formation situation information in the environmental data information sent by the platform layer, the situation analysis module generates the initial overall task strategy of the ship formation according to the expert knowledge rules based on the Bayesian model. The expert knowledge rules are decision-making frameworks designed for multi-agent systems based on expert knowledge for naval combat game scenarios. The decision-making module, based on the initial overall task strategy output by the situation analysis module, generates action instructions for each agent in the formation through a trained strategy decision model (i.e., Figure 1in the task queue). The action module generates an action list and executes specific operations according to the action instructions generated by the decision-making module, and feeds back the execution results to the platform layer for result evaluation.

[0052] The reinforcement learning layer is used to create, train, and store a policy decision-making model to support the decision-making module of the rule layer and optimize the individual decision-making capabilities of each agent.

[0053] A method for offensive and defensive deduction of a maritime ship formation based on multi-agent game and expert knowledge, the method comprising:

[0054] S1. Obtain environmental data information; the environmental data information includes formation situation information for representing the states of friendly and enemy ships.

[0055] Among them, the environmental data information includes formation situation information and combat environment information. The formation situation information is used to represent the states of friendly and enemy agents; such as the states of our operators, opponent operators, whether the enemy agents are visible and their formation distributions and headings, etc., and the current task states and positions of our agents; the combat environment information is the environmental information of the maritime battlefield, such as sea conditions, wind speed, water flow, obstacle distributions, etc.

[0056] Specifically, in this embodiment, the strength values of friendly and enemy ship agents (such as aircraft carriers, cruisers, destroyers, frigates, and models such as large, medium, and small) are introduced as part of the state characteristics.

[0057] Aircraft carrier: strength value H1; cruiser: strength value H2; destroyer: strength value H3; frigate: strength value H4, and so on.

[0058] The strength values for large, medium, and small are T1, T2, and T3 respectively.

[0059] Final strength algorithm: D i = H n × T m , where i is the ship agent number; n is the ship type (such as aircraft carrier, cruiser, destroyer, frigate); m is the ship class (such as large, medium, small).

[0060] The strength is a constant set by the user considering the combat capabilities of various ship agents, and the combat capabilities here can include the ability of the ship to attack enemy targets, the ability to defend against enemy attacks, etc.

[0061] Strength state representation: The strength state D can be represented as a combination of the ship strength values of friendly and enemy sides. For example:

[0062] D = (D 我方 , D 敌方 )

[0063] where D 我方 and D 敌方 are the strength value vectors of the ships of both sides respectively.

[0064] S2. Obtain the likelihood function value for Bayesian estimation based on the formation situation information combined with the pre-designed expert knowledge, and then obtain the initial overall mission strategy of the naval ship formation through Bayesian estimation.

[0065] The initial overall mission strategy includes a search strategy, an attack strategy, and a defense strategy.

[0066] Based on the Bayesian estimation method, expert knowledge can calculate the probability distribution of choosing a certain action given the friendly and enemy strength values. Through Bayesian inference, the action probability distribution is dynamically updated, thereby outputting the final strategy. The expert's prior belief about the action is, where p1, p2, and p3 are custom constants: P(a1) = p1, P(a2) = p2, P(a3) = p3, where p1 + p2 + p3 = 1. (For example, P(a1) = 0.4, P(a2) = 0.3, P(a3) = 0.3).

[0067] For the naval combat game scenario, domain experts designed the likelihood function in Bayesian estimation based on practical experience and combat theory (i.e., expert knowledge).

[0068] Calculation of the likelihood function in the friendly and enemy ship confrontation scenario: Assume that the state s consists of the strength values of the friendly and enemy ships:

[0069] D = (D 我方 , D 敌方 )

[0070] The action space is {a1 = attack, a2 = defense, a3 = search}. Calculation of the likelihood function based on expert knowledge. The expert defines the following rules:

[0071] Likelihood of attack: If our side's strength value is significantly higher than the enemy's, the likelihood of attack is higher.

[0072] Likelihood of defense: If our side's strength value is close to the enemy's, the likelihood of defense is higher.

[0073] Likelihood of search: If our side's strength value is significantly lower than the enemy's, the likelihood of search is higher.

[0074] Posterior probability: Calculate the posterior probability of choosing action a i under strength D through Bayes' theorem:

[0075]

[0076] where P(a i ) is the prior probability (the initial belief of the expert about action a i ). P(D|a i ) is the likelihood function. is the normalization factor, ensuring that the total probability of the posterior distribution is 1.

[0077] Action selection (i.e., initial overall task strategy selection): At this time, actions a1, a2, or a3 are selected according to the posterior probability. (For example, P(a1|D) = 0.5: In the current state, the probability of selecting an attack is 50%. P(a2|D) = 0.1667: In the current state, the probability of selecting a defense is 16.67%. P(a3|D) = 0.1667: In the current state, the probability of selecting a search is 16.67%).)

[0078] For example: Suppose in a scenario of a battle between enemy and friendly ships:

[0079] The action space is {a1 = attack, a2 = defense, a3 = search}.

[0080] D represents the strength value of the enemy and friendly ships.

[0081] The prior distribution is: P(a1) = 0.4, P(a2) = 0.3, P(a3) = 0.3

[0082] The likelihood function is: P(D|a1) = 0.6, P(D|a2) = 0.2, P(D|a3) = 0.2

[0083] Calculate the posterior distribution through Bayes' theorem:

[0084]

[0085] The meaning of the posterior distribution:

[0086] P(a1|D) = 0.5: In the current state, the probability of selecting an attack is 50%.

[0087] P(a2|D) = 0.1667: In the current state, the probability of selecting a defense is 16.67%.

[0088] P(a3|D) = 0.1667: In the current state, the probability of selecting a search is 16.67%.

[0089] Search strategy: The search strategy is used to obtain information about the state of the enemy ships. Agent P adopts the search strategy S p , where S p represents the search task of agent P. The goal of the search strategy is to maximize the battlefield awareness ability, expressed as max∑ p∈N A p , where Ap is the perception area of agent P, and N is the set of all agents in the formation.

[0090] Attack strategy: The selection of the attack target can be represented by the target selection function:

[0091]

[0092] where P target is the priority, and the target with a higher priority is preferentially selected for attack; d target is the distance (unit: meter) between the agent that launches the attack on our side and the target to be attacked; T target is the threat level of the target to be attacked (such as the weapon threat level of the target, whether the enemy ship is equipped with high-threat weapons, etc.), and the larger the value, the greater the threat; C target is the type of the target to be attacked (such as aircraft carrier, cruiser, destroyer, etc.), and different types of targets have different priorities; S target is the ship type of the target to be attacked. The ship type usually refers to the specific model or category of the enemy ship, such as large destroyer, small destroyer, etc.; w1, w2, w3, and w4 are all weight coefficients used to adjust the importance of each factor in target selection.

[0093] In actual combat, different types of targets usually need to be matched with different weapons. The specific weapon allocation rules are as follows:

[0094]

[0095] Among them, Missile (anti-ship missile) is preferentially used to attack large ships (aircraft carrier (i.e., Aircraft Carrier), cruiser (i.e., Cruiser)), especially at long distances. Torpedo is suitable for attacking enemy ships at closer ranges, especially targets for rapid attacks (such as destroyer (i.e., Destroyer) or submarine (i.e., Submarine)). Gun is suitable for smaller targets or when the threat of the enemy ship is low, such as light frigate (i.e., Corvette) or medium frigate (i.e., Frigate). If the enemy target is relatively powerful or has a high defense ability, the firepower is concentrated on one target and powerful weapons are used for attack. When facing multiple targets, the fleet chooses to disperse the firepower to ensure that each target can be attacked, thereby improving the overall mission completion rate.

[0096] S3. Input the environmental data information into the pre-trained policy decision model corresponding to the overall initial task strategy to obtain the action instructions of each agent in the maritime ship formation.

[0097] Corresponding to the initial overall task strategy, there are three strategy decision-making models, namely: the search strategy model corresponding to the search strategy, the attack strategy model corresponding to the attack strategy, and the defense strategy model corresponding to the defense strategy.

[0098] After S2 obtains the posterior probabilities corresponding to different actions, it randomly selects an action according to the posterior probabilities. The higher the posterior probability of an action, the greater the probability of being selected. The environmental data information is input into the strategy decision-making model corresponding to the selected action for fine decision-making to obtain the action instructions of each agent in the naval ship formation.

[0099] In other embodiments, the action with the highest posterior probability can also be directly used as the selected action.

[0100] In the naval combat scenario of multi-agent games, since the ultimate goal is to win, all agents in the ship formation belong to a fully cooperative scenario, and all agents share the same reward function. In this embodiment, a collaborative reward function is introduced into the reward function, and the generated action plan takes into account the cooperative relationship between ships to ensure the efficient completion of the formation task. The collaborative reward function is:

[0101]

[0102] Among them, R coop is the global collaborative reward; w i is the weight of agent i; R i is the local return of agent i.

[0103] The reward function for training the strategy decision-making model corresponding to the search strategy includes R patrol :

[0104] R patrol = R critical + R detect + R fuel

[0105] R critical = r1 × I critical

[0106] R detect = r2 × I detect

[0107] R fuel = -λ1 × (Fuel total - Fuel remain )

[0108] Among them, r1 is a positive reward value, indicating the reward for searching the key area; I criticalis an indicator function, which is 1 when the agent enters the critical area and 0 otherwise; r2 is a positive reward value, representing the reward for discovering an enemy target; I detect is an indicator function, which is 1 when the agent discovers an enemy target and 0 otherwise; λ1 is a positive penalty coefficient, representing the penalty intensity for fuel consumption; Fuel total is the initial fuel quantity; Fuel remain is the remaining fuel quantity;

[0109] The reward function for training the policy decision model corresponding to the attack policy includes R attack :

[0110] R attack = R destroy + R fail + R ammo

[0111] R destroy = r4 × I destroy

[0112] R fail = -r5 × I fail

[0113] R ammo = -λ2 × (Ammo total - Ammo remain )

[0114] where r4 is a positive reward value, representing the reward for destroying an enemy target; I destroy is an indicator function, which is 1 when the agent destroys an enemy target and 0 otherwise; r5 is a positive penalty value, representing the penalty for not destroying an enemy target; I fail is an indicator function, which is 1 when the agent does not destroy an enemy target and 0 otherwise; λ2 is a positive penalty coefficient, representing the penalty intensity for ammo consumption; Ammo total is the initial ammo quantity; Ammo remain is the remaining ammo quantity;

[0115] The reward function for training the policy decision model corresponding to the attack policy includes R defense :

[0116] R defense = R defend + R region-protect + R resource-cost

[0117] R defend = r6 × I defend

[0118] R region-protect = r7 × I region-protect

[0119] R resource-cost = -λ3 × (Resource total - Resource remain )

[0120] Wherein, r6 is a positive reward value, representing the reward for successfully defending against an enemy attack; I defend is an indicator function, which is 1 when the agent successfully defends against an enemy attack and 0 otherwise; r7 is a positive reward value, representing the reward for protecting important areas; I region-protect is an indicator function, which is 1 when the agent protects important areas and 0 otherwise; λ3 is a positive penalty coefficient, representing the penalty intensity for resource consumption; Resource total is the initial resource quantity; Resource remain is the remaining resource quantity.

[0121] Wherein, the above-mentioned R patrol 、R attack and R defense all represent the reward functions for the action execution of a single agent, while R coop represents the reward function for a single agent to have a synergy relationship with others. That is, in R coop , R i represents the above-mentioned R patrol , R attack or R defense .

[0122] Given the high requirements of adversarial games for the decision-making time of agents, it is more suitable to adopt a centralized training and decentralized execution framework (Centralized Training and Decentralized Execution, CTDE). Among them, the value network and the target network are deployed on the central controller, while the policy network is distributed among each agent. That is, in the centralized training stage, all agents share global information for joint training; after the training is completed, the agents use the local policy network for independent decision-making. As the training progresses, the agents gradually optimize their decision-making processes, improving the task completion rate and combat efficiency.

[0123] The training of the policy decision-making model in this embodiment adopts the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm.

[0124] Training the policy decision-making model mainly includes:

[0125] Initializing and defining the feature space (i.e., Figure 1(in the feature space, training model creation, and action system module): Format and store key feature information in the environment, such as enemy dynamics, own status, and battlefield situation. Provide input data for the reinforcement learning model to ensure that during MAPPO training, decision-making strategies can be optimized based on complete battlefield information, optimize state representation, reduce computational complexity, and improve training efficiency.

[0126] The core objective of the MAPPO algorithm is to update the agent's policy by optimizing the loss function, enabling it to make optimal decisions in a multi-agent environment. The loss function of MAPPO includes policy loss, value loss, entropy loss, and global policy coordination loss:

[0127] L(θ) = L clip (θ) - c1L VF (θ) + c2S[π θ (s) + λL coop (θ)

[0128] Among them, L clip (θ) is the policy loss, which mainly reflects the quality of an agent taking a certain action in a specific environment. By adjusting the agent's policy to minimize this loss, the decision-making process can be improved; L VF (θ) is the value loss, which reflects the error between the agent's evaluation of the current state and the true return. In a ship formation, the agent needs to evaluate the value of the current action based on the state (such as fleet distance, enemy position) and optimize it according to the mission objective; S[π θ (s) is the entropy loss, aiming to increase the exploration of the agent and prevent the policy from converging prematurely. In a maritime ship formation, the agent not only has to execute the known optimal policy but also explore new and potentially more efficient tactics. For example, exploring different attack routes; L coop (θ) is the global policy coordination loss. The global coordination loss ensures that each agent's policy is optimized not only based on its own return but also considering interactions with other agents, promoting the coordinated operation of the entire system.

[0129] Specifically, in this embodiment, the policy loss L clip (θ) is:

[0130]

[0131] Among them, r t (θ) represents the ratio of the current policy to the old policy; is the advantage function, used to measure the effect of the current policy relative to the baseline policy. ∈ is the clipping threshold, used to control the magnitude of policy update. The function is used to limit r tThe value range of (θ) is used to avoid excessive policy updates and ensure the stability of the training process. It is defined as: if r t (θ) > 1 + ∈, then clip r t (θ) to 1 + ∈; if r t (θ) < 1 - ∈, then clip r t (θ) to 1 - ∈; if r t (θ) is within the range of [1 - ∈, 1 + ∈], then keep r t (θ) unchanged.

[0132] The value loss L VF (θ) is:

[0133]

[0134] where V θ (s t ) is the agent's value evaluation of the s t state, and R t is the actual return at state s t .

[0135] The entropy loss S[π θ (s) is:

[0136]

[0137] where π θ (a t |s t ) represents the probability of the action a t output by the policy network at state s t .

[0138] The global policy coordination loss L coop (θ) is:

[0139]

[0140] where N is the set of all agents in the formation.

[0141] When using the hybrid decision-making method to solve the intelligent game problem, the strategy of a single opponent is difficult to provide sufficient intelligent support for the intelligent agent. If only one opponent is selected for training, the intelligent agent will often converge to the strategy for that specific opponent. When the opponent's strategy changes, its game level will drop significantly. To solve the problems of poor algorithm generalization ability and strong dependence on the opponent's strategy, in this embodiment, three different types of expert strategy opponents are used to train the intelligent agent to improve the training effectiveness. They are: The herringbone offensive strategy: Attack in a herringbone (V-shaped) formation, using the method of outflanking from both wings and breaking through the center to trap the enemy target within the attack fire range. It is generally applicable when the enemy fleet has strong defenses and the enemy firepower needs to be weakened by surrounding from both wings. The seven-team offensive strategy: The fleet is divided into seven independent squads, and each squad undertakes different tasks. Adopt the method of decentralized attack and flexible maneuver to strike multiple targets simultaneously. It is applicable to the battlefield environment where the enemy defense system is dispersed or the fleet is relatively loose. The left, middle, and right three-team offensive strategy: The fleet is divided into three groups: the left wing, the middle road, and the right wing, and each part performs different tasks. Through zonal operations, the left and right wings are responsible for outflanking, and the middle road troops serve as the main attacking force. This strategy can adapt to different enemy defense forms and is especially suitable for combating enemy fleets with strong central defenses.

[0142] Embodiment of the computer device:

[0143] A computer device includes a processor that is used to execute a computer program to implement the steps of the above-mentioned method for the offensive and defensive deduction of a maritime ship formation based on multi-agent game and expert knowledge. The specific method for the offensive and defensive deduction of a maritime ship formation based on multi-agent game and expert knowledge has been introduced in sufficient detail in the above-mentioned embodiment of the method for the offensive and defensive deduction of a maritime ship formation based on multi-agent game and expert knowledge, and will not be repeated here.

[0144] Specifically, the processor can be a CPU, or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. The processor can also be a processor that supports the advanced reduced instruction set machine (ARM) architecture.

[0145] The present invention has the following characteristics:

[0146] (1) The present invention is a method for offensive and defensive deduction of naval ship formations based on multi-agent game and expert knowledge. By introducing a multi-agent game model to simulate the dynamic interaction and collaborative decision-making process among ships in a ship formation, it can reproduce the complex behavior patterns in a real formation. This method takes the optimization of the game strategies of multi-agents as the core, simulates the task cooperation and resource allocation mechanisms among ships, can not only accurately reflect the impact of individual decisions on the overall formation, but also reflect the constraints and guidance of the overall behavior of the formation on individual strategies, so as to provide intelligent support for the task planning and optimization of ship formations.

[0147] (2) The present invention constructs a hybrid-driven multi-agent game model architecture by combining two methods: expert knowledge rules and reinforcement learning. Expert knowledge rules are used to construct high-level task decisions. According to whether the enemy is visible or not, the system will select different strategies. For example, if the enemy unit is not visible, the decision-making system will choose to perform a reconnaissance task; if the enemy is visible, it will perform an attack task. Reinforcement learning is mainly used for low-level action decisions. This hybrid decision-making framework can effectively combine the advantages of both, can provide clear guidance through rules, and can use reinforcement learning to improve the system's autonomous learning ability and adaptability.

[0148] (3) The present invention introduces a policy equilibrium loss mechanism in model training. During the training process based on the MAPPO algorithm, the policy equilibrium loss mechanism introduces entropy loss and global coordination loss into the loss function, so that during reinforcement learning training: individual agents can retain a certain degree of autonomy to enable ships to independently avoid attacks or adjust firepower distribution. The overall fleet can still maintain effective tactical coordination to ensure a reasonable formation and firepower distribution of the overall formation.

[0149] (4) The formation simulation results generated by the present invention are closer to the dynamic behavior characteristics of real naval formations, and show clear details of the geometric layout, movement trajectories, and resource collaborative allocation of ship formations in visual displays. There is less noise interference in the simulation data, and the model performance is more stable, providing high-value references for the tactical deduction and strategy optimization of ship formations.

[0150] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still make modifications to the technical solutions described in the foregoing embodiments without creative efforts, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for attack and defense simulation of a maritime fleet based on multi-agent game and expert knowledge, characterized in that: The method includes: S1. Acquire environmental data information; the environmental data information includes formation situation information for indicating the status of enemy and friendly ships; S2. Obtaining a likelihood function value for Bayesian estimation based on the formation situation information combined with pre-designed expert knowledge, and then obtaining an initial task overall strategy for the maritime ship formation through Bayesian estimation; S3. Input the environmental data information into a pre-trained strategy decision model corresponding to the overall strategy of the initial task to obtain action instructions for each intelligent agent in the maritime ship formation.

2. The method for attack and defense simulation of a maritime fleet based on multi-agent game and expert knowledge according to claim 1 is characterized in that: The ship status includes strength, and the strength is obtained according to the ship type and ship class.

3. The method for attack and defense simulation of a maritime fleet based on multi-agent game and expert knowledge according to claim 2 is characterized in that: The overall strategy of the initial mission includes an attack strategy, a defense strategy, and a search strategy for obtaining enemy ship status information; the likelihood function value satisfies: If our strength is higher than the enemy's and the difference between our strength and the enemy's strength is higher than the strength difference threshold, then the likelihood function value corresponding to the attack strategy is the highest; If the strength difference between the enemy and us is not higher than the strength difference threshold, the corresponding likelihood function value of the defense strategy is the highest; If our strength is lower than the enemy's strength and the strength difference between the enemy and us is higher than the strength difference threshold, the corresponding likelihood function value of the search strategy is the highest.

4. The method for attack and defense simulation of a maritime fleet based on multi-agent game and expert knowledge according to claim 3 is characterized in that: The attack strategy includes a target selection rule for selecting an attack target and a weapon allocation rule for selecting a weapon.

5. The method for attack and defense simulation of a maritime fleet based on multi-agent game and expert knowledge according to claim 4 is characterized in that: The target selection rule is to preferentially select targets with high priority for attack, and the priority is determined according to the following method: Among them, P target is the priority; d target T is the distance between the attacking agent and the attacked target; target C is the threat level of the target being attacked; target is the type of the target being attacked; S target is the type of the target ship; w1, w2, w3 and w4 are all weight coefficients.

6. The method for attack and defense simulation of a maritime fleet based on multi-agent game and expert knowledge according to claim 4 is characterized in that: The weapon allocation rules are: If the target is an aircraft carrier or a cruiser, and the distance between the attacking agent and the target is greater than the first threshold, the weapon selected is an anti-ship missile; If the target is a destroyer or a submarine, and the distance between the attacking agent and the target is less than or equal to the first threshold, the weapon selected is a torpedo; If the attacked target is a light frigate or a medium frigate, and the threat level of the attacked target is less than or equal to the second threshold, the weapon selected is artillery fire.

7. The method for attack and defense simulation of a maritime fleet based on multi-agent game and expert knowledge according to claim 3 is characterized in that: The strategy decision model is a reinforcement learning model, and the reward function for training the strategy decision model corresponding to the search strategy includes R patrol : R patrol =R critical +R detect +R fuel R critical =r1×I critical R detect =r2×I detect R fuel =-λ1×(Fuel total -Fuel remain ) Among them, r1 is a positive reward value, indicating the reward for searching the key area; I critical is the indicator function, which is 1 when the agent enters the critical area and 0 otherwise; r2 is the positive reward value, indicating the reward for discovering the enemy target; I detect is an indicator function, which is 1 when the agent finds the enemy target and 0 otherwise; λ1 is a positive penalty coefficient, indicating the penalty intensity of fuel consumption; total is the initial fuel amount; remain is the remaining fuel amount; The reward function for training the strategy decision model corresponding to the attack strategy includes R attack : R attack =R destroy +R fail +R ammo R destroy =r4×I destroy R fail =-r5×I fail R ammo =-λ2×(Ammo total -Ammo remain ) Among them, r4 is a positive reward value, indicating the reward for destroying the enemy target; I destroy is an indicator function, which is 1 when the agent destroys the enemy target, otherwise it is 0; r5 is a positive penalty value, indicating the penalty for not destroying the enemy target; I fail is an indicator function, which is 1 when the agent fails to destroy the enemy target, otherwise it is 0; λ2 is a positive penalty coefficient, indicating the penalty intensity of ammunition consumption; Ammo total Initial ammunition amount; Ammo remain The amount of ammunition remaining; The reward function for training the strategy decision model corresponding to the attack strategy includes R defense : R defense =R defend +R region-protect +R resource-cost R defend =r6×I defend R region-protect =r7×I region-protect R resource-cost =-λ3×(Resource total -Resource remain ) Among them, r6 is a positive reward value, indicating the reward for successfully defending against enemy attacks; I defend is an indicator function, which is 1 when the agent successfully defends against enemy attacks, otherwise it is 0; r7 is a positive reward value, indicating the reward for protecting important areas; I region-protect is an indicator function, which is 1 when the agent protects an important area and 0 otherwise; λ3 is a positive penalty coefficient, indicating the penalty intensity for resource consumption; Resource total is the initial resource amount; Resource remain The remaining resources.

8. The method for attack and defense simulation of a maritime fleet based on multi-agent game and expert knowledge according to claim 1 is characterized in that: The strategy decision model is a reinforcement learning model, and the loss function when training the strategy decision model includes at least one of strategy loss, value loss, entropy loss and global strategy coordination loss.

9. The method for attack and defense simulation of a maritime fleet based on multi-agent game and expert knowledge according to claim 1, characterized in that: The strategy decision model is a reinforcement learning model, and the reward function when training the strategy decision model includes a collaborative reward function: Among them, R coop is the global collaboration reward; w i is the weight of agent i; R i is the local reward of agent i; N is the number of agents.

10. A computer device comprising a processor, characterized in that: The processor is used to execute a computer program to implement the steps of the method for attack and defense simulation of a maritime ship formation based on multi-agent game and expert knowledge as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Cooperative multi-agent control method and device based on value function supervision

    CN115755593A

Cited By

  • Aerial formation strike distribution method and related products

    CN120973024A