Ship Missile Target Allocation Method, Device and Equipment Based on Deep Reinforcement Learning

Through the ship missile target allocation method based on deep reinforcement learning, the Transformer model and Markov decision-making process are used to solve the problem of rapid and efficient missile target allocation in ship air defense and anti-missile operations, and high-yield missile target allocation decisions are achieved.

CN115509190BActive Publication Date: 2025-07-08HUNAN DUNYI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211181219.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-27
Publication Date
2025-07-08
Estimated Expiration
2042-09-27

AI Technical Summary

Technical Problem

The existing technology is difficult to quickly and efficiently solve the problem of missile target allocation in air defense and anti-missile operations of maritime ships. Especially in large-scale real-time confrontation scenarios, traditional algorithms have a long calculation time, intelligent optimization algorithms are prone to fall into local optimality, and rule-based methods have unreliable solutions.

Method used

Using a method based on deep reinforcement learning, a ship missile target allocation model is constructed, a Transformer model's fusion attention mechanism and Markov decision-making process is used, and a strategy gradient method is used to train the model to achieve ship missile target allocation decisions.

Benefits of technology

A rapid and efficient generation of high-yield missile target allocation scheme has been achieved, which improves the timeliness and distribution benefits of missile target allocation, and adapts to complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115509190B_ABST
    Figure CN115509190B_ABST
Patent Text Reader

Abstract

The present application relates to a ship missile target allocation method, device and equipment based on deep reinforcement learning in the technical field of weapon target allocation. The method includes: constructing a mathematical model for ship multi-type missile target allocation, and establishing a Markov decision process composed of quadruples based on this mathematical model; constructing a deep reinforcement learning model integrating an attention mechanism based on the Transformer model; this model is used to implement ship missile target allocation decisions according to the battlefield information known in the current ship situation awareness; training the deep reinforcement learning model using the policy gradient method with a baseline; according to the quadruple information at the current time step in the Markov decision process, using the trained deep reinforcement learning model integrating the attention mechanism to allocate ship missiles. This method can quickly and efficiently generate high-yield missile target allocation schemes, improving the missile target allocation yield and allocation timeliness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of weapon target allocation, and particularly to a ship missile target allocation method, device and equipment based on deep reinforcement learning. Background Technique

[0002] At present, naval ships face serious air threats during combat. As the combat styles and weaponry in the maritime battlefield become increasingly complex, various medium-range ballistic missiles and anti-ship missiles (ASM) flying low over the sea can bring huge blows to naval ships and their formations, posing great challenges to the air defense and anti-missile operations of naval ships. As the main weapon resource for ships to strike or intercept offensive targets, shipborne missiles are costly and limited in number. Therefore, efficiently and reasonably coordinating various missile resources, namely weapon-target allocation (WTA), is very important in the air defense and anti-missile operations of ships and is also a key issue in command and control research.

[0003] The missile-target allocation problem (MTA) for naval ship air defense is a typical WTA problem. Domestic and foreign research mainly focuses on the defense side. Generally, based on situation awareness, the optimization goals are mainly to minimize the loss of key areas and minimize resource consumption. Mathematical models are established considering constraints such as resources and spatial relationships to achieve the efficient utilization of ship missile resources. The WTA problem has two different categories: static weapon-target allocation (SWTA) and dynamic weapon-target allocation (DWTA). SWTA only considers static and instantaneous decisions and does not consider the time dimension; while DWTA takes into account the time windows of missiles and targets and the dynamic process of subsequent decisions from making a decision to whether the target is hit. Generally, the "shoot-look-shoot" mode is adopted for interaction. In theory, multiple stages of target strikes can be carried out in the DWTA problem, and each stage can be regarded as an SWTA problem.

[0004] The WTA is a class of NP-hard problems. Some scholars have applied exact methods such as branch and bound, integer programming, and exhaustive search to solve them. However, since the computational complexity of searching the solution space grows exponentially with the problem size, these methods are only applicable to small-scale problems. With the rapid increase in the number of missiles and targets in modern naval ship air defense and anti-missile operations, using traditional exact algorithms to solve the MTA problem requires a long computational time, and the timeliness of decision-making cannot meet the actual needs. Currently, the ideas for solving the weapon-target assignment problem of ship air defense and anti-missile at home and abroad are mainly divided into two categories: intelligent optimization algorithms and learning-based methods. Intelligent optimization algorithms include genetic algorithms, ant colony algorithms, particle swarm algorithms, neighborhood search methods, and various other hybrid optimization algorithms. However, intelligent optimization algorithms need to design corresponding coding or search structures for various specific scenarios, and there are problems such as being easily trapped in local optima and the timeliness of solution finding being difficult to meet the requirements of air defense and anti-missile operations. In addition, there are also some intelligent optimization methods using rule-based strategies to allocate solutions. This method can generate solutions that conform to the rules relatively quickly, but the rule formulation highly depends on empirical knowledge. Therefore, the quality of the solutions cannot be guaranteed and it is difficult to extend to other complex scenarios.

[0005] Learning-based methods have shown outstanding potential in aspects such as unmanned cluster control decision-making, game confrontation, vehicle path planning, and satellite scheduling, while there is relatively little research on target assignment decision-making in air defense and anti-missile.

[0006] To sum up, how to solve the complex WTA problem with large scale and real-time confrontation in real time and efficiently remains a current hot topic and difficulty. Summary of the Invention

[0007] Based on this, it is necessary to provide a ship missile target assignment method, device, and equipment based on deep reinforcement learning for the above technical problems.

[0008] A ship missile target assignment method based on deep reinforcement learning, the method includes:

[0009] Establish a mathematical model for ship missile target assignment.

[0010] Establish a Markov decision process composed of quadruples according to the mathematical model of ship missile target assignment.

[0011] Construct a deep reinforcement learning model integrating an attention mechanism based on the Transformer model; the deep reinforcement learning model integrating the attention mechanism is used to achieve ship missile target assignment decision-making according to the battlefield information known in the current ship situation awareness.

[0012] Train the deep reinforcement learning model integrating the attention mechanism using the policy gradient method with a baseline.

[0013] According to the quadruple information at the current time step in the Markov decision process, a trained deep reinforcement learning model with a fusion attention mechanism is used to allocate ship missile targets.

[0014] A ship missile target allocation device based on deep reinforcement learning, the device includes:

[0015] A mathematical modeling module for establishing a mathematical model for ship missile target allocation.

[0016] A Markov decision process construction module for establishing a Markov decision process composed of quadruples according to the mathematical model for ship missile target allocation.

[0017] A deep reinforcement learning model construction module with a fusion attention mechanism for constructing a deep reinforcement learning model with a fusion attention mechanism based on the Transformer model; the deep reinforcement learning model with a fusion attention mechanism is used to make ship missile target allocation decisions according to the battlefield information known in the current ship situation awareness.

[0018] A deep reinforcement learning model training module with a fusion attention mechanism for training the deep reinforcement learning model with a fusion attention mechanism using the policy gradient method with a baseline.

[0019] A ship missile target allocation module for allocating ship missile targets using the trained deep reinforcement learning model with a fusion attention mechanism according to the quadruple information at the current time step in the Markov decision process.

[0020] A computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements any of the above methods.

[0021] The above-mentioned ship missile target allocation method, device and equipment based on deep reinforcement learning, the method includes: constructing a mathematical model for ship multi-type missile target allocation, establishing a Markov decision process composed of quadruples based on this mathematical model; constructing a deep reinforcement learning model with a fusion attention mechanism based on the Transformer model; this model is used to make ship missile target allocation decisions according to the battlefield information known in the current ship situation awareness; training the deep reinforcement learning model with a fusion attention mechanism using the policy gradient method with a baseline; allocating ship missile targets using the trained deep reinforcement learning model with a fusion attention mechanism according to the quadruple information at the current time step in the Markov decision process. This method can quickly and efficiently generate high-yield missile target allocation plans, improving the missile target allocation benefit and allocation timeliness. Description of the Drawings

[0022] Figure 1Schematic diagram of anti-aircraft and anti-missile for a maritime ship in an embodiment;

[0023] Figure 2 Flow schematic diagram of a ship missile target allocation method based on deep reinforcement learning in an embodiment;

[0024] Figure 3 Schematic diagram of the structure of a deep reinforcement learning model integrating an attention mechanism based on a Transformer model in an embodiment;

[0025] Figure 4 Calculation results of different methods in another embodiment;

[0026] Figure 5 Training convergence of different models in an embodiment, where (a) is the network after replacing the Decoder part in the CAMDRL model with LSTM, (b) is a traditional seq2seq network, 4(c) is the CAMDRL model of the present invention, and (d) is the convergence graph of the first 1000 episodes of the three networks;

[0027] Figure 6 Profit situation of the CAMDRL model at different ratios in another embodiment;

[0028] Figure 7 Structure block diagram of a ship missile target allocation device based on deep reinforcement learning in an embodiment;

[0029] Figure 8 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0030] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0031] The schematic diagram of anti-aircraft and anti-missile for a maritime ship is as Figure 1As shown in the figure. Among them, the defending side is a ship equipped with a Ship to Air Missile (SAM) system. The ship carries various types of SAMs. During the ship's air defense operation, the relative position of the ship remains unchanged. During the air defense and anti-missile operation, the detection radar on the ship is turned on throughout the process, and any missile within the detection range of the ship can be tracked and detected by the ship; the attacking side is the attacking ASM target that comes from the air and enters the interception and detection range of the ship. There are various types of incoming ASMs, and the ASMs fly along a predetermined trajectory until they are intercepted or reach the terminal guidance distance. During the entire air defense and anti-missile process of the ship, the ship can intercept the incoming targets multiple times until all the incoming targets in the scenario are intercepted or the ship is hit, then the battle ends.

[0032] In one embodiment, as Figure 2 shown, a ship missile target allocation method based on deep reinforcement learning is provided, and this method includes the following steps:

[0033] Step 200: Establish a mathematical model for ship missile target allocation.

[0034] Specifically, the mathematical model for ship missile target allocation is used to describe the ship air defense and anti-missile target allocation problem. This model takes into account multi-type missile defenses and combines the dynamic confrontation process in air defense and anti-missile, and considers situations such as interception probability, interception rules, and multiple interceptions to construct the model and scenario of the present invention.

[0035] During the ship's air defense operation, the anti-ship missile flies along a predetermined trajectory until it is intercepted or reaches the terminal guidance distance, while the relative position of the ship remains unchanged. During the air defense and anti-missile operation, the detection radar on the ship is turned on throughout the process, and any missile within the detection range of the ship can be tracked and detected by the ship. During the entire air defense and anti-missile process of the ship, the ship can intercept the incoming targets multiple times until all the incoming targets in the scenario are intercepted or the ship is hit, then the battle ends.

[0036] Step 202: Establish a Markov decision process composed of a quadruple according to the mathematical model for ship missile target allocation.

[0037] Specifically, the air defense task of the ship is to use multi-type shipborne missiles to effectively intercept the attacking side's targets. Considering the adversarial nature during its operation process and the dynamic movement process of the attacking side's targets, the missile-target allocation is regarded as a sequential decision-making process for each target, and then transformed into a Markov decision process in reinforcement learning, so as to use the reinforcement learning method to solve the ship missile target problem.

[0038] Step 204: Construct a deep reinforcement learning model (CAMDRL model) with a fusion attention mechanism based on the Transformer model; the deep reinforcement learning model with a fusion attention mechanism is used to make a decision on ship missile target allocation according to the battlefield information known in the current ship situation awareness.

[0039] Specifically, due to the Transformer model architecture, compared with traditional deep network models such as LSTM and GRU, it can convert sequences into parallel processing through the Attention mechanism, and at the same time make the network attention focus on more important information, so as to better model and process sequence data. Therefore, it is widely used in natural language processing, computer vision, and combinatorial optimization problems. In this embodiment, the Transformer is used to construct a deep reinforcement learning model with a fusion attention mechanism for ship missile target allocation, and its network structure is as Figure 3 shown.

[0040] Step 206: Train the deep reinforcement learning model with a fusion attention mechanism using the policy gradient method with a baseline.

[0041] Specifically, for the reinforcement learning model based on the Transformer architecture, the policy gradient method is usually used for model training, and the commonly used policy gradient method is based on Monte Carlo sampling, which will have a relatively high variance. As a preference, in this embodiment, the idea of fixed Q-targets similar to DQN is used to train the deep reinforcement learning model with a fusion attention mechanism to solve this problem.

[0042] Step 208: Allocate ship missile targets using the trained deep reinforcement learning model with a fusion attention mechanism according to the quadruple information at the current time step in the Markov decision process.

[0043] In the above ship missile target allocation method based on deep reinforcement learning, the method includes: constructing a mathematical model for ship multi-type missile target allocation, establishing a Markov decision process composed of quadruples based on this mathematical model; constructing a deep reinforcement learning model with a fusion attention mechanism based on the Transformer model; the deep reinforcement learning model with a fusion attention mechanism is used to make a decision on ship missile target allocation according to the battlefield information known in the current ship situation awareness; training the deep reinforcement learning model with a fusion attention mechanism using the policy gradient method with a baseline; allocating ship missile targets using the trained deep reinforcement learning model with a fusion attention mechanism according to the quadruple information at the current time step in the Markov decision process. This method can quickly and efficiently generate a high-yield missile target allocation plan, improving the missile target allocation benefit and allocation timeliness.

[0044] In one of the embodiments, step 200 includes: obtaining the number k of types of air defense missiles carried by the current ship, the set D = {D1, D2,..., D k} of the interception distances of each type of air defense missile against the incoming target, the set DP = {dp1, dp2,..., dp k} of the probabilities of each type of air defense missile successfully hitting the incoming target, and the set VS = {vs1, vs2,..., vs k} of the flight speeds of each type of air defense missile; where k is an integer greater than 1, D i in the set D is the interception distance of the i-th type of air defense missile against the incoming target, dp i in the set DP is the probability of the i-th type of air defense missile successfully hitting the incoming target, and vs i in the set VS is the flight speed of the i-th type of air defense missile, i = 1, 2,..., k; obtaining the set N = {1, 2,..., n} of anti-ship missiles of the attacking party that come from the air and enter the ship's interception and detection range, the number g of types of incoming anti-ship missiles, the set VA = {va1, va2,..., va g} of the flight speeds of each type of anti-ship missile, and the set T = {t1, t2,..., t g} of the threat levels of each type of anti-ship missile; where 1 and n in the set D are the first and the n-th incoming anti-ship missiles of the attacking party respectively, va1 in the set D is the flight speed of the first incoming anti-ship missile of the attacking party, and t1 in the set D is the threat level of the first incoming anti-ship missile of the attacking party; during the process of the ship performing the missile target allocation task, taking maximizing the number of successfully intercepted incoming anti-ship missiles and maximizing the resource preservation as the optimization objectives, establishing a ship missile target allocation mathematical model; the ship missile target allocation mathematical model is:

[0045]

[0046] Where: m j is the total number of the j-th type of SAM, pr j is the value coefficient of the j-th type of SAM, r i is the reward corresponding to the successfully intercepted ASM, numerically 10 times the threat level T of the ASM, dp ij is a Boolean variable, which is 1 if the ASM is intercepted and 0 otherwise, a and b are the weighting coefficients corresponding to two sub-goals, and x ij is a decision variable, indicating the number of the j-th type of SAM launched by the ship to intercept the i-th ASM;

[0047] Ship air defense ability constraint model: During the ship's air defense and anti-missile operation process, each interception is restricted by the ship's own missile silo capacity and the number of missiles, and at the same time each type of SAM has its own interception range limit, that is:

[0048]

[0049] where m j is the remaining number of the j - type SAMs on the ship, and α j is the maximum launch capacity of the j - type SAM launch silo on the ship in this wave of interception. is the distance between the ship and the i - th ASM when the ship intercepts the i - th ASM, and D j is the interception distance of the j - type SAM.

[0050] Maximum number of anti - ship missiles to be intercepted constraint model: In the actual air defense and anti - missile process of the ship, excessive missile resources will not be allocated to intercept a single target to avoid excessive waste of resources, that is:

[0051]

[0052] where ω is the maximum number of SAMs intercepted by the ship against the i - th incoming ASM in this interception.

[0053] ASM interception condition constraint model: When the ASM enters within the minimum reaction interception distance of the ship, it cannot be intercepted, and the probability of interception at this time is 0. Otherwise, it can be intercepted, and the interception probability is determined by the SAMs launched, that is:

[0054]

[0055] where p i is the probability of the i - th ASM being intercepted, γ k is the hit probability of the j - type SAM, and D min is the minimum reaction interception distance of the ship.

[0056] In one embodiment, step 202 includes: According to the ship missile target assignment mathematical model, a Markov decision process composed of a quadruple is established as H=(S, A, P, R), where S is the state space, A is the action space, P is the state transition function, and R is the reward function. The Markov decision process is as follows: State space S: The input state information {M, O, V, T, N} of the deep reinforcement learning model integrating the attention mechanism is the battlefield information known in the current ship situation awareness, where M and N are the missile quantity information of both the attacking and defending sides respectively, O is the position matrix of the ship and the incoming targets, V is the flight speed of the missiles of both the attacking and defending sides, and T is the threat degree of the incoming missiles; Action space A: According to the constructed problem model, the dimension d of the action space is designed A ​= ωk + 1, the executable action a ∈ {0, 1, 2..., ωk}, where 0 means not launching SAMs, 1 to ω mean the number of the first type of SAMs launched, and so on; the reward function R: for each successfully intercepted incoming anti - ship missile target, an immediate reward r corresponding to the threat level of the target is obtained, and at the same time, for each consumed missile, a corresponding resource consumption penalty c is also obtained. Thus, the cumulative reward function for intercepting n incoming anti - ship missile targets is obtained: The state - transition function P: The defender selects the action a to be executed through the current policy, that is, s t+1 = π θ (a t |s t ), where t is a certain time step, s t and a t represent the state and action of the defender at the current time step, π θ represents the policy used by the defender when choosing an action, θ is the trainable parameter in the policy network, and as the defender continuously learns, the parameters of the policy network will be optimized accordingly.

[0057] In one embodiment, in step 204, the deep reinforcement learning model based on the fusion attention mechanism of the Transformer model includes an encoder and a decoder.

[0058] In one embodiment, in the encoder, after processing the current state - space information as features, it is used as the initial network input information of the deep reinforcement learning model with the fusion attention mechanism. It is mapped to a high - dimensional space (preferably, the dimension d of the high - dimensional space h = 128) through a linear layer and integrated to obtain the information encoding h i of each target; then, through multiple graph - embedding attention layers, the information encoding h i of each target is processed, and then average pooling is performed on the obtained node encodings of each target to obtain the feature representation of the decoder input data; among them, the graph - embedding attention layer includes: 1 MHA layer, the first Add&Norm layer, 1 feed - forward layer, and the second Add&Norm layer. The feed - forward layer consists of two linear mapping functions with trainable parameters and a ReLU activation function.

[0059] In one embodiment, obtaining the feature representation of the decoder input data is specifically: Let h l be the output of the l - th graph - embedding attention layer. In the l - th graph - embedding attention layer, the MHA layer is based on the multi - head attention mechanism, and according to the input embedding information Generate query vectors, key vectors, and value vectors; calculate the weight information through softmax based on the query vectors, key vectors, and value vectors to obtain the output z of each head of attention. l , and then map the outputs of each head of attention to the output of the MHA layer; the specific calculation process is as follows:

[0060] q l = W l Q h l , k l = W l K h l , v l = W l V h l

[0061]

[0062]

[0063] In the formula, W l Q and W l K are parameter matrices with dimensions of Y * d k * d h , W l V is a parameter matrix with dimensions of Y * d v * d h ; is a trainable parameter matrix with dimensions of d h * d h corresponding to the l-th layer;

[0064] Input the output of the MHA layer into the first Add&Norm layer, and obtain through residual network and normalization processing as the input of the forward feedback layer, and input the obtained information into the second Add&Norm layer for processing to obtain the final target node encoding Input the final target node encoding and the mean encoding obtained by mean pooling of the final target node encoding as the feature representation of the decoder input data; the specific calculation process is as follows:

[0065]

[0066]

[0067]

[0068] Among them, FF(·) is the forward feedback layer function; BN(·) is batch normalization.

[0069] In one embodiment, the encoder includes: a first linear layer, a second linear layer, multiple first forward feedback layers, a Mask layer, and a fully connected layer; in the decoder: the last target node is encoded and the SAM quantity information of the ship in the state space are combined and mapped into a high-order global feature encoding through the first linear layer; when decoding at each time step, the offensive target node encoding information at this time step and the global feature encoding are combined and mapped into a feature vector that fuses global and local node information through the second linear layer; the feature vector is input into the first forward feedback layer, and the obtained output is passed through the Mask layer to delete the solutions that do not meet the constraint conditions, and finally the actionable action encoding u is obtained through the fully connected layer t , so as to obtain the current action probability distribution p at the current time step t t , and the calculation formula of the current action probability distribution is:

[0070] p t = softmax(u t ).

[0071] According to the current action probability distribution p t , the final action is obtained by adopting a greedy or probabilistic sampling strategy. In Figure 3 , Context is the SAM quantity information of the ship in the state space.

[0072] In one embodiment, step 206 includes: constructing a baseline network using greedy measurement; the baseline network is similar in structure to the deep reinforcement learning model integrating the attention mechanism, but with different parameters; generating training samples according to a preset policy; the preset policies include: Policy 1 and Policy 2; Policy 1 is to randomly generate the offensive target positions at different distances from the ship to simulate the battlefield situations under different interception batches; Policy 2 is to generate different training samples according to the principle of long, medium, and short distances of the missiles of both the attacking and defending sides during the movement process; training the baseline network and the deep reinforcement learning model integrating the attention mechanism respectively according to the training samples to obtain the policy gradient, and adjusting the parameters according to the policy gradient; the calculation formula of the policy gradient is:

[0073]

[0074] Among them, is the policy gradient, L(π) is the cumulative reward function of the deep reinforcement learning model integrating the attention mechanism, b(π) is the cumulative reward function of the baseline network, πθ It represents the strategy used by the defending side when choosing actions, and θ is the trainable parameter in the deep reinforcement learning model with a fusion attention mechanism.

[0075] It should be understood that although Figure 2 the steps in the flowchart of Figure 2 are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover,

[0076] In a verification embodiment, to test the performance of the deep reinforcement learning model with a fusion attention mechanism (CAMDRL), this embodiment conducts tests on multiple instances of different scales by constructing a data generator to provide simulation instances of different scales. The generator takes parameters such as the total number N of incoming ASMs and the defense ratio ρ as inputs, and then outputs instances including SAMs, ASMs, and their respective attributes. Among them, ρ refers to the ratio of the number of ASM targets to the number of our ship's SAMs. For example, if N = 20 and ρ = 3, then 20 ASMs and 60 SAMs are generated in the instance. The specific attributes of various ASMs and SAMs are shown in Tables 1 and 2. The simulation experiment parameter settings are shown in Table 3. All simulation experiments are carried out on a Win10 PC configured with an Intel i7-10700K, RTX3070 GPU, and 2 * 32G memory, and are compiled and developed based on Python 3.8 and PyCharm 2020.2.

[0077] Table 1 SAM Attributes

[0078] Type Damage Probability (one / two) Interception Range / km Speed / Mach SAM1 0.75 / 0.92 0-900 2 SAM2 0.7 / 0.95 360-900 2 SAM3 0.62 / 0.84 0-450 2

[0079] Table 2 ASM Attribute Information

[0080] Type Damage Probability Speed / Mach Threat Level ASM1 0.8 3.2 1 ASM2 0.82 3.5 2

[0081] Table 3 Simulation Experiment Parameter Table

[0082]

[0083] In this embodiment, different experimental scenarios are constructed based on the scale of incoming ASMs being 20, 40, 60, 80, and 100 respectively. Among them, the defense ratio ρ = {2, 2.5, 3, 3.5, 4, 4.5} is set for each scenario. At the same time, the Heuristic Rule - Based Algorithm (HRA) and the Neighborhood Search Algorithm Based on Feasible Action (FANS) are used as comparison methods to compare with CAMDRL - S1 and CAMDRL - S2. Among them, HRA first constructs feasible solutions according to the common long - range and short - range interception rules in air defense and antimissile, that is, based on the distance between the incoming target and the ship, it gives priority to intercepting the target with a shorter distance, and then based on the threat level of the target, it gives priority to intercepting the target with a higher threat level. The FANS method generates an initial solution by selecting SAMs for each target in the order of target numbers, adds feasible action constraints in the neighborhood operator, and only generates neighborhood solutions that meet the interception constraints, which speeds up the solution - finding speed of the neighborhood search method. The maximum number of iterations of the FANS method is set to 1000. CAMDRL - S1 and CAMDRL - S2 are deep reinforcement learning methods that incorporate an attention mechanism. Both use the policy network and training method proposed in this paper. The difference lies in the method of generating the target point coordinates of the training samples. The former generates sample targets based on S1, that is, randomly generates target point coordinates within the interceptable range of the ship; the latter generates targets based on S2, generates target points according to the three interception ranges of far, medium, and near respectively, and then trains the three types of samples of far, medium, and near in turn. The parameter settings of the CAMDRL model are shown in Table 5. Based on the above settings, the four models are each run 20 times in each group of experiments. Since it is better to eliminate danger earlier in combat, the reward decay coefficient is set to 0.8, and the decay coefficient is multiplied in the subsequent interception wave rewards. Therefore, the final target value for each instance is:

[0084]

[0085] where ki represents the interception wave number when intercepting the i - th ASM. The average result of the test is used as the final basis for comparison, as shown in Table 4 and Figure 4 as follows.

[0086] Table 4 Comparison of Experimental Results of Methods

[0087] Method FANS HRA CAMDRL-S1 CAMDRL-S2 Method FANS HRA CAMDRL-S1 CAMDRL-S2 N20ρ2 208.71 208.04 211.26 211.66 N60ρ3.5 701.81 768.10 767.53 766.81 N20ρ2.5 215.48 225.02 227.36 227.20 N60ρ4 703.47 765.71 766.60 765.43 N20ρ3 213.34 236.60 237.34 237.19 N60ρ4.5 697.32 765.95 766.08 763.92 N20ρ3.5 196.64 212.35 213.30 212.62 N80ρ2 906.08 884.85 924.86 929.73 N20ρ4 242.89 265.68 266.81 267.28 N80ρ2.5 901.56 948.50 961.71 962.27 N20ρ4.5 224.93 243.31 244.06 243.34 N80ρ3 971.09 1053.02 1055.85 1054.46 N40ρ2 427.99 415.03 428.12 431.80 N80ρ3.5 895.03 974.08 975.84 969.77 N40ρ2.5 429.90 456.95 457.45 457.20 N80ρ4 999.12 1084.43 1081.25 1081.50 N40ρ3 452.75 490.94 491.09 489.83 N80ρ4.5 957.46 1041.05 1041.06 1040.17 N40ρ3.5 478.00 519.41 519.94 519.27 N100ρ2 1066.36 1010.76 1071.28 1078.37 N40ρ4 429.34 463.98 466.28 464.76 N100ρ2.5 1122.13 1186.00 1197.44 1202.85 N40ρ4.5 447.17 487.35 486.34 486.85 N100ρ3 1153.31 1256.25 1260.24 1257.92 N60ρ2 717.84 683.63 745.51 733.04 N100ρ3.5 1259.75 1371.41 1373.06 1371.59 N60ρ2.5 678.28 717.72 721.04 717.22 N100ρ4 1143.07 1249.16 1246.47 1248.56 N60ρ3 646.48 709.06 711.11 709.67 N100ρ4.5 1149.81 1258.61 1258.83 1259.52

[0088] Table 5 Model Parameter Settings

[0089] Parameter Setting Batch size 256 Epoch size 2560000 Learning Rate 10-4 Number of Epochs 3 Optimizer Adam Embedding-layer Dimension 128 Hidden-layer Dimension 128 Output-layer Dimension 7

[0090] The experimental results show that when solving the problem of ship air defense and antimissile, at each scale N, as ρ increases, the gap between the FANS method and other methods gradually becomes obvious, and as N increases, this gap becomes more significant. This is because as the problem scale increases, the solution space expands rapidly, and the optimization performance of the FANS method deteriorates. HRA can obtain relatively optimal solutions in most cases, especially when ρ is 4 and 4.5, it performs optimally several times, but it performs the worst several times when ρ is 2. This rule-based heuristic method can make the decision with the maximum current interception probability when solving each incoming target, but it also causes the method to always allocate more high-value missiles to high-value targets. In an adversarial environment with a low ρ value, that is, relatively insufficient SAMs, HRA seems not very applicable. The CAMDRL model performs better than other methods in the vast majority of cases, among which CAMDRL-S1 accounts for 63.3% and CAMDRL-S2 accounts for 26.7%. In all experiments, the situation where CAMDRL-S1 is better than CAMDRL-S2 accounts for 26.7%, and it is most prominent when solving the problem of ship air defense and antimissile. This may be because the instances of S1 are randomly generated in the whole domain. Compared with the way of generating instances of S2 in the far, medium, and near ranges, the policy network can learn the influence of global factors on the results.

[0091] In addition, the experimental time efficiency of different methods was also analyzed. As shown in Table 6, the average running speeds required for decision-making by different methods under 20 and 200 different sample instance numbers are given. It can be seen that when making decisions with a small number of samples, the speed of HRA is much faster than that of FANS and CAMDRL. This rule-based and heuristic method has great advantages when making small-sample decisions because it does not require too much complex calculation. However, in the case of a large number of samples, since deep reinforcement learning can make batch processing decisions on problems, CAMDRL shows certain efficiency advantages (10% - 13%) in large-sample decisions.

[0092] Table 6 Average running speeds of four models under different sample instance numbers

[0093]

[0094] To further compare the training convergence speed of the policy network structure in this paper, the policy network in this paper is compared with the traditional policy network. In the experiment, all experimental parameters and method model hyperparameters are set the same except for the deep network structure. The convergence of different policy network structures is as Figure 5As shown in the figure, where (a) is the network after replacing the Decoder part in the CAMDRL model with LSTM, denoted as Neta, (b) is the traditional seq2seq network, with both encoding and decoding using LSTM, denoted as Netb; 4(c) is the CAMDRL model of the present invention, denoted as Netc, and (d) is the convergence graph of the first 1000 episodes of the three networks. As Figure 5 As shown in (d) of the figure, the CAMDRL model in this embodiment converges significantly faster than the other two networks. When the episode is about 130, it converges stably from the interception benefit of 230 to about 283. Neta reaches convergence at about 410, and Netb reaches stable convergence at about 520. In addition, the convergence objective function value of the CAMDRL model of the present invention is also better than that of the other two networks.

[0095] Finally, the decision-making benefits of the CAMDRL-S1 algorithm were also analyzed under the condition that the number of incoming missiles N = 20 and different defense ratios ρ. The settings of the ship's missiles and the attacker's target attributes are the same as those in Tables 1 and 2. The decision-making benefits are evaluated through the target rate of return, and the calculation formula of the target rate of return is as follows:

[0096]

[0097] As Figure 6 shown in the figure, as ρ increases, the convergence value of the objective function also continuously increases. Among them, the increase is more obvious when ρ is 2.5 and 3. When it goes further up, the growth rate is not significant, but instead the target rate of return drops significantly. It shows that in the current confrontation environment, when ρ = 2, the sub-objective function value that maximizes the number of successfully intercepted incoming missiles in the objective function is the main body, and at this time, the interception benefit of SAM is the largest. As ρ increases, the proportion of the function value that maximizes the sub-objective of resource preservation gradually increases. When ρ is about 3, all incoming targets can be intercepted. At this time, when ρ increases further, the increase in the objective function value is only the function value part of the SAM resource sub-objective, so the increase rate of the objective function value is small, and the target rate of return drops significantly instead. It can be inferred from this that in the current confrontation environment, for a ship to achieve the effect of saturation interception, it needs to be equipped with at least about 3 times the number of SAM resources as the incoming missiles.

[0098] In one embodiment, as Figure 7 shown in the figure, a ship missile target allocation device based on deep reinforcement learning is provided, including: a mathematical modeling module, a Markov decision process construction module, a deep reinforcement learning model construction module integrating an attention mechanism, a deep reinforcement learning model training module integrating an attention mechanism, and a ship missile target allocation module, where:

[0099] The mathematical modeling module is used to establish a mathematical model for ship missile target allocation;

[0100] A Markov decision process construction module, which is used to establish a Markov decision process composed of quadruples according to the mathematical model of ship missile target allocation;

[0101] A deep reinforcement learning model construction module integrating an attention mechanism, which is used to construct a deep reinforcement learning model integrating an attention mechanism based on the Transformer model; the deep reinforcement learning model integrating an attention mechanism is used to realize the ship missile target allocation decision according to the battlefield information known in the current ship situation awareness;

[0102] A deep reinforcement learning model training module integrating an attention mechanism, which is used to train the deep reinforcement learning model integrating an attention mechanism by using the policy gradient method with a baseline;

[0103] A ship missile target allocation module, which is used to allocate ship missiles to targets by using the trained deep reinforcement learning model integrating an attention mechanism according to the quadruple information at the current time step in the Markov decision process.

[0104] In one embodiment, the mathematical modeling module is further used to obtain the number k of types of air defense missiles carried by the current ship, the set D = {D1, D2,..., D k} of the interception distances of each type of air defense missile against incoming targets, the set DP = {dp1, dp2,..., dp k} of the probabilities of each type of air defense missile successfully hitting incoming targets, and the set VS = {vs1, vs2,..., vs k}; where k is an integer greater than 1, D i in the set D is the interception distance of the i-th type of air defense missile against incoming targets, dp i in the set DP is the probability of the i-th type of air defense missile successfully hitting incoming targets, and vs i in the set VS is the flight speed of the i-th type of air defense missile, i = 1, 2,..., k; obtain the set N = {1, 2,..., n} of anti-ship missiles of the attacking party that come from the air and enter the ship's interception and detection range, the number g of types of incoming anti-ship missiles, the set VA = {va1, va2,..., va g} of the flight speeds of each type of anti-ship missile, and the set T = {t1, t2,..., t g}; where in set D, 1 and n are respectively the 1st and the nth incoming anti-ship missiles of the attacking side, va1 in set D is the flight speed of the 1st incoming anti-ship missile of the attacking side, and t1 in set D is the threat level of the 1st incoming anti-ship missile of the attacking side; during the process of the ship performing the missile target assignment task, a mathematical model of ship missile target assignment is established with the optimization goals of maximizing the number of successfully intercepted incoming anti-ship missiles and maximizing resource preservation; the mathematical model of ship missile target assignment is:

[0105]

[0106] where: m j is the total number of the jth type of SAM, pr j is the value coefficient of the jth type of SAM, r i is the reward corresponding to the successfully intercepted ASM, numerically 10 times the threat level T of the ASM, dp ij is a Boolean variable, which is 1 if the ASM is intercepted, and 0 otherwise. a and b are the weighting coefficients corresponding to two sub-goals, x ij is a decision variable, representing the number of the jth type of SAM launched by the ship to intercept the ith ASM;

[0107] Air defense capability constraint model of the ship:

[0108]

[0109] where, m j is the remaining number of the jth type of SAM on the ship, α j is the maximum launch capability of the jth type of SAM launch well on the ship in this wave of interception. is the distance between the ship and the ASM when the ship intercepts the ith ASM, D j is the interception distance of the jth type of SAM.

[0110] Maximum number of executed incoming anti-ship missiles constraint model:

[0111]

[0112] where, ω is the maximum number of SAMs launched by the ship against the ith incoming ASM in this interception.

[0113] ASM interception condition constraint model:

[0114]

[0115] where, p i is the probability that the ith ASM is intercepted, γ k is the hit probability of the jth type of SAM, D minis the minimum reaction interception distance of the ship.

[0116] In one embodiment, the Markov decision process construction module is further configured to establish a Markov decision process composed of a quadruple H = (S, A, P, R) according to the ship missile target allocation mathematical model, where S is the state space, A is the action space, P is the state transition function, and R is the reward function. The Markov decision process is specifically as follows: State space S: The input state information {M, O, V, T, N} of the deep reinforcement learning model integrating the attention mechanism is the battlefield information known in the current ship situation awareness, where M and N are the missile quantity information of the attacking and defending sides respectively, O is the position matrix of the ship and the incoming target, V is the flight speed of the missiles of the attacking and defending sides, and T is the threat level of the incoming missiles; Action space A: According to the constructed problem model, design the action space dimension d A = ωk + 1, and the executable action a ∈ {0, 1, 2..., ωk}, where 0 means not launching SAM, 1 to ω mean the number of the first type of SAM launched, and so on; Reward function R: Obtain an immediate reward r of the corresponding target threat level for each successfully intercepted incoming anti-ship missile target. At the same time, a corresponding resource consumption penalty c will also be obtained for each consumed missile. Thus, the cumulative reward function for intercepting n incoming anti-ship missile targets is obtained: State transition function P: The defense side selects the action a to be executed through the current policy, that is, s t+1 = π θ (a t |s t ), where t is a certain time step, s t and a t represent the state and action of the defense side at the current time step, π θ represents the policy used by the defense side when selecting actions, θ is the trainable parameter in the policy network, and as continuous learning progresses, the parameters of the policy network will be optimized accordingly.

[0117] In one embodiment, the deep reinforcement learning model integrating the attention mechanism in the deep reinforcement learning model construction module includes an encoder and a decoder based on the Transformer model.

[0118] In one embodiment, in the encoder, the deep reinforcement learning model construction module integrating the attention mechanism is further configured to perform feature processing on the current state space information and use it as the initial network input information of the deep reinforcement learning model integrating the attention mechanism. It is mapped to a high-dimensional space (preferably, the high-dimensional space dimension d h = 128) through a linear layer and integrated to obtain the information encoding h i of each target; Then, through multiple graph embedding attention layers, the information encoding h of each targeti Process them, and then encode the nodes of each obtained target Perform average pooling to obtain the feature representation of the decoder input data; among them, the graph embedding attention layer includes: 1 MHA layer, the first Add&Norm layer, 1 feed-forward layer, and the second Add&Norm layer. The feed-forward layer consists of two linear mapping functions with trainable parameters and a ReLU activation function.

[0119] In one embodiment, the deep reinforcement learning model construction module with a fusion attention mechanism is further configured to make h l Be the output of the l-th graph embedding attention layer. In the l-th graph embedding attention layer, the MHA layer is based on the multi-head attention mechanism and generates query vectors, key vectors, and value vectors according to the input embedding information Generate query vectors, key vectors, and value vectors; calculate weight information through softmax based on the query vectors, key vectors, and value vectors to obtain the output z of each head of attention l , and then map the outputs of each head of attention into the output of the MHA layer; its specific calculation process is as follows:

[0120] q l = W l Q h l , k l = W l K h l , v l = W l V h l

[0121]

[0122]

[0123] In the formula, W l Q and W l K Are parameter matrices with dimensions of Y * d k * d h , W l V Is a parameter matrix with dimensions of Y * d v * d h , Is a trainable parameter matrix with dimensions of d h * d h corresponding to the l-th layer; d q 、d k 、dv are the dimensions of query vector, key vector and value vector respectively. The number of attention network heads of MHA layer is Y=8.

[0124] The output of the MHA layer is input into the first Add&Norm layer and then processed by the residual network and normalization to obtain As the input of the forward feedback layer, the obtained information is input into the second Add&Norm layer for processing to obtain the final target node encoding Encode the final target node And encode the final target node Mean encoding obtained by mean pooling As the feature representation of the decoder input data; its specific calculation process is:

[0125]

[0126]

[0127]

[0128] Among them, FF(·) is the forward feedback layer function; BN(·) is batch normalization.

[0129] In one embodiment, the encoder includes: a first linear layer, a second linear layer, a plurality of first forward feedback layers, a Mask layer, and a fully connected layer; in the decoder: a deep reinforcement learning model building module integrating an attention mechanism, which is also used to encode the final target node After being combined with the SAM quantity information of the ship in the state space, it is mapped into a high-order global feature code through the first linear layer; when decoding at each time step, the attacking target node encoding information h of the time step is i l and the global feature encoding h l The combined and mapped into a feature vector that integrates global and local node information through the second linear layer; the feature vector is input into the first forward feedback layer, and the output is passed through the Mask layer to delete the solutions that do not meet the constraints, and finally the action code u is obtained through the fully connected layer. t , thus obtaining the current action probability distribution p at the current time step t t , the calculation formula of the current action probability distribution is:

[0130] p t =softmax(u t ).

[0131] According to the current action probability distribution p t , using a greedy or probabilistic sampling strategy to get the final action.

[0132] In one embodiment, the deep reinforcement learning model training module integrating the attention mechanism is further configured to construct a baseline network using greedy measurement; the baseline network has a similar structure to the deep reinforcement learning model integrating the attention mechanism, but different parameters; generate training samples according to a preset policy; the preset policy includes: Policy 1 and Policy 2; Policy 1 is to randomly generate the target positions of the attacking party at different distances from the ship to simulate the battlefield situation under different interception batches; Policy 2 is to generate different training samples according to the principle of long, medium, and short distances of the missiles of both the attacking and defending sides during the movement process; train the baseline network and the deep reinforcement learning model integrating the attention mechanism respectively according to the training samples to obtain the policy gradient, and adjust the parameters according to the policy gradient; the formula for calculating the policy gradient is:

[0133]

[0134] wherein, is the policy gradient, L(π) is the cumulative reward function of the deep reinforcement learning model integrating the attention mechanism, b(π) is the cumulative reward function of the baseline network, and π θ represents the policy used by the defending party when selecting an action, and θ is the trainable parameter in the deep reinforcement learning model integrating the attention mechanism.

[0135] The baseline network uses a greedy policy, selects the action with the highest probability as the decision result each time to obtain the reward, and uses the evaluation result of this baseline network as the baseline to reduce the generated variance.

[0136] For the specific limitations of the ship missile target allocation device based on deep reinforcement learning, reference can be made to the limitations of the ship missile target allocation method based on deep reinforcement learning in the above text, which will not be elaborated here. Each module in the above ship missile target allocation device based on deep reinforcement learning can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0137] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 8As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes a method for ship missile target allocation based on deep reinforcement learning. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad set on the shell of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0138] Those skilled in the art can understand that Figure 8 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0139] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it realizes the method steps in the above method embodiment.

[0140] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0141] The above embodiments only represent several implementation manners of this application. Their descriptions are relatively specific and detailed, but they should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of the patent of this application should be subject to the appended claims.

Claims

1. A ship missile target allocation method based on deep reinforcement learning, characterized in that, The method includes: Establishing a mathematical model for ship missile target allocation; Establishing a Markov decision process consisting of quadruples according to the mathematical model for ship missile target allocation; Constructing a deep reinforcement learning model integrating an attention mechanism based on the Transformer model; the deep reinforcement learning model integrating the attention mechanism is used to make ship missile target allocation decisions according to the battlefield information known in the current ship situation awareness; Training the deep reinforcement learning model integrating the attention mechanism by using the policy gradient method with a baseline; Allocating ship missiles to targets by using the trained deep reinforcement learning model integrating the attention mechanism according to the quadruple information at the current time step in the Markov decision process.

2. The method according to claim 1, wherein Establishing a mathematical model for ship missile target allocation includes: Obtain the number k of types of air defense missiles carried by the current ship, the set D = {D1, D2,..., D k} of the interception distances of each type of air defense missile against the incoming target, the set DP = {dp1, dp2,..., dp k} of the probabilities of each type of air defense missile successfully hitting the incoming target, and the set VS = {vs1, vs2,..., vs k}; where k is an integer greater than 1, D i in the set D is the interception distance of the i-th type of air defense missile against the incoming target, dp i in the set DP is the probability of the i-th type of air defense missile successfully hitting the incoming target, and vs i in the set VS is the flight speed of the i-th type of air defense missile, i = 1, 2,..., k; Obtain the set N = {1, 2,..., n} of attacking anti-ship missiles that the ship needs to defend against incoming from the air and enter its interception detection range, the number g of types of incoming anti-ship missiles, the set VA = {va1, va2,..., va g} of flight speeds of each type of anti-ship missile, and the set T = {t1, t2,..., t g} of threat levels of each type of anti-ship missile; where 1 and n in the set D are the first and nth incoming attacking anti-ship missiles respectively, va1 in the set D is the flight speed of the first incoming attacking anti-ship missile, and t1 in the set D is the threat level of the first incoming attacking anti-ship missile; During the process of a ship performing the missile target allocation task, taking maximizing the number of incoming anti-ship missiles successfully intercepted and maximizing resource conservation as the optimization objectives, and establishing a mathematical model for ship missile target allocation; the mathematical model for ship missile target allocation is: where: m j is the total quantity of the j-th type of SAM, pr j is the value coefficient of the j-th type of SAM, r i is the reward corresponding to the successfully intercepted ASM, numerically 10 times the threat level T of the ASM, dp ij is a Boolean variable, which is 1 if the ASM is intercepted and 0 otherwise, a and b are the weighting coefficients corresponding to two sub-goals, x ij is a decision variable, representing the quantity of the j-th type of SAM launched by the ship to intercept the i-th ASM; Air defense capability constraint model of the ship: where m j is the remaining number of the j-th type of SAM on the ship, and α j is the maximum launch capacity of the j-th type of SAM silo on the ship during this wave of interception; is the distance between the ship and the ASM when the ship intercepts the i-th ASM, and D j is the interception distance of the j-th type of SAM; Maximum number of executed incoming anti-ship missiles constraint model: Where ω is the maximum number of SAMs intercepted by the ship for the i-th incoming ASM in this interception; ASM interception condition constraint model: where p i is the probability that the i-th ASM is intercepted, γ k is the hit probability of the j-th type of SAM, and D min is the minimum reaction interception distance of the ship.

3. The method according to claim 1, characterized in that, Establishing a Markov decision process consisting of quadruples according to the mathematical model for ship missile target allocation includes: According to the mathematical model for ship missile target allocation, establishing a Markov decision process consisting of quadruples as H=(S, A, P, R), where S is the state space, A is the action space, P is the state transition function, and R is the reward function. The Markov decision process is specifically as follows: State space S: The input state information of the deep reinforcement learning model integrating the attention mechanism { M,O,V,T,N} is the battlefield information known in the current ship situation awareness, where M and N are the missile quantity information of the attacking and defending sides respectively, O is the position matrix of the ship and the incoming targets, V is the flight speed of the missiles of the attacking and defending sides, and T is the threat degree of the incoming missiles; Action space A: Design the dimension d of the action space according to the constructed problem model A = ωk + 1, and the executable action a ∈ {0, 1, 2..., ωk}, where 0 means not launching SAM, 1 to ω mean the number of the first type of SAM launched, and so on; Reward function R: For each successfully intercepted incoming anti-ship missile target, an immediate reward r corresponding to the threat level of the target is obtained. At the same time, for each consumed missile, a corresponding resource consumption penalty c is also incurred. Thus, the cumulative reward function for intercepting n incoming anti-ship missile targets is obtained: State transition function P: The defender selects the action a to be executed through the current policy, i.e., s t+1 = π θ (a t | s t ), where at a certain time step, s t and a t represent the state and action of the defender at the current time step, π θ represents the policy used by the defender when selecting an action, θ are the trainable parameters in the policy network, and as learning progresses, the parameters of the policy network will be optimized accordingly.

4. The method according to claim 1, wherein Constructing a deep reinforcement learning model integrating an attention mechanism based on the Transformer model. In the steps, the deep reinforcement learning model integrating the attention mechanism based on the Transformer model includes an encoder and a decoder.

5. The method according to claim 4, characterized in that, In the encoder, After processing the current state space information as the initial network input information of the deep reinforcement learning model with a fusion attention mechanism, it is mapped to a high-dimensional space and integrated through a linear layer to obtain the information encoding h of each target. i ; Then, the information encoding h of each target is processed through multiple graph embedding attention layers. i After that, average pooling is performed on the obtained node encodings of each target. to obtain the feature representation of the decoder input data; among them, the graph embedding attention layer includes: 1 MHA layer, the first Add&Norm layer, 1 feed-forward layer, and the second Add&Norm layer, and the feed-forward layer consists of two linear mapping functions with trainable parameters and a ReLU activation function.

6. The method according to claim 5, wherein Obtaining the feature representation of the decoder input data, specifically: Let h l be the output of the l-th graph embedding attention layer. In the l-th graph embedding attention layer, the MHA layer is based on the multi-head attention mechanism and generates query vectors, key vectors, and value vectors according to the input embedding information ; calculates the weight information through softmax based on the query vectors, key vectors, and value vectors to obtain the output z of each head of attention l , and then combines and maps the output of each head of attention into the output of the MHA layer; its specific calculation process is as follows: q l = W l Q h l , k l = W l K h l , v l = W l V h l where, W l Q and W l K are parameter matrices of dimension Y * d k * d h , and W l V is a parameter matrix of dimension Y * d v * d h . is a trainable parameter matrix corresponding to the dimension d h * d h of the l-th layer. d q , d k , and d v are the dimensions of the query vector, key vector, and value vector respectively. Y is the number of attention network heads in the MHA layer, and d h is the dimension of the high-dimensional space; The output of the MHA layer is input into the first Add&Norm layer, and after residual network and normalization processing, we get which is used as the input of the forward feedback layer, and the obtained information is input into the second Add&Norm layer for processing to obtain the final target node encoding The final target node encoding and the mean encoding obtained by performing mean pooling on the final target node encoding are used as the feature representation of the decoder input data; ​ Its specific calculation process is: Where FF(·) is the forward feedback layer function; BN(·) is batch normalization.

7. The method according to any one of claims 4 to 6, characterized in that The encoder includes: a first linear layer, a second linear layer, multiple first forward feedback layers, a Mask layer, and a fully connected layer; In the decoder: Encode the last target node After combining with the SAM quantity information of the ships in the state space, it is mapped into a high-order global feature encoding through the first linear layer; when decoding at each time step, the attacking party target node encoding information at this time step and the global feature encoding are combined and mapped into a feature vector that fuses global and local node information through the second linear layer; the feature vector is input into the first forward feedback layer, and the obtained output is passed through the Mask layer to delete the solutions that do not meet the constraint conditions, and finally the actionable action encoding u t is obtained, so as to obtain the current action probability distribution p t at the current time step t. The calculation formula of the current action probability distribution is: p t = softmax(u t ) According to the current action probability distribution p t , the final action is obtained by using a greedy or probabilistic sampling strategy.

8. The method according to claim 1, wherein Training the deep reinforcement learning model integrating the attention mechanism by using the policy gradient method with a baseline includes: Constructing a baseline network using greedy measurement; the baseline network has a similar structure to the deep reinforcement learning model integrating the attention mechanism, but different parameters; Generating training samples according to a preset policy; the preset policy includes: Policy 1 and Policy 2; Policy 1 is to randomly generate the target positions of the attacking party at different distances from the ship to simulate the battlefield situation under different interception batches; Policy 2 is to generate different training samples according to the principle of long, medium, and short distances of the missiles of both the attacking and defending sides during the movement process; Training the baseline network and the deep reinforcement learning model integrating the attention mechanism respectively according to the training samples to obtain the policy gradient, and adjusting the parameters according to the policy gradient; the policy gradient calculation formula is: Among them, is the policy gradient, L(π) is the cumulative reward function of the deep reinforcement learning model integrating the attention mechanism, b(π) is the cumulative reward function of the baseline network, and π θ represents the policy used by the defender when choosing an action, and θ is the trainable parameter in the deep reinforcement learning model integrating the attention mechanism.

9. A ship missile target allocation device based on deep reinforcement learning, characterized in that, The device includes: A mathematical modeling module for establishing a mathematical model for ship missile target allocation; A Markov decision process construction module for establishing a Markov decision process composed of quadruples according to the mathematical model for ship missile target allocation; A deep reinforcement learning model construction module integrating an attention mechanism for constructing a deep reinforcement learning model integrating an attention mechanism based on the Transformer model; the deep reinforcement learning model integrating an attention mechanism is used to make ship missile target allocation decisions according to the battlefield information known in the current ship situation awareness; A training module for the deep reinforcement learning model integrating an attention mechanism for training the deep reinforcement learning model integrating an attention mechanism by using the policy gradient method with a baseline; A ship missile target allocation module for allocating ship missile targets by using the trained deep reinforcement learning model integrating an attention mechanism according to the quadruple information at the current time step in the Markov decision process.

10. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the method described in any one of claims 1 to 8 is implemented.