Dynamic multi-index optimization heterogeneous weapon target allocation method based on attention mechanism
Through the deep reinforcement learning training model based on attention mechanism, the problems of multi-index optimization and dynamic environmental changes in heterogeneous weapon allocation are solved, and the rapid generation of high-quality multi-index optimization strategies is achieved, which improves the efficiency and response capabilities of weapon allocation.
Patent Information
- Application Number
- CN202510572349.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-08
AI Technical Summary
The existing heterogeneous weapon allocation technology has failed to effectively consider multi-index optimization and dynamic environmental changes. The traditional multi-index optimization iteration algorithm has high time cost and is difficult to apply in large-scale problems.
A dynamic multi-index optimization method based on attention mechanism is adopted, and the attention model is trained through deep reinforcement learning to build a dynamic multi-index optimization model for heterogeneous weapons. The encoder and decoder are used to extract targets and weapon features respectively, generate multi-index optimization strategies, and update parameters to achieve the termination condition when the environment changes.
It realizes the optimization problem of multi-indicator inference in the dynamic battlefield, automatically detects environmental changes and quickly generates effective solutions, and improves the response ability and allocation efficiency of the weapon platform.
Smart Images

Figure CN120450341A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technology in the field of weapon target allocation, specifically a dynamic multi-index optimization heterogeneous weapon target allocation method based on an attention mechanism. Background Art
[0002] Existing heterogeneous weapon allocation technologies do not take into account multi-indicator optimization and dynamic environmental changes. Single-indicator optimization cannot coordinate multi-indicator conflicts. Traditional multi-indicator optimization iterative algorithms require high time costs and are difficult to apply to large-scale problems. Summary of the Invention
[0003] In response to the defects of the existing technologies that all of them are single-index optimization and the lack of non-iterative algorithms for the multi-index weapon target allocation problem, this invention proposes a dynamic multi-index optimization heterogeneous weapon target allocation method based on the attention mechanism. It improves the encoder and decoder in the learning multi-index optimization method, making it suitable for the heterogeneous weapon target allocation problem. It can effectively reduce the target threat while improving the efficiency of weapon target allocation, and is suitable for scenarios with extremely high time requirements.
[0004] The present invention is achieved through the following technical solutions:
[0005] The present invention relates to a dynamic multi-index optimization heterogeneous weapon target allocation method based on an attention mechanism. Minimizing the target threat value and minimizing the weapon cost are dual optimization indicators to construct a dynamic multi-index optimization model for heterogeneous weapons. An attention model is trained offline based on deep reinforcement learning to solve the multi-index optimization model. In the online stage, a multi-index optimization heterogeneous weapon target allocation strategy is generated in real time through the trained attention model. After the inference is completed, if the termination condition of the allocation optimization process is not met, the parameters are updated after environmental detection and input into the attention model until the termination condition is met and the process ends.
[0006] The attention model includes: an encoder and a decoder, wherein: the encoder encodes the weapon and target parameters of the input problem instance to obtain several groups of sub-problem weapon embeddings and target embeddings, and the decoder inputs several groups of sub-problem weapon and target embeddings in parallel and decodes them. Each sub-problem generates a solution through n sequential steps to form a Pareto solution set.
[0007] The encoder includes: a target sub-encoder, a weapon sub-encoder and a routing sub-encoder, wherein: the target sub-encoder extracts the individual information of each target based on the indicator, the weapon sub-encoder extracts the individual information of each weapon based on the indicator, and the routing sub-encoder dynamically and adaptively generates a set of weight vectors, performs weighted aggregation on the indicator embeddings of the target and weapon, thereby obtaining the target embedding and weapon embedding of the sub-problem.
[0008] The termination condition refers to: all targets are destroyed or all weapon resources are exhausted. When the termination condition has not been met, the environment detection mechanism is entered; if an environmental change is detected, the parameters are updated and input into the attention model for solution, otherwise the detection is cyclically performed; when the termination condition is met, the target allocation method ends. Technical Effects
[0009] This invention uses a multi-metric combined optimization attention model trained through deep reinforcement learning. Improved target and weapon sub-encoders extract target and weapon features, respectively. A routing sub-encoder adaptively generates appropriate weight features. Combining weapon and target features, this method decomposes the multi-metric optimization problem into multiple single-objective sub-problems, resolving the conflicting optimization issues inherent in simultaneously optimizing threat and weapon cost. Compared to existing technologies, this invention can infer Pareto solutions to multi-metric optimization problems within seconds. It can automatically detect environmental changes in dynamic battlefields and quickly generate effective new solutions when these changes occur, thereby improving the responsiveness of weapon platforms. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 Flowchart of the present invention;
[0011] Figure 2 This is the overall structure diagram of the attention model of the present invention;
[0012] Figure 3 The encoder structure of the present invention;
[0013] Figure 4 The decoder structure of the present invention;
[0014] Figure 5 Training HV for the model of the present invention n Indicator convergence curve;
[0015] Figure 6 This is a comparison chart of the Pareto solution sets of the present invention and different algorithms. DETAILED DESCRIPTION
[0016] like Figure 1 As shown, this embodiment relates to a dynamic multi-index optimization heterogeneous weapon target allocation method based on an attention mechanism, including:
[0017] S1, based on the heterogeneous weapon target allocation problem, with the minimization of target threat value and weapon cost as dual optimization indicators, a multi-constrained dynamic multi-index optimization model and its constraint conditions are constructed;
[0018] The multi-constrained dynamic multi-index optimization model refers to: Among them: min is the minimization index; the first index function is to minimize the target threat value The second indicator function is to minimize the cost of weapons Where: v j is the value of target j; p ij is the success rate of weapon i destroying target j; m is the number of weapons; n is the number of targets; c i is the cost of weapon i; x ij is a 0-1 integer decision variable for whether to assign the i-th weapon to the target task j, x ij =1 for distribution, x ij =0 means no allocation.
[0019] The constraints of the dynamic multi-index optimization model include: Constraint I: To ensure that each weapon is assigned at most one mission; Constraint II: To ensure that the number of weapons allocated does not exceed the maximum available number; Constraint III: x ij =0 or 1, which is a 0-1 integer variable constraint; Constraint IV: To ensure that the distance between the assigned weapon and the mission does not exceed the weapon's range, where: d ij is the distance from weapon i to task j; D i is the range of weapon i.
[0020] In this embodiment, the number of subproblems is set to 21, the node embedding dimension is 128, the number of attention heads of the multi-head attention (MHA) sublayer is 8, and the number of MHA sublayer stacking layers is 3;
[0021] Table 1 Damage probability matrix
[0022] The parameters of the multi-index weapon target allocation problem in this embodiment are all randomly generated: the weapon value c i With threat value v j Generated from the interval [5,20]. Weapon range D i Sampling from the interval [300, 500], weapon target distance d ij Generated from the interval [250, 400]. Weapon type and target type To distinguish different categories ( represents a fictitious target with a threat value of 0, a destruction probability of 1, and a distance of 0). Weapon reliability p ri Generated uniformly from the interval [0.9, 1]. As shown in Table 1, the baseline destruction probability of different weapon types against various targets is: The final destruction probability is given by Calculated;
[0023] S2 uses deep reinforcement learning to train an attention model for solving the multi-attribute weapon target allocation problem offline, specifically including:
[0024] 2.1 The process of solving the weapon target assignment problem is modeled as a Markov process. The states include: the set of unassigned weapons, the target state, and the assignment saturation (AS); the action is the t-th weapon selecting a target to expand the current solution; the reward is based on the sub-problem objective function. g(·|w j ). The state transfer updates the weapons to be allocated and the saturated feature state. The discount factor is set to 1 to ensure that the long-term benefits are equivalent to the cumulative rewards. The problem size is m×n, where m is the number of weapons and n is the number of targets.
[0025] 2.2 Randomly initialize the attention model parameters θ and generate N uniformly distributed weight vectors {w 1 ,…,w N} for sub-problem construction;
[0026] 2.3 Each training cycle (epoch) performs 125 iterations, and each iteration samples B = 96 multi-indicator weapon target assignment instances {S1,…,S B}, the size of each instance is 20×20, the parameters of each instance are randomly sampled from its distribution, and B is the batch size during one iteration of training;
[0027] 2.4 Input the weapon parameters, target parameters and weight vector of the problem instance into the attention model, and after encoding and decoding, output the multi-index weapon target allocation solution set
[0028] 2.5 Gradient update uses the Top-k baseline algorithm, that is, the average of the top k solutions in the batch is taken as the baseline. In this embodiment, k = 2; combined with the Top-k baseline b(S i ,w j ), the gradient of the parameter θ Among them: g(x|w) is the aggregation function, The learning rate is lr=10 -4 , using the commonly used Adam neural network parameter optimizer.
[0029] In this embodiment, the aggregation function adopts weighted aggregation.
[0030] Based on this, after 20 training cycles, the trained attention model converged and could solve the multi-index optimization problem and obtain the Pareto frontier solution set, that is, obtain the weapon target allocation plan for combat commanders to make decisions;
[0031] like Figure 5 As shown, this is the HV of the model on the test set nThe indicator convergence process diagram shows that the loss function of the model converges quickly at the beginning and stabilizes after a period of time, indicating that the model has converged.
[0032] In the embodiment, the initial number of weapons m=20, the number of targets n=20, and the total number of targets including the virtual targets is 21. The remaining parameters are randomly generated within their respective ranges. The final damage probability p ij Calculated by the formula.
[0033] S3. Use the attention model trained in step S2 to reason about the instance, directly output the approximate Pareto front solution set, and determine whether the preset termination condition of the dynamic algorithm is met, including:
[0034] 3.1 The encoder of the attention model generates an initial weight vector {w 1 ,…,w N}, the three sub-encoders encode the weapon and target parameters and weight vectors, and finally output the target embedding and weapon embedding of N sub-problems;
[0035] 3.2 The decoder of the attention model decodes the target embeddings and weapon embeddings of all sub-problems in parallel. In step t, the assignment probabilities of all n targets are calculated for the t-th weapon, the target with the highest probability is selected as the assignment object, the AS feature is updated, and the decoding continues in step t+1 until the last weapon is decoded. Finally, the solutions of N sub-problems are output, forming an approximate Pareto solution set.
[0036] like Figure 2 As shown, the attention model includes: Figure 3 The encoder shown and Figure 4 The decoder shown in the figure includes a target sub-encoder, a weapon sub-encoder, a routing sub-encoder, and a weighted aggregation unit. The target sub-encoder is used to extract individual information of each target based on the indicator, and the weapon sub-encoder extracts individual information of each weapon based on the indicator. The routing sub-encoder dynamically and adaptively generates a set of weight vectors and performs weighted aggregation on the indicator embeddings of the target and weapon to obtain the target embedding and weapon embedding of the sub-problem. The weighted aggregation unit decomposes the multi-indicator weapon target assignment problem into several sub-problems and generates sub-problem target embeddings and sub-problem weapon embeddings for each of the N sub-problems. These are then input into the decoder in parallel. The decoder obtains the target embedding and weapon embedding of the sub-problem. Each sub-problem generates a solution through n sequential steps.
[0037] like Figure 3 As shown, the target sub-encoder contains L attention layers, for each optimization index f i (i=1,…,q) obtain exclusive indicator features With target value vj , distance d to the weapon target ij and target type As the target feature, the target feature passes through L attention layers to obtain the corresponding target embedding
[0038] In each attention layer, the indicator features are first converted into embeddings through linear projection Through the multi-head attention layer and batch normalization layer, a new embedding is obtained Batch normalization is also used through the feedforward sublayer to obtain a new embedding Enter the next attention layer through skip connection.
[0039] like Figure 3 As shown, the weapon sub-encoder contains L attention layers, according to the indicator characteristics With weapon cost c i , Weapon Range D i , weapon reliability parameter p ri and weapon type W i type As weapon features, weapon features are linearly projected and passed through L attention layers to generate corresponding weapon embeddings.
[0040] like Figure 3 As shown, the routing sub-encoder adopts a two-layer neural network structure with N predefined initial weight vectors As input, output is the optimal weight vector for prediction
[0041] The weighted aggregation unit converts the weight vector Target Embedding and weapon embedding Perform weighted aggregation to decompose the multi-index weapon target allocation problem into several sub-problems and generate sub-problem target embeddings for each of the N sub-problems. and sub-problem weapon embedding
[0042] like Figure 4 As shown, the decoder includes: a linear mapping layer, a multi-head attention layer, a compatibility layer and a softmax layer, wherein: the first linear mapping layer embeds the t-th weapon into W t (L) and weapon embedding mean Concatenate to get context embedding The second linear mapping layer embeds the target and Allocation Saturation (AS) characteristics Concatenate to get the new embedding of the target The multi-head attention layer obtains new context embedding through MHA operation Get the scores of different targets through the compatibility layer The probability of selecting the jth target for weapon i in step t is calculated through the Softmax layer, and the target j with the maximum probability is selected as the assignment object of the tth weapon. The process is cyclical until all weapons are assigned, and the assignment sequence is obtained.
[0043] The score in: are two learnable matrices, and ζ = 10 is a given parameter.
[0044] Preferably, in the process of heterogeneous weapon target allocation, when the target distance exceeds the weapon range, that is, d ij ≥D i , the target is masked, that is, its selection probability is set to 0, so that weapons are allowed to remain unassigned, that is, all unassigned weapons are assigned to the virtual target. The virtual target is added to the end of the target sequence, so there are a total of The process of selecting a virtual target is the same as that of a real target.
[0045] The probability that weapon i selects target j is
[0046] The distribution sequence x1, x2, ..., x m Medium x t Corresponding to the decision variable x in step S1 ij , all sub-problems are decoded in parallel by the decoder, and the solutions of N sub-problems are obtained in parallel, forming the Pareto solution set of the multi-index optimization problem. Therefore, the probability of generating a complete solution x for case s is
[0047] S4. If the termination condition is not reached, perform environmental detection; otherwise, terminate the target allocation and output the result of the multi-index optimization, i.e., the Pareto solution set.
[0048] The termination conditions are: all targets are destroyed or all weapon resources are exhausted.
[0049] The environmental detection mentioned above refers to detecting whether the decision variables or the parameters of the optimization model have changed, specifically including:
[0050] a) Detect the first type of change, i.e., whether the number of weapons m and the number of targets n have changed. If so, exit the detection phase, update the instance parameters, input the attention model, and execute step S3. Otherwise, execute step b.
[0051] b) Detect the second type of change, that is, whether the indicators f1 and f2 have changed. Specifically, the following steps are performed: randomly extract some solutions of the attention model inference, calculate their indicator function values and compare them with the historical indicator function values. When the threshold is exceeded, it is determined that the environment has changed and the instance parameters are updated before inputting into the attention model. Otherwise, the environment change detection cycle continues until the environment changes or the termination condition is reached.
[0052] The example parameters include: weapon and target parameters, namely, weapon quantity, weapon cost, weapon type, weapon reliability, weapon range, target quantity, target threat value, target type, and target distance.
[0053] After specific actual experiments, HV was tested on 200 randomly generated instances. n and C1 R The indicators are shown in Table 2, which are the comparison results, and the standard deviation is in brackets. n and C1 R In terms of indicators, higher HV n The value indicates that the solution set obtained by the algorithm is more diverse and can cover a wider Pareto frontier. The higher the C1 R A value of indicates that the algorithm can generate more competitive solutions. The algorithm with AS features significantly outperforms the version without AS, which verifies the effectiveness of AS features and also shows that the introduction of expert features can enhance the model's ability to understand the problem, thereby generating a better set of solutions.
[0054] Table 2 Comparison of indicators with and without AS characteristics
[0055] In order to illustrate the improvement of the solution quality and solution speed of the present invention, the HV algorithm is compared with the multi-index optimization heuristic algorithms NSGA-II, MOEA / D, SPEA2, and NSGA-III. n and C1 R The indicators and running time are shown in Table 3, which are the average values of different algorithms under three indicators. The running time of the proposed method is the shortest and is significantly lower than that of other heuristic algorithms by an order of magnitude. n A high value indicates that the diversity of Pareto solution sets is better, C1 R A high value indicates that the Pareto solution set has a stronger dominance and a better solution quality. Compared with other iterative algorithms, the present invention not only has a faster solution speed but also a better solution quality.
[0056] Table 3 HV of different algorithms n ,C1 R , average running time
[0057] The Pareto solution set distribution obtained by this invention and other heuristic algorithms is as follows: Figure 6 As shown in the figure, the vertical axis is the target threat value and the horizontal axis is the weapon cost. It can be seen from the figure that the two indicators of the Pareto solution set obtained by the algorithm of the present invention are smaller, indicating that the Pareto solution set obtained by the present invention has better dominance and better solution quality, which is consistent with the indicator results in Table 3, and the calculation time is shorter.
[0058] like Figure 6 As shown, each point corresponds to a non-dominated weapon task allocation scheme; Figure 6 It can be seen that any solution has its own advantages, either low target threat value or low weapon cost; the many solutions solved in this example can assist in the final decision. Decision makers can choose the best solution from many solutions to allocate weapons tasks according to the actual scenario.
[0059] Compared with the existing technology, this invention solves the problem of lacking a non-iterative algorithm for multi-index weapon target allocation by training a multi-index combined optimization attention model based on deep reinforcement learning. The improved target sub-encoder and weapon sub-encoder can extract target features and weapon features respectively, and the routing sub-encoder can adaptively generate appropriate weight features. Combining weapon and target features, the multi-index optimization problem is decomposed into multiple single-target sub-problems, solving the conflict problem of optimizing both threat and weapon cost simultaneously. The decoder improves the context embedding and distribution saturation characteristics Enables more efficient decoding, This prevents over-allocation of weapons and improves the quality of the solution. The decoder solves all subproblems in parallel, minimizing time consumption. Compared to iterative multi-metric optimization algorithms, the attention model trained with deep reinforcement learning can learn the parameter characteristics of weapons and targets, as well as the probability distribution of the optimal weapon allocation. This allows for efficient reasoning and high-quality Pareto solutions without the need for iteration.
[0060] The above-mentioned specific implementation can be partially adjusted in different ways by those skilled in the art without departing from the principles and purpose of the present invention. The scope of protection of the present invention shall be based on the claims and shall not be limited by the above-mentioned specific implementation. All implementation schemes within its scope shall be subject to the constraints of the present invention.
Claims
1. A dynamic multi-index optimization method for heterogeneous weapon target allocation based on attention mechanism, characterized by: Taking minimizing target threat value and minimizing weapon cost as dual optimization indicators, a dynamic multi-index optimization model for heterogeneous weapons is constructed; Offline training of attention model based on deep reinforcement learning is used to solve multi-index optimization model; In the online phase, the trained attention model is used to generate a multi-metric optimization strategy for heterogeneous weapon target allocation in real time; After the inference is completed, if the termination condition of the allocation optimization process is not met, the parameters are updated after the environment is detected and input into the attention model until the termination condition is met and the process ends; The multi-constrained dynamic multi-index optimization model refers to: Among them: min is the minimization index; the first index function is to minimize the target threat value The second indicator function is to minimize the cost of weapons Where: v j is the value of target j; p ij is the success rate of weapon i destroying target j; m is the number of weapons; n is the number of targets; c i is the cost of weapon i; x ij is a 0-1 integer decision variable for whether to assign the i-th weapon to the target task j, x ij =1 for distribution, x ij =0 means no allocation; The constraints of the dynamic multi-index optimization model include: Constraint I: To ensure that each weapon is assigned at most one mission; Constraint II: To ensure that the number of weapons allocated does not exceed the maximum available number; Constraint III: x ij =0 or 1, which is a 0-1 integer variable constraint; Constraint IV: To ensure that the distance between the assigned weapon and the mission does not exceed the weapon's range, where: d ij is the distance from weapon i to task j; D i is the range of weapon i.
2. The dynamic multi-index optimization heterogeneous weapon target allocation method based on the attention mechanism according to claim 1 is characterized in that: The attention model includes an encoder and a decoder, wherein the encoder encodes the weapon and target parameters of the input problem instance to obtain a plurality of sub-problem weapon embeddings and target embeddings, and the decoder inputs the plurality of sub-problem weapon and target embeddings in parallel and decodes them. Each sub-problem generates a solution through n sequential steps, forming a Pareto solution set. The encoder includes: a target sub-encoder, a weapon sub-encoder and a routing sub-encoder, wherein: the target sub-encoder extracts the individual information of each target based on the indicator, the weapon sub-encoder extracts the individual information of each weapon based on the indicator, and the routing sub-encoder dynamically and adaptively generates a set of weight vectors, performs weighted aggregation on the indicator embeddings of the target and weapon, thereby obtaining the target embedding and weapon embedding of the sub-problem.
3. The dynamic multi-index optimization heterogeneous weapon target allocation method based on the attention mechanism according to claim 1 is characterized in that: The termination condition refers to: all targets are destroyed or all weapon resources are exhausted. When the termination condition has not been met, the environment detection mechanism is entered; if an environmental change is detected, the parameters are updated and input into the attention model for solution, otherwise the detection is cyclically performed; when the termination condition is met, the target allocation method ends.
4. The dynamic multi-index optimization heterogeneous weapon target allocation method based on the attention mechanism according to claim 2 is characterized in that: The target sub-encoder contains L attention layers, for each optimization index f i (i=1,…,q) obtain exclusive indicator features With target value v j , distance d to the weapon target ij and target type As the target feature, the target feature passes through L attention layers to obtain the corresponding target embedding In each attention layer, the indicator features are first converted into embeddings through linear projection Through the multi-head attention layer and batch normalization layer, a new embedding is obtained Batch normalization is also used through the feedforward sublayer to obtain a new embedding Enter the next attention layer through skip connection; The weapon sub-encoder contains L attention layers, according to the indicator characteristics With weapon cost c i , Weapon Range D i , weapon reliability parameter p ri and weapon type As weapon features, weapon features are linearly projected and passed through L attention layers to generate corresponding weapon embeddings. The routing sub-encoder adopts a double-layer neural network structure with N predefined initial weight vectors As input, output is the optimal weight vector for prediction The weighted aggregation unit converts the weight vector Target Embedding and weapon embedding Perform weighted aggregation to decompose the multi-index weapon target allocation problem into several sub-problems and generate sub-problem target embeddings for each of the N sub-problems. and subproblem weapons embedded in W i (L) =(W i 1(L) ,…,W i N(L) ) T ; The decoder includes: a linear mapping layer, a multi-head attention layer, a compatibility layer and a softmax layer, wherein: the first linear mapping layer embeds the t-th weapon into W t (L) and weapon embedding mean Concatenate to get context embedding The second linear mapping layer embeds the target and Allocation Saturation (AS) characteristics Concatenate to get the new embedding of the target The multi-head attention layer obtains new context embedding through MHA operation Get the scores of different targets through the compatibility layer The Softmax layer calculates the probability of selecting the jth target for weapon i in step t and selects the target j with the maximum probability as the assignment object of the tth weapon. The process is repeated until all weapons are assigned, and the assignment sequence is obtained. The score in: are two learnable matrices, and ζ = 10 is a given parameter.
5. The dynamic multi-index optimization heterogeneous weapon target allocation method based on the attention mechanism according to claim 1 is characterized in that: The offline training attention model based on deep reinforcement learning specifically includes: 2.1 The process of solving the weapon target assignment problem is modeled as a Markov process. The states include: the set of unassigned weapons, the target state, and the assignment saturation (AS); the action is the t-th weapon selecting a target to expand the current solution; the reward is based on the sub-problem objective function. g(·|w j ). The state transfer updates the weapons to be assigned and the saturated feature state. The discount factor is set to 1 to ensure that the long-term benefits are equivalent to the cumulative rewards. The problem size is m×n, where m is the number of weapons and n is the number of targets. 2.2 Randomly initialize the attention model parameters θ and generate N uniformly distributed weight vectors {w 1 ,…,w N } for sub-problem construction; 2.3 Each training cycle (epoch) performs 125 iterations, and each iteration samples B = 96 multi-indicator weapon target assignment instances {S1,…,S B }, the size of each instance is 20×20, the parameters of each instance are randomly sampled from its distribution, and B is the batch size during one iteration of training; 2.4 Input the weapon parameters, target parameters and weight vector of the problem instance into the attention model, and after encoding and decoding, output the multi-index weapon target allocation solution set 2.5 Gradient update uses the Top-k baseline algorithm, which takes the average of the top k solutions in the batch as the baseline, combined with the Top-k baseline b(S i ,w j ), the gradient of the parameter θ Among them: g(x|w) is the aggregation function, The learning rate is lr=10 -4 , using the commonly used Adam neural network parameter optimizer.
6. The dynamic multi-index optimization heterogeneous weapon target allocation method based on the attention mechanism according to claim 1 is characterized in that: The multi-index optimization strategy for heterogeneous weapon target allocation uses a trained attention model to reason about instances, directly output an approximate Pareto front solution set, and determine whether the dynamic algorithm's preset termination conditions are met. Specifically, the following steps are involved: 3.1 The encoder of the attention model generates an initial weight vector {w 1 ,…,w N }, the three sub-encoders encode the weapon and target parameters and weight vectors, and finally output the target embedding and weapon embedding of N sub-problems; 3.2 The decoder of the attention model decodes the target embeddings and weapon embeddings of all sub-problems in parallel. In the tth step, the assignment probabilities of all n targets are calculated for the tth weapon, the target with the largest probability is selected as the assignment object, the AS feature is updated, and the decoding of the t+1th step is continued until the decoding of the last weapon is completed. Finally, the solutions of the N sub-problems are output to form an approximate Pareto solution set.
7. The dynamic multi-index optimization heterogeneous weapon target allocation method based on the attention mechanism according to claim 1 is characterized in that: In the process of heterogeneous weapon target allocation, when the target distance exceeds the weapon range, that is, d ij ≥D i , the target is masked, that is, its selection probability is set to 0, so that the weapon is allowed to remain unassigned, that is, all unassigned weapons are assigned to the virtual target, and the virtual target is added to the end of the target sequence, so there are a total of targets, the process of selecting virtual targets is the same as that of real targets, that is, the probability of weapon i selecting target j is The distribution sequence x1, x2, ..., x m Medium x t Corresponding to the decision variable x in step S1 ij , all sub-problems are decoded in parallel by the decoder, and the solutions of N sub-problems are obtained in parallel, which constitute the Pareto solution set of the multi-index optimization problem. Therefore, the probability of generating a complete solution x for case s is 8. The dynamic multi-index optimization heterogeneous weapon target allocation method based on the attention mechanism according to claim 1 is characterized in that: The environmental detection mentioned above refers to detecting whether the decision variables or the parameters of the optimization model have changed, specifically including: a) Detect the first type of change, i.e., whether the number of weapons m and the number of targets n have changed. If so, exit the detection phase, update the instance parameters, and input the attention model for inference. Otherwise, execute step b. b) Detecting the second type of change, i.e., whether the indicators f1 and f2 have changed, is done by randomly extracting some solutions inferred by the attention model, calculating their indicator function values, and comparing them with the historical indicator function values. If the threshold is exceeded, it is determined that the environment has changed, and the instance parameters are updated and input into the attention model. Otherwise, the environment change detection cycle continues until the environment changes or the termination condition is met. The example parameters include: parameters of weapons and targets, namely, the number of weapons, the cost of weapons, the type of weapons, the reliability of weapons, the range of weapons, the number of targets, the threat value of targets, the type of targets and the distance of targets.