A target intelligent allocation method and system based on human-computer combination strategy learning

By using a human-machine combined strategy learning method, a target allocation system was established, which solved the problems of inaccurate target allocation models and slow solution speed in existing technologies, and realized fast and accurate decision-making in target allocation of large-scale clusters.

CN114358142BActive Publication Date: 2025-12-19CHINA ACAD OF LAUNCH VEHICLE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111532049.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-14
Publication Date
2025-12-19
Estimated Expiration
2041-12-14

AI Technical Summary

Technical Problem

Existing technologies suffer from problems such as inaccurate model construction, lack of predictability, slow solution speed, and susceptibility to local optima in the target allocation process. In particular, they are difficult to effectively handle multidimensional uncertainties in the target allocation of large-scale clusters.

Method used

A human-machine collaborative strategy learning approach is adopted. By establishing a sample library of human experience-based criterion strategies and an AHP quantization sample library, and using reinforcement learning to train the target allocation criteria and feature quantization model, a target allocation system is constructed, including a criterion model construction module, a quantization model construction module, and a target allocation module, to achieve rapid decision-making.

Benefits of technology

It improves the accuracy and efficiency of target allocation, can converge quickly in dynamic and uncertain environments, is suitable for various task requirements and situational input conditions, and supports human-machine collaborative decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114358142B_ABST
    Figure CN114358142B_ABST
Patent Text Reader

Abstract

The application discloses a target intelligent allocation method and system based on man-machine combination strategy learning, which comprises the following steps: step 1, modeling and training a target allocation criterion model based on an artificial experience criterion strategy sample library; step 2, modeling and training a target characteristic quantification model based on an AHP quantification sample library; and step 3, inputting task requirements and target situation, and utilizing the target allocation criterion model obtained in step 1 and the target characteristic quantification model obtained in step 2 to perform target allocation modeling optimization to obtain a target allocation result. The application can effectively integrate human experience, support machine learning and training of target allocation, effectively exert the respective advantages of man and machine, and explore a target allocation method, so as to promote man-machine combination strategy learning and improve decision-making effect and efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of target intelligent allocation, and particularly relates to a target intelligent allocation method and system based on a man-machine combined strategy learning. BACKGROUND

[0002] With the development of intelligent, network, collaboration and control technology and unmanned platform technology, various unmanned cluster systems develop rapidly. These cluster targets have strong maneuvering capability, flexible configuration, speed advantage, collaboration advantage and quantity advantage. Effective countermeasures can be implemented by using the cluster-to-cluster mode. In the cluster confrontation process, target intelligent allocation is a difficult problem. From the technical method, target allocation has typical complex nonlinear characteristics and belongs to the NP difficult problem.

[0003] The commonly used traditional target allocation algorithms mainly include branch and bound method, implicit enumeration method, dynamic programming method and cut plane method. These algorithms have relatively complicated processes and are difficult to handle large-scale target allocation problems. Heuristic optimization methods provide new methods and new ideas for handling complex problems by simulating natural phenomena or processes, including genetic algorithm (GA), particle swarm optimization (PSO), ant colony optimization (ACO), differential evolution algorithm (DE) and the like.

[0004] Specifically, for example, Xu Kehuo of Armored Forces Engineering College proposed an artificial immune algorithm with global and local updates, adopted optimal antibody inhibition technology to avoid falling into local optimum, and had relatively wide convergence speed and precision. Wang Yi of Air Force Command College generated new decisions from known decisions to reduce repeated search, obtained training samples of allocation schemes by using branch and bound method, deduced target allocation schemes by constructing and running fuzzy K nearest neighbor classifier machine learning method, and realized fast decision. Yang Xiaoling of National University of Defense Technology transformed population initialization, local search, force calculation and particle movement steps of original electromagnetic algorithm to adapt to integer solution space of target problem. By simulating individuals in population as charged particles, attraction and repulsion effect guides individuals to move towards optimal solution direction, and has strong global search capability, and is preliminarily applied to project scheduling and function optimization fields. Wang Zijian of Harbin Institute of Technology researched interceptor interceptor interception capability prediction method, target allocation decision method and cooperative interception strategy decision method of multi-target interceptor, designed a model for decision of interceptor interception strategy, and finally verified effectiveness of the model for cooperative interception strategy decision problem through simulation. BAE company of the United States adopts a control-based method to dynamically allocate weapons to targets in a multi-target situation, and adopts dynamic weapon target allocation (DWTA) based on control. A team of Yonsei University of South Korea adopts a heuristic genetic algorithm for multi-target allocation, introduces heuristic information, effectively accelerates algorithm execution efficiency, and avoids prematureness of genetic algorithm.

[0005] The above-mentioned traditional method mainly obtains an optimal decision scheme through problem modeling, model solving and other links in the process of solving the target allocation problem. However, the model construction of the traditional method is realized according to the experience of experts, the constraint conditions considered are limited, the understanding of the situation and the analysis of the threat of the target are insufficient, the model constructed is inaccurate, and the modeling process lacks global consideration of the mutual influence between different decision times, mainly static decision, and lacks prediction. Dynamic allocation adds modeling of random events that may occur in the process on the basis of the static allocation model, but also increases the complexity of problem solving. In addition, when solving the nonlinear optimization model of the additional constraint, multiple iterations are required for optimization, the solving speed is slow, and the optimization process may fall into a local optimal value or diverge, and a usable target allocation result cannot be obtained. SUMMARY

[0006] The technical problem of the present application is to overcome the shortcomings of the prior art and provide a target intelligent allocation method and system based on human-machine combined strategy learning, which can effectively integrate human experience and support machine learning and training of target allocation, effectively play the respective strengths of human and machine, and explore a target allocation method to promote human-machine combined strategy learning and improve decision-making effect and efficiency.

[0007] In order to solve the above technical problems, the present application discloses a target intelligent allocation method based on human-machine combined strategy learning, comprising:

[0008] Step 1, modeling and training a target allocation criterion model based on an artificial experience criterion strategy sample library;

[0009] Step 2, modeling and training a target characteristic quantification model based on an AHP quantification sample library;

[0010] Step 3, according to the task demand and target situation input, using the target allocation criterion model obtained in step 1 and the target characteristic quantification model obtained in step 2, performing target allocation modeling optimization to obtain a target allocation result.

[0011] In the above-mentioned target intelligent allocation method based on human-machine combined strategy learning, the target allocation criterion model at least includes the following allocation criteria: a maximum damage probability criterion, a maximum threat criterion, a threat degree random allocation criterion, a maximum damage probability minimum unit criterion, a maximum cost-effectiveness ratio criterion, an escape time and remaining balance criterion, and a minimum total time criterion.

[0012] In the above-mentioned target intelligent allocation method based on human-machine combined strategy learning, the target allocation criterion model is modeled and trained based on an artificial experience criterion strategy sample library, comprising:

[0013] An artificial experience criterion strategy sample library is established;

[0014] Each sample in the artificial experience criterion strategy sample is input into the reinforcement learning-based criterion strategy learning model for training, and at the same time, the specific criterion corresponding to the policy selection result is provided by the basic criterion strategy model, the network model parameters of the criterion strategy learning model are obtained through reinforcement learning training, and then the target allocation criterion model is constructed.

[0015] In the target intelligent allocation method based on man-machine combined strategy learning, the artificial experience criterion strategy sample library at least includes: a plurality of task requirements, a plurality of situation input conditions, and artificial strategy selection results corresponding to different task requirements and situation input combination conditions.

[0016] In the target intelligent allocation method based on man-machine combined strategy learning, the target characteristic quantification model is used to determine qualitative and quantitative factors involved in the comprehensive evaluation of the target threat, and at least includes: whether it is assigned by a superior, launch point position, predicted landing point position, range, shutdown point speed, reentry speed, damage type, damage power, damage influence, damage difficulty, survivability, mobility, hit accuracy, remaining flight time, maximum height and target importance.

[0017] In the target intelligent allocation method based on man-machine combined strategy learning, the AHP-based quantification sample library is modeled and trained to obtain the target characteristic quantification model, including:

[0018] An AHP-based quantification sample library is established;

[0019] Each sample in the AHP-based quantification sample library is input into the reinforcement learning-based quantification strategy learning model for training, and at the same time, the corresponding element modeling is provided by the target characteristic quantification modeling, the network model parameters of the quantification strategy learning model are obtained through reinforcement learning training, and then the target characteristic quantification model is constructed.

[0020] In the target intelligent allocation method based on man-machine combined strategy learning, the AHP-based quantification sample library at least includes: a quantitative evaluation element type in the target allocation task, a relative importance score between elements, and an artificial quantification experience result under different combination conditions.

[0021] In the target intelligent allocation method based on man-machine combined strategy learning,

[0022] The model based on the maximum damage probability criterion is represented as follows:

[0023]

[0024] where m represents the target number, n represents the number of fire units, i represents the fire unit number, j represents the target number, i = 1, 2, …, m, j = 1, 2, …, m; x ij represents the allocation decision variable, if the i th fire unit is allocated to attack the j th target, x ij = 1, otherwise x ij = 0; p ij represents the damage probability of the i th fire unit to the j th target; w j represents the threat value of the j th target;

[0025] The model based on the maximum threat criterion is represented as follows:

[0026]

[0027] p ij → w j

[0028] st.

[0029]

[0030] The model based on the threat degree random allocation criterion is represented as follows:

[0031]

[0032] p ij → w i+1

[0033] st.

[0034]

[0035] The model based on the maximum damage probability minimum unit criterion is represented as follows:

[0036]

[0037] where P dj represents the preset damage probability threshold of the j th target, P j represents the joint damage probability of the allocated fire to the j th target, is the mean value of P j ;

[0038] The model based on the maximum cost-effectiveness criterion is represented as follows:

[0039]

[0040] The model based on the maximum escape time and remaining balance criterion is represented as follows:

[0041]

[0042] where t ij denotes the escape time of the jth target from the striking area of the ith fire unit, is the expected damage probability threshold of the jth target;

[0043] The model based on the minimum total time criterion is represented as follows:

[0044]

[0045] where t ij denotes the time for the jth target to reach the far boundary of the killing area of the ith fire unit, t mij denotes the time for the jth target to stay in the killing area of the ith fire unit, t zh denotes the switching time.

[0046] In the above target intelligent allocation method based on human-computer combined strategy learning,

[0047] The model of target importance is represented as follows:

[0048] w(j) = 1 - 0.1I j , 1≤I j ≤3

[0049] where I j denotes the defense priority of the jth target;

[0050] The model of remaining flight time is represented as follows:

[0051]

[0052] where T h1 , T h2 , T h3 , T h4 are four flight time thresholds of different sizes set in advance;

[0053] The model of shutdown point speed is represented as follows:

[0054]

[0055] where T v1 , T v2 , T v3 are three flight speed thresholds of different sizes set in advance;

[0056] The model of maximum height is represented as follows:

[0057]

[0058] where T h1 , T h2, T h3 , T h4 , T h5 , T h6 Six flight height thresholds of different sizes are preset.

[0059] Correspondingly, the application also discloses a target intelligent allocation system based on a man-machine combined strategy learning, which comprises:

[0060] A criterion model construction module is used for modeling and training a target allocation criterion model based on an artificial experience criterion strategy sample library.

[0061] A quantitative model construction module is used for modeling and training a target characteristic quantitative model based on an AHP quantitative sample library.

[0062] A target allocation module is used for target allocation modeling optimization according to a task demand and a target situation input, a target allocation criterion model obtained by the criterion model construction module and a target characteristic quantitative model obtained by the quantitative model construction module, and obtaining a target allocation result.

[0063] The application has the following advantages:

[0064] (1) The application discloses a target intelligent allocation method based on a man-machine combined strategy learning, gives a hierarchical target intelligent allocation method process based on the man-machine combined strategy learning, can effectively integrate artificial priori knowledge rules, artificial decision-making and other human experiences, and improves the accuracy and efficiency of target allocation.

[0065] (2) The application discloses a target intelligent allocation method based on a man-machine combined strategy learning, gives no less than 7 kinds of applicable basic criterion strategy models and no less than 4 kinds of applicable target characteristic quantitative models for the multi-dimensional uncertainty factors of targets and environmental situations in a large-scale cluster target allocation process, and provides basic models for rapid target allocation in the whole process.

[0066] (3) The application discloses a target intelligent allocation method based on a man-machine combined strategy learning, proposes a strategy learning model based on reinforcement learning, the model is applicable to various task demands, various situation input conditions, and the training and learning of artificial strategy selection under different task demand and situation input combination conditions, can quickly converge to obtain applicable network model parameters, and supports man-machine combined decision-making. BRIEF DESCRIPTION OF DRAWINGS

[0067] Figure 1 is a flowchart of a target intelligent allocation method based on a man-machine combined strategy learning in the embodiment of the application;

[0068] Figure 2is a schematic diagram of an artificial experience criterion strategy learning model based on reinforcement learning in an embodiment of the present application;

[0069] Figure 3 is a schematic diagram of a target characteristic artificial quantification strategy learning model. DETAILED DESCRIPTION

[0070] In order to make the purpose, technical solutions and advantages of the present application clearer, the disclosed embodiments of the present application will be described in further detail below with reference to the accompanying drawings.

[0071] The present application is directed to the multi-dimensional uncertainty factors of target quantity, position, category, formation, speed, time and environmental situation in the large-scale cluster target allocation process, and the accumulated human decision-making experience in the cluster game confrontation deduction, and proposes a target intelligent allocation method based on man-machine combined strategy learning, which effectively plays the respective strengths of man and machine, explores the intelligent and rapid decision-making method of target allocation in a dynamic uncertain environment, balances the perfection and complexity of the target allocation problem, and improves the accuracy and efficiency of target allocation.

[0072] As Figure 1 In the present embodiment, the target intelligent allocation method based on man-machine combined strategy learning comprises:

[0073] Step 1, based on the artificial experience criterion strategy sample library, a target allocation criterion model is modeled and trained.

[0074] In the present embodiment, the target allocation criterion model at least includes the following allocation criteria: maximum damage probability criterion, maximum threat criterion, threat degree random allocation criterion, maximum damage probability minimum unit criterion, maximum cost-effectiveness ratio criterion, escape time and remaining balance criterion, and minimum total time criterion.

[0075] Preferably, as Figure 2 The establishment process of the target allocation criterion model is as follows: an artificial experience criterion strategy sample library is established. Each sample in the artificial experience criterion strategy sample is input into the criterion strategy learning model based on reinforcement learning for training, and at the same time, the specific criterion corresponding to the strategy selection result is provided by the basic criterion strategy model above. After reinforcement learning training, the network model parameters of the criterion strategy learning model are obtained, and then the target allocation criterion model is constructed. It can be seen that the present application changes the original direct selection of a specific strategy relying on artificial experience into automatic generation of the most suitable target allocation criterion through intelligent learning model, realizing effective accumulation and application of artificial experience. Among them, the artificial experience criterion strategy sample library at least includes: multiple task requirements, multiple situation input conditions, and artificial strategy selection results under different task requirements and situation input combination conditions.

[0076] Further, the model representation of each distribution criterion is as follows:

[0077] The model representation based on the maximum damage probability criterion is as follows:

[0078]

[0079] wherein m represents the number of targets, n represents the number of fire units, i represents the fire unit number, j represents the target number, i = 1, 2, …, m, j = 1, 2, …, m; x ij represents the distribution decision variable, if the i-th fire unit is assigned to attack the j-th target, x ij = 1, otherwise x ij = 0; p ij represents the damage probability of the i-th fire unit to the j-th target; w j represents the threat value of the j-th target.

[0080] The model representation based on the maximum threat criterion is as follows:

[0081]

[0082] p ij → w j

[0083] st.

[0084]

[0085] The model representation based on the threat degree random distribution criterion is as follows:

[0086]

[0087] p ij → w i+1

[0088] st.

[0089]

[0090] The model representation based on the maximum damage probability minimum unit criterion is as follows:

[0091]

[0092] wherein P dj represents the preset damage probability threshold of the j-th target, P j represents the joint damage probability of the assigned fire to the j-th target, is the mean value of P j .

[0093] The model based on the maximum cost-effectiveness ratio criterion is shown as follows:

[0094]

[0095] The model based on the escape time and remaining balance criterion is shown as follows:

[0096]

[0097] Wherein, t ij represents the escape time of the jth target from the striking area of the i th firepower unit, is the expected damage probability threshold of the jth target.

[0098] The model based on the minimum total time criterion is shown as follows:

[0099]

[0100] Wherein, t ij represents the time when the jth target reaches the far boundary of the i th firepower unit killing area, t mij represents the time when the jth target stays in the i th firepower unit killing area, t zh represents the switching time.

[0101] Step 2, based on the AHP quantitative sample library, modeling and training to obtain the target characteristic quantitative model.

[0102] In this embodiment, the target characteristic quantitative model is used to determine the qualitative and quantitative factors involved in the comprehensive evaluation of the target threat, at least including whether it is designated by the superior, the launch point position, the predicted drop point position, the range, the shutdown point speed, the reentry speed, the damage type, the damage power, the damage influence, the damage difficulty, the survivability, the mobility, the hitting accuracy, the remaining flight time, the maximum height and the target importance.

[0103] Preferably, as Figure 3 The establishment process of the target characteristic quantitative model is as follows: a quantitative sample library based on AHP is established. Each sample in the quantitative sample library based on AHP is input into the quantitative strategy learning model based on reinforcement learning for training, and at the same time, the corresponding element modeling is provided by the above target characteristic quantitative modeling, and after reinforcement learning training, the network model parameters of the quantitative strategy learning model are obtained, and then the target characteristic quantitative model is constructed. It can be seen that the original quantitative result generated by manual importance comparison is changed into the most suitable quantitative strategy and the corresponding quantitative result generated automatically through intelligent learning model, so as to realize the effective accumulation and application of manual experience. Among them, the quantitative sample library based on AHP at least includes: quantitative evaluation element type in target allocation task, element relative importance score between two elements and manual quantitative experience result under different combination conditions.

[0104] Further:

[0105] The model of the target importance is represented as follows:

[0106] w(j) = 1 - 0.1I j , 1≤I j ≤3

[0107] where I j represents the defense priority of the jth target.

[0108] The model of the remaining flight time is represented as follows:

[0109]

[0110] where T h1 , T h2 , T h3 , T h4 are four flight time thresholds of different sizes preset.

[0111] The model of the shutdown point speed is represented as follows:

[0112]

[0113] where T v1 , T v2 , T v3 are three flight speed thresholds of different sizes preset.

[0114] The model of the maximum height is represented as follows:

[0115]

[0116] where T h1 , T h2 , T h3 , T h4 , T h5 , T h6 are six flight height thresholds of different sizes preset.

[0117] Step 3, according to the task requirements and target situation input, using the target assignment criterion model obtained in step 1 and the target characteristic quantification model obtained in step 2, the target assignment modeling optimization is carried out, and the target assignment result is obtained.

[0118] In the embodiment, according to the three levels of criterion strategy, target characteristic quantification and target assignment, steps 1, 2 and 3 are respectively corresponded. Among them, the criterion strategy level is mainly to automatically generate a suitable target assignment criterion according to the task requirement; the target characteristic quantification level is mainly to automatically generate a suitable target characteristic quantification result according to the target situation; and the target assignment level is to use the suitable strategy learned by the criterion strategy level and the target characteristic quantification level to perform target assignment modeling optimization according to the task requirement and the situation input, and obtain a target assignment result.

[0119] Preferably, as described above, the target assignment level mainly completes the target assignment modeling optimization task. Since the target assignment task belongs to a nonlinear combination optimization decision problem, the solution space size increases exponentially with the increase of the number of fire units and the number of targets, so it is necessary to use an intelligent optimization method to improve the convergence speed and quickly realize the target assignment task of large-scale solution.

[0120] In summary, the present application gives an approximate approximation of the artificial experience criterion strategy function by using a deep neural network, defines the relationship between the network model, the state space and the action space. Secondly, the AHP method is used to construct the judgment matrix, and the 1-5 level fuzzy scale method is used to quantize the comparison between two factors in the same layer, and the judgment matrix is formed. The process of obtaining the fuzzy scale by relying on artificial experience scoring is converted into an approximation of a strategy function, and the specific approach can use the deep neural network described above. In addition, a method for target assignment optimization represented by PSO is given (including: 1) PSO initialization; 2) PSO coding; 3) SA-PSO hybrid), which realizes the purpose of solving the target assignment task.

[0121] On the basis of the above embodiment, the following will be described in combination with a specific example.

[0122] The design of the target intelligent assignment method based on man-machine combined strategy learning is as follows:

[0123] (1) Process design of target intelligent assignment method based on man-machine combined strategy learning

[0124] A process of a target intelligent assignment method based on man-machine combined strategy learning is designed, which is developed according to three levels of criterion strategy layer, target characteristic quantification layer and target assignment layer:

[0125] The criterion strategy layer is mainly to automatically generate a suitable target assignment criterion according to the task requirement. The original direct selection of a specific strategy relying on artificial experience is changed into automatic generation of the most suitable target assignment criterion through intelligent learning model, so as to realize effective accumulation and application of artificial experience.

[0126] Target characteristic quantification layer, mainly according to the target situation, automatically generates the corresponding target characteristic quantification result. The original quantification result generated by the importance comparison of artificial is changed to automatically generate the most suitable quantification strategy and obtain the corresponding quantification result through intelligent learning model, so as to realize the effective accumulation and application of artificial experience.

[0127] Target allocation layer, according to the task demand and situation input, uses the corresponding strategy learned by the criterion strategy layer and the target characteristic quantification layer to perform target allocation modeling optimization, and obtains the target allocation result.

[0128] (2) Modeling and learning of criterion strategy layer

[0129] The criterion strategy layer involves two aspects of basic criterion strategy modeling and artificial experience criterion strategy learning:

[0130] Basic criterion strategy modeling:

[0131] Suppose that in a certain task, m targets in the air enter the range of n fire units. According to the problem description, the target allocation model based on different criteria can be established, which can include but is not limited to: according to the problem description, the target allocation model based on different criteria can be established, which can include but is not limited to: based on the maximum damage probability criterion, based on the maximum threat criterion, based on the threat degree random allocation criterion, based on the maximum damage probability minimum unit criterion, based on the maximum cost-effectiveness ratio criterion, based on the escape time and remaining balance criterion, based on the minimum total time criterion, etc.

[0132] Artificial experience criterion strategy learning:

[0133] A sample library of artificial experience criterion strategy is established, which includes: various task requirements, various situation input conditions, and artificial strategy selection results under different task requirement and situation input combination conditions. The above content is input into the reinforcement learning-based criterion strategy learning model for training, and the specific criteria corresponding to the strategy selection results are provided by the basic criterion strategy model in the above, and after reinforcement learning training, the network model parameters learned are generated.

[0134] The learning process can use various general reinforcement learning methods, and only deep neural networks are used as an example to illustrate.

[0135] The sample library continuous state space S is taken as the input of the network, and the continuous action space A is taken as the output of the network. Since deep neural networks can realize the approximation of any continuous function, the deep neural network is used to realize the approximation of the artificial experience criterion strategy function. The network model is denoted as π, and the relationship between the network model π and the state space S and the action space A can be represented by the following formula:

[0136] a = π(s) a∈A, s∈S

[0137] In the network π basic structure, the input is a certain continuous state vector s ∈ S, and the output is the optimal continuous action vector a ∈ A for the state. Then, the learning result can be obtained by using the "continuous-discrete action mapping model". The design of the network π can be adjusted according to the training process. In the example, the network includes an input layer, an output layer, and 2 hidden layers, and the neuron type is Relu type. According to the above parameter structure, assuming that the output of the jth hidden layer of the network π is z j (j = 1, 2), then the action vector q that can be output by the policy network π can be calculated as follows:

[0138]

[0139] q = π(s) = relu(W3relu(W2(relu(W1[s 1] T ))))

[0140] wherein relu(x) = max(0, x) is a rectified linear activation function.

[0141] Then, the obtained after training is:

[0142] Based on the maximum damage probability criterion, the model is represented as follows:

[0143]

[0144] wherein m represents the number of targets, n represents the number of fire units, i represents the fire unit number, j represents the target number, i = 1, 2, …, m, j = 1, 2, …, m; x ij represents the allocation decision variable, if the ith fire unit is allocated to attack the jth target, then x ij = 1, otherwise x ij = 0; p ij represents the damage probability of the ith fire unit to the jth target; w j represents the threat value of the jth target;

[0145] Based on the maximum threat criterion, the model is represented as follows:

[0146]

[0147] p ij → w j

[0148] st.

[0149]

[0150] Based on the threat degree random allocation criterion, the model is represented as follows:

[0151]

[0152] p ij →w i+1

[0153] st.

[0154]

[0155] The model based on the minimum unit criterion of maximum damage probability is shown as follows:

[0156]

[0157] where P dj represents a preset damage probability threshold of the jth target, P j represents a joint damage probability of the jth target by the allocated firepower, is the mean value of P j . When the damage probability of a firepower unit is lower than the preset threshold, the allocation is considered invalid; meanwhile The smaller the value is, the greater the mean value of the damage probability is, thereby ensuring that the target is attacked by smaller firepower resources. In addition, when the number of allocated firepower units is the same , the firepower unit with a greater damage probability is selected to ensure the maximum damage probability of the target.

[0158] The model based on the maximum efficiency-cost ratio criterion is shown as follows:

[0159]

[0160] The model based on the remaining balance criterion of escape time is shown as follows:

[0161]

[0162] where t ij represents the escape time of the jth target from the attack area of the ith firepower unit, is a preset damage probability threshold of the jth target. P j The constraint ensures that the joint damage probability of each target reaches the preset damage probability threshold. If the joint damage probability P j of a target is lower than the damage probability threshold , the allocation of the target is considered invalid. The damage probability threshold is determined by the commander according to the situation.

[0163] The model based on the minimum total time criterion is shown as follows:

[0164]

[0165] Among them, t ij t represents the time it takes for the j-th target to reach the far boundary of the kill zone of the i-th fire unit. mij t represents the time that the j-th target spends in the kill zone of the i-th fire unit. zh Indicates the switching time.

[0166] (3) Target characteristic quantification layer modeling and learning

[0167] The target feature quantification layer involves two stages: target feature quantification modeling and target feature manual quantification strategy learning.

[0168] Quantitative modeling of target characteristics:

[0169] In target assignment tasks, the comprehensive assessment of target threats involves numerous qualitative and quantitative factors, which may include, but are not limited to: whether it is designated by superiors, launch point location, predicted impact point location, range, shutdown velocity, reentry velocity, damage type, damage power, damage impact, degree of damage difficulty, survivability, maneuverability, accuracy, remaining flight time, maximum altitude, and target importance.

[0170] Learning strategies for artificial quantification of target characteristics:

[0171] A quantitative sample library based on AHP is established, which includes: quantitative evaluation element types in target assignment tasks, pairwise relative importance scores between elements, and manual quantitative experience results under different combinations of conditions. The above content is input into a reinforcement learning-based quantitative policy learning model for training. Simultaneously, the target characteristic quantitative modeling described above provides corresponding element modeling. After reinforcement learning training, the learned network model parameters are generated.

[0172] The AHP method is used as an example for illustration:

[0173] The AHP method is used to construct the judgment matrix, and the 1-5 level fuzzy scaling method is applied to quantify the pairwise comparisons of factors in the same layer, forming the judgment matrix A = (a ij ) n×n Where n is the number of factors, a ij Indicates threat factor B i For B j The relative importance of, and satisfying a ij a ji =1, with the following possible values:

[0174] 1: Indicates that two elements are equally important.

[0175] 2: Indicates that when comparing two elements, B i B j Slightly important;

[0176] 3: indicates two elements compared to B i Ratio B j Significant importance;

[0177] 4: indicates two elements compared to B i Ratio B j Strong importance;

[0178] 5: indicates two elements compared to B i Ratio B j Extremely important.

[0179] The above process of obtaining a fuzzy scale by relying on manual experience scoring is converted into an approximation of a strategy function. The process of approximation fitting can use linear fitting, polynomial fitting, etc. It can also use the deep neural network adopted in the above (2) manual experience criterion strategy learning part, and will not be expanded.

[0180] Preferably, the target importance, the remaining flight time, the shutdown point speed, and the maximum height are selected as the elements in the example. The analysis and quantification are as follows:

[0181] Target importance

[0182] Assuming that the importance is divided into k levels, there are m targets in total, satisfying k≤m. Among them, level 1 is the highest, level 2 is the second, and so on. The model of target importance is represented as follows:

[0183] w(j)=1-0.1I j ,1≤I j ≤3

[0184] Where I j represents the priority of the jth target.

[0185] Remaining flight time

[0186] According to the linear difference processing of the target flight time size change characteristics, the model of the remaining flight time is represented as follows:

[0187]

[0188] Where T h1 , T h2 , T h3 , T h4 are four flight time thresholds of different sizes set in advance, which can be adjusted according to experience.

[0189] Shutdown point speed

[0190] According to the linear difference processing of the target shutdown point speed size change characteristics, the model of the shutdown point speed is represented as follows:

[0191]

[0192] wherein T v1 , T v2 , T v3 are three preset flight speed thresholds of different sizes, which are adjusted according to experience.

[0193] Maximum height

[0194] According to the flight height change characteristics, linear difference processing is performed to obtain a model of the maximum height, which is represented as follows:

[0195]

[0196] wherein T h1 , T h2 , T h3 , T h4 , T h5 , T h6 are six preset flight height thresholds of different sizes, which are adjusted according to experience.

[0197] (4) Target allocation layer synthesis

[0198] The target allocation layer mainly completes the target allocation modeling optimization task. Since the target allocation task belongs to a nonlinear combination optimization decision problem, the solution space size increases exponentially with the increase of the number of fire units and the number of targets, and therefore, an intelligent optimization method needs to be used to improve the convergence speed and quickly realize the target allocation task of large-scale solution.

[0199] A PSO algorithm is used to perform target allocation optimization.

[0200] PSO initialization:

[0201] A group of random particles is initialized to obtain an initial solution. Then, an optimal solution is searched through iteration, and the particles are updated according to the individual extreme value and the global extreme value in the iteration process. The optimal solution found by the particle itself is the individual extreme value, and the optimal solution of the current entire population is the global extreme value. The speed and position of the particle can be updated according to the following formula

[0202] v(t+1) = wv(t) + c1r1[pbest(t) - x(t)] + c2r2[gbest(t) - x(t)]

[0203] x(t+1) = x(t) + v(t+1)

[0204] where v(t) and x(t) represent the velocity and position of the particle at time t. w represents the inertia weight, r1 and r2 are random numbers between 0 and 1, and c1 and c2 represent learning factors, which are used to measure the ability of the particle to learn from the excellent particle in the process of approaching the optimal point. c1 adjusts the step length of the particle approaching the individual optimum, and c2 adjusts the step length of the particle approaching the population optimum. When the value of the learning factor is small, the particle moves in the area far from the excellent particle. When the value of the learning factor is large, the particle approaches the excellent particle at a large speed, but when the value is too large, the particle will again move away from the excellent particle.

[0205] PSO encoding:

[0206] A real number-based encoding method is designed, in which the particle position represents a candidate scheme of a target allocation. The dimension of the particle position vector is m (i.e. the number of fire units), and the total number of particles is R. The position vector of the rth particle is X r = [X r1 X r2 …X m ], where X n (i = 1, 2, …, m) is an integer between 0 and n. The particle velocity vector is V r = [v r1 v r2 …v m ], where v n (i = 1, 2, …, m) is an integer between -(n-1) and (n-1).

[0207] X r is converted into a decision variable form suitable for 0-1 integer programming:

[0208] x = [x ij ] m×n (i = 1, 2, …, m, j = 1, 2, …, n)

[0209] The specific conversion formula is shown in the following formula:

[0210]

[0211] The flight speed and position of the rth particle in the i-dimensional subspace are updated as follows:

[0212]

[0213] where c1 and c2 are learning factors, which are normal numbers; r1 and r2 are random numbers between 0 and 1; w is the inertia weight; P r is the optimal position searched by the rth particle, also known as the individual extreme value; P gThe best position found so far for the entire species, also called a global extremum; indicates rounding off.

[0214] SA-PSO hybrid

[0215] Step 1: initialize PSO parameters. Determine the inertia weight w, learning factors c1, c2 and population size R, set the maximum number of iterations k max .

[0216] Step 2: randomly generate a population of R particles, that is, randomly generate R initial populations and R initial velocities where r = 1, 2, …, R.

[0217] Step 3: calculate the fitness g r of each particle and compare it with the individual extremum P r , and update the optimal one as the individual extremum P r .

[0218] Step 4: compare each particle individual extremum P r with the global extremum P g , and update the optimal one as the global extremum P g .

[0219] Step 5: if the termination condition is met, end the program, otherwise, execute Step 6.

[0220] Step 6: calculate the flight speed and position of each particle at the next time according to the specific particle update rule, and limit the speed and position in (v min , v max ) and (X min , X max ) respectively.

[0221] Step 7: execute the SA algorithm

[0222] {

[0223] Step 1: initialize SA parameters, set the initial temperature T and the number of iterations L for each T value.

[0224] Step 2: for k = 1, 2, …, L, execute steps 3-6.

[0225] Step 3: generate a new solution x r '.

[0226] Step 4: calculate e(r) = g r -g r ', where g r ' is the fitness function of the new solution.

[0227] Step 5: If e(r) < 0, accept x r Otherwise accept x with probability exp(-e(r) / T) r

[0228] Step 6: If the termination condition is met, output the current solution as the optimal solution and end the program; otherwise, go to Step 8.

[0229] }

[0230] Step 8: Gradually reduce the temperature at the annealing temperature convergence rate a, that is, T = aT, and if T >= 0, go to Step 3, otherwise end the program.

[0231] On the basis of the above embodiment, the application further discloses a target intelligent allocation system based on a man-machine combined strategy learning, which comprises: a criterion model construction module, which is used for modeling and training to obtain a target allocation criterion model based on an artificial experience criterion strategy sample library; a quantitative model construction module, which is used for modeling and training to obtain a target characteristic quantitative model based on an AHP quantitative sample library; and a target allocation module, which is used for target allocation modeling optimization to obtain a target allocation result according to task requirements and target situation input, the target allocation criterion model obtained by the criterion model construction module and the target characteristic quantitative model obtained by the quantitative model construction module.

[0232] For the system embodiment, since it corresponds to the method embodiment, the description is relatively simple, and the related parts can be referred to the description in the method embodiment part.

[0233] Although the application has been disclosed as above with the preferred embodiments, it is not intended to limit the application, and any person skilled in the art can make possible changes and modifications to the technical solutions of the application by using the disclosed methods and technical contents without departing from the spirit and scope of the application, therefore, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the application, which does not depart from the content of the technical solutions of the application, all belong to the protection scope of the technical solutions of the application.

[0234] The contents not described in detail in the specification of the application belong to the known technology of the person skilled in the art.​​

Claims

1. A target intelligent allocation method based on human-computer combination strategy learning, characterized in that, Comprise: Step 1, based on artificial experience criterion strategy sample library, modeling and training to get target allocation criterion model; comprising: establishing artificial experience criterion strategy sample library; each sample in the artificial experience criterion strategy sample is input into the criterion strategy learning model based on reinforcement learning for training, at the same time, the specific criterion corresponding to the policy selection result provided by the basic criterion strategy model is obtained, after reinforcement learning training, the network model parameters of the criterion strategy learning model are obtained, and then the target allocation criterion model is constructed; wherein, the target allocation criterion model at least includes the following allocation criteria: maximum damage probability criterion, maximum threat criterion, threat degree random allocation criterion, maximum damage probability minimum unit criterion, maximum cost-effective ratio criterion, escape time and remaining balance criterion and minimum total time criterion; the artificial experience criterion strategy sample library at least includes: multiple task requirements, multiple situation input conditions, and artificial strategy selection results corresponding to different task requirements and situation input combination conditions; Step 2, based on AHP quantitative sample library, modeling and training to get target characteristic quantization model; comprising: establishing AHP-based quantitative sample library; each sample in the AHP-based quantitative sample library is input into the quantitative strategy learning model based on reinforcement learning for training, at the same time, the corresponding element modeling is provided by the target characteristic quantization modeling, after reinforcement learning training, the network model parameters of the quantitative strategy learning model are obtained, and then the target characteristic quantization model is constructed; wherein, the target characteristic quantization model is used to determine the qualitative and quantitative factors involved in the comprehensive evaluation of target threat, at least including: whether it is assigned by the superior, launch point position, predicted landing point position, range, shutdown point speed, reentry speed, damage type, damage power, damage influence, damage difficulty, survivability, mobility, hitting accuracy, remaining flight time, maximum height and target importance; the AHP-based quantitative sample library at least includes: quantitative evaluation element type in target allocation task, relative importance score between elements, and artificial quantitative experience results under different combination conditions; Step 3, according to the task requirement and the target situation input, using the target allocation criterion model obtained in step 1 and the target characteristic quantization model obtained in step 2, the target allocation modeling optimization is carried out, and the target allocation result is obtained.

2. The target intelligent allocation method based on man-machine combined strategy learning according to claim 1, wherein the model of the maximum damage probability criterion is represented as follows: The model of the maximum threat criterion is represented as follows: where m represents the target number, n represents the number of fire units, i represents the fire unit number, j represents the target number, i = 1, 2, …, m, j = 1, 2, …, m; x ij represents the allocation decision variable, if the i th fire unit attacks the j th target, x ij = 1, otherwise x ij = 0; p ij represents the damage probability of the i th fire unit to the j th target; w j represents the threat value of the j th target; st. p ij →w j The model of the threat degree random allocation criterion is represented as follows: st. p ij →w i+1 The model of the maximum damage probability minimum unit criterion is represented as follows: The model of the maximum cost-effective ratio criterion is represented as follows: where P dj represents the preset damage probability threshold of the jth target, P j represents the joint damage probability of the allocated firepower to the jth target, is the mean of P j . The model of the escape time and remaining balance criterion is represented as follows: The model of the minimum total time criterion is represented as follows: where t ij represents the escape time of the jth target from the striking area of the ith firepower unit, is the expected damage probability threshold of the jth target; 3. The target intelligent allocation method based on man-machine combined strategy learning according to claim 2, wherein the model of the target importance is represented as follows: Wherein, t mij represents the time of the jth target staying in the i th firepower unit killing area, t zh represents the time of turning firepower. The model of the remaining flight time is represented as follows: ​ w(j) = 1 - 0.1 I j ,1≤I j ≤3 wherein I j represents the guard priority of the jth target; ​ wherein T t1 , T t2 , T t3 , T t4 are four preset time-of-flight thresholds of different sizes; The model of the shutdown point speed is expressed as follows: Wherein, T v1 , T v2 , T v3 are three different flight speed thresholds set in advance The model of the maximum height is expressed as follows: wherein T h1 , T h2 , T h3 , T h4 , T h5 , T h6 are six flight altitude thresholds of different sizes set in advance.

4. A target intelligent allocation system based on human-combined strategy learning for implementing the method of claim 1, characterized in that, The method comprises the following steps: A criterion model construction module is used to model and train a target allocation criterion model based on an artificial experience criterion strategy sample library; An AHP-based quantitative sample library is used to model and train a target characteristic quantitative model; A target allocation module is used to perform target allocation modeling optimization according to a task demand and a target situation input, a target allocation criterion model obtained by the criterion model construction module, and a target characteristic quantitative model obtained by the quantitative model construction module, to obtain a target allocation result.

Citation Information

Patent Citations

  • Unmanned chariot team firepower distribution method based on deep reinforcement learning

    CN112364972A

  • Cooperative target allocation design method for unmanned aerial vehicle cluster in uncertain environment

    CN113031650A