Explainable air combat decision-making methods and devices
Patent Information
- Application Number
- CN202211457546.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-21
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2042-11-21
AI Technical Summary
[0004]传统自主空战机动决策算法的研究可以大致分为基于博弈论的方法、基于优化理论的方法和基于数据驱动的方法,然而上述方法普遍存在准确性较低且模型可解释性较差等问题
[0017]从上面所述可以看出,本公开实施例提供的可解释的空战决策方法及装置,该方法包括:从空战环境中获取当前态势信息;其中,空战环境中存在第一战斗机和与第一战斗机敌对的第二战斗机;根据当前态势信息,从第一战斗机对应的规则库中确定与当前态势信息对应的规则,并根据规则构建得到第一战斗机对应的匹配集;通过第二战斗机行为预测模型,得到第二战斗机的预测行为,根据预测行为和匹配集确定第一战斗机的备选机动动作,并根据备选机动动作和匹配集构建得到第一战斗机对应的机动动作集;根据机动动作集,控制第一战斗机执行机动动作集对应的机动动作。通过应用本公开,不仅能够提高空战决策的对抗性能,还能够获得具有泛化性和可解释性的空战决策规则。
Smart Images

Figure CN116048107B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to an interpretable air combat decision-making method and apparatus. Background Technology
[0002] This section is intended to provide background or context for the embodiments of this disclosure as set forth in the claims. The description herein is not intended to be a prior art simply because it is included in this section.
[0003] Air combat decision-making, due to its highly adversarial and dynamic nature, presents a formidable challenge in the field of adversarial decision-making. From early methods based on statistical decision-making and knowledge reasoning, to model-based optimal decision-making methods, and then to artificial intelligence-based methods, the exploration of autonomous air combat technology for fighter jets is increasingly trending towards novel intelligent approaches.
[0004] Research on traditional autonomous air combat maneuver decision-making algorithms can be broadly categorized into game theory-based methods, optimization theory-based methods, and data-driven methods. However, these methods generally suffer from low accuracy and poor model interpretability. Summary of the Invention
[0005] In view of this, the purpose of this disclosure is to propose an interpretable air combat decision-making method and apparatus, which at least to some extent solves the problems of decision-making efficiency and the interpretability of the decision-making process in related technologies.
[0006] To achieve the above objectives, this exemplary embodiment provides an interpretable air combat decision-making method, including:
[0007] Obtain current situation information from an air combat environment; wherein, the air combat environment contains a first fighter jet and a second fighter jet hostile to the first fighter jet;
[0008] Based on the current situation information, a rule corresponding to the current situation information is determined from the rule base corresponding to the first fighter jet, and a matching set corresponding to the first fighter jet is constructed based on the rule;
[0009] The predicted behavior of the second fighter jet is obtained through the behavior prediction model of the second fighter jet. Based on the predicted behavior and the matching set, the alternative maneuvers of the first fighter jet are determined. Based on the alternative maneuvers and the matching set, the maneuver set corresponding to the first fighter jet is constructed.
[0010] Based on the set of maneuvers, the first fighter jet is controlled to perform the maneuvers corresponding to the set of maneuvers.
[0011] Obtain execution rewards from the air combat environment; update the previous maneuver set and the preparatory maneuver set based on the execution rewards.
[0012] Based on the same inventive concept, an exemplary embodiment of this disclosure also provides an air combat decision-making device, including:
[0013] The situation information acquisition module is configured to acquire current situation information from an air combat environment; wherein, the air combat environment contains a first fighter jet and a second fighter jet hostile to the first fighter jet;
[0014] The matching set construction module is configured to determine the rules corresponding to the current situation information from the rule base corresponding to the first fighter jet based on the current situation information, and construct the matching set corresponding to the first fighter jet based on the rules;
[0015] The maneuver action set construction module is configured to predict the predicted behavior of the second fighter jet through the behavior prediction model of the second fighter jet, determine the candidate maneuvers of the first fighter jet based on the predicted behavior and the matching set, and construct the maneuver action set corresponding to the first fighter jet based on the candidate maneuvers and the matching set.
[0016] The maneuver action set execution module is configured to control the first fighter jet to perform the maneuver actions corresponding to the maneuver action set according to the maneuver action set.
[0017] As can be seen from the above description, the interpretable air combat decision-making method and apparatus provided in this disclosure include: acquiring current situation information from an air combat environment; wherein, the air combat environment contains a first fighter jet and a second fighter jet hostile to the first fighter jet; determining rules corresponding to the current situation information from a rule base corresponding to the first fighter jet based on the current situation information, and constructing a matching set corresponding to the first fighter jet based on the rules; obtaining the predicted behavior of the second fighter jet through a behavior prediction model of the second fighter jet, determining alternative maneuvers of the first fighter jet based on the predicted behavior and the matching set, and constructing a set of maneuvers corresponding to the first fighter jet based on the alternative maneuvers and the matching set; and controlling the first fighter jet to execute the maneuvers corresponding to the set of maneuvers based on the set of maneuvers. By applying this disclosure, not only can the adversarial performance of air combat decision-making be improved, but also air combat decision-making rules with generalizability and interpretability can be obtained. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in this disclosure or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1A flowchart illustrating an interpretable air combat decision-making method provided in an embodiment of this disclosure;
[0020] Figure 2 Another flowchart illustrating the interpretable air combat decision-making method provided in the embodiments of this disclosure;
[0021] Figure 3 A schematic diagram of the situation in an air combat environment provided by embodiments of this disclosure;
[0022] Figure 4 A schematic diagram of the no-escape zone for a short-range air-to-air missile provided in an embodiment of this disclosure;
[0023] Figure 5 A schematic diagram of simulation experiment results provided in the embodiments of this disclosure;
[0024] Figure 6 Another schematic diagram illustrating the simulation experiment results provided in the embodiments of this disclosure;
[0025] Figure 7 Another schematic diagram illustrating the simulation experiment results provided in the embodiments of this disclosure;
[0026] Figure 8 A schematic diagram of the air combat decision-making device provided in an embodiment of this disclosure;
[0027] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this disclosure clearer, the principles and spirit of this disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are given merely to enable those skilled in the art to better understand and implement this disclosure, and are not intended to limit the scope of this disclosure in any way. Rather, these embodiments are provided to make this disclosure more thorough and complete, and to fully convey the scope of this disclosure to those skilled in the art.
[0029] In this article, it is important to understand that any number of elements in the accompanying figures is for illustrative purposes and not for limitation, and any naming is for distinction only and has no limiting meaning.
[0030] The principles and spirit of this disclosure will be explained in detail below with reference to several representative embodiments.
[0031] refer to Figure 1 This is a flowchart illustrating an interpretable air combat decision-making method provided in the embodiments of this disclosure.
[0032] An explainable air combat decision-making method includes the following steps:
[0033] Step S110: Obtain current situation information from the air combat environment; wherein, the air combat environment contains a first fighter jet and a second fighter jet hostile to the first fighter jet.
[0034] In this context, an air combat environment can be defined as an environment in which a first fighter jet engages in aerial combat with the aim of shooting down a second fighter jet and gaining air superiority. That is, the first fighter jet is our aircraft, and the second fighter jet is the enemy aircraft.
[0035] Current situation information can be data that dynamically reflects the offensive and defensive actions of the first and second fighter jets in the air combat environment and has temporal and spatial relationships.
[0036] Step S120: Based on the current situation information, determine the rule corresponding to the current situation information from the rule base corresponding to the first fighter jet, and construct the matching set corresponding to the first fighter jet based on the rule.
[0037] The rule base contains several rules, each of which contains a corresponding precondition and an execution maneuver. Once the precondition is triggered, the execution maneuver is executed.
[0038] In some exemplary embodiments, the rule includes antecedents and executed maneuvers, and the antecedents and executed maneuvers are mapped to a gain prediction vector, prediction error, fitness, maneuver set size, empirical parameter value, and quantity parameter value to achieve preliminary modeling of air combat decision rules.
[0039] As an example: situation information is represented as s, in binary encoding, with each bit being 0 or 1; the rule base is represented as [P]; the rule is represented as cl, the antecedent is represented as c, the execution of the maneuver is represented as a, the profit prediction vector is represented as p, the prediction error is represented as ∈, the fitness is represented as F, the empirical parameter is represented as exp, the size of the maneuver set is represented as as, and the quantity parameter is represented as num.
[0040] Then, mapping the antecedent and the executed maneuver to the revenue prediction vector, prediction error, fitness, maneuver set size, empirical parameter value, and quantity parameter value can be expressed as:
[0041] <c,a>→<p,∈,F,as,exp,num>;
[0042] In this context, the antecedent c of rule cl consists of the triple {0, 1, #}, where # is a wildcard that can match either 0 or 1. Each element p[o] in the profit prediction vector p represents the predicted profit of the first fighter jet after the first fighter jet performs maneuver a and the second fighter jet performs maneuver o.
[0043] In some exemplary embodiments, the step of determining the rule corresponding to the current situation information from the rule base corresponding to the first fighter jet based on the current situation information, and constructing a matching set corresponding to the first fighter jet based on the rule, includes:
[0044] The non-wildcard portion of the antecedent of all the rules in the rule base is compared and matched with the current situation information, and the selected rules are added to the matching set.
[0045] In practice, when situation information s is obtained, the situation information s is compared with the non-wildcard part of the antecedent c of all rules cl in the rule base [P], and the selected rules cl are combined into a matching set [M].
[0046] In this process, the maneuver 'a' in rule cl of the matching set [M] is taken as a candidate maneuver, and the matching set [M] contains the set of maneuvers corresponding to all candidate maneuvers. This matching mechanism ensures that a set of rules with different generalization levels can be obtained for each given situational information 's'.
[0047] In some exemplary embodiments, the execution maneuver actions in the rules of the matching set are used as candidate maneuver actions;
[0048] The step of determining the rule corresponding to the current situation information from the rule base corresponding to the first fighter jet based on the current situation information, and constructing the matching set corresponding to the first fighter jet based on the rule, includes:
[0049] Based on the profit prediction vector of the rule and the fitness, determine the predicted profit for each candidate maneuver in the matching set;
[0050] Based on the second fighter jet behavior prediction model, the expected maneuver prediction corresponding to each candidate maneuver is obtained according to the prediction benefit.
[0051] In response to determining that the maximum value in the expected maneuver prediction is less than or equal to a preset expected maneuver prediction threshold, supplementary rules are generated.
[0052] In the air combat maneuver decision-making problem, the rewards obtained from the environment during the initial learning phase are sparse, and feedback may only be received from the environment upon firing. Therefore, during this period, the elements in the rule's reward prediction vector p and the prediction error ∈ are all initially valued at 0. During the update process, the parameter F representing accuracy in the rule will gradually increase because the rule accurately predicts the rewards obtained from the environment. However, these rules do not really play a role in maneuver strategy formation, and only become effective when their empirical parameter exp exceeds the deletion threshold θ.del Only then can it be deleted from [P]. To reduce the impact of such rules on learning efficiency, this disclosure proposes a rule coverage mechanism and its triggering conditions. Specifically, firstly, each candidate maneuver a in [M] is calculated. i Forecasted returns
[0053]
[0054] Where cl.p and cl.F represent the payoff prediction vector and fitness of rule cl, respectively. elements in Representing the first fighter jet in performing maneuvers a i The second fighter jet followed the estimated strategy. When making a decision, a i The predicted gains that can be obtained from the corresponding set of maneuvers.
[0055] Based on this, and combined with the current second fighter jet behavior prediction model, each candidate maneuver a in [M] i The expected maneuver prediction can be expressed as:
[0056]
[0057] This can be understood as: considering the second fighter jet behavior prediction model Then, the maneuvering motion value function in (s,a). When At that time, a rule overriding mechanism is used to generate new rules, in which... The threshold is a predefined value that can be determined based on the specific task.
[0058] In summary, when the matching set [M] does not cover the predefined minimum set of maneuvers or the maximum value of the expected maneuver prediction is less than θ... EAP The rule overriding mechanism is triggered as needed. The rule overriding mechanism is invoked on demand to supplement the rules corresponding to the missing maneuvers in the matching set [M].
[0059] Step S130: Using the second fighter behavior prediction model, predict the predicted behavior of the second fighter, determine the candidate maneuvers of the first fighter based on the predicted behavior and the matching set, and construct the maneuver set corresponding to the first fighter based on the candidate maneuvers and the matching set.
[0060] In some exemplary embodiments, the method for constructing a second fighter jet behavior prediction model includes:
[0061] A sample set is constructed, comprising several samples; wherein the samples include: sample data and label data; the sample data includes training potential information; and the label data includes the training prediction behavior of the second fighter jet corresponding to the training potential information.
[0062] Based on the sample set, a second fighter jet behavior prediction model is constructed and trained using a predetermined machine learning algorithm.
[0063] The predetermined machine learning algorithm can be selected from one or more of the following: Naive Bayes algorithm, decision tree algorithm, support vector machine algorithm, KNN algorithm, neural network algorithm, deep learning algorithm, and logistic regression algorithm.
[0064] In some exemplary embodiments, the method for updating the second fighter jet behavior prediction model includes:
[0065] Acquire the data set corresponding to the second fighter jet, including situational information and behavioral information;
[0066] The second fighter jet behavior prediction model is updated based on the data components.
[0067] The data set, which includes situational information and behavioral information, characterizes the behavioral information of the second fighter jet under different situational information conditions.
[0068] As an example: the behavior of the second fighter jet is represented by o.
[0069] Among them, situational-behavioral information (SBI) of the second fighter jet is collected on a round-by-round basis. i ,o i The second fighter behavior prediction model is updated at the end of each round using the following formula:
[0070]
[0071] in, For the predicted second fighter jet in the situation Execution Action The probability, η e ∈[0,1] is the information entropy constant.
[0072] Based on the foregoing, the constructed second fighter jet behavior prediction model can predict the behavior of a second fighter jet under situational information with similar features to the visit history. Another advantage of this method is that the machine learning algorithm model can learn the latent behavioral characteristics of the second fighter jet. Furthermore, the constructed second fighter jet behavior prediction model can simultaneously participate in maneuver selection and rule parameter updates.
[0073] In some exemplary embodiments, the step of predicting the predicted behavior of the second fighter jet using a second fighter jet behavior prediction model, determining candidate maneuvers for the first fighter jet based on the predicted behavior and the matching set, and constructing a set of maneuvers corresponding to the first fighter jet based on the candidate maneuvers and the matching set includes:
[0074] Obtain the maneuver value function and fuse the second fighter behavior prediction model into the maneuver value function to obtain the fused maneuver value function;
[0075] Based on a specific greedy algorithm, the candidate maneuvers are determined through the fused maneuver value function;
[0076] Add the rule containing the candidate maneuver from the matching set to the maneuver set.
[0077] In specific implementation, the second fighter jet behavior prediction model is integrated to select maneuvers. Optionally, a specific greedy algorithm is the ∈-greedy strategy. Using the ∈-greedy strategy, the first fighter jet's maneuvers are selected and executed based on the expected maneuver prediction of the maneuver value function integrated with the second fighter jet behavior prediction model. This process can be formally represented as:
[0078]
[0079] Where prob is a random variable drawn from a uniform distribution in [0,1], and a random Indicates from the maneuvering space Randomly selected maneuvers.
[0080] Step S140: According to the set of maneuvers, control the first fighter jet to perform the maneuvers corresponding to the set of maneuvers.
[0081] In some exemplary embodiments, after controlling the first fighter jet to perform the maneuver corresponding to the maneuver set according to the maneuver set, the method further includes:
[0082] Obtain execution rewards from the aforementioned air combat environment;
[0083] The set of maneuvers from the previous moment is updated based on the execution reward.
[0084] Specifically, this includes: updated empirical parameters, maneuver set size parameters, quantity parameters, profit prediction vector, prediction error, and fitness.
[0085] In practice, update the set of maneuvers [A]. -1 The parameters of the rules. Select the selected maneuver a according to the above.s Then, the rules in the matching set [M] that contain this maneuver constitute the maneuver set [A]. s After execution, the first fighter jet receives feedback r from the environment and uses it to update the set of maneuvers from the previous moment [A]. -1 The rules in [A]. -1 For each rule cl in the set, since it is selected into the maneuver set, its empirical parameter cl.exp is first incremented by 1, and the maneuver set size parameter cl.as is updated as follows:
[0086]
[0087] In the formula, x represents [A]. -1 In the rule, β1∈(0,1] is the learning rate.
[0088] Target predicted value P t (s)(abbreviated as P) t () indicates that the first fighter jet performed a maneuver in the previous s-1 situation. -1 Then, within the subsequent time period, the expected total discount future return generated by choosing the optimal maneuver strategy. Therefore, P t This can be formally represented as:
[0089]
[0090] Where, r -1 The reward is the value from the previous time step, and γ∈[0,1) is the discount factor. This is the value function estimation of the situation s. Since the reward obtained by the first fighter jet is also related to the behavior of the second fighter jet, it is only related to the maneuvering action o of the second fighter jet at the previous moment. -1 The relevant profit prediction vectors have been updated:
[0091] cl.p[o -1 ]←cl.p[o -1 ]+β2(P t -cl.p[o -1 ])
[0092] Where β2∈(0,1] is the learning rate.
[0093] Prediction error ∈ describes the target prediction P t With the first fighter jet to execute a -1 The difference in expected returns. Due to the second fighter jet behavior prediction model The profit prediction vector p is continuously updated online, and ∈ is updated accordingly. Let cl.ζ represent the profit of rule cl under the current second fighter behavior prediction model, using the expected rule prediction cl.ζ, which can be expressed as:
[0094]
[0095] in, This indicates that the second fighter jet behavior prediction model is in the previous situation s -1 The predicted value. Then the prediction error of the rule can be updated according to the following formula:
[0096] cl.∈←cl.∈+β3(|P t -cl.ζ|-cl.∈)
[0097] Where β3∈(0,1] is the learning rate.
[0098] Fitness F reflects the accuracy of the rule, and its update method can be expressed as:
[0099]
[0100] Where x is the set of maneuvers [A] -1 In the rules, β4∈(0,1] is the learning rate, k(cl) represents the fitness weighted by the accuracy of rule cl, and is a function of the prediction error:
[0101]
[0102] Where α∈(0,1] and Used to further differentiate the accuracy of the rules. ∈0 is a predefined maximum tolerance error threshold.
[0103] In some exemplary embodiments, after updating the previous time-to-time maneuver action set according to the execution reward, the method further includes:
[0104] Determine the average time span from the last time a new rule was generated in the previous maneuver action set to the current time.
[0105] In response to determining that the average time span is greater than a preset average time span threshold, based on the fitness of the rule, two rules are selected from the set of maneuvers at the previous moment as parent rules. Based on the antecedent of the parent rules, an antecedent of a child rule matching the situation information is created, and the new rule is generated based on the antecedent of the child rule.
[0106] In practice, new rules are generated based on a genetic algorithm (GA). After the parameters are updated, the movement action set [A] is... -1 New rules are generated using GA. First, [A] is calculated. -1 The rule specifies the average time span from the last application of GA to the current time. When the calculated result exceeds a predefined threshold θ, the rule is applied. GAThe GA will only be triggered again at a certain time. This triggering condition can be formally described as:
[0107]
[0108] Where cl.ts represents the timestamp when the GA was last applied to the action set where rule cl (formerly) is located, cl.num is the number of rules in the rule base [P] that are consistent with the cl parameter, and t is the current time.
[0109] Once the triggering condition is met, based on the fitness F of the rule, use the roulette wheel method to select from [A]. -1 Two rules are selected as parent rules. Then, based on the antecedents of the two parent rules, crossover and mutation operators are used to create a situation that can still match the situation. -1 The antecedent of the offspring rule. Specifically, using the two-point crossover operator with probability χ, the rule's maneuvering action remains unchanged. Using the single-point mutation operator with probability μ, this mutation operation can only change the mutated bit to # or still be able to interact with s. -1 The specific values that match. Finally, initialize the other parameters of the generated child rules.
[0110] In some exemplary embodiments, after generating the new rule based on the antecedent of the child rule, the method further includes:
[0111] In response to determining that the prediction error of the parent rule is less than a preset prediction error threshold and the number of times it is selected by the set of maneuvers is greater than a preset number threshold, the child rule is merged into the parent rule, and the merged child rule is deleted.
[0112] In some exemplary embodiments, the method further includes:
[0113] Construct a data set including the situational information, the maneuvers of the first fighter jet, and the actions of the second fighter jet;
[0114] The rules that match the data group are selected from the rule base to construct a preliminary set of maneuvers, and the preliminary set of maneuvers is updated according to the eligibility value of the data group and the accuracy of the rules.
[0115] In practice, new rules are integrated. The GA fusion mechanism checks whether the generated child rules can be fused with their parent rules. Specifically, when the parent rule is sufficiently accurate (cl. ∈ < ∈ 0) and has been selected into the maneuver action set a sufficient number of times (cl. exp > θ), the fusion is performed. sub θ subWhen a predefined fusion threshold is met, newly generated child rules will be fused, and the parent rule count parameter cl.num will be incremented by 1. Alternatively, if a child rule has not been fused by any parent rule, and its predecessor already exists in the rule base [P], then the existing rule count parameter cl.num will be incremented by 1, and the newly generated child rule will be fused. Otherwise, the newly generated child rule will be added to [P].
[0116] Another fusion mechanism employed is the set of maneuvering actions [A]. -1 The rules are integrated. First, in [A]... -1 Find the one with the strongest generalization ability (i.e., the antecedent c contains the most #), sufficient accuracy (cl. ∈ < ∈ 0), and a sufficient number of times it is selected into the set of maneuvers (cl. exp > θ). sub The rule is as follows. Secondly, try to fuse [A] using this rule. -1 The remaining rules in [A] are then calculated, and the quantity parameter cl.num is accumulated. The rules to be merged will then be derived from [A]. -1 The rules are removed from the rule base [P]. This fusion mechanism effectively controls the number of macro rules in [P], ensuring rule quality and the matching speed of the rule base.
[0117] In practice, the trace set and eligibility values are updated. The method maintains a historical trace set Ξ and a trace e during the learning process to record the information needed for rule updates. Specifically, the trace set records multiple joint state-maneuver triples (s, a, o), where... As a situation, and These represent the maneuvers of the first and second fighter jets, respectively. The trace e records the qualification value corresponding to the triple (s, a, o) in Ξ.
[0118] The first and second fighter jets respectively perform maneuvers at s and a. s and o s Afterwards, if Ξ does not include (s,a) s ,o s If the condition is met, then add it to Ξ. Then, the eligibility value corresponding to the situation s is updated using the following formula:
[0119]
[0120] To improve algorithm efficiency, only non-repeating and sufficiently important tuples are stored in Ξ. Specifically, if the eligibility value corresponding to (s,a,o) is less than the predefined threshold (e(s,a,o)<θ), then... et If (s,a,o) is not found in the trace set Ξ, then (s,a,o) will be removed from the trace set Ξ. To achieve the above function, after each round of learning, the qualification value corresponding to (s,a,o) is decayed as follows:
[0121] e(s,a,o)=λ·γe ·e(s,a,o)
[0122] Where, λ·γ e The proximity of the tuple (s, a, o) is defined, γ. e ∈[0,1) is the decay factor, and λ∈(0,1) is the trace decay parameter, which determines the decay rate of the trace. Through the above mechanism, the space complexity of each iteration will not increase linearly with the appearance of a new tuple.
[0123] In practice, update the preparatory maneuver action set [A]. et Parameters of the rules. Motion set [A] -1 After the rules in the rule base [P] are updated, the rules that can match the historical trace set are updated to different degrees according to their accuracy and corresponding qualification value. This can be understood as the contribution of the historical trace to the agent's current state.
[0124] First, for each record (s, a, o) in Ξ, search in [P] for rules that match the situation s and the first fighter maneuver a, and combine the matching results into a set of preliminary maneuvers [A]. et Then, based on the eligibility value pair [A] in e... et The rules in the Chinese system have been updated.
[0125] When allocating TD errors, the accuracy of the rule is considered; the higher the accuracy, the greater the degree of update. Since the fitness parameter cl.F of the rule represents its accuracy, the normalized fitness is used as the relative accuracy of the rule. Therefore, the update methods for the profit prediction p and the prediction error ∈ are as follows:
[0126]
[0127]
[0128] Where x represents [A] et In the rules, β5∈(0,1] and β6∈(0,1] are the learning rates.
[0129] First, the set of maneuvers from the previous moment is updated based on environmental feedback, information on the behavior of the second fighter jet, and predictions of expected maneuvers [A]. -1 The parameters of the rule, and in [A] -1 GA is used to generate and evolve new rules.
[0130] Then, based on the accuracy-based eligibility trace mechanism, rules that match historical traces are selected from the rule base [P] to form a set [A]. et It will be updated based on eligibility scores and rule accuracy.
[0131] refer to Figure 2 In this disclosure, data modules containing rules are represented by rounded rectangles, while main functional modules are represented by rectangles. This indicates the maneuver library of our aircraft (the first fighter jet). This indicates the maneuver library of enemy aircraft (secondary fighter).
[0132] The scheme can be divided into two stages: maneuver selection and rule updating. Thin solid arrows represent the data flow for maneuver selection, while thick solid arrows represent the data flow for rule updating. The maneuver selection stage is primarily responsible for interacting with the environment. It receives situational information s and uses matching and rule coverage mechanisms in the rule base [P] to filter out rules that match s, forming a matching set [M]. Then, it comprehensively considers online enemy aircraft behavior information to select maneuver a. s Then, construct the corresponding set of maneuvers [A]. Finally, execute the selected maneuver and obtain a reward r from the environment.
[0133] The rule update phase is mainly responsible for parameter updates and rule evolution, and it consists of three main parts. First, the set of maneuvers from the previous moment is updated based on environmental feedback, enemy aircraft behavior information, and expected maneuver predictions [A]. -1 The parameters of the rule, and in [A] -1 GA is used to generate and evolve new rules. Then, based on an accuracy-based eligibility trace mechanism, rules that match historical traces are selected from the rule base [P] to form a set [A]. et It will be updated based on eligibility scores and rule accuracy.
[0134] In related technologies, air combat maneuver decision-making models based on deep reinforcement learning are a "black box," lacking interpretability and security in their strategies. Autonomous air combat maneuver decision-making algorithms, however, take situational information from both sides as input, combining our aircraft's maneuverability, available airborne weapons, and combat targets. Through mathematical modeling, online reasoning, rule matching, intelligent optimization, and deep neural networks, they select the optimal decision-making scheme from a maneuver action library or tactical library. Traditional research on autonomous air combat maneuver decision-making algorithms can be broadly categorized into game theory-based methods, optimization theory-based methods, and data-driven methods. However, these methods generally suffer from: strong reliance on expert rules or labeled data; the need for accurate and complete air combat situation assessment models; and complex solutions and low real-time performance. Furthermore, the state space of the autonomous air combat maneuver decision-making environment is vast, rewards are sparse, and the deep reinforcement learning-based strategy model is a "black box," making it difficult for humans to understand the rule mapping relationships. This limits its application in fields such as finance and the military, where security and interpretability are paramount. The field of air combat maneuver decision-making is complex, and the interpretability of decision-making models is a problem that must be solved in future research in order to establish a reliable human-machine mutual trust mechanism.
[0135] As described above, this disclosure firstly redesigns the model in terms of rule representation, maneuver selection, and rule evolution. Secondly, enemy aircraft behavior information is incorporated into the rule representation, and a neural network-based enemy aircraft behavior prediction model is constructed. Combined with the rule antecedent representation mechanism, the enemy aircraft behavior prediction model can guide our aircraft's maneuver selection under relative enemy-friendly situations with similar feature representations. Then, a maneuver set parameter update mechanism is designed based on the enemy aircraft behavior prediction model, and an accuracy-based qualification trace mechanism is proposed to accelerate rule evolution. Furthermore, this disclosure can obtain a maneuver decision rule base with accuracy, generalization, and interpretability while ensuring our aircraft's combat performance.
[0136] In some exemplary embodiments, the effectiveness of this disclosure is verified using an in-visual-range 1v1 air combat maneuver environment employing short-range air-to-air missiles as an example.
[0137] Environment settings
[0138] Example of air combat maneuver decision-making at a certain moment in the adversarial environment: Figure 3 As shown, the fighter jet adopts a six-degree-of-freedom model and is confined to a three-dimensional space of 200km×200km×20km. In this experiment, the fighter jet's onboard weapon is a short-range air-to-air missile, and the relevant parameter settings are shown in Table 1, based on existing research. The initial situation is that two enemy and friendly aircraft are flying towards each other, with initial position coordinates of [95km, (100km, (5km]) and [110km, (100km, (5km]). The initial speed of both aircraft is 200m / s, and the initial roll angle and track inclination angle are 0°. The enemy aircraft's track deflection angle is 0°, and the friendly aircraft's track deflection angle is 180°. Under this initial situation, the two sides are in a state of equilibrium. The round ends when one side gains the opportunity to fire, stalls (<80m / s) or overspeeds (>400m / s), exceeds the maximum allowable altitude (18km) or falls below the minimum altitude (0.2km), or reaches the maximum time constraint (500 time steps). The positions and attitude angles of both aircraft are then reset.
[0139] Table 1 (Parameter settings for fighter jets equipped with short-range air-to-air missiles)
[0140]
[0141]
[0142] The maneuver space of a fighter jet comprises seven basic fighter jet maneuvers proposed by NASA: constant speed flight (a0), maximum G-force left turn (a1), maximum G-force right turn (a2), maximum G-force climb (a3), maximum G-force dive (a4), maximum acceleration flight (a5), and maximum deceleration flight (a6). The trajectory and attitude control of a fighter jet can be translated into control of the tangential overload η.x Normal overload η f and roll angle μ p Control.
[0143] Relative situation construction
[0144] Selected relative situation information [q r ,q b ,β,d,v R ,Δv,h R ,Δh] serves as the situational input for the decision-making model, such as Figure 3 As shown. Where q r Let q be the target azimuth angle. b Let β be the target heading angle, β be the angle between the two aircraft's speeds, d be the distance between the two aircraft, and v be the target heading angle. R Let h be the speed of our aircraft, Δv be the speed difference between our aircraft and the enemy aircraft, and h be the speed of our aircraft. R Let Δh be the altitude of our aircraft, and Δh be the altitude difference between our aircraft and the enemy aircraft. The parameters in the relative situation information can be calculated using the following formula:
[0145]
[0146]
[0147]
[0148]
[0149] Δv=v R -v B
[0150] Δh=z R -z B
[0151] h R =z R
[0152] In the formula, R represents the relevant parameters of our aircraft, B represents the relevant parameters of the enemy aircraft, (x,y,z) represents the position coordinates of the fighter's center of mass in the coordinate system, and ψ p γ is the track deflection angle. p This refers to the track inclination angle.
[0153] Taking into account, such as Figure 4 The diagram illustrates the no-escape zone for short-range air-to-air missiles and the situational relationship between the enemy and friendly forces. Taking our aircraft as an example, if q is satisfied... r <30°,q b If the angle is greater than 120°, β is less than 45°, and 1.5km is less than d and 4.0km, then our aircraft has the opportunity to fire.
[0154] Experimental setup
[0155] Table 2 (q) r and q b Discretization and Encoding
[0156] q r , q b ]]> [0°,30°) [30°,60°) [60°,90°) [90°,120°) [120°,150°) [150°,180°]
[0157] Table 3 (β, d, v) R ,Δv,h R Discretization and encoding of Δh
[0158] β [0°,45°) [45°,90°) [90°,135°) [135°,180°] d [0km, 1km) [1km, 4km) [4km, 7.5km) [7.5km, +∞) <![CDATA[v R ]]> [0m / s, 200m / s) [200m / s, 300m / s) [300m / s, 400m / s) [400m / s,+∞) Δv (-∞, 50m / s) [-50m / s, 0m / s) [0m / s, 50m / s) [50m / s,+∞) <![CDATA[h R ]]> [0km, 1km) [1km, 5km) [5km, 8km) [8km,+∞) Δh (-∞, -3km) [-3km, 0km) [0km, 3km) [3km,+∞)
[0159] First, the continuous situational information is discretized and encoded. As shown in Tables 2 and 3, q r and q b The initial situation information is represented by 3 binary bits, and the remaining situation information is represented by 2 binary bits. Therefore, after discretizing and encoding the relative situation information, the state information consists of 18 binary bits. The initial situation of this environment can be represented as 000 000 11 11 01 10 1010.
[0160] Secondly, the reward function is designed. The reward function consists of two parts: environmental reward and situational advantage reward. When our aircraft gains a firing opportunity, it receives a positive reward; when our aircraft stalls or overspeeds, exceeds altitude limits, or when an enemy aircraft gains a firing opportunity, it receives a penalty from the environment, and the round ends. Therefore, the environmental reward function f... env for:
[0161]
[0162] To accelerate strategy learning and mitigate the impact of sparse rewards in air combat maneuvering environments, this paper references existing methods for assessing the air combat situation of fighter jets equipped with air-to-air missiles, and incorporates relative situational advantage (angle advantage f) into the assessment. a Distance advantage f d Speed advantage v High altitude advantage h Introducing a reward function design, the reward derived from the composite advantage function can be expressed as:
[0163] f situation =ω a f a +ω d f d +ω v f v +ω h f h
[0164] Where, ω a +ω d +ω v +ω h =1, ω aω d ω v and ω h The weights for angle advantage, distance advantage, speed advantage, and altitude advantage are respectively assigned. Then, the adversarial environment reward f is comprehensively considered. env and situational advantage reward f situation The real-time reward obtained by my machine can be represented as:
[0165]
[0166] Among them, the situational advantage coefficient ρ s =0.1. The enemy aircraft selects the maneuver corresponding to gaining the maximum combined situational advantage based on the relative situational advantage function.
[0167] Experimental results
[0168] Figure 5 , Figure 6 and Figure 7 The learned decision model is presented for use in a 1v1 air combat maneuver decision-making environment, including the dual-aircraft combat trajectory and relative situation information (target azimuth q) closely related to firing conditions. r Target heading angle q b The curves showing the changes in the angle β between the two aircraft's speeds and the distance d between them are shown. Initially, the two aircraft are flying towards each other at a distance of 15 km, thus the engagement begins beyond visual range. In the early stages of the engagement, both sides employ accelerated straight-line maneuvers to quickly shorten the distance between the two aircraft. During this phase, q... r =q b =0°, β=180°. Then, the engagement transitioned to close-range air combat. To avoid a collision, our aircraft adjusted its nose direction by making a left turn. The enemy aircraft, seeking maximum situational advantage, selected maneuvers based on the current situation (e.g., Figure 5 The aircraft then entered close-range dogfighting, with each aircraft taking turns ascending and descending due to angular advantages. Ultimately, the Chinese aircraft gained the opportunity to fire, and the corresponding relative situation information is shown in Table 4, satisfying the firing conditions. Statistically, in this combat scenario, when the Chinese aircraft used the air combat model provided in this disclosure for decision-making, the combat win rate was approximately 83%, while the enemy aircraft's win rate was 11.5%. After 500,000 rounds of training, the air combat rule base [P] provided in this disclosure retained 677 macro rules, of which 620 were previously used for maneuver selection (i.e., cl.exp > 0).
[0169] Table 4 (Situational Information at the End of the Confrontation)
[0170]
[0171] Table 5 (Examples of rules generated after learning)
[0172]
[0173]
[0174] Table 5 provides the implementation details. Figure 5 , Figure 6 and Figure 7 Examples of rules corresponding to the adversarial trajectory. Combining the encoding information in Tables 2 and 3, the first rule can be interpreted as: if q r ∈[30°,60°), q b ∈[150°,180°], β∈[45°,90°), d∈[1km, 4km), h R If the distance is ∈ [5km, 8km), then our aircraft should execute the maneuver "maximum overload left turn". According to the aforementioned angular positional relationship, in this situation, our aircraft is in a superior position behind the enemy aircraft. In the framework of a zero-sum game, from our aircraft's decision-making perspective, the enemy aircraft, in order to minimize the gains our aircraft gain, should execute the maneuver a0 "uniform speed flight" according to the gain prediction vector of this rule. At this time, our aircraft's gain prediction is 0. Therefore, Rule 1 recommends that our aircraft should turn left at this time to further establish angular advantage, thereby reaching the firing condition as soon as possible.
[0175] One interpretation that matches rule 2 is: if q r ∈[150°,180°], q b ∈[60°,90°), β∈[0°,45°), d∈[1km,4km), v R If ∈[200m / s,300m / s) and Δh>0, then our aircraft should perform the maneuver "maximum overload right turn". At this time, our aircraft is being pursued by the enemy aircraft, and the distance and speed angle meet the enemy aircraft's firing conditions. Since our aircraft has an altitude advantage at this time, the rule suggests that our aircraft should turn right to get rid of the enemy aircraft's pursuit and destroy the enemy's angle and distance advantage.
[0176] Similarly, one interpretation that matches rule 3 is: if q r ∈[0°,30°), q b If ∈[120°,180°], d>4km, Δh>0, then our aircraft should execute the maneuver "maximum overload dive". At this time, our aircraft is tailing the enemy aircraft, and the distance between the two aircraft has exceeded the constraints of the firing conditions. At the same time, our aircraft has an altitude advantage. Therefore, our aircraft can convert potential energy into kinetic energy through the dive maneuver, gain speed, and further reduce the distance with the enemy aircraft, thereby gradually achieving the firing conditions. The maneuver corresponding to rule 4 is "maximum overload right turn", and the situation angle relationship described is: q r ∈[60°,90°), q b ∈[150°,180°], at this point, our machine can further reduce q through the maneuvers recommended by this rule. rThis establishes an advantage in terms of perspective.
[0177] Furthermore, generalization rules with high accuracy were learned, such as rules 5 through 9, whose antecedents are basically composed of wildcards #. Rule 7, in particular, achieved a fitness of 0.997, indicating high accuracy. Rule 9 describes the situational relationship as follows: the distance between the two aircraft, d, is greater than 7.5 km, far exceeding the prescribed firing distance, and our aircraft's speed, v... R The speed range is [0 m / s, 300 m / s), which does not meet the optimal flight speed specified in Table 1. Therefore, our aircraft can shorten the distance between the two aircraft by accelerating and increasing its speed, thereby gaining a speed advantage and accumulating energy to complete the next maneuver.
[0178] Note that the above explanations of the rules are all intuitive and are only intended to illustrate the potential for obtaining interpretable decision-making models in air combat maneuver decision-making problems. In the application of the rule-based [P] model, each decision cannot be made by a single rule, but requires comprehensive consideration of all rules that match the current situation, combined with parameters such as the fitness and payoff prediction vector of each rule, to select and execute the maneuver.
[0179] In summary, by leveraging the designed rule representation, maneuver selection, and rule evolution mechanisms, while ensuring our aircraft's combat performance and learning efficiency, we can obtain a maneuver decision rule base that is accurate, generalizable, and interpretable.
[0180] It should be noted that the method of this disclosure embodiment can be executed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method of this disclosure embodiment, and the multiple devices will interact with each other to complete the method described.
[0181] It should be noted that the above description describes some embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in a different order than those in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0182] Based on the same inventive concept, corresponding to any of the above-described embodiments, this disclosure also provides an air combat decision-making device.
[0183] refer to Figure 8 The air combat decision-making device includes:
[0184] The situation information acquisition module 810 is configured to acquire current situation information from an air combat environment; wherein, the air combat environment contains a first fighter jet and a second fighter jet hostile to the first fighter jet;
[0185] The matching set construction module 820 is configured to determine the rules corresponding to the current situation information from the rule base corresponding to the first fighter jet based on the current situation information, and construct the matching set corresponding to the first fighter jet based on the rules;
[0186] The maneuver action set construction module 830 is configured to predict the predicted behavior of the second fighter jet through the second fighter jet behavior prediction model, determine the candidate maneuvers of the first fighter jet based on the predicted behavior and the matching set, and construct the maneuver action set corresponding to the first fighter jet based on the candidate maneuvers and the matching set.
[0187] The maneuver action set execution module 840 is configured to control the first fighter jet to perform the maneuver actions corresponding to the maneuver action set according to the maneuver action set.
[0188] In some exemplary embodiments, the rule includes an antecedent;
[0189] Match set building module 820 is specifically configured as follows:
[0190] The non-wildcard portion of the antecedent of all the rules in the rule base is compared and matched with the current situation information, and the selected rules are added to the matching set.
[0191] In some exemplary embodiments, the rule further includes performing maneuver actions, mapping the antecedent and the performed maneuver actions to a profit prediction vector, prediction error and fitness, maneuver action set size, empirical parameter value, and quantity parameter value, and using the performed maneuver actions in the rules of the matching set as candidate maneuver actions; then, the matching set construction module 820 is specifically configured as follows:
[0192] Based on the profit prediction vector of the rule and the fitness, determine the predicted profit for each candidate maneuver in the matching set;
[0193] Based on the second fighter jet behavior prediction model, the expected maneuver prediction corresponding to each candidate maneuver is obtained according to the prediction benefit.
[0194] In response to determining that the maximum value in the expected maneuver prediction is less than or equal to a preset expected maneuver prediction threshold, supplementary rules are generated.
[0195] In some exemplary embodiments, the motion set construction module 830 is specifically configured as follows:
[0196] Obtain the maneuver value function and fuse the second fighter behavior prediction model into the maneuver value function to obtain the fused maneuver value function;
[0197] Based on a specific greedy algorithm, the candidate maneuvers are determined through the fused maneuver value function;
[0198] Add the rule containing the candidate maneuver from the matching set to the maneuver set.
[0199] In some exemplary embodiments, the air combat decision-making device is further configured to:
[0200] Obtain execution rewards from the aforementioned air combat environment;
[0201] The set of maneuvers from the previous moment is updated based on the execution reward.
[0202] In some exemplary embodiments, the air combat decision-making device is further configured to:
[0203] Determine the average time span from the last time a new rule was generated in the previous maneuver action set to the current time.
[0204] In response to determining that the average time span is greater than a preset average time span threshold, based on the fitness of the rule, two rules are selected from the set of maneuvers at the previous moment as parent rules. Based on the antecedent of the parent rules, an antecedent of a child rule matching the situation information is created, and the new rule is generated based on the antecedent of the child rule.
[0205] In some exemplary embodiments, the air combat decision-making device is further configured to:
[0206] In response to determining that the prediction error of the parent rule is less than a preset prediction error threshold and the number of times it is selected by the set of maneuvers is greater than a preset number threshold, the child rule is merged into the parent rule, and the merged child rule is deleted.
[0207] In some exemplary embodiments, the air combat decision-making device is further configured to:
[0208] Construct a data set including the situational information, the maneuvers of the first fighter jet, and the actions of the second fighter jet;
[0209] The rules that match the data group are selected from the rule base to construct a preliminary set of maneuvers, and the preliminary set of maneuvers is updated according to the eligibility value of the data group and the accuracy of the rules.
[0210] In some exemplary embodiments, the air combat decision-making device is further configured to:
[0211] A sample set is constructed, comprising several samples; wherein the samples include: sample data and label data; the sample data includes training potential information; and the label data includes the training prediction behavior of the second fighter jet corresponding to the training potential information.
[0212] Based on the sample set, a second fighter jet behavior prediction model is constructed and trained using a predetermined machine learning algorithm.
[0213] For ease of description, the above apparatus is described in terms of its functions, divided into various modules. Of course, in implementing this disclosure, the functions of each module can be implemented in one or more software and / or hardware.
[0214] The apparatus of the above embodiments is used to implement the corresponding interpretable air combat decision-making method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0215] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the interpretable air combat decision-making method described in any of the above embodiments.
[0216] Figure 9 This embodiment illustrates a more specific hardware structure of an electronic device, which may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.
[0217] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0218] The memory 1020 can be implemented in the form of ROM (Read-Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0219] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.
[0220] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0221] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.
[0222] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.
[0223] The electronic devices described above are used to implement the corresponding interpretable air combat decision-making methods in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0224] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the interpretable air combat decision-making method as described in any of the above embodiments.
[0225] The aforementioned non-transitory computer-readable storage media can be any available medium or data storage device that a computer can access, including but not limited to magnetic storage (e.g., floppy disks, hard disks, magnetic tapes, magneto-optical disks (MOs), etc.), optical storage (e.g., CDs, DVDs, BDs, HVDs, etc.), and semiconductor storage (e.g., ROMs, EPROMs, EEPROMs, non-volatile memory (NAND (FLASH), solid-state drives (SSDs)), etc.).
[0226] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the interpretable air combat decision-making method as described in any of the embodiments in the exemplary method section above, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0227] Those skilled in the art will recognize that embodiments of this disclosure can be implemented as a system, method, or computer program product. Therefore, this disclosure can be implemented as entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as a "circuit," "module," or "system." Furthermore, in some embodiments, this disclosure can also be implemented as a computer program product contained in one or more computer-readable media, which includes computer-readable program code.
[0228] Any combination of one or more computer-readable media may be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example,, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (not exhaustive) of a computer-readable storage medium may include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device.
[0229] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0230] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0231] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0232] It should be understood that each block of a flowchart and / or block diagram, as well as combinations of blocks in a flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine that, when executed by a computer or other programmable data processing device, creates means for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram.
[0233] These computer program instructions may also be stored in a computer-readable medium that enables a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable medium produce a product comprising an instruction apparatus that implements the functions / operations specified in the boxes of a flowchart and / or block diagram.
[0234] Computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, such that the instructions that execute on the computer or other programmable apparatus can provide a process for implementing the functions / operations specified in the boxes of a flowchart and / or block diagram.
[0235] Furthermore, although the operations of the methods of this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the operations shown must be performed to achieve the desired result. Rather, the steps depicted in the flowcharts may be executed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0236] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this disclosure should have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar terms used in the embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0237] While the spirit and principles of this disclosure have been described with reference to several specific embodiments, it should be understood that this disclosure is not limited to the disclosed specific embodiments, and the division of aspects does not imply that features in these aspects cannot be combined for benefit; such division is merely for convenience of expression. This disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims. The scope of the appended claims is to be interpreted in the broadest sense, thereby encompassing all such modifications and equivalent structures and functions.
Claims
1. An interpretable air combat decision-making method, characterized in that, include: Obtain current situation information from an air combat environment; wherein, the air combat environment contains a first fighter jet and a second fighter jet hostile to the first fighter jet; Based on the current situation information, a rule corresponding to the current situation information is determined from the rule base corresponding to the first fighter jet, and a matching set corresponding to the first fighter jet is constructed based on the rule; The rule includes an antecedent. and perform maneuvers ; The step of determining the rule corresponding to the current situation information from the rule base corresponding to the first fighter jet based on the current situation information, and constructing a matching set corresponding to the first fighter jet based on the rule, includes: Compare and match the non-wildcard portion of the antecedent of all the rules in the rule base with the current situation information, and add the filtered rules to the matching set. The preceding element and the aforementioned execution of maneuvering actions Mapping to the revenue prediction vector Prediction error Adaptability , scale of mobile action Empirical parameter values and quantity parameter values Profit prediction vector Each element in This indicates that the first fighter jet is performing a maneuver. The second fighter jet performed a maneuver. Subsequently, the predicted gain of the first fighter jet will be determined by selecting the executed maneuvers from the rules in the matching set as candidate maneuvers. The step of determining the rule corresponding to the current situation information from the rule base corresponding to the first fighter jet based on the current situation information, and constructing the matching set corresponding to the first fighter jet based on the rule, includes: Based on the profit prediction vector of the rule and the fitness, determine the predicted profit for each candidate maneuver in the matching set. ; ; in, and Representing rules respectively The profit prediction vector and fitness, elements in Representing the first fighter jet in performing maneuvers The second fighter jet followed the estimated strategy. When making decisions, The predicted gains that can be obtained from the corresponding set of maneuvers; Based on the second fighter jet behavior prediction model, and according to the predicted gains, the expected maneuver prediction for each candidate maneuver is obtained; expressed as: ; To consider the second fighter jet behavior prediction model Afterwards, The function of the maneuvering action value; when At that time, a rule overriding mechanism is used to generate new rules, in which... For a predefined threshold; In response to determining that the maximum value in the expected maneuver prediction is less than or equal to a preset expected maneuver prediction threshold, supplementary rules are generated; when the matching set [M] does not cover the predefined minimum maneuver set or the maximum value of the expected maneuver prediction is less than... The rule overriding mechanism is triggered in a timely manner; the rule overriding mechanism is invoked as needed to supplement the rules corresponding to the missing maneuvers in the matching set [M]. The predicted behavior of the second fighter jet is obtained through a second fighter jet behavior prediction model. Based on the predicted behavior and the matching set, candidate maneuvers for the first fighter jet are determined. A set of maneuvers corresponding to the first fighter jet is constructed based on the candidate maneuvers and the matching set, including: obtaining a maneuver value function and fusing the second fighter jet behavior prediction model into the maneuver value function to obtain a fused maneuver value function; based on The strategy involves determining the candidate maneuvers using the fused maneuver value function, and adding the rules containing the candidate maneuvers from the matching set to the maneuver set. Based on the set of maneuvers, the first fighter jet is controlled to perform the maneuvers corresponding to the set of maneuvers.
2. The method according to claim 1, characterized in that, After controlling the first fighter jet to perform the maneuvers corresponding to the maneuver set according to the maneuver set, the method further includes: Obtain execution rewards from the aforementioned air combat environment; The previous action set is updated based on the execution reward, specifically including: updating the experience parameter, action set size parameter, quantity parameter, profit prediction vector, prediction error, and fitness. In practice, the set of maneuvers will be updated. The parameters of the rules are selected according to the above-mentioned maneuvering actions. Then, the rules in the matching set [M] that contain the maneuver constitute the maneuver set [A]. After execution, the first fighter jet received feedback from the environment. And used to update the set of maneuvers from the previous moment. The rules in; for Each rule in Since it was selected into the set of maneuvers, its empirical parameters are as follows: Increment by 1, the size parameter of the motion set Update as follows: In the formula, Substitute The rules in The learning rate; Target predicted value ,Right now This indicates that the first fighter jet was in the previous situation. Execute maneuvers Then, within the subsequent timeframe, the expected total discount future return generated by selecting maneuver actions according to the optimal strategy; therefore, Formal representation: in, As a reward for the previous moment. As a discount factor, In the situation The value function is estimated; since the reward obtained by the first fighter jet is also related to the behavior of the second fighter jet, it is only related to the maneuver of the second fighter jet at the previous moment. The relevant profit prediction vectors have been updated: in, The learning rate; Prediction error Describes target prediction With the first fighter jet Differences in expected returns; due to the second fighter jet behavior prediction model and profit prediction vector Continuously updated online The prediction is updated accordingly; using the expectation rule for forecasting. Representation rules The benefit under the current second fighter jet behavior prediction model is expressed as: in, This indicates that the second fighter jet behavior prediction model is in the previous situation. The predicted value, then the prediction error of the rule. Updates can be performed using the following formula: in, The learning rate; fitness This reflects the accuracy of the rules, and its update method is expressed as follows: in, It is a collection of motor action sequences The rules in For learning rate, Representation rules The accuracy-weighted fitness is a function of the prediction error: in, and Used to further differentiate the accuracy of the rules. This is a predefined maximum tolerance error threshold; After updating the previous time-to-time maneuver action set based on the execution reward, the method further includes: Determine the average time span from the last time a new rule was generated in the previous maneuver action set to the current time. In response to determining that the average time span is greater than a preset average time span threshold, based on the fitness of the rule, two rules are selected from the set of maneuvers at the previous moment as parent rules, and based on the antecedent of the parent rules, an antecedent of a child rule matching the situation information is created, and the new rule is generated based on the antecedent of the child rule. After generating the new rule based on the antecedent of the child rule, the method further includes: In response to determining that the prediction error of the parent rule is less than a preset prediction error threshold and the number of times it is selected by the set of maneuvers is greater than a preset number threshold, the child rule is merged into the parent rule, and the merged child rule is deleted. The method further includes: Construct a data set including the situational information, the maneuvers of the first fighter jet, and the behavior of the second fighter jet; The rules that match the data group are filtered in the rule base to construct a preliminary maneuver action set, and the preliminary maneuver action set is updated according to the eligibility value of the data group and the accuracy of the rules. In practice, the trace set and eligibility values are updated; the method maintains the historical trace set during the learning process. Heji Used to record information required for rule updates; specifically, the trace set records multiple joint state-maneuver triples. ,in As a situation, and These represent the maneuvers of the first and second fighter jets, respectively; tracks Recorded middle The qualification value corresponding to a triple; The first and second fighter jets were in Perform maneuvers separately and Afterwards, if Not included Then add it to Then, the situation The corresponding qualification value is updated according to the following formula: To improve algorithm efficiency Only non-repeating and sufficiently important tuples are stored; specifically, if The corresponding qualification value is less than the predefined threshold. ),but set of traces Remove from the middle; to achieve the above function, after each round of learning, The corresponding qualification value decays according to the following formula: in, Tuples are defined Recentness As the attenuation factor, The trace decay parameter determines the decay rate of the trace; through the above mechanism, the space complexity of each iteration does not increase linearly with the appearance of a new tuple. In practice, the preparatory maneuver action set should be updated. Parameters of the rules; set of motion actions After the rules in the rule base [P] are updated, the rules that can match the historical trace set are updated to different degrees according to their accuracy and corresponding qualification value, which is the contribution of the historical trace to the agent reaching the current state. First of all, for Each record in Searching in [P] can be related to the situation. and the first fighter maneuver The matching rules are used to assemble the matching results into a set of preliminary maneuvers. Then, according to The qualification value pair The rules in the Chinese system have been updated. When allocating TD errors, the accuracy of the rules is considered; the higher the accuracy, the greater the degree of update. This is due to the fitness parameter of the rules. This represents its accuracy; normalized fitness is used as the relative accuracy of the rule; then, the profit prediction... and prediction error The update method is as follows: in, express The rules in and The learning rate; First, the set of maneuvers from the previous moment is updated based on environmental feedback, information on the behavior of the second fighter jet, and predictions of expected maneuvers. The parameters of the rules, and in GA is used to generate and evolve new rules.
3. The method according to claim 1, characterized in that, The method further includes: A sample set is constructed, comprising several samples; wherein the samples include: sample data and label data; the sample data includes training potential information; and the label data includes the training prediction behavior of the second fighter jet corresponding to the training potential information. Based on the sample set, a second fighter jet behavior prediction model is constructed and trained using a predetermined machine learning algorithm. The predetermined machine learning algorithm is selected from one or more of the following: Naive Bayes algorithm, decision tree algorithm, support vector machine algorithm, KNN algorithm, neural network algorithm, deep learning algorithm, and logistic regression algorithm; The update methods for the second fighter jet behavior prediction model include: Acquire the data set corresponding to the second fighter jet, including situational information and behavioral information; The second fighter jet behavior prediction model is updated based on the data components; The data set, which includes situational information and behavioral information, characterizes the behavioral information of the second fighter jet under different situational information. The behavior of the second fighter jet is represented as follows: ; Among these, situational-behavioral information of the second fighter jet is collected on a round-by-round basis. And at the end of each round, the second fighter behavior prediction model is updated using the following formula: in, For the predicted second fighter jet in the situation Execution Action The probability, Let be the information entropy constant. .
4. An air combat decision-making device, characterized in that, include: The situation information acquisition module is configured to acquire current situation information from an air combat environment; wherein, the air combat environment contains a first fighter jet and a second fighter jet hostile to the first fighter jet; The matching set construction module is configured to determine the rules corresponding to the current situation information from the rule base corresponding to the first fighter jet based on the current situation information, and construct the matching set corresponding to the first fighter jet based on the rules; The rule includes an antecedent. and perform maneuvers ; The step of determining the rule corresponding to the current situation information from the rule base corresponding to the first fighter jet based on the current situation information, and constructing a matching set corresponding to the first fighter jet based on the rule, includes: Compare and match the non-wildcard portion of the antecedent of all the rules in the rule base with the current situation information, and add the filtered rules to the matching set. The preceding element and the aforementioned execution of maneuvering actions Mapping to the revenue prediction vector Prediction error Adaptability , scale of mobile action Empirical parameter values and quantity parameter values Profit prediction vector Each element in This indicates that the first fighter jet is performing a maneuver. The second fighter jet performed a maneuver. Subsequently, the predicted gain of the first fighter jet will be determined by selecting the executed maneuvers from the rules in the matching set as candidate maneuvers. The step of determining the rule corresponding to the current situation information from the rule base corresponding to the first fighter jet based on the current situation information, and constructing the matching set corresponding to the first fighter jet based on the rule, includes: Based on the profit prediction vector of the rule and the fitness, determine the predicted profit for each candidate maneuver in the matching set. ; ; in, and Representing rules respectively The profit prediction vector and fitness, elements in Representing the first fighter jet in performing maneuvers The second fighter jet followed the estimated strategy. When making decisions, The predicted gains that can be obtained from the corresponding set of maneuvers; Based on the second fighter jet behavior prediction model, and according to the predicted gains, the expected maneuver prediction for each candidate maneuver is obtained; expressed as: ; To consider the second fighter jet behavior prediction model Afterwards, The function of the maneuvering action value; when At that time, a rule overriding mechanism is used to generate new rules, in which... For a predefined threshold; In response to determining that the maximum value in the expected maneuver prediction is less than or equal to a preset expected maneuver prediction threshold, supplementary rules are generated; when the matching set [M] does not cover the predefined minimum maneuver set or the maximum value of the expected maneuver prediction is less than... The rule overriding mechanism is triggered in a timely manner; the rule overriding mechanism is invoked as needed to supplement the rules corresponding to the missing maneuvers in the matching set [M]. The maneuver set construction module is configured to predict the predicted behavior of the second fighter jet using a second fighter jet behavior prediction model, determine candidate maneuvers for the first fighter jet based on the predicted behavior and the matching set, and construct a maneuver set corresponding to the first fighter jet based on the candidate maneuvers and the matching set. This includes: obtaining a maneuver value function and fusing the second fighter jet behavior prediction model into the maneuver value function to obtain a fused maneuver value function; based on... The strategy involves determining the candidate maneuvers using the fused maneuver value function, and adding the rules containing the candidate maneuvers from the matching set to the maneuver set. The maneuver action set execution module is configured to control the first fighter jet to perform the maneuver actions corresponding to the maneuver action set according to the maneuver action set.
Citation Information
Patent Citations
Unmanned plane short-range dogfight decision-making method based on rehearsal maneuvering rule system
CN107390706A
Unmanned aerial vehicle cooperative air combat decision-making method based on genetic fuzzy tree
CN111240353A