Air game confrontation decision inversion method based on integrated fuzzy decision tree
Through the integrated fuzzy decision tree method, the problem of category imbalance in air games against data sets is solved, the accuracy of behavior inversion and the fuzzification efficiency of fuzzy features is improved, and the rationality and compliance of decision actions are ensured.
Patent Information
- Application Number
- CN202411893176.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-13
AI Technical Summary
The existing technology does not consider the category imbalance characteristics of the aerial game adversarial data set when establishing the behavior inversion model of the decision tree/fuzzy decision tree, which leads to the model tending to identify a large number of behavior categories and ignores a small number of behavior categories, which makes it difficult to ensure the overall inversion accuracy of the behavior inversion model.
A method of air game confrontation decision inversion based on integrated fuzzy decision trees is proposed. By training multiple fuzzy decision trees, the relative features are fuzzyed by the number of fuzzy features, the training subset pool is constructed, and the branch frequency parameters and leaf node confidence parameters of the decision tree are optimized. Finally, the decision action is determined through the confidence accumulation value and normalization processing.
The accuracy of behavioral inversion of air game adversarial data with unbalanced characteristics is improved, the efficiency of fuzzy characteristics is enhanced, and the rationality and compliance of decision-making actions are ensured.
Smart Images

Figure CN119990345A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent decision-making, and in particular relates to an air game confrontation decision inversion method based on an integrated fuzzy decision tree. Background Art
[0002] In recent years, deep reinforcement learning technology has received high attention and in-depth research in the field of enabling intelligent aerial game confrontation. However, the intelligent decision-making algorithm based on deep reinforcement learning has a deep neural network with "black box" properties as its decision kernel. The opacity of its decision-making process makes it difficult to understand the decision-making mechanism and judge the rationality and compliance of the decision results, which greatly limits the application of intelligent decision-making algorithms in the field of aerial game confrontation.
[0003] In order to clarify the internal decision logic of intelligent decision-making algorithms and invert the decision-making behavior of intelligent decision-making algorithms, the methods that are currently being studied more are decision trees and their derived tree models. Among the derived tree models, the fuzzy decision tree based on fuzzy theory, with its fuzzy decision logic, is highly consistent with human thinking and decision-making habits, that is, it can receive fuzzy language and fuzzy information and make correct identification and judgment. However, when the above tree models perform behavioral inversion on intelligent decision-making algorithms, they do not consider the unbalanced characteristics of game confrontation data (there are significant differences in the number of data in different behavioral categories). Therefore, the trained models tend to identify a larger number of behavioral categories and ignore a smaller number of behavioral categories, which makes it difficult to ensure the overall inversion accuracy of the behavioral inversion model.
[0004] The existing patent CN111353606A discloses a deep reinforcement learning air combat game interpretation method and system based on a fuzzy decision tree. For example, the existing patent CN116384480A discloses a deep reinforcement learning decision interpretation system. In the prior art, when establishing a behavior inversion model of a decision tree / fuzzy decision tree, the category imbalance characteristics of the air game confrontation data set have not been considered, resulting in the trained model tending to identify a larger number of behavior categories and ignore a smaller number of behavior categories, which in turn makes it difficult to ensure the overall inversion accuracy of the behavior inversion model. Summary of the invention
[0005] The purpose of the present invention is to solve the problems raised in the background technology and to propose an air game confrontation decision inversion method based on an integrated fuzzy decision tree.
[0006] To achieve the above object, the technical solution adopted by the present invention is:
[0007] The present invention proposes an air game confrontation decision inversion method based on an integrated fuzzy decision tree, comprising:
[0008] Step 1: During the training process, a set of standardized first relative state data of the red and blue agents and a set of action data categories of the red agent are obtained, and the action data categories of the same category and the corresponding first relative state data are stored as a whole in the corresponding data pool;
[0009] Step 2: construct a training subset pool based on each data pool;
[0010] Step 3: for each subset in the training subset pool, randomly set the number of fuzzy features for the relative features corresponding to each first eigenvalue of the first relative state data, and perform fuzzification processing on each first eigenvalue according to the number of fuzzy features to obtain a fuzzy data subset;
[0011] Step 4: Take each fuzzy data subset as the root node of the fuzzy decision tree, determine the intermediate nodes and leaf nodes of the fuzzy decision tree, and obtain the fuzzy decision tree corresponding to each fuzzy data subset;
[0012] Step 5: During the decision-making process, the standardized second relative state data of the red and blue agents at the current moment is obtained, and after fuzzification, it is input into each fuzzy decision tree, and each fuzzy decision tree outputs the confidence of all action data categories;
[0013] Step 6: Calculate the cumulative confidence value of each action data category, and normalize each confidence cumulative value. Finally, the action data category corresponding to the largest confidence cumulative value after normalization is used as the final decision action of the red agent at the current moment.
[0014] Preferably, each of the relative state data is a vector value formed by relative characteristics of the red and blue agents, and the relative characteristics include at least relative distance, relative height, relative angle and relative speed.
[0015] Preferably, in the training process, a set of standardized first relative state data of the red and blue agents and a set of action data categories of the red agent are obtained during the game, and the action data categories of the same category and the corresponding first relative state data are stored as a whole in a corresponding data pool, including:
[0016] The set of the first relative state data after standardization of the red and blue agents is used Indicates that represents the i-th first relative state data, and in represents the jth first eigenvalue in the i-th first relative state data, df represents the number of first eigenvalues, and the set of action data categories of the red agent is denoted by Indicates that, y i represents the i-th action data category, N represents the number of first relative state data or action data categories, and the first relative state data in the first relative state data set correspond one-to-one to the action data categories in the action data category set in sequence;
[0017] All data pools are represented as Among them, D k represents the data pool of the kth action data category, K represents the number of action data categories, and N k represents the number of action data categories belonging to the kth category, represents the i-th first relative state data belonging to the k-th action data category, Indicates that it belongs to the i-th action data category in the k-th action data category.
[0018] Preferably, when constructing the training subset pool based on each data pool, the construction of each subset in the training subset pool includes:
[0019] Step 2.1: Set the size of each subset in the training subset pool to N t ;
[0020] Step 2.2, generate labels for corresponding action data categories based on uniform distribution sampling;
[0021] Step 2.3: Based on the generated labels, select the corresponding data pool, randomly sample a data from it, and store it in the subset;
[0022] Step 2.4: Count the number of data in the subset, if it is less than N t , then jump to step 2.2, if it is equal, then output the subset H l represents the lth subset in the training subset pool, represents the i-th first relative state data in the l-th subset, represents the i-th action data category in the l-th subset;
[0023] Step 2.5: Repeat steps 2.1 to 2.4 to form L subsets as the training subset pool.
[0024] Preferably, for each subset in the training subset pool, the number of fuzzy features is randomly set for the relative features corresponding to each first eigenvalue of the first relative state data, and each first eigenvalue is fuzzified according to the number of fuzzy features to obtain the corresponding fuzzy data subset, including:
[0025] Step 3.1: For each subset in the training subset pool, the relative feature corresponding to the jth first feature value of the first relative state data in the subset is expressed as A j , then the relative feature set corresponding to all the first eigenvalues of the first relative state data in the subset is represented by the set Indicates that the number of fuzzy features is set for each relative feature in set A, recorded as B j Relative feature A j The number of fuzzy features;
[0026] The first eigenvalue corresponding to each relative feature is fuzzified according to the number of fuzzy features to obtain the corresponding fuzzy data subset, including:
[0027] Step 3.2: For discrete features, the K-means clustering algorithm is used for fuzzy processing, and the steps are as follows:
[0028] Step 3.2.1: Based on the number of fuzzy features B j , from the subset H l Randomly select B from the first relative state data j The first eigenvalue of a discrete relative feature As the initial category center, and recorded as Represents the subset H l The e-th category center when the j-th first eigenvalue is fuzzy processed;
[0029] Step 3.2.2: Calculate subset H l The Euclidean distance from the first eigenvalue of each discrete relative feature to the center of each category, and the formula is as follows:
[0030]
[0031] in, Represents the subset H l The first eigenvalue of the j'th discrete relative feature of the i-th first relative state data Euclidean distance to the center of the e-th cluster;
[0032] Subset H l The first eigenvalue of each discrete relative feature in is assigned to the nearest category center, and the data assigned to each category center is represented by the set express, Represents the subset H l Relative feature A j The data set assigned to the center of the e-th category when the corresponding first eigenvalue is fuzzified;
[0033] Step 3.2.3, calculate the data mean in the data set of each category center as the new category center of the corresponding category center, and the formula is as follows:
[0034]
[0035] in, Represents the subset H l Relative feature A j The corresponding first eigenvalue is fuzzified and the new category center of the e-th category center is obtained. Representing a collection of data The number of data in Indicates assigning the new category center to the corresponding category center;
[0036] Step 3.2.4: Repeat steps 3.2.2-3.2.3 until the center of each category no longer changes, and the optimal center of each category is obtained. Whether it belongs to the data set in each optimal category, to output the first eigenvalue of the discrete relative feature The fuzzy value of And satisfy:
[0037]
[0038] in, Represents the subset H l The fuzzy value of the first eigenvalue of the j'th discrete relative feature of the ith first relative state data relative to the eth optimal category center;
[0039] Step 3.3: For continuous features, mixed Gaussian distribution is used for fuzzy processing, and the steps are as follows:
[0040] Step 3.3.1: Based on the number of fuzzy features B j , let subset H l The first eigenvalue of each continuous relative characteristic in Obey Include B j Gaussian mixture model with Gaussian distributions, Represents the subset H l The jth first relative state data in the ith * The first eigenvalue of the continuous relative feature, and the probability density function of the Gaussian mixture model is as follows:
[0041]
[0042] in, The first eigenvalue representing the relative feature of the continuous type The probability density function of the Gaussian mixture model is π w , μ w and σ w Respectively represent the mixing coefficient, mean and standard deviation of the w-th Gaussian distribution in the Gaussian mixture model, and π w >0, represents the probability density function of the Gaussian distribution;
[0043] Among them, the parameter set of the Gaussian mixture model is defined as Indicates B j The mixing coefficient of the Gaussian distribution, Indicates B j The mean of a Gaussian distribution, Indicates B j The standard deviation of each Gaussian distribution, each Gaussian distribution corresponds to the fuzzy feature one by one;
[0044] Step 3.3.2, solve the parameters of the Gaussian mixture model through the log-likelihood function, and the formula is as follows:
[0045]
[0046] Where L() represents the log-likelihood function;
[0047] Step 3.3.3: Iterate and solve the parameters of the Gaussian mixture model through the expectation maximization algorithm: First, calculate the first eigenvalue of each continuous relative feature The posterior probability of belonging to each Gaussian distribution is as follows:
[0048]
[0049] in, Represents the first eigenvalue of the tth iteration of the Gaussian mixture model The posterior probability of belonging to the w-th Gaussian distribution, and They represent the mixing coefficient, mean and standard deviation of the w-th Gaussian distribution of the t-th iteration of the Gaussian mixture model respectively;
[0050] Update the parameters of the Gaussian mixture model at the t+1th iteration according to the posterior probability, and the formula is as follows:
[0051]
[0052]
[0053]
[0054] in, and They represent the mixing coefficient, mean and standard deviation of the w-th Gaussian distribution of the t+1-th iteration of the Gaussian mixture model, respectively, and T represents transposition;
[0055] Step 3.3.4, continue to iterate until the log-likelihood function L(θ) of the Gaussian mixture model converges. When convergence occurs, the parameters of the corresponding Gaussian mixture model are the optimal parameters, denoted as
[0056] Calculate the fuzzy value of the first eigenvalue of each continuous relative feature and record it as in:
[0057]
[0058] in, Represents the subset H l The jth first relative state data in the ith * The fuzzy value of the first eigenvalue of the continuous relative feature relative to the wth Gaussian distribution;
[0059] Step 3.4: For subset H l The fuzzy value of each first eigenvalue in is calculated using the fuzzy data subset Indicates that, F l,i Represents the subset H l The set of fuzzy values of all first eigenvalues of the i-th first relative state data in , and Represents the subset H l The fuzzy value of the jth first eigenvalue of the i-th first relative state data in
[0060] Then for the L subsets in the training subset pool, L fuzzy data subsets are obtained.
[0061] Preferably, the method of using each fuzzy data subset as the root node of a fuzzy decision tree, determining the intermediate nodes and leaf nodes of the fuzzy decision tree, and obtaining the fuzzy decision tree corresponding to each fuzzy data subset comprises:
[0062] Step 4.1, for each fuzzy data subset, take the fuzzy data subset as the root node of the fuzzy decision tree, set the branch frequency parameter threshold α and the leaf node confidence parameter threshold β, where the branch frequency parameter represents the percentage of data on the branch of the fuzzy decision tree to the total data, and the leaf node confidence parameter represents the probability of classifying the leaf node as a certain action data category;
[0063] Step 4.2, traverse each node;
[0064] Step 4.3: For the currently traversed node, record it as node q, and calculate the branch frequency parameter α of node q q and leaf node confidence parameter β q , and β q =[β q,1 ,β q,2 ,…,β q,k ,…,β q,K ] T , where β q,k It represents the confidence that the leaf node of the currently traversed node q belongs to the action data category k, and the set of fuzzy features of the relative features of the previously split nodes on the branch where the currently traversed node q is located is recorded as in, Represents the relative characteristics of the p-1th split node on the branch where the currently traversed node q is located No. fuzzy features, and the fuzzy features of the relative features of the currently traversed node q on the branch after splitting are expressed as The branch frequency parameter α of the currently traversed node q q and leaf node confidence parameter β q The calculation formula is as follows:
[0065]
[0066]
[0067] Where V( ) represents potential, Depend on The fuzzy value products of each fuzzy feature in are summed up to get: Depend on The product of the fuzzy values of the fuzzy features belonging to action data category k is obtained, and for the root node, the value of its branch frequency parameter is 1, and the value of the leaf node confidence parameter is
[0068] Step 4.4: Determine whether the currently traversed node q is a leaf node, and the judgment condition is: the branch frequency parameter α of the currently traversed node q q Is it less than or equal to the preset branch frequency parameter threshold α, or is the leaf node of the currently traversed node q affiliated to the category confidence parameter β q Is it greater than or equal to the preset leaf node confidence parameter threshold β? If the condition is met, the currently traversed node q belongs to a leaf node, then return to step 4.2, that is, continue to traverse the next node, otherwise go to step 4.5;
[0069] Step 4.5: Calculate the cumulative mutual information value CMI and cumulative conditional entropy value CCE of all candidate relative features in the currently traversed node q, and the formula is as follows:
[0070] The formula for the cumulative mutual information value CMI is as follows:
[0071]
[0072] in,
[0073]
[0074]
[0075]
[0076]
[0077]
[0078] in, Indicates the number of fuzzy features in the currently traversed node q, represents the sequence number of the fuzzy feature in the currently traversed node q, CMI1, CMI2, CMI3 and CMI4 all represent the cumulative joint information, V k Indicates intermediate parameters;
[0079] The formula for the cumulative conditional entropy value CCE is as follows:
[0080]
[0081] Step 4.6: Calculate the ratio of the cumulative mutual information value to the cumulative conditional entropy value of all candidate relative features in the currently traversed node q, and use the candidate relative feature corresponding to the largest ratio as the split feature of the currently traversed node q Then according to the split feature split, and The corresponding calculation formula is as follows:
[0082]
[0083] Among them, arg max represents the maximum value function, and C represents the set of all candidate relative features in the currently traversed node q;
[0084] Then return to step 4.2 until all nodes are traversed and there are no nodes to be split, then output the fuzzy decision tree, then output the fuzzy decision trees corresponding to all fuzzy data subsets, all fuzzy decision trees are expressed as Among them, FDT l Represents the lth fuzzy decision tree.
[0085] Preferably, in the decision-making process, the standardized second relative state data of the red and blue agents at the current moment is obtained, fuzzified, and input into each fuzzy decision tree, and each fuzzy decision tree outputs the confidence of all action data categories, including:
[0086] Step 5.1, fuzzifying each second eigenvalue in the second relative state data;
[0087] Step 5.1.1, for the relative feature type corresponding to the second eigenvalue being discrete, calculate the Euclidean distance between the second eigenvalue of each discrete relative feature and the center of each optimal category, and assign the second eigenvalue of each discrete relative feature to the center of the optimal category closest to it, and output the fuzzy value of the second eigenvalue of each discrete relative feature based on whether the second eigenvalue of each discrete relative feature belongs to the data set in each optimal category;
[0088] Step 5.1.2: When the relative feature type corresponding to the second eigenvalue is continuous, the fuzzy value of the second eigenvalue of each continuous relative feature is calculated according to the optimal parameters of the Gaussian mixture model;
[0089] Step 5.2, input the fuzzy values of all second eigenvalues as a whole into each fuzzy decision tree;
[0090] Step 5.3: Each fuzzy decision tree outputs the confidence of all action data categories, and the confidence of all fuzzy decision trees outputting all action data categories is expressed as where β l,k It represents the confidence that the lth fuzzy decision tree belongs to action data category k.
[0091] Preferably, the confidence cumulative values of each action data category are calculated, and each confidence cumulative value is normalized, and finally the action data category corresponding to the maximum confidence cumulative value after normalization is used as the final decision action of the red agent at the current moment, including:
[0092] Step 6.1: The formula for calculating the cumulative confidence value of each action data category is as follows:
[0093]
[0094] Among them, cβ k Represents the cumulative confidence value for action data category k;
[0095] Step 6.2: Normalize each confidence cumulative value, and the corresponding calculation formula is as follows:
[0096]
[0097] in, represents the normalized cumulative confidence value of action data category k, and m represents the sequence number of the action data category;
[0098] Step 6.3: Select the maximum confidence cumulative value after normalization, and the corresponding calculation formula is as follows:
[0099]
[0100] Among them, k * represents the final decision action, and arg max represents the maximum value function.
[0101] Compared with the prior art, the present invention has the following beneficial effects:
[0102] This aerial game confrontation decision inversion method based on integrated fuzzy decision trees first trains multiple fuzzy decision trees, then obtains the standardized second relative state data of the red and blue agents at the current moment, and inputs it into each fuzzy decision tree. Each fuzzy decision tree outputs the confidence of all action data categories, calculates the cumulative confidence value of each action data category, and normalizes each confidence cumulative value. Finally, the action data category corresponding to the largest confidence cumulative value after normalization is used as the final decision action of the red agent at the current moment, and then the decision in the confrontation process is obtained, thereby improving the accuracy of behavioral inversion of aerial game confrontation data with unbalanced characteristics; during fuzzy processing, different fuzzy processing methods are designed for discrete relative features and continuous relative features, thereby improving the fuzzification efficiency of relative features. BRIEF DESCRIPTION OF THE DRAWINGS
[0103] Figure 1 It is a flow chart of the air game confrontation decision inversion method based on integrated fuzzy decision tree of the present invention;
[0104] Figure 2 Flowchart of the decision-making process of the present invention. DETAILED DESCRIPTION
[0105] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0106] like Figure 1-Figure 2 As shown, an air game confrontation decision inversion method based on an integrated fuzzy decision tree is provided, including:
[0107] Step 1: During the training process, obtain the set of standardized first relative state data of the red and blue agents and the set of action data categories of the red agent during the game, and store the action data categories of the same category and the corresponding first relative state data as a whole into the corresponding data pool, specifically:
[0108] The set of the first relative state data of the red and blue agents after standardization is used Indicates that Represents the i-th first relative state data (where The i in can be understood as different moments), and in represents the jth first eigenvalue in the i-th first relative state data (e.g., the first eigenvalue is that the relative distance between the red and blue sides is 100 km), df represents the number of first eigenvalues, and the set of action data categories of the red agent is denoted by Indicates that, y i represents the i-th action data category, N represents the number of first relative state data or action data categories, and the first relative state data in the first relative state data set correspond one-to-one to the action data categories in the action data category set in sequence;
[0109] All data pools are represented as Among them, D k represents the data pool of the kth action data category, K represents the number of action data categories, and N k represents the number of action data categories belonging to the kth category, represents the i-th first relative state data belonging to the k-th action data category, Indicates that it belongs to the i-th action data category in the k-th action data category.
[0110] It should be noted that during the training process, a set consisting of the standardized first relative state data of the red and blue agents and a set consisting of the action data categories of the red agent can be taken within a preset period of time, wherein the standardized first relative state data is used to standardize the relative state data of the red and blue agents, wherein the relative state data of the red and blue agents are obtained by comparing the state data of the red agent with the state data of the blue agent (or the state data of the blue agent with the state data of the red agent), and each relative state data is a vector value composed of the relative characteristics of the red and blue agents, and the relative characteristics at least include relative distance, relative height, relative angle and relative speed, and the action data categories at least include left-cut flight, right-cut flight, pursuit of enemy aircraft, guided flight, tail-end avoidance and other actions.
[0111] Step 2: construct a training subset pool based on each data pool. The construction of each subset in the training subset pool includes:
[0112] Step 2.1: Set the size of each subset in the training subset pool to N t ;
[0113] Step 2.2, generate labels for corresponding action data categories based on uniform distribution sampling (set a label for each data pool (such as 1 represents the data pool corresponding to the first action data category, 2 represents the data pool corresponding to the second action data category, etc.), generate labels through uniform distribution sampling, such as the label generated by uniform distribution sampling is 1, which represents the data pool corresponding to the first action data category label);
[0114] Step 2.3: Based on the generated labels, select the corresponding data pool, randomly sample a data from it, and store it in the subset;
[0115] Step 2.4: Count the number of data in the subset, if it is less than N t , then jump to step 2.2, if it is equal, then output the subset H l represents the lth subset in the training subset pool, represents the i-th first relative state data in the l-th subset, represents the i-th action data category in the l-th subset;
[0116] Step 2.5: Repeat steps 2.1 to 2.4 to form L subsets as the training subset pool.
[0117] Step 3: For each subset in the training subset pool, randomly set the number of fuzzy features for the relative features corresponding to each first eigenvalue of the first relative state data, and perform fuzzification processing on each first eigenvalue according to the number of fuzzy features to obtain a fuzzy data subset, specifically:
[0118] Step 3.1: For each subset in the training subset pool, the relative feature corresponding to the jth first feature value of the first relative state data in the subset is expressed as A j (such as the relative distance between red and blue), then the relative feature set corresponding to all the first feature values of the first relative state data in the subset is represented by the set Indicates that the number of fuzzy features is set for each relative feature in set A, recorded as B j Relative feature A jThe number of fuzzy features (for example, the number of fuzzy features of the relative distance between red and blue is 3, and the three fuzzy features are respectively close distance, medium distance, and far distance);
[0119] The first eigenvalue corresponding to each relative feature is fuzzified according to the number of fuzzy features to obtain the corresponding fuzzy data subset, including:
[0120] Step 3.2: For discrete features, the K-means clustering algorithm is used for fuzzy processing, and the steps are as follows:
[0121] Step 3.2.1: Based on the number of fuzzy features B j , from the subset H l Randomly select B from the first relative state data j The first eigenvalue of a discrete relative feature As the initial category center, and recorded as Represents the subset H l The e-th category center when the j-th first eigenvalue is fuzzy processed;
[0122] Step 3.2.2: Calculate subset H l The Euclidean distance from the first eigenvalue of each discrete relative feature to the center of each category, and the formula is as follows:
[0123]
[0124] in, Represents the subset H l The first eigenvalue of the j'th discrete relative feature of the i-th first relative state data Euclidean distance to the center of the e-th cluster;
[0125] Subset H l The first eigenvalue of each discrete relative feature in is assigned to the nearest category center, and the data assigned to each category center is represented by the set express, Represents the subset H l Relative feature A j The data set assigned to the center of the e-th category when the corresponding first eigenvalue is fuzzified;
[0126] Step 3.2.3, calculate the data mean in the data set of each category center as the new category center of the corresponding category center, and the formula is as follows:
[0127]
[0128] in, Represents the subset H lRelative feature A j The corresponding first eigenvalue is fuzzified and the new category center of the e-th category center is obtained. Representing a collection of data The number of data in Indicates assigning the new category center to the corresponding category center;
[0129] Step 3.2.4: Repeat steps 3.2.2-3.2.3 until the center of each category no longer changes, and the optimal center of each category is obtained. Whether it belongs to the data set in each optimal category, to output the first eigenvalue of the discrete relative feature The fuzzy value of And satisfy:
[0130]
[0131] in, Represents the subset H l The fuzzy value of the first eigenvalue of the j'th discrete relative feature of the ith first relative state data relative to the eth optimal category center, such as B j =5 o'clock,
[0132] Step 3.3: For continuous features, mixed Gaussian distribution is used for fuzzy processing, and the steps are as follows:
[0133] Step 3.3.1: Based on the number of fuzzy features B j , let subset H l The first eigenvalue of each continuous relative characteristic in Obey Include B j Gaussian mixture model with Gaussian distributions, Represents the subset H l The jth first relative state data in the ith * The first eigenvalue of the continuous relative feature, and the probability density function of the Gaussian mixture model is as follows:
[0134]
[0135] in, The first eigenvalue representing the relative feature of the continuous type The probability density function of the Gaussian mixture model is π w , μ w and σ w Respectively represent the mixing coefficient, mean and standard deviation of the w-th Gaussian distribution in the Gaussian mixture model, and πw >0, represents the probability density function of the Gaussian distribution;
[0136] Among them, the parameter set of the Gaussian mixture model is defined as Indicates B j The mixing coefficient of the Gaussian distribution, Indicates B j The mean of a Gaussian distribution, Indicates B j The standard deviation of each Gaussian distribution, each Gaussian distribution corresponds to the fuzzy feature one by one;
[0137] Step 3.3.2, solve the parameters of the Gaussian mixture model through the log-likelihood function, and the formula is as follows:
[0138]
[0139] Where L() represents the log-likelihood function;
[0140] Step 3.3.3, iteratively solve the parameters of the Gaussian mixture model through the expectation-maximization algorithm (EM algorithm): first calculate the first eigenvalue of each continuous relative feature The posterior probability of belonging to each Gaussian distribution is as follows:
[0141]
[0142] in, Represents the first eigenvalue of the tth iteration of the Gaussian mixture model The posterior probability of belonging to the w-th Gaussian distribution, and They represent the mixing coefficient, mean and standard deviation of the w-th Gaussian distribution of the t-th iteration of the Gaussian mixture model respectively;
[0143] Update the parameters of the Gaussian mixture model at the t+1th iteration according to the posterior probability, and the formula is as follows:
[0144]
[0145]
[0146]
[0147] in, and They represent the mixing coefficient, mean and standard deviation of the w-th Gaussian distribution of the t+1-th iteration of the Gaussian mixture model, respectively, and T represents transposition;
[0148] Step 3.3.4, continue to iterate until the log-likelihood function L(θ) of the Gaussian mixture model converges. When convergence occurs, the parameters of the corresponding Gaussian mixture model are the optimal parameters, denoted as
[0149] Calculate the fuzzy value of the first eigenvalue of each continuous relative feature and record it as in:
[0150]
[0151] in, Represents the subset H l The jth first relative state data in the ith * The first eigenvalue of the continuous relative feature is relative to the fuzzy value of the wth Gaussian distribution, such as B j =3,
[0152] Step 3.4: For subset H l The fuzzy value of each first eigenvalue in is calculated using the fuzzy data subset Indicates that, F l,i Represents the subset H l The set of fuzzy values of all first eigenvalues of the i-th first relative state data in , and Represents the subset H l The fuzzy value of the jth first eigenvalue of the i-th first relative state data in
[0153] Then for the L subsets in the training subset pool, L fuzzy data subsets are obtained.
[0154] Step 4: Take each fuzzy data subset as the root node of the fuzzy decision tree, determine the intermediate nodes and leaf nodes of the fuzzy decision tree, and obtain the fuzzy decision tree corresponding to each fuzzy data subset, specifically:
[0155] Step 4.1: For each fuzzy data subset, take the fuzzy data subset as the root node of the fuzzy decision tree, set the branch frequency parameter threshold α and the leaf node confidence parameter threshold β (according to experience, the branch frequency parameter threshold α is 0.05, and the leaf node confidence parameter threshold β is 0.9), where the branch frequency parameter represents the percentage of data on the branch of the fuzzy decision tree to the total data, and the leaf node confidence parameter represents the probability of classifying the leaf node as a certain action data category;
[0156] Step 4.2, traverse each node;
[0157] Step 4.3: For the currently traversed node, record it as node q, and calculate the branch frequency parameter α of node q q and leaf node confidence parameter β q , and β q =[β q,1 ,β q,2 ,…,β q,k ,…,β q,K ] T , where β q,k It represents the confidence that the leaf node of the currently traversed node q belongs to the action data category k, and the set of fuzzy features of the relative features of the previously split nodes on the branch where the currently traversed node q is located is recorded as in, Represents the relative characteristics of the p-1th split node on the branch where the currently traversed node q is located No. fuzzy features, and the fuzzy features of the relative features of the currently traversed node q on the branch after splitting are expressed as The branch frequency parameter α of the currently traversed node q q and leaf node confidence parameter β q The calculation formula is as follows:
[0158]
[0159]
[0160] Among them, V() represents potential, Depend on The fuzzy value products of each fuzzy feature in are summed up to get: Depend on The product of the fuzzy values of the fuzzy features belonging to action data category k is obtained, and for the root node, the value of its branch frequency parameter is 1, and the value of the leaf node confidence parameter is
[0161] Step 4.4: Determine whether the currently traversed node q is a leaf node, and the judgment condition is: the branch frequency parameter α of the currently traversed node q q Is it less than or equal to the preset branch frequency parameter threshold α, or is the leaf node of the currently traversed node q affiliated to the category confidence parameter β q Is it greater than or equal to the preset leaf node confidence parameter threshold β? If the condition is met, the currently traversed node q belongs to a leaf node, then return to step 4.2, that is, continue to traverse the next node, otherwise go to step 4.5;
[0162] Step 4.5: Calculate the cumulative mutual information value CMI and cumulative conditional entropy value CCE of all candidate relative features in the currently traversed node q, and the formula is as follows:
[0163] The formula for the cumulative mutual information value CMI is as follows:
[0164]
[0165] in,
[0166]
[0167]
[0168]
[0169]
[0170] in, Indicates the number of fuzzy features in the currently traversed node q, represents the sequence number of the fuzzy feature in the currently traversed node q, CMI1, CMI2, CMI3 and CMI4 all represent the cumulative joint information, V k Indicates intermediate parameters;
[0171] The formula for the cumulative conditional entropy value CCE is as follows:
[0172]
[0173] Step 4.6: Calculate the ratio of the cumulative mutual information value to the cumulative conditional entropy value of all candidate relative features in the currently traversed node q, and use the candidate relative feature corresponding to the largest ratio as the split feature of the currently traversed node q Then according to the split feature split, and The corresponding calculation formula is as follows:
[0174]
[0175] Among them, argmax represents the maximum value function, and C represents the set of all candidate relative features in the currently traversed node q;
[0176] Then return to step 4.2 until all nodes are traversed and there are no nodes to be split, then output the fuzzy decision tree, then output the fuzzy decision trees corresponding to all fuzzy data subsets, all fuzzy decision trees are expressed as Among them, FDT l Represents the lth fuzzy decision tree.
[0177] like Figure 2As shown, in step 5, during the decision-making process, the standardized second relative state data of the red and blue agents at the current moment is obtained, and after fuzzification, it is input into each fuzzy decision tree. Each fuzzy decision tree outputs the confidence of all action data categories, specifically:
[0178] Step 5.1, fuzzifying each second eigenvalue in the second relative state data;
[0179] Step 5.1.1, for the relative feature type corresponding to the second eigenvalue is discrete, calculate the Euclidean distance between the second eigenvalue of each discrete relative feature and the center of each optimal category, and assign the second eigenvalue of each discrete relative feature to the optimal category center closest to it, and output the fuzzy value of the second eigenvalue of each discrete relative feature according to whether the second eigenvalue of each discrete relative feature belongs to the data set in each optimal category (the specific process is the same as the corresponding step in step 3.2);
[0180] Step 5.1.2: For the relative feature type corresponding to the second eigenvalue being continuous, calculate the fuzzy value of the second eigenvalue of each continuous relative feature according to the optimal parameters of the Gaussian mixture model (the specific process is the same as the corresponding step in step 3.3);
[0181] Step 5.2, input the fuzzy values of all second eigenvalues as a whole into each fuzzy decision tree;
[0182] Step 5.3: Each fuzzy decision tree outputs the confidence of all action data categories, and the confidence of all fuzzy decision trees outputting all action data categories is expressed as where β l,k It represents the confidence that the lth fuzzy decision tree belongs to action data category k.
[0183] Step 6: Calculate the confidence cumulative value of each action data category, and normalize each confidence cumulative value. Finally, the action data category corresponding to the largest confidence cumulative value after normalization is used as the final decision action of the red agent at the current moment. Specifically:
[0184] Step 6.1: The formula for calculating the cumulative confidence value of each action data category is as follows:
[0185]
[0186] Among them, cβ k Represents the cumulative confidence value for action data category k;
[0187] Step 6.2: Normalize each confidence cumulative value, and the corresponding calculation formula is as follows:
[0188]
[0189] in, represents the normalized cumulative confidence value of action data category k, and m represents the sequence number of the action data category;
[0190] Step 6.3: Select the maximum confidence cumulative value after normalization, and the corresponding calculation formula is as follows:
[0191]
[0192] Among them, k * represents the final decision action, and arg max represents the maximum value function.
[0193] in Figure 2 Examples of the first fuzzy decision tree and the Lth fuzzy decision tree are given in, but not limited to them, and examples of the confidence of the outputs of the first fuzzy decision tree and the Lth fuzzy decision tree are given in, but not limited to them, where Figure 2 The probability in the bar chart in the lower right corner of represents the meaning of the normalized cumulative confidence value.
[0194] This aerial game confrontation decision inversion method based on integrated fuzzy decision trees first trains multiple fuzzy decision trees, then obtains the standardized second relative state data of the red and blue agents at the current moment, and inputs it into each fuzzy decision tree. Each fuzzy decision tree outputs the confidence of all action data categories, calculates the cumulative confidence value of each action data category, and normalizes each confidence cumulative value. Finally, the action data category corresponding to the largest confidence cumulative value after normalization is used as the final decision action of the red agent at the current moment, and then the decision in the confrontation process is obtained, thereby improving the accuracy of behavioral inversion of aerial game confrontation data with unbalanced characteristics; during fuzzy processing, different fuzzy processing methods are designed for discrete relative features and continuous relative features, thereby improving the fuzzification efficiency of relative features.
[0195] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. An air game confrontation decision inversion method based on integrated fuzzy decision tree, characterized by: The air game confrontation decision inversion method based on integrated fuzzy decision tree includes: Step 1: During the training process, a set of standardized first relative state data of the red and blue agents and a set of action data categories of the red agent are obtained, and the action data categories of the same category and the corresponding first relative state data are stored as a whole in the corresponding data pool; Step 2: construct a training subset pool based on each data pool; Step 3: for each subset in the training subset pool, randomly set the number of fuzzy features for the relative features corresponding to each first eigenvalue of the first relative state data, and perform fuzzification processing on each first eigenvalue according to the number of fuzzy features to obtain a fuzzy data subset; Step 4: Take each fuzzy data subset as the root node of the fuzzy decision tree, determine the intermediate nodes and leaf nodes of the fuzzy decision tree, and obtain the fuzzy decision tree corresponding to each fuzzy data subset; Step 5: During the decision-making process, the standardized second relative state data of the red and blue agents at the current moment is obtained, and after fuzzification, it is input into each fuzzy decision tree, and each fuzzy decision tree outputs the confidence of all action data categories; Step 6: Calculate the cumulative confidence value of each action data category, and normalize each confidence cumulative value. Finally, the action data category corresponding to the largest confidence cumulative value after normalization is used as the final decision action of the red agent at the current moment.
2. The air game confrontation decision inversion method based on integrated fuzzy decision tree according to claim 1 is characterized in that: Each of the relative state data is a vector value formed by the relative characteristics of the red and blue intelligent agents, and the relative characteristics at least include relative distance, relative height, relative angle and relative speed.
3. The air game confrontation decision inversion method based on integrated fuzzy decision tree according to claim 1 is characterized in that: During the training process, a set of standardized first relative state data of the red and blue agents and a set of action data categories of the red agent are obtained during the game, and the action data categories of the same category and the corresponding first relative state data are stored as a whole in a corresponding data pool, including: The set of the first relative state data after standardization of the red and blue agents is used Indicates that represents the i-th first relative state data, and in represents the jth first eigenvalue in the i-th first relative state data, df represents the number of first eigenvalues, and the set of action data categories of the red agent is denoted by Indicates that, y i represents the i-th action data category, N represents the number of first relative state data or action data categories, and the first relative state data in the first relative state data set correspond one-to-one to the action data categories in the action data category set in sequence; All data pools are represented as Among them, D k represents the data pool of the kth action data category, K represents the number of action data categories, and N k represents the number of action data categories belonging to the kth category, represents the i-th first relative state data belonging to the k-th action data category, Indicates that it belongs to the i-th action data category in the k-th action data category.
4. The air game confrontation decision inversion method based on integrated fuzzy decision tree as claimed in claim 3 is characterized by: When constructing the training subset pool based on each data pool, the construction of each subset in the training subset pool includes: Step 2.1: Set the size of each subset in the training subset pool to N t ; Step 2.2, generate labels for corresponding action data categories based on uniform distribution sampling; Step 2.3: Based on the generated labels, select the corresponding data pool, randomly sample a data from it, and store it in the subset; Step 2.4: Count the number of data in the subset, if it is less than N t , then jump to step 2.2, if it is equal, then output the subset H l represents the lth subset in the training subset pool, represents the i-th first relative state data in the l-th subset, represents the i-th action data category in the l-th subset; Step 2.5: Repeat steps 2.1 to 2.4 to form L subsets as the training subset pool.
5. The air game confrontation decision inversion method based on integrated fuzzy decision tree according to claim 4 is characterized in that: For each subset in the training subset pool, the number of fuzzy features is randomly set for the relative features corresponding to each first eigenvalue of the first relative state data, and each first eigenvalue is fuzzified according to the number of fuzzy features to obtain a corresponding fuzzy data subset, including: Step 3.1: For each subset in the training subset pool, the relative feature corresponding to the jth first feature value of the first relative state data in the subset is expressed as A j , then the relative feature set corresponding to all the first eigenvalues of the first relative state data in the subset is represented by the set Indicates that the number of fuzzy features is set for each relative feature in set A, recorded as B j Relative feature A j The number of fuzzy features; The first eigenvalue corresponding to each relative feature is fuzzified according to the number of fuzzy features to obtain the corresponding fuzzy data subset, including: Step 3.2: For discrete features, the K-means clustering algorithm is used for fuzzy processing, and the steps are as follows: Step 3.2.1: Based on the number of fuzzy features B j , from the subset H l Randomly select B from the first relative state data j The first eigenvalue of a discrete relative feature As the initial category center, and recorded as Represents the subset H l The e-th category center when the j-th first eigenvalue is fuzzy processed; Step 3.2.2: Calculate subset H l The Euclidean distance from the first eigenvalue of each discrete relative feature to the center of each category, and the formula is as follows: in, Represents the subset H l The first eigenvalue of the j'th discrete relative feature of the i-th first relative state data Euclidean distance to the center of the e-th cluster; Subset H l The first eigenvalue of each discrete relative feature in is assigned to the nearest category center, and the data assigned to each category center is represented by the set express, Represents the subset H l Relative feature A j The data set assigned to the center of the e-th category when the corresponding first eigenvalue is fuzzified; Step 3.2.3, calculate the data mean in the data set of each category center as the new category center of the corresponding category center, and the formula is as follows: in, Represents the subset H l Relative feature A j The corresponding first eigenvalue is fuzzified and the new category center of the e-th category center is obtained. Representing a collection of data The number of data in Indicates assigning the new category center to the corresponding category center; Step 3.2.4: Repeat steps 3.2.2-3.2.3 until the center of each category no longer changes, and the optimal center of each category is obtained. Whether it belongs to the data set in each optimal category, to output the first eigenvalue of the discrete relative feature The fuzzy value of And satisfy: in, Represents the subset H l The fuzzy value of the first eigenvalue of the j'th discrete relative feature of the ith first relative state data relative to the eth optimal category center; Step 3.3: For continuous features, mixed Gaussian distribution is used for fuzzy processing, and the steps are as follows: Step 3.3.1: Based on the number of fuzzy features B j , let subset H l The first eigenvalue of each continuous relative characteristic in Obey Include B j Gaussian mixture model with Gaussian distributions, Represents the subset H l The jth first relative state data in the ith * The first eigenvalue of the continuous relative feature, and the probability density function of the Gaussian mixture model is as follows: in, The first eigenvalue representing the relative feature of the continuous type The probability density function of the Gaussian mixture model is π w , μ w and σ w Respectively represent the mixing coefficient, mean and standard deviation of the w-th Gaussian distribution in the Gaussian mixture model, and π w >0, represents the probability density function of the Gaussian distribution; Among them, the parameter set of the Gaussian mixture model is defined as Indicates B j The mixing coefficient of the Gaussian distribution, Indicates B j The mean of a Gaussian distribution, Indicates B j The standard deviation of each Gaussian distribution, each Gaussian distribution corresponds to the fuzzy feature one by one; Step 3.3.2, solve the parameters of the Gaussian mixture model through the log-likelihood function, and the formula is as follows: Where L() represents the log-likelihood function; Step 3.3.3: Iterate and solve the parameters of the Gaussian mixture model through the expectation maximization algorithm: First, calculate the first eigenvalue of each continuous relative feature The posterior probability of belonging to each Gaussian distribution is as follows: in, Represents the first eigenvalue of the tth iteration of the Gaussian mixture model The posterior probability of belonging to the w-th Gaussian distribution, and They represent the mixing coefficient, mean and standard deviation of the w-th Gaussian distribution of the t-th iteration of the Gaussian mixture model respectively; Update the parameters of the Gaussian mixture model at the t+1th iteration according to the posterior probability, and the formula is as follows: in, and They represent the mixing coefficient, mean and standard deviation of the w-th Gaussian distribution of the t+1-th iteration of the Gaussian mixture model, respectively, and T represents transposition; Step 3.3.4, continue to iterate until the log-likelihood function L(θ) of the Gaussian mixture model converges. When convergence occurs, the parameters of the corresponding Gaussian mixture model are the optimal parameters, denoted as Step 3.3.5: Calculate the fuzzy value of the first eigenvalue of each continuous relative feature and record it as in: in, Represents the subset H l The jth first relative state data in the ith * The fuzzy value of the first eigenvalue of the continuous relative feature relative to the wth Gaussian distribution; Step 3.4: For subset H l The fuzzy value of each first eigenvalue in is calculated using the fuzzy data subset Indicates that, F l,i Represents the subset H l The set of fuzzy values of all first eigenvalues of the i-th first relative state data in , and Represents the subset H l The fuzzy value of the jth first eigenvalue of the i-th first relative state data in Then for the L subsets in the training subset pool, L fuzzy data subsets are obtained.
6. The air game confrontation decision inversion method based on integrated fuzzy decision tree according to claim 5, characterized in that: The method of taking each fuzzy data subset as the root node of the fuzzy decision tree, determining the intermediate nodes and leaf nodes of the fuzzy decision tree, and obtaining the fuzzy decision tree corresponding to each fuzzy data subset includes: Step 4.1, for each fuzzy data subset, take the fuzzy data subset as the root node of the fuzzy decision tree, set the branch frequency parameter threshold α and the leaf node confidence parameter threshold β, where the branch frequency parameter represents the percentage of data on the branch of the fuzzy decision tree to the total data, and the leaf node confidence parameter represents the probability of classifying the leaf node as a certain action data category; Step 4.2, traverse each node; Step 4.3: For the currently traversed node, record it as node q, and calculate the branch frequency parameter α of node q. q and leaf node confidence parameter β q , and β q =[β q,1 ,β q,2 ,…,β q,k ,…,β q,K ] T , where β q,k It represents the confidence that the leaf node of the currently traversed node q belongs to the action data category k, and the set of fuzzy features of the relative features of the previously split nodes on the branch where the currently traversed node q is located is recorded as in, Represents the relative characteristics of the p-1th split node on the branch where the currently traversed node q is located No. fuzzy features, and the fuzzy features of the relative features of the currently traversed node q on the branch after splitting are expressed as The branch frequency parameter α of the currently traversed node q q and leaf node confidence parameter β q The calculation formula is as follows: Among them, V() represents potential, Depend on The fuzzy value products of each fuzzy feature in are summed up to get: Depend on The product of the fuzzy values of the fuzzy features belonging to action data category k is obtained, and for the root node, the value of its branch frequency parameter is 1, and the value of the leaf node confidence parameter is Step 4.4: Determine whether the currently traversed node q is a leaf node, and the judgment condition is: the branch frequency parameter α of the currently traversed node q q Is it less than or equal to the preset branch frequency parameter threshold α, or is the leaf node of the currently traversed node q affiliated to the category confidence parameter β q Is it greater than or equal to the preset leaf node confidence parameter threshold β? If the condition is met, the currently traversed node q belongs to a leaf node, then return to step 4.2, that is, continue to traverse the next node, otherwise go to step 4.5; Step 4.5: Calculate the cumulative mutual information value CMI and cumulative conditional entropy value CCE of all candidate relative features in the currently traversed node q, and the formula is as follows: The formula for the cumulative mutual information value CMI is as follows: in, in, Indicates the number of fuzzy features in the currently traversed node q, represents the sequence number of the fuzzy feature in the currently traversed node q, CMI1, CMI2, CMI3 and CMI4 all represent the cumulative joint information, V k Indicates intermediate parameters; The formula for the cumulative conditional entropy value CCE is as follows: Step 4.6: Calculate the ratio of the cumulative mutual information value to the cumulative conditional entropy value of all candidate relative features in the currently traversed node q, and use the candidate relative feature corresponding to the largest ratio as the split feature of the currently traversed node q Then according to the split feature split, and The corresponding calculation formula is as follows: Among them, argmax represents the maximum value function, and C represents the set of all candidate relative features in the currently traversed node q; Then return to step 4.2 until all nodes are traversed and there are no nodes to be split, then output the fuzzy decision tree, then output the fuzzy decision trees corresponding to all fuzzy data subsets, all fuzzy decision trees are expressed as Among them, FDT l Represents the lth fuzzy decision tree.
7. The air game confrontation decision inversion method based on integrated fuzzy decision tree according to claim 1 is characterized by: In the decision-making process, the standardized second relative state data of the red and blue agents at the current moment is obtained, fuzzified, and input into each fuzzy decision tree. Each fuzzy decision tree outputs the confidence of all action data categories, including: Step 5.1, fuzzifying each second eigenvalue in the second relative state data; Step 5.1.1, for the relative feature type corresponding to the second eigenvalue being discrete, calculate the Euclidean distance between the second eigenvalue of each discrete relative feature and the center of each optimal category, and assign the second eigenvalue of each discrete relative feature to the center of the optimal category closest to it, and output the fuzzy value of the second eigenvalue of each discrete relative feature based on whether the second eigenvalue of each discrete relative feature belongs to the data set in each optimal category; Step 5.1.2: When the relative feature type corresponding to the second eigenvalue is continuous, the fuzzy value of the second eigenvalue of each continuous relative feature is calculated according to the optimal parameters of the Gaussian mixture model; Step 5.2, input the fuzzy values of all second eigenvalues as a whole into each fuzzy decision tree; Step 5.3: Each fuzzy decision tree outputs the confidence of all action data categories, and the confidence of all fuzzy decision trees outputting all action data categories is expressed as where β l,k It represents the confidence that the lth fuzzy decision tree belongs to the action data category k.
8. The air game confrontation decision inversion method based on integrated fuzzy decision tree according to claim 7 is characterized by: The confidence cumulative values of each action data category are calculated, and each confidence cumulative value is normalized, and finally the action data category corresponding to the maximum confidence cumulative value after normalization is used as the final decision action of the red agent at the current moment, including: Step 6.1: The formula for calculating the cumulative confidence value of each action data category is as follows: Among them, cβ k Represents the cumulative confidence value for action data category k; Step 6.2: Normalize each confidence cumulative value, and the corresponding calculation formula is as follows: in, represents the normalized cumulative confidence value of action data category k, and m represents the sequence number of the action data category; Step 6.3: Select the maximum confidence cumulative value after normalization, and the corresponding calculation formula is as follows: Among them, k * represents the final decision action, and arg max represents the maximum value function.