Short-distance air combat maneuver decision-making method based on Monte Carlo tree search and reinforcement learning
By building a virtual air combat environment and neural network, combined with Monte Carlo tree search and reinforcement learning, the problems of high-dimensional state space and long-term delay feedback are solved, and the rapid, global optimality and stability of the aircraft's autonomous maneuver decision-making are achieved.
Patent Information
- Application Number
- CN202311095814.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-28
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-08-28
AI Technical Summary
In the prior art, under the problems of high-dimensional state space and long-term delay feedback, traditional reinforcement learning methods cannot effectively search and learn, and it is difficult to find a balance between exploring unknown states and known states, resulting in unstable maneuvering decision making and local optimal solutions.
Using Monte Carlo tree search and reinforcement learning methods, we use air combat virtual environments, strategy deep neural networks and value neural networks to build air combat virtual environments, strategy deep neural networks and value neural networks, and use Monte Carlo tree search to optimize maneuver decision-making.
It improves the speed and global optimization of maneuvering decisions, can adapt to high-dimensional state space and environmental uncertainty, quickly find optimization strategies, and avoids the local optimal solutions of traditional methods.
Smart Images

Figure CN120277980A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of aircraft maneuver decision-making, and more specifically, to a close-range air combat maneuver decision-making method based on Monte Carlo tree search and reinforcement learning. Background Art
[0002] Autonomous air combat maneuver decision-making, as one of the most critical capabilities of modern aircraft in air combat, is playing an increasingly important role with the improvement of airborne computing power and the complexity of the air combat environment. Due to the improvement of modern air detection capabilities, the intelligence advantages such as situation extraction that can be achieved by each aircraft during air combat are reduced, and the degree of intelligence and optimality of the aircraft's autonomous maneuver decision-making ability will largely determine the outcome of air combat. The aircraft's autonomous maneuver decision-making needs to use a fast and reasonable decision-making method to always seize the advantageous position of the air combat situation in the air combat environment, and use the situation advantage to attack the enemy in advance with airborne weapons. Although pilots have good combat experience, they cannot guarantee the optimality of maneuver decisions and may have situations such as human operation errors or inaccurate operations. Therefore, researching the autonomous air combat ability of aircraft has become an inevitable trend in the development of aircraft air combat technology.
[0003] The methods for aircraft air combat autonomous maneuver decision-making mainly include methods based on game theory, methods based on optimization, and methods based on artificial intelligence. The maneuver decisions based on game theory and optimization methods strongly depend on the matching degree between the mathematical modeling problem and the actual air combat problem, and are only applicable to simple air combat scenarios. The air combat decision-making algorithms based on reinforcement learning have the following problems: (1) They will face the challenge of the curse of dimensionality when dealing with high-dimensional state space problems, that is, the dimension of the state space is very large, resulting in the inability of traditional reinforcement learning methods to effectively search and learn; (2) They will face difficulties when dealing with long-time delay feedback problems, because in this case, the agent needs to wait for feedback at a very long time interval, resulting in a slow and unstable learning process; (3) Reinforcement learning needs to find a balance between exploring unknown states and exploiting known states in order not to fall into local optimal solutions during the learning process, but setting the exploration space too large or too small will affect the algorithm convergence. Summary of the Invention
[0004] The present invention aims to provide a close-range air combat maneuver decision-making method based on Monte Carlo tree search and reinforcement learning, which can solve the above problems.
[0005] To solve the above problems, the technical solution adopted by the present invention is as follows:
[0006] A close-range air combat maneuver decision-making method based on Monte Carlo tree search and reinforcement learning, comprising
[0007] S1. Construct an air combat virtual environment, including an aircraft guidance model, a maneuver library, maneuverability limitations, and a single-step reward for air combat decisions. The maneuver library records the maneuvers of several aircraft. The single-step reward for air combat decisions is obtained based on the aircraft guidance model and the outcome of the air combat victory or defeat of the aircraft after performing the maneuver selected from the maneuver library. Maneuvers that do not meet the maneuverability limitations need to be deleted from the maneuver library. The maneuverability limitations include the speed limit, climb angle limit, and altitude position limit of the aircraft.
[0008] S2. Construct a policy deep neural network, a value neural network, and a Monte Carlo tree search.
[0009] S3. Based on the Monte Carlo tree search, the air combat virtual environment, and the historical data of the aircraft, calculate the data of the experience sample pool through Selfplay, and perform offline training on the policy deep neural network and the value neural network with the data of the experience sample pool. During the training process, the probability output by the Monte Carlo tree search needs to be corrected according to the noise.
[0010] S4. According to the real-time data of the aircraft, the air combat virtual environment, the Monte Carlo tree search, and the trained policy deep neural network and value neural network, obtain the maneuvers of the aircraft and the corresponding probabilities of the maneuvers, and select the maneuver corresponding to the maximum probability as the short-range air combat maneuver decision of the aircraft.
[0011] In a preferred embodiment of the present invention, the aircraft guidance model is:
[0012]
[0013] The maneuver decision input a(t) = [n x , n, μ] T are the longitudinal overload n x , the normal overload n, and the speed roll angle μ respectively, and the state quantity x(t) = [V, χ, γ, x, y, z] T are the flight speed, yaw angle, climb angle, and spatial position respectively.
[0014] The state x(k) at the current k moment is discretely updated by the decision quantity. When the aircraft guidance model updates the state x(k + 1) at the next moment, according to the discrete difference equation
[0015] x(k + 1) = x(k) + f(x(k), u(k))Δt
[0016] perform approximate state update, where Δt is the update step size.
[0017] In a preferred embodiment of the present invention, the maneuver actions in the maneuver action library include straight flight at a constant speed, straight flight with acceleration, straight flight with deceleration, maneuver left turn, maneuver right turn, maneuver climb, and maneuver dive, and each maneuver action has a corresponding action control amount.
[0018] In a preferred embodiment of the present invention, the aircraft maneuverability limit is:
[0019] V min ≤V≤V max ,γ min ≤γ≤γ max ,z min ≤z≤z max
[0020] Wherein, V min ,V max are the minimum and maximum flight speeds respectively, γ min ,γ max are the minimum and maximum climb angles respectively, z min ,z max are the minimum and maximum height positions respectively.
[0021] In a preferred embodiment of the present invention, the air combat decision reward function is
[0022]
[0023] Wherein, T is the number of steps for updating the air combat state, and R t is the single-step reward for the air combat decision before this reward update;
[0024] R t =a d T d +a h T h +a φq T φq
[0025]
[0026] Wherein, T d ,T h ,T φq are the relative distance, relative height, and angle advantage functions of the aircraft respectively, a d ,a h ,a φq are the corresponding weight coefficients, satisfying 1>a φq >a d >a h >0, σ d ,σ h1 ,σ h2 are the slope coefficients, and are all greater than 1, and d is the distance between the two aircraft, d* is the maximum attack distance of the airborne weapon, Δh is the relative altitude between the two aircraft, Δh min , Δh max are the lower and upper bounds of the optimal altitude difference respectively, and φ, q are the azimuth angle of our aircraft and the approach angle of the enemy aircraft respectively.
[0027] In a preferred embodiment of the present invention, the state inputs of the policy deep neural network and the value neural network are both:
[0028] s = [p r , v r , p b , v b , d, Δh, Δv, φ, q] T
[0029] Wherein, p r , v r are the position and velocity vectors of our aircraft respectively, p b , v b are the position and velocity vectors of the enemy aircraft respectively, and Δv, φ, q are the relative velocity, azimuth angle and approach angle of our aircraft respectively;
[0030] For the policy deep neural network, its output P NN is the selection probability of each action selected under the current aircraft state,
[0031] P NN = log_softmax(s p )
[0032] Wherein, s p is the output value of the last linear network of the policy neural network;
[0033] For the value neural network, its output v NN is the value of the current state,
[0034] v NN = tanh(s v )
[0035] Wherein, s v is the output value of the last linear network of the value neural network.
[0036] In a preferred embodiment of the present invention, the algorithm process of Monte Carlo tree search includes:
[0037] 1) Create a node according to the current aircraft state, and the node attributes include: (s, a, p, N, UCB), which represent the state, maneuver action, action probability, access times and UCB value of the aircraft in the air combat respectively. The calculation formula of UCB is
[0038] UCB(s,a) = Q(s,a) + U(s,a)
[0039]
[0040] Wherein, c is a weight coefficient, representing the proportion of the probability value output by the policy network, ΣN j is the sum of the visit counts of all child nodes of the current node. For all newly created nodes, Q(s,a) is initially 0. For the root node, the action a in its attributes is empty, and the probability p is 1. For non-root nodes, the action and probability value are initially defined by their parent nodes;
[0041] 2) In each playout of Monte Carlo tree search, the following operations are performed in sequence: Select the child node with the largest UCB value among all child nodes until the current node is a leaf node; Evaluate the action probability P NN and value v NN in the current state by the policy neural network and the value neural network. If there is a winner in the air combat state corresponding to the current leaf node, then according to the win-loss result, set v NN = 1 or v NN = -1, and then perform backtracking to update the attributes of each layer of tree nodes. If there is no winner, after node expansion according to P NN update the attributes of each layer of tree nodes according to the original v NN ;
[0042] After several playout operations, obtain the action a i in the attributes of each child node of the root node corresponding to the current aircraft state through the softmax function
[0043]
[0044] Wherein, N i is the visit count of each child node of the root node, ∈1 = 1.0e -10 , ∈2 ∈ (0,1] is the temperature coefficient.
[0045] In a preferred embodiment of the present invention, when offline training the policy deep neural network and the value neural network, the policy deep neural network and the value neural network are spliced into an overall deep neural network, and by the Loss function
[0046] Loss = (v NN - v) 2 - (P MCTS ) T lnP NN + c θ ‖θ‖ 2
[0047] Perform backward gradient optimization of the neural network. In the formula, v NN , P NN is a vector composed of the values and probabilities corresponding to each state of the training samples in the experience pool under the current neural network. v is the value of the air combat result corresponding to each state in the experience sample pool data, and P MCTS is the action probability output by the Monte Carlo tree search corresponding to each state in the experience sample pool data. θ is the parameter of the current neural network layer, and c θ > 0 is the weight of the network layer update parameter.
[0048] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0049] 1) Based on the Monte Carlo tree search (MCTS) algorithm, a reinforcement learning offline training architecture is added. Through the policy deep neural network and the value neural network, the node value is directly evaluated, avoiding the most time-consuming simulation link in the classical algorithm, improving its decision-making speed and global optimality of the decision-making. It also enables the Monte Carlo tree search to adapt to the high-dimensional state space, quickly find a better action strategy in the high-dimensional state space, and enable the reinforcement learning neural network to obtain feedback faster and optimize the strategy;
[0050] 2) Reinforcement learning may be affected when facing environmental uncertainties because it usually learns and plans based on a model, and the model may have errors. The Monte Carlo tree search algorithm can better adapt to environmental uncertainties by exploring unknown states and actions in a limited number of steps.
[0051] To make the above objects, features, and advantages of the present invention more obvious and understandable, specific embodiments of the present invention are hereinafter given, and in conjunction with the accompanying drawings, the following detailed description is provided. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0053] Figure 1 is a maneuver decision-making offline training framework based on MCTS and reinforcement learning;
[0054] Figure 2 is the Selplay training data collection process;
[0055] Figure 3It is an online maneuver decision-making framework based on MCTS and reinforcement learning. Detailed implementation manners
[0056] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.
[0057] Please refer to Figure 1 , Figure 2 and Figure 3 , the present invention provides a close-range air combat maneuver decision-making method based on Monte Carlo tree search and reinforcement learning. The cooperative air combat scenario considered by this method is close-range dogfight air combat. Since the combat distance is relatively close, the battlefield situation awareness level is global transparent situation, that is, both sides of the air combat can obtain the position and velocity vector information of the opponent. The two sides of the air combat are the our aircraft and the enemy aircraft. The maneuver decision-making model constructed by the present invention includes an air combat virtual environment, a policy deep neural network, a value neural network, Monte Carlo tree search and Selfplay.
[0058] 1. Design of air combat virtual environment:
[0059] The air combat virtual environment established by the present invention mainly realizes three functions: updating the aircraft state according to the maneuver, feedbacking available maneuvers according to the current aircraft state, and feedbacking action rewards according to the current state of our aircraft and the enemy aircraft. The air combat virtual environment established by the present invention includes an aircraft guidance model, a maneuver library, aircraft maneuverability limitations and single-step air combat decision rewards.
[0060] Considering the commonly used "guidance first and then tracking" control structure and decision-making real-time performance in actual engineering, the present invention selects the following three-degree-of-freedom aircraft guidance model:
[0061]
[0062] The maneuver decision input a(t) = [n x , n, μ] T are the longitudinal overload n x , the normal overload n and the speed roll angle μ respectively, and the state quantity x(t) = [V, χ, γ, x, y, z] T are the flight speed, yaw angle, climb angle and spatial position respectively.
[0063] In the MCTS (Monte Carlo tree search) algorithm and the reinforcement learning decision algorithm, the state x(k) at the current k moment is discretely updated by the decision quantity. Therefore, when the present invention updates the state x(k + 1) at the next moment according to the aircraft guidance model (1), according to the following discrete difference equation,
[0064] x(k + 1) = x(k) + f(x(k), a(k))Δt (2)
[0065] Perform state approximate update. When the update step size Δt is small, the discrete approximation error of this method is small.
[0066] After establishing the aircraft guidance model, the present invention gives the following maneuver action library and the maneuverability limit of the aircraft (each maneuver action can be expanded, and each threshold can be selected according to different maneuverabilities):
[0067] Table 1 Maneuver Action Library and Corresponding Control Quantities
[0068] Optional Maneuvering Actions Corresponding Action Control Quantities Constant Velocity Straight Flight <![CDATA[a(t) = [0, 1, 0] T > Accelerated Straight Flight <![CDATA[a(t) = [1, 1, 0] T > Decelerated Straight Flight <![CDATA[a(t) = [-1, 1, 0] T > Maneuvering Left Turn <![CDATA[a(t) = [0, 3, 30°] T > Maneuvering Right Turn <![CDATA[a(t) = [0, 3, -30°] T > Maneuvering Climb <![CDATA[a(t) = [0, 3, 0] T > Maneuvering Dive <![CDATA[a(t) = [0, -3, 0] T >
[0069] The maneuverability limit of the aircraft is
[0070] V min ≤ V ≤ V max , γ min ≤ γ ≤ γ max , z min ≤ z ≤ z max (3)
[0071] Wherein, V min , V max are the minimum and maximum flight speeds respectively, γ min , γ max are the minimum and maximum climb angles respectively, z min , z max are the minimum and maximum height positions respectively.
[0072] The aircraft maneuver decision is different maneuver actions in Table 1, belonging to the discrete action space.
[0073] In the training stage, the single-step reward function of the air combat decision is shown in Equation (4), that is, if there is a winner in the air combat between two aircraft (our aircraft and the enemy aircraft), the actual reward is ±1. If there is no victory or defeat within the limited number of state update times, the value is assigned to the process state according to the cumulative advantage of the two aircraft during the air combat process.
[0074]
[0075] Wherein, T is the number of steps for updating the air combat state, and R t is the single-step reward of the air combat decision before this reward update, and is defined as follows:
[0076] R t = a d T d + a h T h+a φq T φq
[0077]
[0078] wherein, T d , T h , T φq is the relative distance, relative altitude and angle advantage function of the aircraft, a d , a h , a φq is the corresponding weight coefficient, satisfying 1 > a φq > a d > a h > 0, σ d , σ h1 , σ h2 is the slope coefficient, and all are greater than 1, d is the distance between the two aircraft, d * is the maximum attack distance of the airborne weapon, Δh is the relative altitude between the two aircraft, Δh min , Δh max are the lower bound and upper bound of the optimal altitude difference respectively, φ and q are the azimuth angle of our aircraft and the approach angle of the enemy aircraft respectively.
[0079] It should be noted that the above air combat virtual environment needs to feedback the available maneuvering actions in the current flight state. For the available maneuvering actions in the current flight state, if they do not meet the maneuverability limit (3), they need to be deleted from the maneuvering action library in Table 1. That is, assuming the current flight state is x(k), through different maneuvering actions a in Table 1, different next moment states x(k + 1) can be obtained in turn according to (2). If the state x(k + 1) meets the state limit of (3), the corresponding maneuvering action a is added to the available maneuvering action library, otherwise it is not added.
[0080] 2. Design of the policy deep neural network and the value neural network:
[0081] For the intermediate layer of the neural network of the policy deep neural network and the value neural network, the commonly used deep linear layer in the prior art is selected, and the activation function between the linear layer networks is selected as the commonly used Relu function. The state inputs of the policy network and the value network are both:
[0082] s = [p r , v r , p b , v b , d, Δh, Δv, φ, q] T (6)
[0083] In the formula, p r , v r are the position and velocity vectors of our aircraft respectively, pb , v b are the position and velocity vectors of the enemy aircraft, respectively, and Δv, φ, q are the relative velocity, azimuth angle, and approach angle of our aircraft, respectively. Since the order of magnitude of the position vector, the distance d between the two aircraft, and the relative altitude Δh between the two aircraft is relatively large compared to other state variables, their units are scaled to the kilometer unit.
[0084] For the policy deep neural network, its output P NN is the selection probability of each action under the current aircraft state. It is necessary to map the output value of the last linear network of the policy deep neural network to the probability space through the common log_softmax function. Therefore, the activation function of its output layer is:
[0085] P NN = log_softmax(s p ) (7)
[0086] where s p is the output value of the last linear network of the policy deep neural network.
[0087] For the value neural network, its output v NN is the value of the current state. Considering the reward setting, that is, the value output by Equation (4) belongs to the interval [-1, 1]. The state value is mapped to the [-1, 1] space through the tanh function, that is
[0088] v NN = tanh(s v ) (8)
[0089] where s v is the output value of the last linear network of the value neural network.
[0090] 3. Monte Carlo tree search playout process:
[0091] The Monte Carlo tree search algorithm guesses the actions in a finite number of future steps to update the state and detect the node value. After several playouts, the "optimal" action and the corresponding probability under the current state can be obtained. Its process is roughly similar to the classic Monte Carlo tree search algorithm, mainly divided into three steps: selection, expansion, and backtracking. The specific process is as follows:
[0092] 1) Create a node according to the current aircraft state to form a Monte Carlo tree search. The node attributes include: (s, a, p, N, UCB), which represent the state, maneuver action, action probability, visit count, and UCB (a commonly used professional term in the classic Monte Carlo tree search, which is a calculation method of the node value) value of the aircraft in air combat. The calculation formula of UCB is
[0093] UCB(s,a) = Q(s,a) + U(s,a)
[0094]
[0095] Where c is the weight coefficient, representing the proportion of the probability value output by the policy network, ΣN j is the sum of the visit counts of all child nodes of the current node. For all newly created nodes, Q(s,a) is initially 0. For the root node, the action a in its attributes is empty (or assigned arbitrarily), the probability p is 1, while for non-root nodes, the action and probability value are initially defined by their parent nodes.
[0096] 2) In each playout of Monte Carlo tree search, the following operations are performed in sequence:
[0097] Select the child node with the largest UCB value among all child nodes until the current node is a leaf node; evaluate the action probability P NN and value v NN at the current state by the policy neural network and value neural network. If there is a winner in the air combat state corresponding to the current leaf node, then according to the win-loss result, set v NN = 1 or v NN = -1, and then perform backtracking to update the attributes of each layer of tree nodes. If there is no winner, after node expansion according to P NN update the attributes of each layer of tree nodes according to the original v NN .
[0098] After several playout operations as described above, a Monte Carlo tree with a certain depth and number of nodes can be obtained. Through the softmax function, the action a i in the attributes of each child node of the root node corresponding to the current aircraft state i MCTS and its corresponding probability P
[0099]
[0100] can be obtained: i where N -10 is the visit count of each child node of the root node, ∈1 = 1.0e and ∈2 ∈ (0,1] is the temperature coefficient. From (10), the current optimal action i MCTS and its corresponding probability P i can be selected according to the probability of each action. It should be noted that in the offline training stage, a noise probability value P i Dirichlet generated by Dirichlet noise needs to be added to the probability P To increase a certain exploration possibility on the basis of the action probabilities calculated by the Monte Carlo tree search method; in the online phase, no noise needs to be added, and the selection probabilities of each action are completely determined by the Monte Carlo tree.
[0101] 4. Collect experience sample pool data through Selfplay:
[0102] By combining the Monte Carlo search tree algorithm of the policy deep neural network and the value neural network, the maneuver with the highest selection probability and its probability in the current aircraft state can be achieved. Therefore, through the self-play strategy, two aircraft can be created in the air combat virtual environment, and each maneuver and its probability in the current state can be obtained through the Monte Carlo search tree algorithm respectively, as Figure 2 shown. By updating the states of the aircraft from different perspectives, the corresponding probability values P of the maneuvers with the highest probability in each state are collected. MCTS . After obtaining the air combat results through a limited number of Selfplays, according to the states of the two aircraft recorded during the air combat, the final value v of the two aircraft is determined by equation (4). p1 , v p2 . Starting from the perspectives of the two aircraft respectively, value data is added to their corresponding state and probability samples, and after merging, it is put into the sample pool. The following specific steps can be iterated to fill the experience pool with data:
[0103] Step1: Set the maximum number of iterations K of Selfplay, initialize the iteration number k = 1, initialize the air combat virtual environment, and initialize the states x r (k) and x b (k) of the red and blue aircraft in the environment, and initialize the aircraft identifier flag = 1.
[0104] Step2: Based on the states x r (k) and x b (k) of the two aircraft, if falg = 1 and the current aircraft perspective is the red side, then calculate the root node state s i from the perspective of the red aircraft; if flag = 0, then calculate the root node state s i from the perspective of the blue aircraft.
[0105] Step3: Create a Monte Carlo tree based on the root node state s i , and after a limited number of playouts, output each action and its probability (a i , P i MCTS ), record and save the current (s i , P i MCTS ), and record the aircraft state x r(k), x b (k) is saved. According to each action and its probability (a i , P i MCTS ), the optimal action is obtained After that, the state of the aircraft in the current perspective is updated by (2), that is
[0106] If flag = 1, then update x r (k) to x r (k) + f(x r (k), a * )Δt;
[0107] If flag = 0, then update x b (k) to x b (k) + f(x b (k), a * )Δt.
[0108] Step4: Determine whether k = K. If so, go to Step5; if not, perform a negation operation on flag, set k = k + 1, and jump to Step2.
[0109] Step5: The air combat process ends. According to the saved states of the two aircraft x r (k), x b (k), k = 1, 2,..., K, calculate the process rewards from the perspective of our aircraft (red aircraft) and the enemy aircraft (blue aircraft) respectively by (5) and the process rewards from the perspective of the enemy aircraft (blue aircraft) k = 1, 2,..., K, and according to the air combat result, obtain the final value v of the red aircraft p1 and the final value v of the blue aircraft p2 , that is
[0110]
[0111] According to the saved (s i , P i MCTS ) data of the two aircraft respectively, expand them to (s i , P i MCTS , v p1 ), (s i , P i MCTS , v p2 ) and put them into the experience pool, and the overall process ends.
[0112] 5. Parameter Update of Policy Deep Neural Network and Value Neural Network:
[0113] As can be seen from Figure 1 , the policy deep neural network and the value neural network can be concatenated into an overall deep neural network. The backward gradient optimization of the neural network is performed by the following Loss function (12). After the update, it is equivalent to refreshing the overall network. Since the network structure remains unchanged, the relationship between the updated parameters and the policy deep neural network and the value neural network is the same as before.
[0114] Loss = (v NN - v) 2 - (P MCTS ) T lnP NN + c θ ‖θ‖ 2 (13)
[0115] In the formula, v NN , P NN is a vector composed of the values and probabilities corresponding to each state of the training samples in the experience pool under the current neural network. From the acquisition process of the Selfplay sample pool, v is the value of the air combat result corresponding to each state of the training samples in the experience pool, and P MCTS is the action probability output by the Monte Carlo tree search corresponding to each state of the training samples in the experience pool. θ is the parameter of the current neural network layer, and c θ > 0 is the weight of the network layer update parameter.
[0116] In summary, the present invention provides a complete optimization scheme for maneuver decision-making based on Monte Carlo tree search and reinforcement learning, mainly for the specific implementation process of offline training. The online use process is based on the trained policy deep neural network and value neural network. By using the limited-step playout iteration in the Monte Carlo tree search algorithm introduced above, the selection probability of each action under the current aircraft state can be obtained online. The optimal maneuver action can be selected according to the probability, that is
[0117]
[0118] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A close-range air combat maneuver decision-making method based on Monte Carlo tree search and reinforcement learning, characterized in that including S1. Construct an air combat virtual environment, including an aircraft guidance model, a maneuver action library, maneuverability limitations, and single-step rewards for air combat decisions. The maneuver action library records the maneuver actions of several aircraft; the single-step reward for air combat decisions is obtained based on the aircraft guidance model and the outcome of the air combat victory or defeat of the aircraft after performing the maneuver action selected from the maneuver action library; for maneuver actions that do not meet the maneuverability limitations, they need to be deleted from the maneuver action library; the maneuverability limitations include the speed limitation, climb angle limitation, and altitude position limitation of the aircraft; S2. Construct a policy deep neural network, a value neural network, and Monte Carlo tree search; S3. Based on Monte Carlo tree search, the air combat virtual environment, and the historical data of the aircraft, calculate the data of the experience sample pool through Selfplay, and perform offline training on the policy deep neural network and the value neural network with the data of the experience sample pool. During the training process, the probability output by Monte Carlo tree search needs to be corrected according to the noise; S4. According to the real-time data of the aircraft, the air combat virtual environment, Monte Carlo tree search, and the trained policy deep neural network and value neural network, obtain the maneuver actions of the aircraft and the probabilities corresponding to the maneuver actions, and select the maneuver action corresponding to the maximum probability as the close-range air combat maneuver decision of the aircraft.
2. The close-range air combat maneuver decision-making method based on Monte Carlo tree search and reinforcement learning according to claim 1, characterized in that The aircraft guidance model is: The input of maneuver decision-making \(a(t)=[n x ,n,\mu] T are respectively the longitudinal overload \(n x \), the normal overload \(n\) and the speed roll angle \(\mu\), and the state variables \(x(t)=[V,\chi,\gamma,x,y,z] T are respectively the flight speed, the yaw angle, the climb angle and the spatial position; The state x(k) at the current k moment is discretely updated by the decision-making quantity. When the aircraft guidance model updates the state x(k + 1) at the next moment, according to the discrete difference equation x(k + 1) = x(k) + f(x(k), u(k))Δt perform approximate state update, where Δt is the update step size.
3. The close-range air combat maneuver decision-making method based on Monte Carlo tree search and reinforcement learning according to claim 2, characterized in that The maneuver actions in the maneuver action library include uniform straight flight, accelerated straight flight, decelerated straight flight, maneuver left turn, maneuver right turn, maneuver climb, and maneuver dive, and each maneuver action has a corresponding action control quantity.
4. The method for making close-range air combat maneuver decisions based on Monte Carlo tree search and reinforcement learning according to claim 3, characterized in that The maneuverability limitations of the aircraft are: V min V ≤ V ≤ V max , γ min γ ≤ γ ≤ γ max , z min z ≤ z ≤ z max where, V min , V max are the minimum and maximum flight speeds, γ min , γ max are the minimum and maximum climb angles, z min , z max are the minimum and maximum altitude positions respectively.
5. The method for making close-range air combat maneuver decisions based on Monte Carlo tree search and reinforcement learning according to claim 4, characterized in that, The single-step reward function for air combat decisions is where T is the number of steps for updating the air combat state, and R t is the single-step reward for the air combat decision before this reward update; Among them, T d , T h , T φq are respectively the relative distance, relative height, and angular advantage functions of the aircraft. a d , a h , a φq are the corresponding weight coefficients, satisfying 1 > a φq > a d > a h > 0. σ d , σ h1 , σ h2 are slope coefficients, and all are greater than 1. d is the distance between the two aircraft, d * is the maximum attack distance of the airborne weapon. Δh is the relative height between the two aircraft, Δh min , Δh max are respectively the lower bound and upper bound of the optimal height difference. φ and q are respectively the azimuth angle of our aircraft and the approach angle of the enemy aircraft.
6. The method for making a close air combat maneuver decision based on Monte Carlo tree search and reinforcement learning according to claim 5, characterized in that The state inputs of the policy deep neural network and the value neural network are both: s = [p r , v r , p b , v b , d, Δh, Δv, φ, q] T where p r , v r are the position and velocity vectors of our aircraft, respectively, and p b , v b are the position and velocity vectors of the enemy aircraft, respectively. Δv, φ, and q are the relative velocity, azimuth angle, and approach angle of our aircraft; For the policy deep neural network, its output P NN is the selection probability of choosing each action in the current aircraft state, P NN = log_softmax(s p ) where s p is the output value of the linear network of the last layer of the policy deep neural network; For the value neural network, its output v NN is the value of the current state, v NN = tanh(s v ) where s v is the output value of the last linear network of the value neural network.
7. The close-range air combat maneuver decision-making method based on Monte Carlo tree search and reinforcement learning according to claim 6, characterized in that, The algorithm flow of Monte Carlo tree search includes: 1) Create a node according to the current aircraft state. The node attributes include: (s, a, p, N, UCB), which respectively represent the state, maneuver action, action probability, number of visits, and UCB value of the aircraft in air combat. The calculation formula of UCB is UCB(s, a) = Q(s, a) + U(s, a) where c is the weight coefficient, representing the proportion of the probability value output by the policy network, and ΣN j is the sum of the visit counts of all child nodes of the current node. For all newly created nodes, Q(s,a) is initially 0. For the root node, the action a in its attributes is empty, and the probability p is 1. For non-root nodes, the action and probability value are initially defined by their parent nodes; 2) In each playout of Monte Carlo Tree Search, the following operations are performed in sequence: Select the child node with the largest UCB value among all child nodes until the current node is a leaf node; Evaluate the action probability P NN and the value v NN at the current state by the policy neural network and the value neural network. If there is a winner in the air combat state corresponding to the current leaf node, then set v NN = 1 or v NN = -1 according to the win-lose result, and then perform backtracking to update the attributes of tree nodes at each layer. If there is no winner, after expanding the node according to P NN , update the attributes of tree nodes at each layer according to the original v NN . After several playout operations, action a in the attributes of each child node corresponding to the current aircraft state is obtained through the softmax function i and its corresponding probability P i MCTS : where N i is the access times of each child node of the root node, ∈1 = 1.0e -10 , ∈2 ∈ (0, 1] is the temperature coefficient.
8. The method for making a short-range air combat maneuver decision based on Monte Carlo tree search and reinforcement learning according to claim 7, wherein When performing offline training on the policy deep neural network and the value neural network, splice the policy deep neural network and the value neural network into an overall deep neural network, and use the Loss function Loss=(v NN -v) 2 -(P MCTS ) T lnP NN +c θ ‖θ‖ 2 Perform backward gradient optimization of the neural network, where v NN , P NN is a vector composed of the values and probabilities corresponding to each state of the training samples in the experience pool under the current neural network. v is the air combat result value corresponding to each state in the experience sample pool data, and P MCTS is the action probability output by Monte Carlo tree search corresponding to each state in the experience sample pool data. θ is the parameter of the current neural network layer, and c θ > 0 is the weight of the network layer update parameter.
Citation Information
Patent Citations
Unmanned aerial vehicle obstacle avoidance and path planning method
CN113110592A
Method for realizing real-time determination of optimal decision-making action by intelligent real-time decision-making system
CN114462566A
Curling stone decision-making method of deep reinforcement learning based on Monte Carlo tree search
CN114581834A
Automatic driving longitudinal decision-making method based on Monte Carlo tree search
CN116341662A
Chess self-learning method and device based on machine learning
US20220379224A1