Methods for close-range search and precision strike of loitering munitions against maneuvering targets under conditions of information opacity
By combining the A* algorithm and deep learning networks, the trajectory planning and strike strategy of loitering munitions are optimized, solving the problem of trajectory planning and strike in environments with opaque information, and enabling rapid and accurate strikes against mobile targets.
Patent Information
- Application Number
- CN202410481569.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-22
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-04-22
AI Technical Summary
Existing loitering munitions face difficulties in online adjustment of trajectory planning and engagement of maneuvering targets in environments with opaque information, making it impossible to quickly respond to changes in targets and enemy maneuvers, leading to an increased miss rate.
By combining the A* algorithm with a deep learning network, the guidance law is optimized through interactive learning, generating autonomous search and strike strategies. The reward function of the A* algorithm and the reward function of the deep learning network are used to optimize the trajectory planning and strike method of the loitering munition.
In environments with opaque information, loitering munitions can quickly generate autonomous search plans, reconstruct trajectories, and achieve precise positioning and strikes against maneuvering targets, meeting the end-to-end time sensitivity requirements online.
Smart Images

Figure CN118794303B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of guidance technology, and in particular relates to a method for close-range search and precision strike of maneuvering targets by a loitering munition under conditions of information opacity. Background Technology
[0002] Loitering munitions are a new type of guided weapon that organically integrates advanced UAV technology and munitions technology. They can reach designated target areas at low altitudes and low speeds to perform various combat missions such as "cruising flight," "standby," "reconnaissance," and "strike." They can be launched from different delivery platforms like traditional munitions, and can also conduct reconnaissance of combat areas like UAVs. They are less expensive than cruise missiles, can quickly enter target areas, and feature high launch speed, long loiter time, and tactical flexibility. They can perfectly integrate battlefield reconnaissance, target identification, and damage assessment, possessing rapid response and precision strike capabilities. Existing methods for generating precision strike trajectories mainly include heuristic algorithms such as genetic algorithms and particle swarm optimization, traditional guidance law algorithms, and traditional graph search algorithms such as the flight extension node method. However, these studies rarely consider the problem of generating trajectories for precise strikes against maneuvering targets by loitering munitions in environments with opaque information.
[0003] The current method for generating trajectories for precise strikes against maneuvering targets by loitering munitions has the following shortcomings:
[0004] 1) In trajectory planning and trajectory reconstruction due to mission changes, the heuristic algorithms that are currently widely researched and applied are simple to implement and stable to run, but they rely on global scene information. The computational load increases exponentially with the increase of the exploration space and the complexity of the influencing factors. For changes in the enemy's situation in complex battlefield environments, the algorithm must be run again, making it difficult to balance online real-time performance and global optimization.
[0005] 2) Regarding precision strikes, key enemy targets may possess maneuverability and evade our loitering munitions through serpentine maneuvers, increasing the miss rate of our loitering munitions. Existing algorithms for correcting traditional guidance rates (bias ratio guidance law, optimal guidance law, etc.) are mostly designed for stationary targets or rely on stringent assumptions, lacking long-term decision-making capabilities and unable to cope with game-theoretic scenarios where we are moving and the enemy is moving. Trajectory generation methods based on optimal control or heuristic algorithms have long optimization times and rely on precise information about the scene and the enemy's maneuver parameters, making it difficult to meet the real-time response requirements for precision strikes against maneuvering targets.
[0006] In summary, current loitering munition trajectory dynamic planning primarily relies on a "man-in-the-loop" command and control mode, resulting in a low level of automation in weapon command and control. There is an urgent need to develop intelligent loitering munition precision strike technology for maneuvering targets. In opaque operational environments, enemy target information cannot be accurately obtained, and enemy targets may possess maneuverability to evade attacks. During the target approach search phase, loitering munitions have poor adaptability to incomplete information environments, and existing trajectory planning algorithms face difficulties in online adjustment, failing to meet the demands for speed and accuracy in online trajectory reconstruction under target changes or mission alterations. In the precision strike phase, targets evade attacks through maneuvering maneuvers such as sudden movements, stops, and turns; existing target precision strike algorithms lack long-term decision-making capabilities, leading to an increase in miss rates. Summary of the Invention
[0007] The purpose of this invention is to provide a method for loitering munitions to conduct close-range search and precision strikes against maneuvering targets under conditions of information opacity. This addresses the shortcomings of loitering munitions in adapting to incomplete information environments during the target close-range search phase. Existing trajectory planning algorithms face difficulties in online adjustment and cannot meet the requirements for speed and accuracy in online trajectory reconstruction under target changes or mission alterations. Furthermore, during the target precision strike phase, targets evade attacks through maneuvering patterns such as sudden movements, stops, and turns. Existing target precision strike algorithms lack long-term decision-making capabilities, leading to an increase in misses.
[0008] This invention employs the following technical solution: a method for close-range search and precision strike of maneuvering targets by a loitering munition under information opacity, characterized in that it includes:
[0009] Step 1: Calculate the trajectory of the loitering munition to the target using the A* algorithm, and smooth the trajectory.
[0010] Step 2: Calculate the guidance law based on the smoothed trajectory, and use a deep learning network to interactively learn the guidance law as the action and the loitering munition's vector as the state, thereby optimizing the guidance law so that the loitering munition can accurately strike the target.
[0011] The reward function of the A* algorithm in step 1 is:
[0012] g = α × q + β × b;
[0013]
[0014]
[0015] In the formula, α and β represent the weight values of the two, respectively; q is the normalized value of the importance probability; b is the normalized value of the distance length; p is the original importance probability, q + p = 1; L long L represents the longest distance traveled. short L represents the shortest distance;now This represents the current distance traveled.
[0016] Furthermore, the reward function of the deep learning network in step 2 is:
[0017]
[0018] In the formula, r represents the single-step reward after correction at the i-th time step; i The single-step reward before the correction at the i-th time step is the sum of the dense event-driven reward and the sparse final-state reward; γ a τ is the decay factor, τ is the total number of time steps for the smart ammunition in this round, and G is the final state reward obtained at the end of this round;
[0019] The reward function for intensive events is as follows:
[0020] r exp =w1×(r1+r2)+w2×r3,
[0021] In the formula, w1 and w2 represent the weighting coefficients of the above-mentioned dense returns, respectively;
[0022] in,
[0023] r1 = dist mt_old -dist mt_now r2 = v m ×cosε;r3=dist mo_now -dist mo_old In the formula, dist mt_old Dist indicates the distance between the smart munition and the target at the previous moment. mt_now This indicates the distance between the smart munition and the target at the current moment; v m The velocity of the smart munition is represented by ε, where ε represents the angle between the smart munition's velocity vector and the target's line of sight; dist mo_now Dist indicates the current distance between the smart munition and the no-fly zone. mo_old This indicates the distance between the smart munition and the no-fly zone at the previous moment.
[0024] Furthermore, in step 2, the vector of the loitering munition is:
[0025]
[0026] In the formula, dist mt For the distance between the target and the target, y m For the elevation angle of the ammunition, ψ m For the ammunition yaw angle, ε mt For the line-of-sight tilt angle of the bullet, β mt For the angle of view of the bullet, v mFor ammunition speed, D T For time constraints, D V For terminal speed, D A From the terminal angle.
[0027] The beneficial effects of this invention are:
[0028] This invention relates to a loitering munition technology for precision strikes against maneuvering targets in environments with opaque information. Addressing the challenges of deception and misdirection by enemy targets during combat, as well as the limitations of our sensors leading to inaccurate target location information, the loitering munition can rapidly generate an autonomous search plan, reconstruct its trajectory, and achieve precise target localization by combining constraints on insurmountable areas and flight conditions. Upon target detection, the loitering munition can generate a forward-looking maneuvering strategy to predict future target routes and improve strike accuracy, particularly in response to target maneuvering and escape attempts.
[0029] This invention can transfer most of the trajectory optimization computation to the offline training stage. When the target information is inaccurate, the loitering munition can generate maneuver strategies online end-to-end, meet the time sensitivity requirements of the precision strike algorithm, and achieve rapid closure of the OODA loop.
[0030] This invention utilizes an interactive self-learning method to enable loitering munitions to learn spontaneously in an offline environment, responding to the maneuvering strategies of various escape modes of enemy targets. It rapidly generates precision-guided trajectories for maneuvering targets online end-to-end, meeting the requirements of operational time sensitivity. In response to enemy targets exhibiting behaviors such as concealment, deception, and position shifting, it rapidly covers and searches the possible operational areas of the enemy based on their potential maneuverability, thus solving the problem of precision-targeting maneuvering targets after they have been located. Attached Figure Description
[0031] Figure 1 This is a combat simulation diagram of Embodiment 1 of the present invention;
[0032] Figure 2 The search track after smoothing in Embodiment 1 of the present invention;
[0033] Figure 3 This refers to the change in the viewing angle of the loitering munition in Embodiment 1 of the present invention;
[0034] Figure 4 The lateral overload variation of the loitering munition during the entire search process of the trajectory tracking missile in Embodiment 1 of the present invention;
[0035] Figure 5 The simulation results are for the traditional proportional guidance method;
[0036] Figure 6 The simulation results are for the traditional proportional guidance method;
[0037] Figure 7The simulation results show the results of generating guidance commands for striking maneuvering targets using the intelligent precision strike algorithm designed in Embodiment 1 of the present invention.
[0038] Figure 8 The simulation results show the results of generating guidance commands for striking maneuvering targets using the intelligent precision strike algorithm designed in Embodiment 1 of the present invention. Detailed Implementation
[0039] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0040] Based on the key nodes of the strike mission, this invention can be divided into a close-range search phase and a precision strike phase on mobile targets:
[0041] (1) Close-range search phase: Based on the detection capabilities of the loitering munition, the optimal search scheme and optimal search trajectory are generated. The coordinates of all possible enemy targets are traversed, and the reconstructed trajectory is generated by combining various factors such as the probability of enemy target existence, path obstacle constraints, flight capability constraints, and target distance.
[0042] (2) Precision strike phase for maneuvering targets: In response to the potential maneuvering characteristics of enemy targets, an autonomous decision-making framework based on knowledge-assisted deep reinforcement learning is adopted to quickly generate response strategies and complete the precision strike mission.
[0043] Specifically, the method for close-range search and precision strike of maneuvering targets by a loitering munition under opaque information disclosed in this invention includes the following steps:
[0044] Step 1: Calculate the trajectory of the loitering munition to the target using the A* algorithm, and then smooth the trajectory.
[0045] Because loitering munitions have a limited turning radius, this mission becomes virtually impossible when the angle between two waypoints is too small. To address this issue, considering the varying probability of target presence during close-range searches and the different distances from the starting point in different areas, a new reward function is proposed.
[0046] By considering the probability of the target's existence in the area and its distance, normalization is performed first, and then different weighting processes are used to obtain the target search route with the highest comprehensive value, thereby achieving better search efficiency.
[0047] Step 1 includes:
[0048] Step 101: Initialize the node linked list and set the true cost of the target point (i.e. the length of the initial optimal trajectory) to infinity g(start) = ∞.
[0049] Step 102: Calculate the value between (start, goal) using the cost function. During the iterative search process, simultaneously maintain the three linked lists: OPEN, CLOSED, and INCONS. The general form of the cost function can be expressed as:
[0050] f(s) = g(s) + ε × h(s)
[0051] In the formula, g(s) represents the actual cost from the starting point to the current node s; h(s) represents the estimated cost from the current node to the target node; and ε is the heuristic weight coefficient (ε≥1).
[0052] Step 103: When the algorithm repeatedly expands to a node in CLOSED, it decreases the g value of that node, updates its cost value, and adds the node to INCONS. Nodes in INCONS are not immediately expanded; instead, they are merged with the OPEN list and reordered for expansion in the next iteration.
[0053] Step 104: After completing a single trajectory planning, if the planning time does not exceed the limit, update the heuristic weights and node linked list, and use this as the basis for the next weighted planning. The heuristic weights adopt a linear reduction method, represented as:
[0054] ε = ε - Δε
[0055] In the formula, the empirical value of Δε is taken as 0.1.
[0056] The linked list update includes: emptying the CLOSED list; moving nodes from the INCONS list to the OPEN list and emptying the INCONS list; and recalculating the node values in the OPEN list according to the new heuristic weights.
[0057] Step 105: In node expansion, the minimum track length, maximum turning angle, maximum climb and glide angle, and other aircraft dynamic constraints, as well as environmental constraints such as flight altitude, no-fly zones and terrain restrictions, are used as the basis. If a node does not meet the aircraft dynamic constraints or conflicts with the environmental constraints, the node is considered an invalid node and will not be expanded.
[0058] Step 106: The exit condition for a single search is that the target point is expanded. Since the cost of the OPEN table nodes is updated according to the new heuristic weights each time the weighted algorithm is executed, the exit condition for each execution of the weighted SAS algorithm is:
[0059] f(s goal ) <min s∈OPEN (f(s))
[0060] In the formula, s goal Represents the target node; f(s) goal) represents the cost of the current path; min s∈OPEN(f(s)) This represents the minimum cost of all nodes in the OPEN table. During a single weighted SAS algorithm search, if the cost of all nodes in the OPEN table is greater than the current path cost, the search is terminated for that path.
[0061] Step 107: Reorder the search targets using the reward function. First, calculate the normalized values of probability and distance using the importance probability normalization formula and the distance length normalization formula, respectively. Then, substitute these values into the reward function formula to obtain the reward values of the search targets and sort them.
[0062] The definition of the reward function is:
[0063] g = α × q + β × b
[0064] In the formula, q and b are the normalized values of probability and distance, respectively, and α and β represent the weight values of the two.
[0065] Importance probability normalization formula:
[0066]
[0067] In the formula, q is the normalized value of the importance probability; p is the original importance probability, and it is necessary to ensure that the sum of the importance probabilities is 1, that is, q+p=1;
[0068] Normalized formula for distance length:
[0069]
[0070] In the formula, b is the normalized value of the distance; L long L represents the longest distance traveled, the longest distance traveled among all targets. short L represents the shortest distance, which is the shortest distance among all objectives. now This represents the current distance traveled.
[0071] Since the curves generated by the track points under the grid are not smooth enough and there are too many broken lines, a combination of second-order and third-order Bézier curves is used to optimize the path smoothness.
[0072] Step 2: Calculate the guidance law based on the smoothed trajectory, and use a deep learning network to interactively learn the guidance law as the action and the loitering munition's vector as the state, and optimize the guidance law so that the loitering munition can accurately strike the target.
[0073] Step 2 includes the following steps:
[0074] Step 201: State Design
[0075] The acquired raw battlefield information is processed to obtain situational information, which is then normalized and used as the state of the Markov Decision Process (MDP) model, including the distance between the missile and the target. mt ammunition velocity v m Ammunition yaw angle ψ m , Ammunition pitch angle y m β, the angle of deflection of the bullet's line of sight mt ε, the angle of view of the bullet mt It also includes the output of the higher-level decision-making module, i.e., constraint information, including time constraint D. T Terminal speed D V Terminal angle D A The state design of MDP is a vector S composed of these scalar normalized state variables:
[0076] Step 202: Motion Design
[0077] Reinforcement learning algorithms plan flight paths by outputting actions, which are considered as guidance laws. This invention uses a traditional proportional guidance algorithm as prior knowledge to design the guidance law of the munition. That is, based on the yaw and pitch overloads calculated by the proportional guidance method, it learns the command sequence and finally uses the combined overloads to guide the loitering munition to strike.
[0078] Step 203: Reward Function Design
[0079] Considering the characteristics of the object, the purpose of the task, and the constraint information, the reward function is designed into two types: sparse final state reward function and dense event-guided reward function.
[0080] Sparse final state reward refers to the reward value given to smart ammunition based on the mission completion result at the end of the round. According to the constraints specified in the input scheme, the final state result includes 6 cases, and the result description and reward value are as follows:
[0081] 1. Smart ammunition hits target: 200*k;
[0082] 2. Fly away from the combat zone: -200;
[0083] 3. The angle of the intelligent ammunition terminal does not meet the constraint: -100;
[0084] 4. Failure to reach the target area within a certain time: -100;
[0085] 5. Collision into no-fly zones, such as terrain that cannot be crossed: -200;
[0086] 6. The speed of the intelligent ammunition terminal does not meet the constraint: -100.
[0087] In the formula, k is the discount factor, defined as: Where `max` represents the maximum number of steps allowed in a single scene, and `used` represents the current number of steps taken. The shorter the time taken to complete the task, the larger the discount factor and the higher the reward, which is used to guide the smart munitions in planning the shortest possible trajectory.
[0088] Based on whether the battlefield situation at the next moment is more conducive to mission completion after a single time step update, a dense event-guided reward function is constructed to guide the learning direction of the intelligent ammunition strategy, which is an adjustment to the sparse final state reward.
[0089] This invention extracts four rules for designing reward functions for intensive events.
[0090] Rule 1: The planned trajectory must be close to the target in terms of distance.
[0091] The reward is positive if the distance between the smart munition and the target decreases, and negative otherwise. As shown in the following formula:
[0092] r1 = dist mt_old -dist mt_now
[0093] In the formula, dist mt_old Dist indicates the distance between the smart munition and the target at the previous moment. mt_now This indicates the distance between the smart munition and the target at the current moment.
[0094] Rule 2: The planned trajectory velocity direction should be consistent with the target's line-of-sight angle.
[0095] The lead angle is defined as the angle between the velocity vector direction and the target's line-of-sight angle. When the velocity magnitudes are the same, the reward is negatively correlated with the lead angle magnitude; while when the lead angle magnitudes are the same, the reward is positively correlated with the velocity magnitude. This is expressed as:
[0096] r2 = v m ×cosε
[0097] In the formula, v m ε represents the velocity of the smart munition, and ε represents the angle between the smart munition's velocity vector and the target's line of sight.
[0098] Rule 3: The planned trajectory of smart munitions must be within the flyable area.
[0099] Smart munitions cannot enter no-fly zones defined by terrain constraints, operational constraints, etc. If the distance between the smart munition and the no-fly zone increases, the reward is positive, with the positive reward increasing as the distance increases; conversely, if the distance between the smart munition and the no-fly zone decreases, the penalty is negative, with the negative penalty increasing as the distance increases. This is represented as:
[0100] r3 = distmo_now -dist mo_old
[0101] In the formula, dist mo_now Dist indicates the current distance between the smart munition and the no-fly zone. mo_old This indicates the distance between the smart munition and the no-fly zone at the previous moment.
[0102] Rule 4: Smart munitions priority strategies should differ depending on the situation.
[0103] By weighted summing the above dense rewards, we obtain the reward function based on real-time exploration as shown in the following equation:
[0104] r exp =w1×(r1+r2)+w2×r3
[0105] In the formula, w1 and w2 represent the weighting coefficients of the aforementioned dense rewards, and their specific values are dynamically adjusted during the interaction. r1 and r2 both guide the smart munitions to fly towards the target, therefore they share the weighting coefficient w1.
[0106] The single-step reward for autonomous trajectory planning of intelligent munitions includes two types: dense event reward and sparse final state reward. In each training round, only the last time step receives a sparse reward reflecting the mission outcome. In reality, the trajectory planning result is influenced by the decisions made at each time step in that round, not just the decision made at the last time step. Using the final state reward only to evaluate the decision made at the last time step ignores the influence of previous time steps on the result.
[0107] Because the final state reward should be related to all actions taken in this round, and the closer the step is to the final state, the stronger its correlation with the task outcome, the single-step reward for each sample is adjusted before storing it in the experience pool, for example:
[0108]
[0109] In the formula, r i The single-step reward before the correction at the i-th time step, i.e., the reward for intensive event bootstrapping r. exp The sum of the sparse final state rewards, γ represents the corrected single-step reward at the i-th time step. a τ is the decay factor, τ is the total number of time steps for the smart ammunition in this round, and G is the final state reward obtained at the end of this round.
[0110] Example 1
[0111] The combat simulation diagram constructed by this invention is as follows: Figure 1As shown, the entire simulation scene is processed into a grid, with a size of 30km × 30km. The purple area in the figure represents insurmountable mountain peaks, and the three square areas represent potential target locations provided by the grid information network, labeled as target points A, B, and C. The different brightness colors of the target points distinguish different probability of target presence; the blue brightness, from brightest to darkest, corresponds to a probability of 70%, 50%, and 30% respectively (corresponding to target points B, C, and A). Based on the search order calculation method designed above, considering the distance between the loitering munition and the target location, and the probability of target presence, α = 0.7 and β = 0.3 are selected to calculate the priority score, generating a search order of C→B→A. The smoothed search trajectory is shown below. Figure 2 As shown, this track is suitable for actual operation.
[0112] In the simulation scenario of loitering munition trajectory tracking control, the available overload is set to n∈[-5g,5g], K p =10.0, K i =0.01, the loitering munition's viewpoint changes as follows Figure 3 As shown. The lateral overload changes of the loitering munition during its trajectory tracking process throughout the entire search process are as follows. Figure 4 As shown, in Figure 4 During the flight, the overall lateral overload variation was relatively stable, fluctuating cyclically with two plateau periods. During these plateau periods, the loitering munition was in a circling and patrolling phase, conducting a comprehensive search of areas where targets might be present. The entire curve shows that the lateral overload of the loitering munition remained below the set value of 5g throughout the entire flight. This indicates that after smoothing, the generated path allows the loitering munition to track the trajectory well, and the generated trajectory meets the minimum turning radius and other specifications of the loitering munition. (Comparison...) Figure 5 , 6 and Figure 7 , 8 As can be seen, the loitering munition clearly switches between flight phases. During loitering flight, the trajectory angle changes gradually, facilitating the munition's safe arrival at the search area. During hovering flight, the munition rapidly maneuvers around the target area, improving search efficiency. The trajectory generated by this invention and the parameter settings of the tracking algorithm both meet the flight constraints of the loitering munition, satisfying the search mission requirements in battlefield environments with incomplete information. In 100 Monte Carlo simulations conducted in a randomly generated scenario at the target location, the average algorithm performance was calculated, with an average running time of 0.072 seconds and an average path length of 35 kilometers. The approach search algorithm designed in this invention meets the requirements for fast online computation, and the planned trajectory meets the set range constraints.
[0113] To address the shortcomings of strategy optimization, such as its failure to prioritize tasks, a single-step reward correction mechanism that prioritizes tasks is designed. Furthermore, a strategy exploration mechanism and a neural network structure are designed to generate and optimize the global optimal trajectory of intelligent munitions under task changes.
[0114] The actions output by the policy network (Actor) in DDPG are deterministic, eliminating the need for sampling from a distribution and significantly reducing computational cost. During training, random noise is added to the actions output by the policy network to enhance the exploration capabilities of the smart munitions and diversify the experience. For example: a t =μ(s) t |θ t )+N, where N follows a zero-mean global distribution, i.e. K e is a constant coefficient, and i is the number of training iterations.
[0115] In the early stages of training, a traditional guidance method is introduced, transforming the constraint information output by the autonomous decision-making module into sparse final-state rewards and dense event-guided rewards. An actor-critic framework is adopted, where the evaluation network receives feedback rewards from environmental inputs to update the evaluation network, and the output action evaluation is fed back to the actor policy network to update the action policy. The policy network outputs a policy, and after policy exploration, a control overload 'a' is obtained. After adding exploration noise, the guidance law of the ammunition is obtained. A single-step reward correction is introduced to obtain a new reward 'r'. The current state 's', the ammunition overload 'a', the new reward 'r', and the next state 's' are stored in the experience pool to complete the single-step interaction process.
[0116] In the simulation test scenario, it is set that when the target detects that the distance between the target and the missile is less than 500m, it will start to perform serpentine maneuvers to escape the attack. That is, two overloads are generated in the y direction, and the overload direction is changed every 2 to 3 seconds to perform reverse maneuvers.
[0117] The reinforcement learning training parameters are set as follows:
[0118] 1. Remaining Time Return Coefficient K tgo =1;
[0119] 2. Remaining time threshold t threshold =0.11;
[0120] 3. Policy Network Actor Learning Rate α θ =0.0001;
[0121] 4. Final state reward correction coefficient γ a =0.98;
[0122] 5. Evaluate the learning rate α of the network Critic. ω =0.001;
[0123] 6. Maximum and minimum available overload n max =±5;
[0124] 7. Reward discount factor γ = 0.9;
[0125] 8. Action exploration variance σ 2=15;
[0126] 9. Experience pool size capacity = 20000;
[0127] 10. Variance decay rate K e =0.9995;
[0128] 11. The target network update coefficient τ = 0.01;
[0129] 12. Single-step update duration Δt = 0.1;
[0130] 13. Batch training sample size: batch = 128;
[0131] 14. Total number of training rounds M = 10000.
[0132] During training round 3700, the loitering munition demonstrated good strike effectiveness. The network parameters were saved and used to generate guidance commands for the loitering munition during simulation. The initial missile and target information were set as follows: loitering munition (650±50, 2000±50, 900, 100) and target point (1000±50, 2000±50, 50, 20±4).
[0133] Using the same parameter settings, the loitering munitions were instructed to generate guidance commands for attack using both the traditional proportional guidance method and the intelligent precision strike algorithm for maneuvering targets designed in this invention. The simulation results are as follows: Figure 8 As shown, the accuracy of the algorithm designed in this invention in striking the target can be clearly observed. The simulation scenario was changed to increase the randomness of the target's initial position, velocity magnitude, velocity direction, and maneuver timing. 100 simulations were conducted using both proportional guidance and reinforcement learning guidance. The success rates were statistically analyzed as follows: proportional guidance had a success rate of 6%; reinforcement learning guidance had a success rate of 100%, fully verifying the effectiveness of the precision strike algorithm of this invention in a target serpentine maneuvering combat scenario.
[0134] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for close-range search and precision strike of maneuvering targets by a loitering munition under conditions of information opacity, characterized in that: include: Step 1: Calculate the trajectory of the loitering munition to the target using the A* algorithm, and smooth the trajectory. Step 2: Calculate the guidance law based on the smoothed trajectory, and use a deep learning network to interactively learn the guidance law as the action and the loitering munition's vector as the state, thereby optimizing the guidance law so that the loitering munition can accurately strike the target. The reward function of the A* algorithm in step 1 is: g = α × q + β × b; In the formula, α and β represent the weight values of the two, respectively; q is the normalized value of the importance probability; b is the normalized value of the distance length; p is the original importance probability, q + p = 1; L long L represents the longest distance traveled. short L represents the shortest distance; now This represents the current distance. The reward function for the deep learning network in step 2 is: In the formula, r represents the single-step reward after correction at the i-th time step; i The single-step reward before the correction at the i-th time step is the sum of the dense event-driven reward and the sparse final-state reward; γ a τ is the decay factor, τ is the total number of time steps for the smart ammunition in this round, and G is the final state reward obtained at the end of this round; The reward function for intensive events is as follows: r exp =w1×(r1+r2)+w2×r3, In the formula, w1 and w2 represent the weighting coefficients of the above-mentioned dense returns, respectively; in, r1 = dist mt_old -dist mt_now r2 = v m ×cosε;r3=dist mo_now -dist mo_old In the formula, dist mt_old Dist indicates the distance between the smart munition and the target at the previous moment. mt_now This indicates the distance between the smart munition and the target at the current moment; v m The velocity of the smart munition is represented by ε, where ε represents the angle between the smart munition's velocity vector and the target's line of sight; dist mo_now Dist indicates the current distance between the smart munition and the no-fly zone. mo_old This indicates the distance between the smart munition and the no-fly zone at the previous moment; In step 2, the vector of the loitering munition is: In the formula, dist mt For the distance between the target and the target, y m For the elevation angle of the ammunition, ψ m For the ammunition yaw angle, ε mt For the line-of-sight tilt angle of the bullet, β mt For the angle of view of the bullet, v m For ammunition speed, D T For time constraints, D V For terminal speed, D A From the terminal angle.
Citation Information
Patent Citations
Unmanned aerial vehicle air combat autonomous avoidance maneuvering decision-making method based on deep reinforcement learning
CN116185059A
Method for controlling the weaponry of multifunctional tactical aircrafts and system for implementation thereof
RU2759057C1