Trajectory planning method based on improved sampling reinforcement learning

This patented method for trajectory planning based on improved sampling reinforcement learning, through the construction of a dynamic environment model and the definition of a learning agent and Markov decision processes, generates thinned trajectories and updates strategies using adaptive step-size thinning sampling technology. It is applied to the field of aerospace guidance and control, particularly in this area. Specifically, it involves a trajectory planning method based on improved sampling reinforcement learning. By constructing a dynamic environment model and learning agent and Markov decision processes, it enhances the penetration and attack capabilities of aircraft in complex adversarial environments. This solves the problems of low sampling efficiency, poor training effects, and insufficient adaptability of penetration strategies in existing technologies, achieving efficient strategy optimization and stable training results.

CN121302889APending Publication Date: 2026-01-09HARBIN ENG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511458476.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-13
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

In modern information warfare, when aircraft are carrying out penetration and strike missions and facing multi-layered and multi-type air defense systems, existing deep reinforcement learning algorithms have low sampling efficiency and difficulty in convergence during training in long-track, multi-stage ballistic integration scenarios, resulting in insufficient strategy optimization effects and adaptability.

Method used

An improved sampling reinforcement learning trajectory planning method is adopted. By constructing a dynamic environment model, defining the learning agent and Markov decision process, and using adaptive step size sparsification sampling technology, sparsified trajectories are generated and policies are updated, thereby improving training stability and policy generalization ability.

Benefits of technology

It effectively balances the sampling accuracy and efficiency of deep reinforcement learning training in aerospace guidance and control, improves strategy generalization ability and training stability, and is suitable for engineering scenarios with high requirements for real-time performance and reliability, such as missile penetration and target interception.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121302889A_ABST
    Figure CN121302889A_ABST
Patent Text Reader

Abstract

The invention relates to a trajectory planning method based on improved sampling reinforcement learning. The method comprises the following steps: constructing a dynamic environment model; based on the dynamic environment model, defining a learning agent and a Markov decision process; performing algorithm parameter initialization based on the defined learning agent and a Markov decision process; based on the initialized algorithm, generating an original bulletproof trajectory through interaction between the intelligent agent and the environment; carrying out thinning processing on the original trajectory to generate a thinned trajectory; the thinning track is stored in an experience playback buffer pool, and strategy updating is carried out by using the thinning track; and based on the updated strategy, executing a defense penetration strike task. According to the method, the sampling precision and efficiency of deep reinforcement learning training in aerospace guidance control are effectively balanced, the strategy generalization ability and the training stability are improved, and a feasible scheme is provided for intelligent game strategy optimization in a complex dynamic scene; the method is especially suitable for engineering scenes with high requirements on real-time performance and reliability, such as missile penetration and target interception.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of aerospace guidance and control technology, and in particular to a trajectory planning method based on improved sampling reinforcement learning. Background Technology

[0002] In modern information warfare, aircraft performing penetration strike missions face the formidable challenge of multi-layered and multi-type enemy air defense systems. These systems typically integrate radar detection, infrared detection, and various types of anti-aircraft missiles, forming a comprehensive defense network from long-range early warning to short-range interception. Traditional penetration strategies often rely on pre-planned fixed flight paths, but this approach has significant limitations in the highly dynamic battlefield environment. When enemy air defense systems adjust their deployments based on real-time situations or new threats emerge, fixed-track aircraft are easily detected and intercepted, leading to mission failure.

[0003] In recent years, deep reinforcement learning technology has demonstrated significant application potential in aerospace guidance and control due to its autonomous learning and strategy optimization capabilities in complex decision-making problems. However, the direct application of this type of algorithm in penetration and attack missions still faces several challenges. One of these challenges is the low sampling efficiency and difficulty in convergence during training when facing mission scenarios such as long trajectories and multi-stage ballistic integration, which restricts the effectiveness and adaptability of strategy optimization. To address these issues, this invention provides a penetration and attack trajectory planning method based on improved sampling reinforcement learning. By improving the sampling strategy and optimizing the reinforcement learning algorithm, it enhances the penetration and attack capabilities of aircraft in complex adversarial environments, solving the problems of low sampling efficiency, poor training results, and insufficient adaptability of penetration strategies in existing technologies. Summary of the Invention

[0004] The purpose of this invention is to provide a trajectory planning method based on improved sampling reinforcement learning to solve the technical problems of low sampling efficiency and data redundancy in deep reinforcement learning training during penetration attacks, thereby improving training stability and policy generalization ability.

[0005] To achieve the above objectives, the present invention provides the following solution:

[0006] A trajectory planning method based on improved sampling reinforcement learning includes:

[0007] Based on the combat scenario of penetrating missiles striking stationary ground targets, a dynamic environment model of penetrating missiles, interceptor missiles, and targets is constructed.

[0008] Based on the dynamic environment model, a learning agent and a Markov decision process are defined.

[0009] Based on the defined learning agent and Markov decision process, the algorithm parameters are initialized.

[0010] Based on the initialized algorithm, the original trajectory of the penetrating bullet is generated through the interaction between the agent and the environment;

[0011] The original trajectory is thinned out to generate a thinned trajectory;

[0012] The thinning trajectory is stored in the experience replay buffer pool, and the policy is updated using the thinning trajectory;

[0013] Based on the updated strategy, execute penetration strike missions.

[0014] Optionally, constructing dynamic environment models for penetrating missiles, interceptor missiles, and targets includes:

[0015] The combat environment is extracted into a geometric representation in an inertial coordinate system, a combat geometry model is constructed, and the global state of the combat environment is determined.

[0016] Using speed magnitude and velocity angle To describe the kinematic relationships of each aircraft and construct a dynamic model;

[0017] Introducing a first-order inertial element to accelerate the command of the penetrating missile. As input, the actual normal acceleration is output to construct the control and response model.

[0018] Optionally, the engagement geometry model includes three participants: penetration missiles... The target that the penetration missile is intended to attack. and interceptor missiles to protect targets. Each spacecraft is modeled as a point mass, and its instantaneous state is described by a set of core variables, including its position coordinates in a two-dimensional plane. Velocity tilt angle Flight speed and actual acceleration ;

[0019] The global state of the combat environment is the set of state vectors of all participating entities;

[0020] The dynamic model is as follows:

[0021] ;

[0022] in, and These are the rates of change of position in the horizontal and vertical directions, respectively. It is the rate of change of velocity, determined by the thrust of the aircraft. ,resistance and quality Joint decision, It is the rate of change of the velocity tilt angle, which is the acceleration in the direction normal to the aircraft's velocity vector. and current speed Decide;

[0023] The first-order inertial element is:

[0024] ;

[0025] in, It is the response time constant of the aircraft control system. It is the actual normal acceleration.

[0026] Optionally, the definition of the learning agent and the Markov decision process includes:

[0027] The proximal policy optimization algorithm is adopted as the core policy update framework for the agent.

[0028] The observation space is designed based on the relative kinematics between the penetration missile, the target, and the interceptor missile;

[0029] Aircraft-based command normal acceleration Design the action space;

[0030] A policy network and a value network are designed using a multilayer perceptron structure.

[0031] Design a composite reward function.

[0032] Optionally, the reward function comprises four parts;

[0033] The first part is based on the relative distance between the penetration missile and the target. An exponentially dense reward is used to encourage the agent to continuously approach the target throughout the flight;

[0034] The second part is a tiered terminal reward based on the distance between the penetrating missile and the target terminal. When the final miss distance is less than the preset kill radius threshold, a preset positive reward is given to enhance the hit behavior.

[0035] The third part is based on the relative distance between the penetrating missile and the interceptor missile. Exponential negative rewards are used to guide agents to actively avoid the threat of interceptor missiles;

[0036] The fourth part is a tiered terminal penalty based on the relative distance between the penetrating missile and the interceptor missile. When the final miss distance is less than the preset kill radius threshold, a preset negative reward is given to enhance the maneuvering penetration behavior.

[0037] Optionally, initializing the algorithm parameters includes:

[0038] Initialize the weights and bias parameters of the policy network and value network;

[0039] Initialize a program with a preset capacity Experience replay buffer pool ;

[0040] Select and initialize an adaptive step-size numerical integrator. ;

[0041] Define and initialize a sparsification processor .

[0042] Optionally, based on the initialized algorithm, the original trajectory of the penetration projectile is generated through the interaction between the agent and the environment, including:

[0043] At the beginning of each training round, the initial conditions of the dynamic environment model are randomized within a predefined distribution range; at the same time, a temporary list structure is created to store all state transition samples generated in this round in chronological order.

[0044] At each point in the round, the agent interacts with the environment once; the agent bases its actions on the policy network and the current observations. Output an action command Action commands It is passed to the dynamic environment model as input for the next state evolution;

[0045] After receiving the action command, the environment calls the pre-initialized adaptive step integrator. To calculate the next state; Adaptive step-size numerical integrator Based on its internal error control mechanism, it automatically performs variable-step integration and evolves the state at the next time step. The state transition samples generated in this process That is, the strategy data is recorded in a temporary list structure, including the current state. Sampling actions in the current policy network Rewards from environmental feedback after an action is taken and the state at the next moment This cycle continues until the preset termination condition for the round is triggered, and the resulting temporary list structure is a complete original trajectory.

[0046] Optionally, the original trajectory is thinned to generate a thinned trajectory, including:

[0047] Using the initialized sparsification processor The original trajectory is processed to generate a thinned trajectory; the processing includes: based on a preset sample retention rate. Random sampling is performed, and a preset mechanism is used to ensure that the starting and ending points of the trajectory are unconditionally preserved.

[0048] Optionally, storing the thinning trajectory in an experience replay buffer and using the thinning trajectory for policy updates includes:

[0049] Store the sample points in the thinning trajectory into the experience replay buffer pool;

[0050] Batch policy data is extracted from the experience replay buffer, and the advantage function is calculated for the samples at each time step. This is used for subsequent strategy updates;

[0051] Based on the dominant function value The PPO strategy update refers to performing multiple rounds of stochastic gradient ascent on the same batch of sample data stored in the buffer pool to optimize an objective function.

[0052] After completing multiple rounds of optimization on the current batch of data, discard that batch of data and update the optimized policy network using the PPO policy. (Network parameters are) The original trajectory of the penetration bullet is regenerated through the interaction between the agent and the environment, and a new round of sample generation, thinning and training begins.

[0053] The closed-loop process of sample generation, sparsification, and training is iterated continuously until the performance indicators of the agent converge or the preset total number of training steps is reached, thus obtaining the final guidance strategy.

[0054] The beneficial effects of this invention are as follows:

[0055] This invention effectively balances the sampling accuracy and efficiency of deep reinforcement learning training in aerospace guidance and control through adaptive step-size sparsification sampling technology, improves the strategy generalization ability and training stability, and provides a feasible solution for intelligent game strategy optimization in complex dynamic scenarios. It is especially suitable for engineering scenarios with high requirements for real-time performance and reliability, such as missile penetration and target interception. Attached Figure Description

[0056] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0057] Figure 1 This is a schematic diagram of a trajectory planning method based on improved sampling reinforcement learning according to an embodiment of the present invention;

[0058] Figure 2 This is a schematic diagram of an air combat scenario model according to an embodiment of the present invention;

[0059] Figure 3 This is a schematic diagram comparing the training and learning curves of an agent under adaptive step-size sparsification sampling according to an embodiment of the present invention;

[0060] Figure 4 This is a schematic diagram of the off-target amount as a function of training in an embodiment of the present invention using sparse sampling;

[0061] Figure 5 This is a schematic diagram comparing the number of downsampling rounds with the same number of sampling steps in an embodiment of the present invention;

[0062] Figure 6 This is a schematic diagram of the penetration attack scenario verification after adaptive step size thinning sampling in an embodiment of the present invention; wherein, (a) is the engagement trajectory, (b) is the acceleration of the penetration missile and the interceptor missile, (c) is the zero-control miss distance of the two engagement processes, and (d) is the Monte Carlo simulated ballistic trajectory. Detailed Implementation

[0063] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0064] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0065] like Figure 1 As shown, this embodiment proposes a trajectory planning method based on improved sampling reinforcement learning, including:

[0066] Step 1: Based on the confrontation scenario of penetrating missiles striking stationary ground targets, construct a dynamic environment model of penetrating missiles, interceptor missiles, and targets;

[0067] Step 2: Based on the dynamic environment model, define the learning agent and the Markov decision process;

[0068] Step 3: Initialize the algorithm parameters based on the defined learning agent and Markov decision process;

[0069] Step 4: Based on the initialized algorithm, the original trajectory of the penetration bullet is generated through the interaction between the agent and the environment;

[0070] Step 5: Perform thinning processing on the original trajectory to generate a thinned trajectory;

[0071] Step 6: Store the thinning trajectory in the experience replay buffer pool, and use the thinning trajectory to update the strategy;

[0072] Step 7: Based on the updated strategy, execute the penetration strike mission.

[0073] Furthermore, constructing a dynamic environment model for the penetrating missile, interceptor missile, and target includes:

[0074] The combat environment is extracted into a geometric representation in an inertial coordinate system, a combat geometry model is constructed, and the global state of the combat environment is determined.

[0075] Using speed magnitude and velocity angle To describe the kinematic relationships of each aircraft and construct a dynamic model;

[0076] Introducing a first-order inertial element to accelerate the command of the penetrating missile. As input, the actual normal acceleration is output to construct the control and response model.

[0077] Specifically, in this embodiment, step 1, establishing a dynamic environment model, includes:

[0078] The implementation of this invention first requires the construction of a dynamic simulation environment that can accurately reflect the laws of the physical world. This environment is the foundation for reinforcement learning agents to learn and iterate policies. In this embodiment, a typical adversarial scenario of a penetrating missile striking a stationary ground target is taken as an example. This scenario includes three entities: the penetrating missile (agent), the interceptor missile, and the target.

[0079] S1. Description of the combat environment and determination of state variables;

[0080] Figure 2 The battle is distilled into an inertial coordinate system. The geometric representation in the model includes three participants: penetration missiles. The target that the penetration missile is intended to attack. and interceptor missiles to protect targets. Each aircraft in the system (penetration missile and interceptor missile) is modeled as a point mass, and its instantaneous state is described by a set of core variables, including its position coordinates in a two-dimensional plane. Velocity tilt angle Flight speed and actual acceleration Use subscripts respectively , and This indicates penetration missiles, targets, and interceptor missiles. Therefore, any aircraft The state vector can be represented as .

[0081] The global state of the entire combat environment is the set of state vectors of all participating entities. The combat process is divided into two interrelated but independent parts: the attack of the target by the penetration missile (…). ) and interceptor missiles intercepting penetrating missiles to protect the target ( In these kinds of battles, The distance between each aircraft Indicates line of sight. The line-of-sight angle is denoted by the subscript 0, and the initial conditions at the start of the engagement are represented by the subscript 0.

[0082] S2: Establish the dynamic equations;

[0083] The trajectory of each aircraft is determined by a set of nonlinear ordinary differential equations. The magnitude of the velocity is used as the basis for this determination. and velocity angle To describe the kinematic relationship, the specific set of dynamic equations is as follows:

[0084] ;

[0085] in, and These are the rates of change of position in the horizontal and vertical directions, respectively. It is the rate of change of velocity, determined by the thrust of the aircraft. ,resistance and quality Jointly determined. In the simplified analysis of many guidance problems, velocity can be assumed. It remained unchanged during the main combat phase. It is the rate of change of the velocity tilt angle, which is directly caused by the acceleration in the direction normal to the aircraft's velocity vector. and current speed Decision. This normal acceleration. It is the key physical quantity that intelligent agents or guidance laws need to control.

[0086] S3: Control and Response Model;

[0087] The decision command output by the intelligent agent is the normal acceleration. In a real physical system, this command cannot be instantaneously converted into the aircraft's actual acceleration because the aircraft's control system (including the autopilot and servos) has response delays and dynamic characteristics. To simulate this process, a first-order inertial element is introduced into the model. This element converts the commanded acceleration... As input, the output is the actual normal acceleration. Its mathematical relationship is:

[0088] ;

[0089] in, This is the response time constant of the aircraft control system, representing how quickly the system reaches the command value. This model reflects the dynamic delay between the issuance of a command and the actual maneuver, and is a crucial step in ensuring a successful transition from simulation to reality.

[0090] Furthermore, the definitions of learning agents and Markov decision processes include:

[0091] The proximal policy optimization algorithm is adopted as the core policy update framework for the agent.

[0092] The observation space is designed based on the relative kinematics between the penetration missile, the target, and the interceptor missile;

[0093] Aircraft-based command normal acceleration Design the action space;

[0094] A policy network and a value network are designed using a multilayer perceptron structure.

[0095] Design a composite reward function.

[0096] Specifically, in this embodiment, step 2, the definition of the learning agent and the Markov decision process includes:

[0097] After the environmental model is established, it is necessary to define the key elements of the reinforcement learning agent, which serves as the core of decision-making, and its learning process (i.e., Markov decision process).

[0098] S1: The policy update algorithm framework is established;

[0099] This invention employs the Proximal Policy Optimization (PPO) algorithm as the core policy update framework for the agent. PPO is an advanced policy gradient algorithm, belonging to the on-policy method. Its core advantage lies in introducing a "clipping" objective function, which limits the change in policy between the old and new policies to a small "trust region" during each update. This effectively avoids the policy collapse problem caused by excessively large update steps in traditional policy gradient algorithms, ensuring the stability of the training process and monotonic performance improvement. This characteristic is crucial for complex aerospace missions that require stable learning and avoid catastrophic forgetting.

[0100] S2: Observation space design;

[0101] Because the complete state space of the environment has too high a dimension and contains information that the agent cannot directly obtain, a low-dimensional but information-rich observation space was designed. This space consists of the relative kinematic quantities between the agent (penetrating missile) and other key entities (target and interceptor missile).

[0102] An effective observation vector Includes two sets of relative distances: the penetration missile relative to the target (MT) and the interceptor missile (MD). Approach speed and line-of-sight angular rate Specifically:

[0103] ;

[0104] Among them, relative distance Approach speed Line-of-sight angular rate The observations simultaneously considered two engagement processes, with distance, approach speed, and line-of-sight angular rate treated as policy inputs. These quantities directly reflect the situation of the interception or evasion mission, providing sufficient basis for the agent to make high-quality decisions.

[0105] S3: Motionspace Design

[0106] The agent's action space is designed to output a scalar value corresponding to the aircraft's command normal acceleration. This means that by controlling the acceleration perpendicular to the velocity vector, the curvature of the flight trajectory is changed, thereby adjusting the flight direction. The range of values ​​for this command acceleration is limited by the aircraft's maximum maneuverability.

[0107] S4: Network Structure Design

[0108] The agent's decision-making and evaluation functions are carried out by two separate but structurally similar deep neural networks: the policy network (Actor) and the value network (Critic). In this embodiment, both networks employ a multilayer perceptron (MLP) structure. The network structure is configured such that the input layer receives a 6-dimensional observation vector, followed by three hidden layers containing 128 neurons each, using the ReLU nonlinear activation function. The output layer of the policy network determines its dimension based on the definition of the action space (i.e., the mean and variance of a Gaussian distribution), while the output layer of the value network is a single node used to output the value estimate of the state.

[0109] S5: Reward Function Design:

[0110] The reward function is the sole signal guiding the agent's learning, and its design directly determines the performance of the final policy. This invention employs a composite reward function, aiming to balance process guidance and the final goal. To guide the agent's learning, a composite reward function is used:

[0111] ;

[0112] ;

[0113] ;

[0114] in, For dense reward functions used in process guidance; A sparse ladder reward function for terminal judgment. , and Hyperparameters that shape the reward function , and This represents different levels of strike accuracy, meaning that higher strike accuracy... The greater the reward or penalty value, the better. This function effectively guides the agent to learn complex game strategies by rewarding the target of the attack and penalizing the risk of being intercepted.

[0115] The reward function mainly consists of four parts: the first part is based on the relative distance between the penetrating missile and the target. The first part consists of an exponentially dense reward system to encourage the agent to continuously approach the target throughout the flight; the second part is a tiered terminal reward based on the distance between the penetrating missile and the target terminal, where a larger positive reward is given when the final miss distance is less than a preset kill radius threshold to reinforce the hit behavior; the third part is based on the relative distance between the penetrating missile and the interceptor missile. The first part is an exponential negative reward (or penalty) used to guide the agent to actively avoid the threat of interceptor missiles; the second part is a tiered terminal penalty based on the relative distance between the penetration missile and the interceptor missile. When the final miss distance is less than the preset kill radius threshold, a larger negative reward is given to enhance the maneuver penetration behavior.

[0116] Furthermore, the algorithm parameter initialization includes:

[0117] Initialize the weights and bias parameters of the policy network and value network;

[0118] Initialize a program with a preset capacity Experience replay buffer pool ;

[0119] Select and initialize an adaptive step-size numerical integrator. ;

[0120] Define and initialize a sparsification processor .

[0121] Specifically, in this embodiment, step 2, algorithm parameter initialization, includes:

[0122] Before training begins, each component of the algorithm needs to be properly initialized to ensure that the training process can start smoothly and proceed efficiently.

[0123] S1: Neural network parameter initialization:

[0124] Initialize the weights and biases of the policy network and value network. Using standard initialization methods such as Xavier or He can ensure that the network has good gradient propagation properties in the early stages of training, avoiding gradient vanishing or exploding problems, thereby accelerating network convergence.

[0125] S2: Initialize the experience replay buffer pool:

[0126] Initialize a program with a preset capacity Experience replay buffer pool At the start of training, the buffer pool is empty. This invention aims to fill this buffer pool with higher efficiency through subsequent thinning sampling, allowing it to contain more diverse trajectory information within the same capacity.

[0127] S3: Adaptive step-size integrator initialization:

[0128] Select and initialize an adaptive step-size numerical integrator. This embodiment uses the Runge-Kutta-Fairberg (RKF 45) method. The key to initialization is setting a reasonable absolute or relative error tolerance. This parameter determines the balance between the accuracy of the integration and computational efficiency: a smaller value... This will result in higher accuracy and denser sampling points, while larger... Conversely.

[0129] S4: Spallation processor initialization:

[0130] Define and initialize a sparsification processor Its core parameter is the sample retention rate. This value, between 0 and 1, represents the probability that each sample point in the original trajectory will be retained during the sparsity process. The choice of this parameter directly affects the sparsity of the samples and needs to be weighed based on the complexity of the specific task and the information requirements.

[0131] Furthermore, based on the initialized algorithm, the initial trajectory of the penetration projectile is generated through the interaction between the agent and the environment, including:

[0132] At the beginning of each training round, the initial conditions of the dynamic environment model are randomized within a predefined distribution range; at the same time, a temporary list structure is created to store all state transition samples generated in this round in chronological order.

[0133] At each point in the round, the agent interacts with the environment once; the agent bases its actions on the policy network and the current observations. Output an action command Action commands It is passed to the dynamic environment model as input for the next state evolution;

[0134] After receiving the action command, the environment calls the pre-initialized adaptive step integrator. To calculate the next state; Adaptive step-size numerical integrator Based on its internal error control mechanism, it automatically performs variable-step integration and evolves the state at the next time step. The state transition samples generated in this process The strategy data is recorded in a temporary list structure; this loop continues until the preset termination condition of the round is triggered, and the final temporary list structure is a complete original trajectory.

[0135] Specifically, in this implementation, step 4, the trajectory generation stage based on adaptive step size, includes:

[0136] This stage is one of the core components of this embodiment, generating raw, high-fidelity training data through the interaction between the agent and the environment.

[0137] S1: Round Start and Environment Reset:

[0138] At the start of each training round, the simulation environment is first reset. To enhance the generalization ability of the strategy, the initial conditions of the environment (such as the initial positions, velocities, and tilt angles of the penetrating and intercepting missiles) are randomized within a predefined distribution range. Simultaneously, a temporary list structure is created. It is used to store all state transition samples generated in the current round in chronological order.

[0139] S2: The interaction loop between the agent and the environment:

[0140] At each point in time during a round, the agent interacts with the environment once. The agent then determines the interaction based on its policy network. and current observations Output an action command This instruction is passed to the environment model as input for the next state evolution.

[0141] S3: Adaptive Integral and State Evolution:

[0142] After receiving the action command, the environment calls the pre-initialized adaptive step integrator. The integrator then calculates the next state. Based on its internal error control mechanism, it automatically performs variable-step integration. It calculates an optimal time step that satisfies the preset accuracy requirements while maximizing computational efficiency. And evolve into the state of the next moment. The state transition samples generated in this process Completely recorded The loop continues until the end condition of the round is triggered (including hitting the target, being intercepted, or reaching the maximum simulation time).

[0143] Further, the original trajectory is thinned to generate a thinned trajectory, including:

[0144] Using the initialized sparsification processor The original trajectory is processed to generate a thinned trajectory; the processing includes: based on a preset sample retention rate. Random sampling is performed, and a preset mechanism is used to ensure that the starting and ending points of the trajectory are unconditionally preserved.

[0145] Specifically, in this embodiment, step 5, the trajectory sample thinning process, includes:

[0146] This stage is performed after each round is completed. It aims to reduce the dimensionality and remove redundancy from the original data, and is a key step in improving sample quality.

[0147] S1: Obtain the original trajectory:

[0148] When a round ends, the result obtained from step four This is a complete original trajectory. The characteristic of this trajectory is that the data point density is uneven. The data points are dense during the phase of drastic changes in the aircraft's state (such as when performing high-G maneuvers), while they are relatively sparse during the stable flight phase.

[0149] S2: Perform thinning treatment:

[0150] Call the pre-initialized sparsification processor right The process involves processing. The core of this processing is based on a preset sample retention rate. Random sampling is performed. Specifically, the system iterates through each sample point in the original trajectory and uses a random number generator to determine whether the sample point should be retained. To ensure that critical temporal information of the trajectory is not lost, a mechanism is usually designed to ensure that the start and end points of the trajectory are unconditionally retained.

[0151] S3: Generate a thinned trajectory:

[0152] After the aforementioned data forgetting process, a new, significantly shorter, thinning trajectory emerges. This new trajectory, while retaining the key dynamic features of the original trajectory, significantly reduces the number of data points and lowers the correlation between samples, laying the foundation for efficient storage and training in the future.

[0153] Furthermore, storing the thinned trajectory in the experience replay buffer and using the thinned trajectory for policy updates includes:

[0154] A general dominance estimation method is used to calculate the dominance function value for the samples at each time step in the buffer pool. ;

[0155] Calculate the advantage function by extracting batch strategy data from the experience replay buffer pool;

[0156] Based on the dominant function value The PPO strategy update is performed; the PPO strategy update refers to performing multiple rounds of stochastic gradient ascent on the same batch of data to optimize an objective function.

[0157] After completing multiple rounds of optimization on the current batch of data, discard that batch of data and use the optimized policy network. The original trajectory of the penetration bullet is regenerated through the interaction between the intelligent agent and the environment, and a new round of sample generation, thinning and training begins.

[0158] The closed-loop process of sample generation, sparsification, and training is iterated continuously until the performance indicators of the agent converge or the preset total number of training steps is reached, thus obtaining the final guidance strategy.

[0159] Specifically, in this embodiment, step 6, the storage and training phase, includes:

[0160] This stage is the final step in the agent's learning and policy iteration using the processed high-quality samples.

[0161] S1: Sample storage:

[0162] The thinning trajectory generated in step 5 All sample points are stored in the global experience replay buffer. In online policy algorithms like PPO, this data is typically stored in a temporary "rollout" buffer for upcoming single policy updates. Steps four and five are then executed continuously until a sufficient amount (e.g., a batch size) has accumulated in the buffer. (sample).

[0163] S2: Calculation of the dominance function:

[0164] Before updating the policy, the dominance function value needs to be calculated for the samples at each time step in the buffer pool. This invention employs the Universal Advantage Estimation (GAE) method. By exponentially weighting the temporal difference (TD) errors over multiple future steps, GAE effectively balances the prediction bias and variance, thereby providing a more stable and accurate gradient direction signal for policy updates.

[0165] From the experience replay buffer pool Extracting batches of old strategy data Calculate the advantage function:

[0166] ;

[0167] in For the reward of the intelligent agent, As a discount factor, For state The value function below.

[0168] S3: PPO Strategy Update:

[0169] This is the core of agent learning. PPO optimizes a special objective function by performing stochastic gradient ascent over multiple epochs on the same batch of data. This objective function, the PPO-Clip objective, takes the following form:

[0170] ;

[0171] In the formula, This represents the parameters of the current policy network. Let be the dominant function at time t. Expressing expected experience, For the new and old strategies in action The probability ratio is defined by the first term, which is the core objective of the post-clipping policy gradient. The clip function constrains the probability ratio to a certain value. Within the range, to prevent excessively large single update increments, This is the trimming parameter (usually set to 0.1-0.3). The operation ensures the conservatism of the updates. The second term is the mean squared error loss of the value function, used to train the Critic network. Its weight hyperparameter, For parameters Value network in state The estimated value, The first term is the target value of the value network, typically the actual reward or the target calculated using the TD method. The third term is the policy entropy reward, used to encourage exploration and prevent the policy from prematurely converging to a suboptimal solution. This is a weight hyperparameter for the entropy term, used to adjust the degree of exploration.

[0172] S4: Iteration and Convergence

[0173] After performing multiple rounds of optimization on the current batch of data, the batch is discarded (a characteristic of the On-Policy algorithm). Then, the newly updated, higher-performing policy network is used. Returning to step four, a new round of sample generation, spalling, and training begins. This closed-loop process of "sampling-sparsening-training" iterates continuously until the agent's performance metrics (such as average round reward) converge or reach the preset total number of training steps, ultimately resulting in a robust and efficient guidance policy.

[0174] The specific training environment settings for the intelligent agent are shown in Table 1 below:

[0175] Table 1 Specific Training Environment Settings for the Agent

[0176] Horizontal position (km) Vertical height (km) Flight speed (m / s) Maximum overload ( ) Response time constant (s) interceptor missile 0 1000 30 0.3 Penetration missile 0 1800 20 0.3

[0177] Learning curves of agents under different training conditions, for example Figure 3As shown, it was observed that, under the same reward function settings, the agent with sparse sampling scored significantly higher than the agent without sparse sampling. The agent with sparse sampling achieved a cumulative reward of around 200 after approximately 5e5 steps, then steadily increased, gradually rising to around 500 after 2e6 steps, until finally obtaining a stable convergent policy. In contrast, the agent without sparse sampling faced significantly greater difficulties during training, achieving a maximum cumulative reward of only around 200, eventually hitting a training bottleneck and failing to complete the penetration attack mission.

[0178] When using adaptive step-size sparsification sampling technology, the miss distance of the interceptor missile when intercepting a penetrating missile is... Miss distance of anti-tank missiles striking ground targets Changes throughout the training process, such as Figure 4 As shown. To more clearly demonstrate the training effect, the vertical axis is set to logarithmic format. It can be seen that in the initial stage of training, the agent does not yet possess the ability to penetrate defenses, therefore... Initially, the values ​​were very small. However, as the agent's strategy was updated, the miss distance of the interceptor missile increased rapidly, while the miss distance of the penetration missile did not change significantly. This meant that the penetration missile was only maneuvering to avoid interception, but could not simultaneously engage ground targets. Subsequently, as training progressed, the miss distance of the interceptor missile stabilized at... The magnitude of the interceptor missile's flight path indicates that it can assess the threat level of the interceptor missile and avoid interception by flying a relatively fixed distance. This avoids adding additional difficulties to the engagement process of striking ground targets due to excessive maneuverability. At the same time, the miss distance of the interceptor missile when striking ground targets can be observed. As training progresses, the value is continuously reduced until it becomes accurate to... That is, within the meter range. This means that the intelligent agent has successfully acquired a strategy for stable penetration and precise target engagement.

[0179] Figure 5 This section compares the number of sampling rounds with the same number of sampling steps. It can be observed that the training process using adaptive step-size sparsification sampling can undergo more rounds with the same number of steps, thus effectively improving sample diversity.

[0180] Specifically, in this embodiment, step 7, simulating and verifying the strategy obtained through reinforcement learning training under specified battlefield conditions, includes:

[0181] The simulation conditions are listed in Table 2 below:

[0182] Table 2 Simulation Condition Settings

[0183] Horizontal position (km) Vertical height (km) Flight speed (m / s) Initial ballistic inclination angle (deg) interceptor missile 25 0 1000 40 Penetration missile 0 20 1800 -19

[0184] Simulation results are as follows Figure 6 As shown. Figure 6 (a)-(c) represent the results of a single simulation under specified conditions. Figure 6 (d) represents the Monte Carlo simulation results under continuously changing initial positions of the penetration missile.

[0185] Figure 6 (a) shows the engagement trajectory in a single simulation. It can be seen that the interceptor missile missed the target by 237.13 meters at 11.79 seconds, while the penetration missile successfully hit the target at 18.00 seconds with a miss of 0.10 meters. The accelerations of the two missiles in this scenario are as follows: Figure 6 As shown in (b), the penetrating missile increased its maneuverability as it approached the interceptor, using maximum maneuverability to evade the interceptor. The interceptor, however, experienced acceleration saturation due to the interference of the penetrating maneuver, ultimately resulting in a miss. It can be seen that the penetrating maneuver resembled a bang-bang pattern, which exacerbated the range of acceleration changes in the interceptor, making interception more difficult. After successfully penetrating, the penetrating missile rapidly changed the direction of its acceleration, adjusting its flight path to aim at the ground target. This process was achieved through… Figure 6 The results of zero-control miss distance shown in (c) also corroborate this. Zero-control miss distance of the interceptor missile. Even when it missed the target, it was still about 200 meters away, while the zero-control miss distance of the anti-tank missile was... If the missile quickly stabilizes near zero after penetrating the defenses, it means that the penetrating missile has entered the interception triangle for striking ground targets.

[0186] Figure 6 (d) represents the trajectory of the penetrating missile obtained through Monte Carlo simulation with its initial position continuously changed within the range of [15, 25] km. It was observed that the penetrating missile initially corrected the target's heading error using high overload. However, the penetrating missile actively overcorrected the heading error, thus creating adjustment space for the interceptor missile to switch maneuvering direction and miss the target when it approaches. After successful penetration, the trajectory was readjusted until it hit the target with a near-straight trajectory. This process simultaneously considered both maneuvering penetration and precision strike.

[0187] This embodiment discloses an adaptive step-size thinning sampling method and system. By dynamically adjusting the integration step size through an adaptive step-size integrator and combining it with a data forgetting mechanism to delete redundant samples, it solves the problems of low sampling efficiency and data redundancy in deep reinforcement learning training in the aerospace field. Experiments show that this method can significantly improve the training efficiency of intelligent agents and achieve high-precision strategy optimization in scenarios such as missile penetration, demonstrating good engineering application value.

[0188] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A trajectory planning method based on improved sampling reinforcement learning, characterized in that, include: Based on the combat scenario of penetrating missiles striking stationary ground targets, a dynamic environment model of penetrating missiles, interceptor missiles, and targets is constructed. Based on the dynamic environment model, a learning agent and a Markov decision process are defined. Based on the defined learning agent and Markov decision process, the algorithm parameters are initialized. Based on the initialized algorithm, the original trajectory of the penetrating bullet is generated through the interaction between the agent and the environment; The original trajectory is thinned out to generate a thinned trajectory; The thinning trajectory is stored in the experience replay buffer pool, and the policy is updated using the thinning trajectory; Based on the updated strategy, execute penetration strike missions.

2. The trajectory planning method based on improved sampling reinforcement learning according to claim 1, characterized in that, Constructing dynamic environment models for penetrating missiles, interceptor missiles, and targets includes: The combat environment is extracted into a geometric representation in an inertial coordinate system, a combat geometry model is constructed, and the global state of the combat environment is determined. Using speed magnitude and velocity angle To describe the kinematic relationships of each aircraft and construct a dynamic model; Introducing a first-order inertial element to accelerate the command of the penetrating missile. As input, the actual normal acceleration is output to construct the control and response model.

3. The trajectory planning method based on improved sampling reinforcement learning according to claim 2, characterized in that, The engagement geometry model includes three participants: penetration missiles. The target that the penetration missile is intended to attack. and interceptor missiles to protect targets. Each spacecraft is modeled as a point mass, and its instantaneous state is described by a set of core variables, including its position coordinates in a two-dimensional plane. Velocity and tilt angle Flight speed and actual acceleration ; The global state of the combat environment is the set of state vectors of all participating entities; The dynamic model is as follows: ; in, and These are the rates of change of position in the horizontal and vertical directions, respectively. It is the rate of change of velocity, determined by the thrust of the aircraft. ,resistance and quality Joint decision, It is the rate of change of the velocity tilt angle, which is the acceleration in the direction normal to the aircraft's velocity vector. and current speed Decide; The first-order inertial element is: ; in, It is the response time constant of the aircraft control system. It is the actual normal acceleration.

4. The trajectory planning method based on improved sampling reinforcement learning according to claim 1, characterized in that, The definitions of learning agents and Markov decision processes include: The proximal policy optimization algorithm is adopted as the core policy update framework for the agent. The observation space is designed based on the relative kinematics between the penetration missile, the target, and the interceptor missile; Aircraft-based command normal acceleration Design the action space; A policy network and a value network are designed using a multilayer perceptron structure. Design a composite reward function.

5. The trajectory planning method based on improved sampling reinforcement learning according to claim 4, characterized in that, The reward function consists of four parts; The first part is based on the relative distance between the penetration missile and the target. An exponentially dense reward is used to encourage the agent to continuously approach the target throughout the flight; The second part is a tiered terminal reward based on the distance between the penetrating missile and the target terminal. When the final miss distance is less than the preset kill radius threshold, a preset positive reward is given to enhance the hit behavior. The third part is based on the relative distance between the penetrating missile and the interceptor missile. Exponential negative rewards are used to guide agents to actively avoid the threat of interceptor missiles; The fourth part is a tiered terminal penalty based on the relative distance between the penetrating missile and the interceptor missile. When the final miss distance is less than the preset kill radius threshold, a preset negative reward is given to enhance the maneuvering penetration behavior.

6. The trajectory planning method based on improved sampling reinforcement learning according to claim 4, characterized in that, Algorithm parameter initialization includes: Initialize the weights and bias parameters of the policy network and value network; Initialize a program with a preset capacity Experience replay buffer pool ; Select and initialize an adaptive step-size numerical integrator. ; Define and initialize a sparsification processor .

7. The trajectory planning method based on improved sampling reinforcement learning according to claim 6, characterized in that, Based on the initialized algorithm, the original trajectory of the penetration projectile is generated through the interaction between the agent and the environment, including: At the beginning of each training round, the initial conditions of the dynamic environment model are randomized within a predefined distribution range; at the same time, a temporary list structure is created to store all state transition samples generated in this round in chronological order. At each point in the round, the agent interacts with the environment once; the agent bases its actions on the policy network and the current observations. Output an action command Action commands It is passed to the dynamic environment model as input for the next state evolution; After receiving the action command, the environment calls the pre-initialized adaptive step integrator. To calculate the next state; Adaptive step-size numerical integrator Based on its internal error control mechanism, it automatically performs variable-step integration and evolves the state at the next time step. The state transition samples generated in this process That is, the strategy data is recorded in a temporary list structure, including the current state. Sampling actions in the current policy network Rewards from environmental feedback after an action is taken and the state at the next moment This cycle continues until the preset termination condition for the round is triggered, and the resulting temporary list structure is a complete original trajectory.

8. The trajectory planning method based on improved sampling reinforcement learning according to claim 6, characterized in that, The original trajectory is thinned to generate a thinned trajectory, which includes: Using the initialized sparsification processor The original trajectory is processed to generate a thinned trajectory; the processing includes: based on a preset sample retention rate. Random sampling is performed, and a preset mechanism is used to ensure that the starting and ending points of the trajectory are unconditionally preserved.

9. The trajectory planning method based on improved sampling reinforcement learning according to claim 7, characterized in that, Storing the thinned trajectory in the experience replay buffer and using the thinned trajectory for policy updates includes: Store the sample points in the thinning trajectory into the experience replay buffer pool; Batch policy data is extracted from the experience replay buffer, and the advantage function is calculated for the samples at each time step. This is used for subsequent strategy updates; Based on the dominant function value The PPO strategy update refers to performing multiple rounds of stochastic gradient ascent on the same batch of sample data stored in the buffer pool to optimize an objective function. After completing multiple rounds of optimization on the current batch of data, discard that batch of data and update the optimized policy network using the PPO policy. The original trajectory of the penetration bullet is regenerated through the interaction between the intelligent agent and the environment, and a new round of sample generation, thinning and training begins. The closed-loop process of sample generation, sparsification, and training is iterated continuously until the performance indicators of the agent converge or the preset total number of training steps is reached, thus obtaining the final guidance strategy.

Citation Information

Cited By

  • On-satellite PWM-MDP combined high-precision temperature control system and method thereof

    CN122152006A