Incremental trajectory tracking strategy reinforcement learning method under high-frequency decision

By building state space, designing reward function and asymmetry incremental strategy optimization module, combined with deep deterministic strategy gradient algorithm, the accuracy, energy consumption and field of view constraints of trajectory tracking under high-frequency decisions are solved, and efficient trajectory tracking strategy optimization is achieved.

CN120523014APending Publication Date: 2025-08-22NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510652198.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

The existing trajectory tracking methods are difficult to meet the requirements of high precision, low energy consumption and field-of-view angle constraints in high-frequency decision-making scenarios, especially when facing high maneuvering targets, the learning efficiency is low and the strategy performance is limited, making it difficult to achieve optimized trajectory tracking.

Method used

The state space and proportional guidance method based on pure angle measurement are constructed, the basic and real-time feedback reward function are designed, and the progressive incremental strategy optimization module is adopted, combined with the deep deterministic strategy gradient algorithm, and the trajectory tracking strategy is optimized through the transition from low-frequency to high-frequency decisions.

Benefits of technology

It significantly improves interception accuracy and energy efficiency, ensures stable interception within the field of view angle constraints, improves strategy training efficiency, adapts to complex dynamic scenarios, and meets high-precision and low-energy consumption requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120523014A_ABST
    Figure CN120523014A_ABST
Patent Text Reader

Abstract

The invention particularly relates to an incremental trajectory tracking strategy reinforcement learning method under a high-frequency decision, which is used for solving the problem of intercepting a maneuvering target by an interceptor under the constraint of a field angle. According to the method, an initial strategy is provided for deep reinforcement learning by constructing a state space based on pure angle measurement and an action space based on a proportional guidance method, and a basic reward function related to an optimization target and a real-time feedback process reward function are designed, so that the strategy training efficiency is improved. In order to improve the strategy performance, a high-frequency decision-making mode is adopted, a progressive incremental strategy optimization module is provided, low-frequency decision-making is transited to high-frequency decision-making, and the training efficiency and the strategy performance are both considered. Experimental results show that the method achieves better strategy performance in the aspects of trajectory tracking precision, average energy consumption and total energy consumption while ensuring high-frequency decision training efficiency, and shows good generalization ability and adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of interceptor trajectory tracking and control, and in particular to a reinforcement learning method for an incremental trajectory tracking strategy under high-frequency decision-making, which is used to solve the problem of intercepting a maneuvering target by an interceptor under field of view angle constraints. Background Art

[0002] The demand for precise tracking and interception of moving targets is widespread in modern civilian applications, such as aircraft docking in the aerospace sector and drone aerial capture missions. These applications place stringent demands on the relevant technologies in terms of accuracy, resource consumption, and observation range limitations. For example, during aircraft docking, the two aircraft must precisely approach and dock in a complex spatial environment, while also controlling energy consumption to avoid excessive adjustments to the posture and resulting in energy waste. When drones perform aerial capture missions, they must accurately track and intercept target objects within the limited field of view of their cameras, optimizing energy usage to extend their flight time.

[0003] Currently, traditional trajectory tracking methods mainly include proportional navigation (PN), sliding mode control (SM), and optimal control (OC). Although these methods have high accuracy and stability in theory, they have the following shortcomings in practical applications: poor adaptability to target maneuvers. Traditional methods generally assume that the target is stationary or moving at a constant speed, making it difficult to cope with the complex motion characteristics of highly maneuverable targets; difficulty in handling field of view angle constraints. Under field of view angle constraints, traditional methods cannot ensure that the target remains within the seeker's field of view, which can easily lead to target loss; and insufficient energy consumption optimization. Traditional methods have limitations in energy consumption optimization, making it difficult to achieve low-energy consumption designs during interception.

[0004] In recent years, trajectory tracking methods based on deep reinforcement learning (DRL) have gradually attracted attention. DRL methods optimize trajectory tracking strategies by training intelligent agents, and have higher flexibility and adaptability. However, existing trajectory tracking methods based on DRL still face the following challenges in high-frequency decision-making scenarios: low learning efficiency. High-frequency decisions lead to long decision sequences and low sample discrimination, which increases learning complexity and reduces learning efficiency. Strategy performance is limited. Directly training intelligent agents in high-frequency decision-making environments may make strategy optimization difficult and it is difficult to achieve the ideal performance ceiling. Therefore, existing technologies cannot simultaneously meet the requirements of high precision, low energy consumption, and field of view constraints when dealing with the interception of highly maneuverable targets. Developing a method that can efficiently optimize trajectory tracking strategies in high-frequency decision-making environments is an important development direction for current interceptor trajectory tracking technology.

[0005] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present invention, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention

[0006] The present invention provides an incremental trajectory tracking strategy reinforcement learning method under high-frequency decision-making, which is used to solve the problem of intercepting a maneuvering target by an interceptor under field of view angle constraints.

[0007] To achieve this goal, this method provides an initial strategy for deep reinforcement learning by constructing a state space based on pure angle measurement and an action space based on proportional guidance. It also designs a basic reward function related to the optimization objective and a process reward function with real-time feedback to improve the efficiency of strategy training. At the same time, to improve strategy performance, a high-frequency decision-making method is adopted, and a progressive incremental strategy optimization module is proposed to transition from low-frequency decision-making to high-frequency decision-making, taking into account both training efficiency and strategy performance.

[0008] The technical solution adopted in the present invention is:

[0009] First, a state space based on pure angle measurement is constructed. Based on the relative motion equation between the interceptor and the target, the line of sight angle, line of sight angle change rate, interceptor lead angle, and interceptor lead angle change rate are defined. The relative motion equation is:

[0010]

[0011] Among them, V m and V t They represent the speed of the interceptor and the target respectively, r represents the relative distance between the interceptor and the target, q represents the line of sight angle, γ m and γ tdenote the track angles of the interceptor and the target respectively, and the lead angle of the interceptor is defined as λ m =γ m -q.

[0012] These state variables are then normalized to eliminate the differences between different physical dimensions and ensure that each state variable has the same scale, which facilitates effective search in the sample space. The normalized state space can be defined as:

[0013]

[0014] in, They represent the normalized sight angle, sight angle rate, interceptor lead angle and their changing rate respectively.

[0015] Next, we design an action space based on the proportional navigation method. We design an action space module based on the proportional navigation method (PN) and introduce a bias term as a training action. This bias term is used to adjust the normal overload instruction of the interceptor to optimize the trajectory tracking strategy. The bias equation is:

[0016]

[0017] Among them, A m Represents the normal overload instruction of the interceptor, which is used to control the acceleration of the interceptor. N represents the proportional coefficient, and ε represents the biased overload instruction. Represents the line of sight angular rate. During the training process, the agent optimizes the strategy by adjusting the bias term, thereby improving the overall performance of the trajectory tracking system.

[0018] Considering the time delay caused by the missile-borne autopilot during actual flight, the interceptor's autopilot delay characteristic is modeled as a first-order link, thereby more accurately describing the interceptor's response behavior in the dynamic process. The first-order delay characteristic of its autopilot can be expressed as:

[0019]

[0020] Where τ represents the time constant, Indicates trajectory tracking instructions.

[0021] Redesign the basic reward function related to the optimization goal. When the target is successfully intercepted, a positive reward signal is provided. Its value is in a decreasing relationship with the miss distance. The smaller the miss distance, the higher the reward. It is also in a decreasing relationship with the accumulated acceleration. The smaller the accumulated acceleration, the higher the reward. When the line of sight angle exceeds the maximum allowable field of view angle constraint, a large negative reward is applied to punish the behavior of the strategy deviating from the expected constraints. The specific reward function is set as:

[0022]

[0023] Among them, R end represents the reward associated with the combinatorial optimization objective, r min and A max are weight parameters, N represents the total number of steps, when the distance r between the interceptor and the target is less than r min , it indicates that the interception is successful and the agent will receive a positive reward signal; in addition, a negative reward is set for situations where the field of view constraint is exceeded:

[0024]

[0025] Among them, R rep Indicates the negative reward for exceeding the previous angle constraint, and B is a large positive value.

[0026] Redesign the reward function to provide real-time feedback, including the reward function R for dynamically correcting the angular velocity of the line of sight q , whose expression is:

[0027]

[0028] Among them, b1 is the weight parameter, σ1 is the scaling coefficient;

[0029] Reward function R for real-time correction of the leading angle λ , whose expression is:

[0030]

[0031] Among them, b2 is the weight parameter and σ2 is the scaling coefficient;

[0032] Reward function R for optimizing energy consumption A , whose expression is:

[0033]

[0034] Among them, b3 is the weight parameter and σ3 is the scaling coefficient;

[0035] Process Reward R guiding The sum of is:

[0036] R guiding =R q +R λ +R A ,

[0037] Combined with the basic reward function, a complete reward mechanism is formed to drive the strategy learning of the intelligent agent.

[0038] Afterwards, a high-frequency decision-making method is adopted, and a progressive incremental policy optimization (PIPO) module is introduced. In the initial stage of policy learning, low-frequency decisions are used to quickly generate the initial policy, helping the agent to become familiar with the environment and reducing learning complexity. As training progresses, the decision-making frequency is gradually increased, transitioning from low-frequency decisions to high-frequency decisions, so that the policy is gradually optimized in a high-frequency decision-making environment. As the decision-making frequency gradually increases, the weight of the process reward function is adaptively adjusted. The adjustment method is as follows:

[0039] a 1(i) ·f i =a1·f1,

[0040] Among them, a 1(i) is the weight of the reward for the i-th stage process, f i is the decision frequency of the i-th stage, f1 is the decision frequency of the initial stage, and a1 is the process reward weight of the initial stage, i.e., when i=1. Through the above-mentioned progressive incremental strategy optimization, it is ensured that the intelligent agent can achieve stable strategy optimization in a high-frequency decision-making environment, while taking into account both training efficiency and strategy performance.

[0041] Finally, policy learning is implemented based on the Deep Deterministic Policy Gradient (DDPG) algorithm. First, the state space, action space, basic reward function, and process reward function are initialized, and a Markov Decision Process (MDP) model is constructed. Policy learning is then performed using the DDPG algorithm's policy network (Actor Network) and value network (Critic Network). The policy network is used to select the optimal action, and the value network is used to evaluate the expected reward of state-action pairs. In the initial stage of policy learning, low-frequency decisions are used to quickly generate the initial policy, helping the agent familiarize itself with the environment and reducing learning complexity. As training progresses, the decision frequency is gradually increased, transitioning from low-frequency to high-frequency decisions, allowing the policy to be gradually optimized in this high-frequency decision environment. In each training stage, experience replay is used to sample from the experience pool, and the parameters of the value network are updated using the Bellman equation. Simultaneously, the parameters of the policy network are updated using the policy gradient method. During policy learning, the weights of the process reward function are adaptively adjusted to prevent excessive accumulation of process rewards due to increased decision frequency. In the final stage of policy learning, the policy is optimized through high-frequency decisions to ensure stable convergence in this high-frequency decision environment, ultimately obtaining an optimized trajectory tracking strategy.

[0042] According to a second aspect of the present invention, a storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the incremental trajectory tracking strategy reinforcement learning method under high-frequency decision-making described in the first aspect is implemented.

[0043] According to a third aspect of the present invention, a computer program product is provided, on which a computer program is stored. When the computer program is executed by a processor, the incremental trajectory tracking strategy reinforcement learning method under high-frequency decision-making described in the first aspect is implemented.

[0044] According to a fourth aspect of the present invention, there is provided an electronic device, comprising:

[0045] processor; and

[0046] a memory for storing executable instructions of the processor;

[0047] Wherein, the processor is configured to implement the incremental trajectory tracking strategy reinforcement learning method under high-frequency decision-making described in the first aspect above by executing the executable instructions.

[0048] The incremental trajectory tracking strategy reinforcement learning method under high-frequency decision-making provided by the embodiments of the present invention has the following beneficial effects:

[0049] In terms of trajectory tracking accuracy, through the unique design of state space, action space and reward function, compared with traditional trajectory tracking methods such as BPN and SM algorithms, the mean and variance of the miss distance are significantly better, which effectively improves the interception accuracy of maneuvering targets and meets the high-precision requirements of modern civil and other fields; in terms of energy consumption, the basic and process reward functions have a significant effect on the optimization of energy consumption, and have outstanding advantages in average energy consumption and total energy consumption. While ensuring the interception effect, it can reduce the energy consumption of the interceptor, improve energy utilization efficiency, and extend the operating range; in the field of view angle constraint problem, the constraints are cleverly transformed and a penalty mechanism is set. In the experiment, the interceptor lead angle is always controlled The PIPO module overcomes the difficulties of high-frequency decision-making learning, making the learning process stable and the strategy optimization effective, thus avoiding training failures. In addition, the trajectory tracking strategy has strong generalization ability and can accurately hit the target with stable performance in multiple rounds of Monte Carlo tests and different target maneuvers, adapting to complex dynamic scenarios and diverse mission requirements.

[0050] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The accompanying drawings are incorporated into and constitute a part of this specification, illustrate embodiments consistent with the present invention, and together with the description, serve to explain the principles of the present invention. Obviously, the drawings described below are only some embodiments of the present invention, and it is clear that those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0052] Figure 1 is a two-dimensional interceptor and target geometric relationship diagram;

[0053] Figure 2 Learn a flowchart for the complete trajectory tracking strategy based on the IGPL algorithm;

[0054] Figure 3 It is a single-stage learning flow chart;

[0055] Figure 4 This is the IGPL-DDPG training curve;

[0056] Figure 5 This is the Monte Carlo test result diagram;

[0057] Figure 6 This is the result diagram of interceptor target interception test (target sinusoidal maneuver);

[0058] Figure 7 This is the result of the interceptor target interception test (target square wave maneuver).

[0059] Figure 8 This is a test diagram of the effect of the process reward module;

[0060] Figure 9 Flow chart of the method of the present invention. DETAILED DESCRIPTION

[0061] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0062] In addition, the accompanying drawings are merely schematic illustrations of the present invention and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the blocks shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0063] like Figure 1 The two-dimensional geometric relationship diagram between the interceptor and the target is shown. Based on this, the state space is constructed. According to the relative motion equation of the interceptor and the target, the state variables such as the relative position, relative velocity, sight angle and its rate of change of the interceptor and the target are determined. The relative motion equation of the interceptor and the target can be defined as:

[0064]

[0065] Among them, V m and V t They represent the speed of the interceptor and the target respectively, r represents the relative distance between the interceptor and the target, q represents the line of sight angle, γ m and γ t denote the track angles of the interceptor and target, respectively.

[0066] The interceptor's track angle to the target is further defined by the following formula:

[0067]

[0068] Among them, A m represents the normal acceleration of the interceptor, A t represents the normal acceleration of the target, V t and V m are the speeds of the target and interceptor respectively.

[0069] According to the above formula, the lead angle of the interceptor can be defined as the angle between the interceptor's flight path and the line of sight, that is:

[0070] λ m =γ m -q,

[0071] In the embodiment of the present application, under normal circumstances, assuming that the flight attack angle is small, the lead angle λ m It can represent the target sight angle measured by the seeker in real time. Therefore, the field of view constraint of the seeker can be converted into a constraint on the interceptor's lead angle.

[0072] Furthermore, these state variables are normalized to have the same scale, which facilitates the subsequent effective search in the sample space. The normalized state space can be defined as:

[0073]

[0074] in, They represent the normalized sight angle, sight angle rate, interceptor lead angle and their changing rate respectively.

[0075] Furthermore, the motion space based on the proportional navigation method (PN) is designed, such as Figure 2 As shown, the action space module is designed based on PN, and the bias term is introduced as the training action. In the embodiment of the present application, the bias equation is set as:

[0076]

[0077] Among them, N represents the proportional coefficient, ε represents the bias overload instruction, Represents the line of sight angular rate. During the training process, the agent optimizes the strategy by continuously adjusting the bias term, thereby improving the overall performance of the trajectory tracking system.

[0078] Furthermore, considering the time delay characteristics of the interceptor autopilot during actual flight, it is modeled as a first-order link:

[0079]

[0080] Where τ is the time constant, The trajectory tracking instructions are used to more accurately describe the response behavior of the interceptor in the dynamic process.

[0081] Redesign the reward function. In terms of the basic reward function, when the target is successfully intercepted, according to the formula:

[0082]

[0083] Among them, r min and A max are all weight parameters, and N represents the total number of steps.

[0084] In actual scenarios, when the distance r between the interceptor and the target is less than r min When , the agent receives positive rewards, and the smaller the miss distance, the higher the reward, and the smaller the cumulative acceleration, the higher the reward; when the sight angle exceeds the maximum allowable field of view angle constraint, according to the formula:

[0085]

[0086] Apply a large negative reward, where B is a large positive value, to punish the policy for deviating from the expected constraints.

[0087] For a process reward function such as Figure 3 As shown, including the reward function for dynamically correcting the angular velocity of the line of sight In each time step, the agent dynamically corrects the line of sight angular velocity according to this function, where σ1 is the scaling coefficient and b1 is the weight parameter.

[0088] Reward function for real-time correction of lead angle The agent is made to maintain the smallest possible leading angle during the learning process to meet the field of view angle constraint, where σ2 is the scaling factor and b2 is the weight parameter.

[0089] and a reward function for optimizing energy consumption Guide the agent to optimize the strategy in each time step to pursue the lowest energy consumption, where σ3 is the scaling coefficient and b3 is the weight parameter.

[0090] The sum of the process rewards is:

[0091] R guiding =R q +R λ +R A ,

[0092] Overall reward R total Reward R by process guiding Together with the basic reward, it is composed of:

[0093] R total =α1·R guiding +α2·(R end +Rrep),

[0094] Furthermore, a high-frequency decision-making method is adopted and a progressive incremental policy optimization (PIPO) module is introduced. In the initial stage of policy learning, such as Figure 2 and Figure 3 As shown in the figure, low-frequency decision-making is adopted. At this stage, the state, action and reward samples in the experience pool are all short sequence samples. The agent quickly becomes familiar with the environment through these samples, reducing learning complexity and quickly generating initial strategies.

[0095] In some embodiments, the initial decision frequency is set low, and the intelligent agent explores and learns in a relatively simple environment. As training progresses, the decision frequency is gradually increased, transitioning from low-frequency decision-making to high-frequency decision-making.

[0096] In this process, Figure 2As shown in the figure, the number of samples gradually increases, and the agent continuously optimizes based on the existing strategy. At the same time, in order to ensure the effectiveness of strategy learning, the weight of the process reward is adaptively adjusted according to the formula:

[0097] a 1(i) ·f i =a1·f1,

[0098] Among them, a 1(i) is the weight of the reward for the i-th stage process, f i is the decision frequency in stage i, f1 is the decision frequency in the initial stage, and a1 is the process reward weight in the initial stage, i.e., when i = 1. Through this gradual incremental strategy optimization, we ensure that the product of decision frequency and process reward weight remains constant, thus avoiding excessive accumulation of process rewards due to increased decision frequency.

[0099] Finally, policy learning is implemented based on the Deep Deterministic Policy Gradient (DDPG) algorithm. First, the state space, action space, basic reward function, and process reward function are initialized, and a Markov Decision Process (MDP) model is constructed. Policy learning is then performed using the DDPG algorithm's policy network (Actor Network) and value network (Critic Network). The policy network is used to select the optimal action, and the value network is used to evaluate the expected reward of state-action pairs. In the initial stage of policy learning, low-frequency decisions are used to quickly generate the initial policy, helping the agent familiarize itself with the environment and reducing learning complexity. As training progresses, the decision frequency is gradually increased, transitioning from low-frequency to high-frequency decisions, allowing the policy to be gradually optimized in this high-frequency decision environment. In each training stage, experience replay is used to sample from the experience pool, and the parameters of the value network are updated using the Bellman equation. Simultaneously, the parameters of the policy network are updated using the policy gradient method. During policy learning, the weights of the process reward function are adaptively adjusted to prevent excessive accumulation of process rewards due to increased decision frequency. In the final stage of policy learning, the policy is optimized through high-frequency decisions to ensure stable convergence in this high-frequency decision environment, ultimately obtaining an optimized trajectory tracking strategy.

[0100] A device for implementing the incremental trajectory tracking strategy reinforcement learning method under high-frequency decision-making, comprising:

[0101] Processor: an algorithm for executing the incremental trajectory tracking policy reinforcement learning method under the high-frequency decision-making, including the construction of state space and action space, the calculation of reward function, and the training of policy network and value network;

[0102] Memory: used to store the parameters and models required by the method and the data generated during the training process;

[0103] Input interface: used to receive the initial parameters of the interceptor and target, including initial position, velocity, acceleration, etc., as well as field of view constraints;

[0104] Output interface: used to output the optimized trajectory tracking strategy and related performance indicators, including miss distance, energy consumption and intercept time;

[0105] Communication module: used to interact with external systems for data and to implement real-time updates and applications of strategies.

[0106] In terms of experimental verification, a simulation experimental environment for interception trajectory tracking based on random initial conditions is constructed, such as Figure 5 、 Figure 6 and Figure 7 The experimental results are shown. The interception simulation scenario parameters are set, and the mission goal is for the interceptor to pursue a high-speed target under field-of-view constraints, ensuring that the interceptor can successfully intercept the target while maximizing interception accuracy and minimizing energy consumption.

[0107] In actual operation, the initialization environment parameters are randomly generated according to the parameter range, such as the distance between the interceptor and the target, the line of sight angle, the target speed and other parameters.

[0108] We selected the DDPG algorithm as the baseline algorithm and set the dimensions, activation functions, and hyperparameters for each layer of the AN and CN networks. Throughout the training process, the decision frequency gradually increased, with corresponding decision step sizes of 0.1, 0.05, 0.02, and 0.01, respectively.

[0109] The effectiveness of the proposed method can be verified by multiple Monte Carlo (MC) tests, such as Figure 5 In the Monte Carlo test, the spatial trajectories of the interceptor and the target, the changes in the distance between the interceptor and the target, the interceptor's acceleration and lead angle, etc. are observed. The performance of the strategy is analyzed from qualitative and quantitative perspectives, and the advantages of the proposed method in terms of trajectory tracking accuracy, energy consumption and field of view constraints are demonstrated.

[0110] In summary, through the above specific implementation methods, the present invention can effectively realize the reinforcement learning of incremental trajectory tracking strategies under high-frequency decision-making, improve the interception performance of the interceptor against maneuvering targets, and meet the needs of practical applications, especially in scenarios with field of view constraints, energy consumption limitations, and high precision requirements.

[0111] It should be noted that any process or method description in the flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and that the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.

[0112] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0113] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0114] In addition, the functional units in the various embodiments of the present application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into one module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0115] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present application. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A reinforcement learning method for incremental trajectory tracking strategy under high-frequency decision-making, characterized by: The following steps are involved: Step 1: Construct a state space based on pure angle measurement, which includes the line of sight angle, line of sight angular rate, interceptor lead angle, and interceptor lead angle change rate to fully reflect the dynamic state during trajectory tracking. Step 2: Design an action space based on the proportional guidance method and introduce a bias term as a training action to provide an initial strategy for deep reinforcement learning. Step 3: Design a basic reward function related to the optimization goal. The basic reward function is related to the miss distance, energy consumption and field of view constraints, and is used to provide global goal feedback during the policy training process. Step 4: Design a process reward function that provides real-time feedback. The process reward function includes reward terms for gaze angular velocity, lead angle, and energy consumption, which are used to adjust the strategy in real time during strategy training. Step 5: Adopt a high-frequency decision-making approach to improve the policy's responsiveness by shortening the time step, and introduce a progressive incremental policy optimization (PIPO) module. This module transitions from low-frequency decision-making to high-frequency decision-making, gradually improving policy performance while balancing training efficiency and policy performance. Step 6: Implement policy learning based on the deep deterministic policy gradient (DDPG) algorithm, gradually increase the decision frequency through multi-stage learning, and ultimately optimize the trajectory tracking strategy in a high-frequency decision environment.

2. The incremental trajectory tracking strategy reinforcement learning method under high-frequency decision-making according to claim 1 is characterized by: The method for constructing the state space comprises the following steps: Step 1: Based on the relative motion equation between the interceptor and the target, define the state variables such as the sight angle, sight angle rate, interceptor lead angle, and interceptor lead angle change rate; The relative motion equation between the interceptor and the target can be defined as: Among them, V m and V t They represent the speed of the interceptor and the target respectively, r represents the relative distance between the interceptor and the target, q represents the line of sight angle, γ m and γ t denote the track angles of the interceptor and target, respectively; Step 2: Normalize the state variables to eliminate the differences between different physical dimensions and ensure that all state variables have the same scale to facilitate effective search in the sample space; The normalized state space can be defined as: in, They represent the normalized sight angle, sight angle rate, interceptor lead angle and their changing rate respectively.

3. The incremental trajectory tracking strategy reinforcement learning method under high-frequency decision-making according to claim 1 is characterized in that The design method of the action space includes: The motion space module is designed based on the proportional navigation method (PN), and a bias term is introduced as a training action. The bias term is used to adjust the normal overload instruction of the interceptor to optimize the trajectory tracking strategy. The bias equation of the action space module is: Among them, N represents the proportional coefficient, ε represents the bias overload instruction, represents the line of sight angular rate; During the training process, the agent optimizes the strategy by adjusting the bias term ξ, thereby improving the overall performance of the trajectory tracking system.

4. The incremental trajectory tracking strategy reinforcement learning method under high-frequency decision-making according to claim 1 is characterized in that The design method of the reward function includes: Design a basic reward function related to the optimization objective, the basic reward function includes: When the target is successfully intercepted, a positive reward signal is provided, and its value is in a decreasing relationship with the miss distance. The smaller the miss distance, the higher the reward. It has a decreasing function relationship with the cumulative sum of acceleration. The smaller the cumulative sum of acceleration, the higher the reward. When the viewing angle exceeds the maximum allowed field of view constraint, a large negative reward is applied to punish the policy for deviating from the expected constraints. Design a process reward function that provides real-time feedback, which includes: Reward function R for dynamically correcting the angular velocity of the line of sight q , whose expression is: Among them, b1 is the weight parameter, σ1 is the scaling coefficient; Reward function R for real-time correction of the leading angle λ , whose expression is: Among them, b2 is the weight parameter and σ2 is the scaling coefficient; Reward function R for optimizing energy consumption A , whose expression is: Among them, b3 is the weight parameter and σ3 is the scaling coefficient; The sum of the reward functions is: R guiding =R q +R λ +R A , Combined with the basic reward function, a complete reward mechanism is formed to drive the strategy learning of the intelligent agent.

5. The incremental trajectory tracking strategy reinforcement learning method under high-frequency decision-making according to claim 1 is characterized in that The progressive incremental strategy optimization PIPO module comprises the following steps: Step 1: In the initial stage of policy learning, low-frequency decisions are used to quickly generate initial policies, helping the agent become familiar with the environment and reducing learning complexity. Step 2: As training progresses, gradually increase the decision frequency, transitioning from low-frequency decision-making to high-frequency decision-making, so that the strategy is gradually optimized in a high-frequency decision-making environment; Step 3: As the decision frequency gradually increases, the weight of the process reward function is adaptively adjusted to prevent excessive accumulation of process rewards due to the increase in decision frequency. The specific adjustment method is as follows: a 1(i) ·f i =a1·f1, Among them, a 1(i) is the weight of the reward for the i-th stage process, f i is the decision frequency of the i-th stage, f1 is the decision frequency of the initial stage, and a1 is the process reward weight of the initial stage when i=1; Through the above-mentioned gradual incremental strategy optimization, it is ensured that the intelligent agent can achieve stable strategy optimization in a high-frequency decision-making environment, while taking into account both training efficiency and strategy performance.

6. The incremental trajectory tracking strategy reinforcement learning method under high-frequency decision-making according to claim 1 is characterized in that The implementation method based on the deep deterministic policy gradient DDPG algorithm includes the following steps: Step 1: Initialize the state space, action space, basic reward function and process reward function, and build a Markov decision process MDP model; Step 2: Use the DDPG algorithm’s policy network ActorNetwork and value network Critic Network for policy learning, where the policy network is used to select the optimal action and the value network is used to evaluate the expected return value of the state-action pair; Step 3: In the initial stage of policy learning, low-frequency decisions are used to quickly generate initial policies, helping the agent become familiar with the environment and reducing learning complexity. Step 4: As training progresses, gradually increase the decision frequency, transitioning from low-frequency decision-making to high-frequency decision-making, so that the strategy is gradually optimized in a high-frequency decision-making environment; Step 5: In each training phase, the experience replay mechanism is used to sample from the experience pool and the Bellman equation is used to update the parameters of the value network. At the same time, the policy network parameters are updated using the policy gradient method. Step 6: During the policy learning process, the weights of the process reward function are adaptively adjusted to prevent excessive accumulation of process rewards due to the increase in decision frequency.

7. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the incremental trajectory tracking strategy reinforcement learning method under high-frequency decision-making according to any one of claims 1 to 6 is implemented.

8. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the incremental trajectory tracking strategy reinforcement learning method under high-frequency decision-making according to any one of claims 1 to 6 is implemented.

9. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to execute the incremental trajectory tracking strategy reinforcement learning method under high-frequency decision-making according to any one of claims 1 to 6 by executing the executable instructions.