Aircraft trajectory optimization method and device

By constructing a three-dimensional nonlinear relative motion model and a nonlinear interference observer, and combining it with an adaptive dynamic programming method, composite guidance commands were designed. This solved the problems of robustness and trajectory tracking accuracy in intercepting highly maneuverable targets, and enabled the aircraft to intercept targets accurately, quickly, and stably in highly dynamic environments.

CN121995955APending Publication Date: 2026-05-08INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INST OF AUTOMATION CHINESE ACAD OF SCI
Filing Date
2026-01-27
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies lack robustness, have low trajectory tracking accuracy, slow convergence speed, and are prone to overload saturation when intercepting highly maneuverable targets, making it difficult to achieve accurate, fast, and stable interception in three-dimensional space.

Method used

By constructing a three-dimensional nonlinear relative motion model, combining a nonlinear interference observer and an adaptive dynamic programming method, a composite guidance command is designed to achieve real-time estimation and feedforward compensation of target maneuvering interference. The weight update law ensures the rapid and stable convergence of the value function estimation network, and the control is performed considering the overload constraints of the aircraft.

Benefits of technology

It significantly improves the interception accuracy and response speed of aircraft in highly dynamic and highly interference environments, enhances anti-interference robustness, ensures that control commands are within the safe range of physical actuators, avoids the risk of runaway caused by overload saturation, and meets the requirements for millisecond-level real-time control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121995955A_ABST
    Figure CN121995955A_ABST
Patent Text Reader

Abstract

The invention provides an aircraft trajectory optimization method and device, and relates to the technical field of aircraft control, and the method comprises the steps: determining interference estimation data and feedback control data based on relative motion state data, carrying out linear superposition and overload limit processing, determining a composite guidance instruction, and controlling an aircraft. According to the method and the device provided by the invention, the three-dimensional nonlinear relative motion model is constructed, and the feedforward interference compensation signal and the feedback optimal control signal are cooperatively fused, so that the interference influence is eliminated from the source, and the optimality of the flight path is ensured; through the synergistic effect of a plurality of links such as model construction, interference estimation, optimal guidance law design, instruction fusion and trajectory control, the interception precision and response speed of the aircraft to a high-maneuvering target in a high-dynamic and strong-interference environment are remarkably improved, and the method has extremely high engineering application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of aircraft control technology, and in particular to an aircraft trajectory optimization method and apparatus. Background Technology

[0002] In modern warfare, the trajectory control performance of precision-guided weapons directly determines the success rate of target interception. This is especially true when facing targets with high maneuverability and rapid trajectory change capabilities, which places extremely high demands on the robustness, real-time performance, and optimality of aircraft trajectory optimization methods. The motion state of highly maneuverable targets is highly nonlinear, time-varying, and uncertain. Their maneuvering acceleration, as an unknown disturbance, can severely impair the trajectory tracking accuracy of the aircraft and even lead to the failure of the interception mission.

[0003] Traditional aircraft trajectory optimization methods mainly include proportional guidance law, sliding mode control method, and adaptive control method. When intercepting highly maneuverable targets, these methods suffer from problems such as insufficient robustness, low trajectory tracking accuracy, slow convergence speed, and susceptibility to overload saturation.

[0004] Therefore, how to achieve accurate, rapid, and stable interception of highly maneuverable targets in three-dimensional space has become a technical problem that the industry urgently needs to solve. Summary of the Invention

[0005] This invention provides a method and apparatus for optimizing aircraft trajectory, which addresses the shortcomings of existing technologies in intercepting highly maneuverable targets, such as insufficient robustness, low trajectory tracking accuracy, slow convergence speed, and susceptibility to overload saturation. It enables accurate, rapid, and stable interception of highly maneuverable targets in three-dimensional space, thereby improving the engineering practicality and reliability of aircraft trajectory optimization.

[0006] This invention provides a method for optimizing aircraft trajectory, comprising: Acquire relative motion data between the aircraft and the tracked target; The relative motion state data is input into a nonlinear interference observer to determine the interference estimation data; the nonlinear interference observer is constructed based on the relative motion model between the aircraft and the tracked target. A value function estimation network is constructed, and the relative motion state data is input into the value function estimation network to determine the feedback control data; the value function estimation network adjusts the weights of the value function estimation network using a weight update law; Based on the interference estimation data and the feedback control data, a composite guidance command is determined to control the aircraft based on the composite guidance command.

[0007] In some embodiments, the nonlinear disturbance observer includes internal state variables and an observer gain matrix; the structure and parameters of the observer gain matrix are determined based on the Lyapunov stability criterion.

[0008] In some embodiments, inputting the relative motion state data into a nonlinear disturbance observer to determine disturbance estimation data includes: Based on the relative motion state data, update the internal state variables in the nonlinear disturbance observer; The disturbance estimation data is output based on a linear combination of the internal state variables and the relative motion state data.

[0009] In some embodiments, constructing a value function estimation network, by inputting the relative motion state data into the value function estimation network to determine feedback control data, includes: The value function estimation network is constructed based on a single hidden layer neural network; The relative motion state data is input into the value function estimation network to obtain the value function estimate; Based on the estimated value of the value function, combined with the weight matrix of the control input energy consumption term in the performance index function and the control input matrix of the relative motion model, the feedback control data is determined.

[0010] In some embodiments, the weight update law includes an error feedback term; The error feedback term is determined by the performance index function and the error function determined by the value function estimation network; The error feedback term is used to dynamically adjust the update rate of the weights in the value function estimation network.

[0011] In some embodiments, determining the composite guidance command based on the interference estimation data and the feedback control data includes: The interference estimation data and the feedback control data are linearly superimposed to generate an initial composite guidance command; Based on the preset maximum acceleration threshold of the aircraft, the initial composite guidance command is subjected to amplitude limiting processing to obtain the composite guidance command.

[0012] This invention provides an aircraft trajectory optimization device, comprising: The data acquisition module is used to acquire relative motion state data between the aircraft and the target being tracked. The feedforward module is used to input the relative motion state data into the nonlinear interference observer to determine the interference estimation data; the nonlinear interference observer is constructed based on the relative motion model between the aircraft and the tracked target; The feedback module is used to construct a value function estimation network, inputting the relative motion state data into the value function estimation network to determine feedback control data; the value function estimation network adjusts the weights of the value function estimation network using a weight update law; The control module is used to determine composite guidance commands based on the interference estimation data and the feedback control data, so as to control the aircraft based on the composite guidance commands.

[0013] The present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the aircraft trajectory optimization method.

[0014] The present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the aircraft trajectory optimization method.

[0015] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the aircraft trajectory optimization method.

[0016] The aircraft trajectory optimization method and apparatus provided by this invention accurately characterize the strong coupling characteristics between the line-of-sight azimuth angle and the pitch angle by constructing a three-dimensional nonlinear relative motion model, effectively avoiding the truncation error caused by traditional linearized models and laying a model foundation for high-precision control. Through a nonlinear disturbance observer, real-time and accurate estimation and feedforward compensation of target maneuver disturbances are achieved without requiring prior information on target maneuvers or measurement of target acceleration, significantly enhancing the system's robustness against large target maneuvers. Based on an adaptive dynamic programming method, a feedback optimal guidance law is designed, featuring innovative weight updates. The optimal guidance law ensures the rapid and stable convergence of the value function estimation network, meeting the millisecond-level real-time control requirements. Through the linear superposition of feedforward and feedback, and overload processing, the interference effect is eliminated at the source, the optimality of the flight trajectory is guaranteed, and the control commands are kept within the safe range of the aircraft's physical actuators, avoiding the risk of runaway due to overload saturation. Through the synergistic effect of multiple links such as model building, interference estimation, optimal guidance law design, command fusion and trajectory control, the interception accuracy and response speed of the aircraft against highly maneuverable targets in high-dynamic and highly interference environments are significantly improved, which has extremely high engineering application value. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0018] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating the aircraft trajectory optimization method provided by the present invention.

[0020] Figure 2 This is a block diagram illustrating the principle of aircraft trajectory optimization in the aircraft trajectory optimization method provided by this invention.

[0021] Figure 3 This is a curve showing the rate of change of the single network weights used in the adaptive dynamic programming algorithm of the aircraft trajectory optimization method provided by this invention.

[0022] Figure 4 This is a diagram showing the trajectory optimization / guidance command changes of the aircraft trajectory optimization method provided by this invention.

[0023] Figure 5 This is a dynamic graph showing the convergence of the line-of-sight rotation rate of the aircraft trajectory optimization method provided by this invention.

[0024] Figure 6 This is a motion trajectory diagram of an aircraft intercepting a maneuvering target using the aircraft trajectory optimization method provided by this invention.

[0025] Figure 7 This is a schematic diagram of the aircraft trajectory optimization device provided by the present invention.

[0026] Figure 8 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0027] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0028] It should be noted that the terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that comprises a series of steps, units, or modules is not necessarily limited to those explicitly listed, but may include other steps, units, or modules not explicitly listed or inherent to such processes, methods, products, or devices.

[0029] Figure 1 This is a flowchart illustrating the aircraft trajectory optimization method provided by the present invention, as shown below. Figure 1 As shown, the method includes steps 110, 120, 130 and 140.

[0030] Step 110: Obtain the relative motion state data between the aircraft and the tracked target.

[0031] Specifically, the aircraft trajectory optimization method provided in this embodiment of the invention is executed by an aircraft trajectory optimization device. This device can be implemented in software, such as an aircraft trajectory optimization program running on a computer; or it can be implemented in hardware, such as a computer or server that executes the aircraft trajectory optimization method.

[0032] Traditional methods for optimizing aircraft trajectories mainly include proportional guidance laws, sliding mode control, and adaptive control, but these methods all have significant technical drawbacks. The proportional guidance law is designed based on a linearized missile-target relative motion model, ignoring the strong nonlinear coupling characteristics in three-dimensional space. When the target performs high-maneuvering motion, the guidance accuracy drops significantly, making it difficult to meet the requirements of precise interception. The sliding mode control method suppresses interference by designing sliding surfaces and switching control laws, but traditional sliding mode control suffers from severe chattering, which leads to increased wear on the aircraft's actuators. Furthermore, the selection of control gain depends on prior information about the interference boundary, resulting in insufficient robustness in scenarios where the target's maneuvering acceleration is unknown. Traditional adaptive control methods adapt to interference changes by estimating system parameters online, but the convergence speed is slow, and it is difficult to simultaneously guarantee the optimal control performance of the system. In complex interference environments, trajectory divergence problems are prone to occur.

[0033] Adaptive Dynamic Programming (ADP), as an advanced algorithm integrating reinforcement learning and optimal control, solves the optimal control problem of nonlinear systems through iterative learning, without requiring precise solution to the complex Hamilton-Jacobi-Bellman (HJB) equations, thus providing a new technical path for the design of optimal guidance laws for nonlinear systems. However, the pure adaptive dynamic programming method has poor robustness to unknown disturbances. Its value function estimation network is susceptible to disturbances, resulting in slow convergence and increased estimation bias, leading to a decline in the performance of the optimal control law and failing to meet the strong robustness requirements for intercepting highly maneuverable targets.

[0034] Nonlinear Disturbance Observer (NDO), as an effective tool for disturbance estimation and compensation, can significantly improve the system's ability to suppress disturbances by real-time monitoring of the system state, estimating unknown disturbances online, and performing feedforward compensation. Combining the NDO with adaptive dynamic programming is expected to fully leverage the advantages of both technologies: the NDO can estimate and compensate for target maneuvering disturbances in real time, eliminating the adverse effects of disturbances on the system; while adaptive dynamic programming can ensure optimal control performance of the system, enabling the aircraft trajectory to reach its optimal state.

[0035] However, existing fusion schemes still have several shortcomings: First, the coordination between the interference observer and adaptive dynamic programming is poor, and the interference estimation signal cannot be effectively integrated into the value function estimation and weight update process of adaptive dynamic programming, resulting in limited improvement in control performance; second, the design of the weight update law of adaptive dynamic programming lacks specificity, has a slow convergence speed, and is prone to weight oscillation, affecting the real-time performance of the optimal control law; third, the fusion strategy of composite guidance commands is simple, does not consider the overload constraints of the aircraft, and is prone to overload saturation, reducing its engineering practicality. Therefore, this invention provides an aircraft trajectory optimization method to achieve accurate and stable interception of highly maneuverable targets.

[0036] An aircraft is an airborne vehicle capable of controlled flight within or outside the atmosphere and possessing the ability to perform trajectory adjustment and target interception missions. In this embodiment of the invention, "aircraft" primarily refers to interceptor missiles, such as air-to-air missiles, surface-to-air missiles, and ship-to-air missiles.

[0037] A tracking target refers to an object that moves in three-dimensional space, and whose trajectory is typically unpredictable, nonlinear, and highly maneuverable for an aircraft. In this embodiment of the invention, the tracking target mainly refers to highly maneuverable enemy aerial targets, including but not limited to fighter jets, cruise missiles, and tactical ballistic missiles.

[0038] Relative motion state data refers to a set of physical quantities used in a three-dimensional coordinate system to describe the relative positional relationship and dynamic changing trend between an aircraft, such as a missile, and its tracked target. In this embodiment of the invention, the relative motion state data includes, but is not limited to, line-of-sight azimuth angle, line-of-sight pitch angle, rate of change of line-of-sight azimuth angle, and rate of change of line-of-sight pitch angle. Preferably, it also includes the missile-target distance and the rate of change of missile-target distance. This set of data together constitutes the input state vector of the missile-target relative motion model, which can fully characterize the geometric motion characteristics of the aircraft during the target tracking process.

[0039] After constructing the relative motion model between the aircraft and the tracked target, it is necessary to acquire accurate relative motion states of the missile and target in real time as the input basis for subsequent interference observation and guidance law calculation. In this embodiment of the invention, the aircraft, such as an air-to-air missile, uses a multi-sensor unit to synchronously collect motion information of the target and itself. Furthermore, in order to remove measurement noise and outliers from the original state data and obtain a smooth and reliable relative motion state signal, a Kalman filter algorithm is used to process the data, ultimately obtaining the relative motion state data between the aircraft and the tracked target.

[0040] In this embodiment of the invention, the real-time motion state data of the target aircraft and itself are synchronously collected by the sensor unit. After filtering, noise reduction and data standardization, a reliable relative motion state signal between the projectile and the target is output.

[0041] In this embodiment of the invention, the sensor unit employs a multi-sensor fusion scheme, including a radar sensor, an inertial measurement unit (IMU), and a global positioning system (GPS). The radar sensor measures the missile-target distance, line-of-sight angle, and corresponding rate of change; the IMU collects motion state parameters such as the missile's acceleration and angular velocity; and the GPS acquires the absolute position information of the missile and the target, assisting in calculating their relative motion state. In engineering implementation, the sensor sampling frequency, the computing power of the embedded processor, and the control cycle must be considered to ensure the real-time performance of the method.

[0042] The collected raw state data includes sensor measurement noise, and abnormal values ​​caused by external environmental interference such as electromagnetic interference and airflow interference, which need to be processed by the data preprocessing module.

[0043] In this embodiment of the invention, a Kalman filter algorithm is used for noise suppression. Its core principle is to establish state equations and observation equations, and then iteratively update the optimal state estimate at the current moment using the state estimate from the previous moment and the observed value at the current moment, thereby eliminating noise and abnormal data. Through the iterative update of the state equations and observation equations, abnormal data caused by sensor measurement noise and external environmental interference are eliminated, ensuring the smoothness and reliability of the relative motion state signal.

[0044] The specific processing steps are as follows: First, based on the relative motion model between the missile and the target, i.e., the relative motion model between the aircraft and the target being tracked, a Kalman filter state equation is established to clarify the evolution law of the system state. Second, based on the measurement characteristics of the sensor, an observation equation is established to describe the relationship between the observed values ​​and the system state. Then, through iterative calculations of the prediction step and the update step, a smooth and reliable relative motion state signal between the missile and the target, i.e., the final relative motion state data, is obtained, providing high-quality input data for subsequent interference observation and guidance law design.

[0045] In addition, if faced with a complex electromagnetic interference environment, anti-interference algorithms such as adaptive filtering and wavelet denoising can be added during data preprocessing to further improve the reliability of state data.

[0046] Step 120: Input the relative motion state data into the nonlinear interference observer to determine the interference estimation data; the nonlinear interference observer is constructed based on the relative motion model between the aircraft and the tracked target.

[0047] Specifically, in actual interception processes, tracking targets such as fighter jets or highly maneuverable missiles often employ unknown maneuvering strategies, such as sinusoidal maneuvers or step maneuvers, to evade interception. The acceleration generated by such maneuvers constitutes an unknown external disturbance for the aircraft's guidance system. To eliminate the impact of this disturbance on guidance accuracy, this invention proposes a nonlinear disturbance observer that estimates the target's maneuvering acceleration in real time without needing to predict the target's maneuvering patterns or measure the target's acceleration.

[0048] In this embodiment of the invention, the nonlinear interference observer is constructed based on a relative motion model. A relative motion model is a set of mathematical equations based on rigid body dynamics and kinematics principles, used to mathematically describe the evolution of the relative position, relative velocity, and relative attitude between the aircraft and the tracked target in three-dimensional space over time.

[0049] To achieve precise interception in three-dimensional space, this invention, based on the relative kinematics of an aircraft (taking a missile as an example) and a target in three-dimensional space, and combined with a line-of-sight angle parameterization method, integrates key parameters such as the line-of-sight angle, the rate of change of the line-of-sight angle, the missile-target distance, and the rate of change of the missile-target distance. It establishes a compact missile-target relative motion model that includes system state variables, nonlinear coupling terms, a control input matrix, an interference matrix, and a target maneuvering acceleration interference term. This model clarifies the logical relationships between the parameters and comprehensively reflects the relative motion between the missile and the target, i.e., between the aircraft and the tracked target. The system state variables in the missile-target relative motion model include the line-of-sight azimuth angle, the line-of-sight pitch angle, and their corresponding rates of change. The nonlinear coupling terms are composed of the combination of the rate of change of the missile-target distance and the rate of change of the line-of-sight angle, accurately reflecting the strong nonlinear coupling characteristics of the missile-target relative motion in three-dimensional space.

[0050] First, the core parameters of the relative motion model are defined. Based on the principles of relative kinematics in three-dimensional space, the line-of-sight azimuth, line-of-sight elevation, missile-target distance, and their derivatives are selected as key parameters describing the system's motion. The line-of-sight azimuth and elevation angles, as key angular parameters describing the relative position of the missile and target, directly reflect the trend of their relative motion. The missile-target distance is the straight-line distance between the missile and the target, and its rate of change is the rate of change of distance over time, used to characterize the speed at which the missile approaches or moves away from the target. The missile acceleration serves as a control input, used to adjust the missile's flight trajectory. The target's maneuvering acceleration, as an unknown disturbance, is a core factor affecting trajectory tracking accuracy.

[0051] Secondly, based on the fundamental equations of relative kinematics, and considering the influence of the rate of change of missile-target distance and the rate of change of line-of-sight angle on relative motion, the fundamental equations of missile-target relative motion are derived. In the derivation process, the coupling characteristics of three-dimensional space are fully considered, and the motion equations of azimuth and elevation angles are coupled and modeled to avoid the accuracy loss caused by neglecting coupling terms in traditional linearized models.

[0052] Finally, through variable substitution and matrix simplification, the basic equations are transformed into a compact form of the projectile-target relative motion model, which is expressed as follows: in , Indicates the angle of inclination of the line of sight; Indicates the angle of deflection of the line of sight; For gaze conversion rate; ,in Indicates the rate of change of the distance between the projectile and the target. Indicates the distance between the target and the target; The missile's acceleration control effectiveness on line-of-sight acceleration is described; The control acceleration input represents the aircraft (missile), where, and These are the azimuth and pitch acceleration components of the missile in the line-of-sight coordinate system, respectively. The effect of the target's maneuvering acceleration on the line-of-sight angular acceleration was described; The acceleration representing the tracking target is considered an unknown disturbance term of the system in this embodiment of the invention. and These represent the azimuth and pitch acceleration components of the target in the line-of-sight coordinate system.

[0053] This model clarifies the relationships between system state variables, nonlinear coupling terms, control input matrix, disturbance matrix, and target maneuver acceleration disturbance terms, providing a precise model foundation for subsequent disturbance observation and guidance law design.

[0054] The core advantage of this model lies in its ability to accurately capture the strong nonlinear coupling characteristics of the relative motion between the projectile and the target in three-dimensional space, without the need for linear approximation of the system, thus providing a reliable guarantee for the accuracy of interference estimation and the design of the optimal guidance law.

[0055] It should be noted that the parameters of the relative motion model between the missile and the target need to be calibrated according to the actual type of aircraft, such as air-to-air missiles, surface-to-air missiles, and combat scenarios. The parameter values ​​can be optimized through live-fire tests or high-precision hardware-in-the-loop simulation platforms.

[0056] In this embodiment of the invention, the target maneuvering acceleration is regarded as an unknown disturbance of the system. A nonlinear disturbance observer with fast convergence characteristics is designed based on the relative motion model to achieve real-time and accurate estimation of the disturbance.

[0057] Interference estimation data refers to a set of physical quantity vectors used to characterize the current maneuvering state of the tracked target, calculated in real time by a nonlinear interference observer. In this embodiment of the invention, the interference estimation data is an estimated vector of target maneuvering acceleration, including estimated components of target maneuvering acceleration in the line-of-sight azimuth direction and line-of-sight pitch direction. This data aims to quantify the unknown disturbances caused by the target's maneuvering flight to the relative motion trajectory of the missile and the target, and is used for feedforward compensation in the subsequent guidance command generation.

[0058] In this embodiment of the invention, based on relative motion state data, a nonlinear disturbance observer is used to determine disturbance estimation data.

[0059] Step 130: Construct a value function estimation network. Input the relative motion state data into the value function estimation network to determine the feedback control data. The value function estimation network uses a weight update law to adjust the weights of the value function estimation network.

[0060] Specifically, after solving the "unknown interference" problem using a nonlinear interference observer, the remaining control task can be equivalent to performing optimal trajectory planning for a nominal system without interference. In this embodiment of the invention, based on a single-network adaptive dynamic programming method, a feedback optimal guidance law is designed for the nominal system, i.e., an ideal system without target maneuvering interference. An innovative weight update law ensures the optimality and fast convergence of the guidance law.

[0061] The weight update law refers to a set of differential equations or iterative rules used in adaptive dynamic programming algorithms to guide the dynamic evolution of the weight vector of the value function estimation network over time.

[0062] Feedback control data refers to a set of optimal control command vectors calculated based on an adaptive dynamic programming algorithm for the nominal relative motion system of an aircraft after external disturbances have been removed. In this specific embodiment, the feedback control data is the optimal feedback control command, which is derived from the gradient information obtained by differentiating the output of the value function estimation network (i.e., the value function estimate) with respect to the system state, combined with the control energy consumption weight matrix in the performance index function and the system control input matrix. This data aims to ensure that the system meets preset performance indicators under nominal conditions, such as minimizing tracking error and minimizing energy consumption, and is used to control the aircraft to fly along the optimal trajectory.

[0063] In this embodiment of the invention, the value function estimation network adopts a single-hidden-layer neural network structure. The hidden layer activation function is selected as the sigmoid function or a polynomial function, and the input layer and hidden layer are dynamically adjusted through a weight update law, so that the estimated value function asymptotically approaches the optimal value function. Based on relative motion state data, the feedback control data can be determined using this value function estimation network.

[0064] Step 140: Based on the interference estimation data and the feedback control data, determine the composite guidance command, and control the aircraft based on the composite guidance command.

[0065] Specifically, Figure 2 This is a block diagram illustrating the principle of aircraft trajectory optimization in the aircraft trajectory optimization method provided by this invention, as shown below. Figure 2 As shown, in order to leverage the synergistic advantages of feedforward compensation and feedback control, a composite guidance command fusion strategy is designed, and overload limitation processing is implemented considering the overload constraints of the aircraft.

[0066] During trajectory optimization, the combined effect of feedforward compensation and feedback control rapidly converges the rate of change of the line-of-sight angle to near zero, satisfying the guidance requirements of the parallel approach method. The core of the parallel approach method is to make the rate of change of the line-of-sight angle between the missile and the target approach zero, ensuring that the missile approaches the target along the optimal trajectory and improving interception accuracy. Furthermore, the missile-target relative motion model is compact, and the computational overhead for data preprocessing, interference observation, and guidance law design is low, making it easy to implement in embedded processors, greatly improving the engineering practicality and adaptability of this method.

[0067] Simultaneously, the distance between the missile and the target is monitored in real time by the missile-target distance monitoring module. When the distance is less than the set interception threshold, the interception is considered successful, guidance control is stopped, and the target interception mission is completed. It should be noted that the interception threshold can be adjusted according to the actual interception mission requirements. In this embodiment of the invention, the interception threshold is set to 30 meters.

[0068] The aircraft trajectory optimization method provided in this invention constructs a three-dimensional nonlinear relative motion model, accurately characterizing the strong coupling characteristics between the line-of-sight azimuth angle and the pitch angle, effectively avoiding the truncation error caused by traditional linearized models, and laying a model foundation for high-precision control. Through a nonlinear disturbance observer, real-time and accurate estimation and feedforward compensation of target maneuver disturbances are achieved without requiring prior information on target maneuvers or measuring target acceleration, significantly enhancing the system's robustness against large target maneuvers. Furthermore, an innovative weight update method is used to design a feedback-optimal guidance law based on adaptive dynamic programming. The optimal guidance law ensures the rapid and stable convergence of the value function estimation network, meeting the millisecond-level real-time control requirements. Through the linear superposition of feedforward and feedback, and overload processing, the interference effect is eliminated at the source, the optimality of the flight trajectory is guaranteed, and the control commands are kept within the safe range of the aircraft's physical actuators, avoiding the risk of runaway due to overload saturation. Through the synergistic effect of multiple links such as model building, interference estimation, optimal guidance law design, command fusion and trajectory control, the interception accuracy and response speed of the aircraft against highly maneuverable targets in high-dynamic and highly interference environments are significantly improved, which has extremely high engineering application value.

[0069] In some embodiments, the nonlinear disturbance observer includes internal state variables and an observer gain matrix; the structure and parameters of the observer gain matrix are determined based on the Lyapunov stability criterion.

[0070] Specifically, in this embodiment of the invention, following the core idea of ​​"state tracking-error correction", a nonlinear interference observer is designed, which includes internal state variables and a diagonal matrix form of observer gain. The preprocessed relative motion state signal is used as input, and the target maneuver acceleration interference is estimated in real time through internal state dynamic evolution and feedback correction, and the interference estimation signal is output. Among them, the internal state variables are used to dynamically track the changing trend of the target maneuver acceleration, and the observer gain matrix is ​​used to adjust the convergence speed and stability of the interference estimation.

[0071] It should be noted that the design of the observer gain matrix is ​​crucial to ensuring the accuracy of interference estimation, directly affecting its stability and speed. In this embodiment of the invention, the structure and parameters of the gain matrix are determined based on the Lyapunov stability criterion, ensuring rapid convergence of interference estimation. First, a Lyapunov candidate function is constructed, which satisfies positive definiteness. Then, the derivative of the Lyapunov candidate function is calculated, and the constraints of the gain matrix are derived based on the stability requirement that the derivative is less than zero. Finally, within the constraint range, combined with the target maneuver characteristics such as maximum maneuver acceleration, maneuver frequency, and system response requirements, the gain parameters in diagonal matrix form are determined, achieving rapid convergence and high-precision tracking of interference estimation.

[0072] The aircraft trajectory optimization method provided in this embodiment of the invention designs a nonlinear disturbance observer based on the Lyapunov stability criterion, which ensures the rapid convergence of disturbance estimation.

[0073] In some embodiments, inputting the relative motion state data into a nonlinear disturbance observer to determine disturbance estimation data includes: Based on the relative motion state data, update the internal state variables in the nonlinear disturbance observer; The disturbance estimation data is output based on a linear combination of the internal state variables and the relative motion state data.

[0074] Specifically, in this embodiment of the invention, the observer's input is a preprocessed relative motion state signal, including the rate of change of the line-of-sight angle and the rate of change of the projectile-target distance. The internal state values ​​are updated in real time through the dynamic evolution equation of the internal state variables. Its core principle is to feed back the system state error to the internal state evolution process through the observer gain matrix, making the internal state values ​​asymptotically approximate the actual target maneuvering acceleration. Then, through a linear combination of the internal state values ​​and the system state variables, an interference estimation signal is output.

[0075] In this embodiment of the invention, the target acceleration is considered as interference, and a nonlinear interference observer is used to estimate this interference. Based on the constructed relative motion model between the aircraft and the tracked target, the nonlinear interference observer is constructed as follows: in, It is an estimate of the target's acceleration disturbance; It is the internal state of the nonlinear observer; It is the observer gain.

[0076] The aircraft trajectory optimization method provided in this invention, by designing a nonlinear interference observer, can estimate the target maneuver acceleration interference in real time without prior information about the target maneuver, providing accurate interference information for feedforward compensation, thereby offsetting the adverse effects of unknown interference on the system and solving the problem of insufficient anti-interference robustness of the simple adaptive dynamic programming method.

[0077] In some embodiments, constructing a value function estimation network, by inputting the relative motion state data into the value function estimation network to determine feedback control data, includes: The value function estimation network is constructed based on a single hidden layer neural network; The relative motion state data is input into the value function estimation network to obtain the value function estimate; Based on the estimated value of the value function, combined with the weight matrix of the control input energy consumption term in the performance index function and the control input matrix of the relative motion model, the feedback control data is determined.

[0078] Specifically, after feedforward compensation of the target maneuver through the interference observer, the remaining system can be considered as an interference-free nominal system. In order to achieve the optimal balance between "trajectory tracking accuracy" and "energy consumption" on this nominal system, this embodiment of the invention uses an adaptive dynamic programming method to design the feedback optimal guidance law.

[0079] The performance index function is used to evaluate the quality of the control signal, taking into account both trajectory tracking accuracy and energy consumption. It is defined as a quadratic form function containing a system state error term and a control input energy consumption term. The system state error term penalizes state deviations from the ideal trajectory, and its importance is adjusted by a weight matrix; a larger weight coefficient indicates a higher requirement for trajectory tracking accuracy. The control input energy consumption term limits the amplitude of the control input to avoid excessive energy consumption of the aircraft, and its importance is also adjusted by a weight matrix.

[0080] The design principle of the performance index function is to minimize control input energy consumption while ensuring trajectory tracking accuracy, achieving a balance between "high accuracy and low energy consumption," and providing an evaluation standard for the derivation of the optimal control law. The performance index function includes a control input energy consumption term, and through weight adjustment, it achieves a balance between trajectory tracking accuracy and energy consumption. Compared to traditional methods, this reduces energy consumption to a certain extent, improves the aircraft's endurance, and extends its effective combat radius.

[0081] The value function is used to characterize the cumulative sum of the performance indicators of a system at all future times, starting from the current state, and is a core concept of adaptive dynamic programming. Since the analytical solution to the value function is difficult to obtain directly, this embodiment of the invention uses a single-hidden-layer neural network to construct a value function estimation network to approximately solve for the optimal value function.

[0082] In this embodiment of the invention, a value function estimation network is constructed based on a single-network adaptive dynamic programming method. A dedicated neural network weight update law containing positive learning rate parameters and error feedback terms is designed. The network receives relative motion state signals and disturbance estimation signals, and outputs the optimal feedback control signal for the nominal system.

[0083] The value function estimation network is designed as follows: the number of input layer nodes equals the dimension of the system state variables, used to receive the preprocessed projectile-object relative motion state signal; the number of hidden layer nodes is determined through simulation testing and optimization, and the Sigmoid function is selected as the activation function, which has strong nonlinear mapping ability and smooth gradient, and can effectively fit the nonlinear characteristics of the value function; the output layer is linear, and the output value is the estimated value of the value function. The initial weights of the network are determined through random initialization, and the range of initial weight values ​​is set based on simulation experience to avoid slow network convergence caused by initial weights being too large or too small.

[0084] Then, the relative motion state data is input into the value function estimation network to obtain the value function estimate; based on the value function estimate, combined with the weight matrix of the control input energy consumption term in the performance index function and the control input matrix of the relative motion model, the feedback control data is determined.

[0085] In other words, based on the relationship between the optimal value function and the performance index function, the optimal feedback guidance law is derived through the optimal control principle: the output of the value function estimation network, i.e. the value function estimate, is differentiated with respect to the system state variables to obtain the gradient information of the value function; then, combined with the weight matrix of the control input energy consumption term in the performance index function and the control input matrix of the projectile-target relative motion model, the optimal feedback control signal is derived. This signal can minimize the system performance index function and ensure the optimality of the trajectory.

[0086] The aircraft trajectory optimization method provided in this invention is based on an adaptive dynamic programming method to design a feedback optimal guidance law. The innovative weight update law ensures the fast and stable convergence of the value function estimation network. By constructing the value function estimation network and determining the feedback control data, the convergence speed is greatly improved, meeting the real-time requirements of aircraft trajectory optimization.

[0087] In some embodiments, the weight update law includes an error feedback term; The error feedback term is determined by the performance index function and the error function determined by the value function estimation network; The error feedback term is used to dynamically adjust the update rate of the weights in the value function estimation network.

[0088] Specifically, in the process of constructing a value function estimation network to solve the optimal control law, the convergence speed and stability of the network weights directly determine the real-time performance and reliability of the guidance system. Traditional weight update methods, such as the standard descent method based on gradients, usually use a fixed learning rate, which makes it difficult to balance convergence speed and oscillation suppression: too small a learning rate leads to slow convergence and cannot adapt to the rapidly changing battlefield environment; too large a learning rate can easily cause weight oscillations or even divergence.

[0089] The design of the weight update law directly affects the convergence speed and stability of the value function estimation network. This invention innovatively designs a dedicated weight update law that includes a positive learning rate parameter and an error feedback term, solving the problems of slow convergence and oscillation in traditional weight update laws. The core logic of the weight update law is to dynamically adjust the weight update rate through the error feedback term: the error feedback term consists of the deviation between the system performance index function and the estimated value of the value function; in other words, the error feedback term is determined by the performance index function and the error function determined by the value function estimation network.

[0090] It should be noted that the error feedback term is used to dynamically adjust the update rate of the weights of the value function estimation network; when the deviation is large, the weight update rate is increased to speed up network convergence; when the deviation is small, the weight update rate is decreased to avoid weight oscillation and ensure network stability.

[0091] In addition, the learning rate parameter is a positive constant, which is determined by simulation test optimization. It needs to be dynamically adjusted according to the target's maximum maneuvering acceleration, maneuvering frequency and other characteristics. Its value needs to take into account both convergence speed and stability. If the learning rate is too large, it will cause weight oscillation. If the learning rate is too small, it will cause slow convergence speed.

[0092] The aircraft trajectory optimization method provided in this invention balances convergence speed and stability by designing a dedicated weight update law, enabling the value function estimation network to converge quickly and stably to the optimal weights, and making the value function estimate accurately approximate the optimal value function.

[0093] In some embodiments, determining the composite guidance command based on the interference estimation data and the feedback control data includes: The interference estimation data and the feedback control data are linearly superimposed to generate an initial composite guidance command; Based on the preset maximum acceleration threshold of the aircraft, the initial composite guidance command is subjected to amplitude limiting processing to obtain the composite guidance command.

[0094] Specifically, in this embodiment of the invention, the fusion of composite guidance commands adopts a linear superposition strategy, which uses the interference estimation signal output by the nonlinear interference observer as a feedforward compensation signal and linearly superimposes it with the optimal feedback control signal output by the adaptive dynamic programming to generate the initial composite guidance command.

[0095] The feedforward compensation signal directly counteracts the effects of target maneuvering interference. Based on the amplitude and direction of the interference estimation signal, a compensation signal with the opposite direction and matching amplitude to the interference is generated, suppressing the interference's disruption to trajectory tracking at its source. The optimal feedback control signal ensures the system's optimal tracking performance, enabling the aircraft trajectory to converge quickly to the optimal trajectory.

[0096] The advantages of the linear superposition strategy lie in its simple structure, low computational overhead, ease of engineering implementation, and ability to balance anti-interference performance and optimal control performance by adjusting the weight coefficients of feedforward compensation and feedback control. In this embodiment of the invention, the weight coefficients are determined based on simulation tests. When the target maneuvering interference is strong, the weight of feedforward compensation is increased; when the interference is weak, the weight of feedback control is increased, achieving dynamic adaptive adjustment.

[0097] Furthermore, the aircraft's actuators, such as thrusters and servos, have maximum acceleration limits. If the acceleration corresponding to the composite guidance command exceeds this limit, overload saturation will occur, leading to damage to the actuators or loss of trajectory control. Therefore, it is necessary to limit the amplitude of the initial composite guidance command through overload limiting.

[0098] The preset maximum acceleration threshold for overload limiting is determined based on actual engineering parameters such as the aircraft's structural strength and propulsion system performance. A saturation function is used for amplitude limiting: when the acceleration corresponding to the initial composite guidance command is less than or equal to the maximum acceleration threshold, the command is directly output; when the acceleration exceeds the maximum acceleration threshold, the command corresponding to the maximum acceleration is output, ensuring that the aircraft's acceleration is controlled within a safe range. For different aircraft actuator characteristics, the saturation function type of the overload limiting module can be adjusted, such as hard saturation or soft saturation, to further optimize the smoothness of control commands.

[0099] The following example of air-to-air missile interception of highly maneuverable fighter jets will be used to elaborate on the specific implementation process, parameter configuration and simulation verification results of the method of the present invention, and fully verify the effectiveness and superiority of the method.

[0100] First, the implementation scenario is set. The interception scenario is a three-dimensional airspace, the target is a highly maneuverable fighter jet capable of various maneuvering modes such as sinusoidal maneuvers, step maneuvers, and random maneuvers; the interceptor is an air-to-air missile equipped with the trajectory optimization method described in this invention, which collects status data through sensors such as radar, inertial measurement unit, and GPS to achieve precise interception of the target.

[0101] Next, the core parameters are configured. Based on the actual engineering scenario and simulation test optimization, the parameter configurations for each core module are determined as follows: Projectile-target relative motion model parameters: System state variables: line of sight azimuth, line of sight pitch, rate of change of line of sight azimuth, rate of change of line of sight pitch; Initial target distance: 5500 meters; Maximum missile acceleration: 40g (g=9.8m / s²) 2 ); Target's maximum acceleration: 2g.

[0102] Data preprocessing module parameters: Iteration step size: 10 milliseconds.

[0103] Nonlinear disturbance observer parameters: Initial internal state values: [0, 0]ᵀ; Observer gain matrix: l = diag(1, -1); Observer iteration step size: 10 milliseconds.

[0104] Adaptive dynamic programming parameters: The performance index function weight matrix is: Q=diag (10000, 100), R=0.000001I2, where I2 is a two-dimensional identity matrix.

[0105] Composite guidance and overload limiting parameters: The weighting coefficients for feedforward compensation and feedback control are both 1 (linearly superimposed). Overload limit threshold: 40g.

[0106] Interception threshold settings: If the distance between the missile and the target is less than 30m, the interception is considered successful.

[0107] Finally, the specific implementation steps are as follows: First, construct the relative motion model between the projectile and the target. According to the method described in the technical solution of this invention, based on the principle of relative kinematics, derive and construct a compact relative motion model between the projectile and the target, clarify the correlation between system state variables, nonlinear coupling terms, control input matrix, interference matrix, and target maneuver acceleration interference terms, and the model can accurately reflect the strong nonlinear coupling characteristics of the relative motion between the projectile and the target in three-dimensional space.

[0108] Secondly, state data acquisition and preprocessing are performed. The radar sensors on the missile measure the missile-target distance, line-of-sight angle, and rate of change in real time. The inertial measurement unit collects the missile's acceleration and angular velocity, and GPS obtains the absolute position information of the missile and the target. The collected raw data is input into the data preprocessing module, where noise is suppressed by the Kalman filter algorithm, and a smooth and reliable missile-target relative motion state signal (line-of-sight azimuth angle, line-of-sight elevation angle, and their rate of change) is output.

[0109] Next, interference estimation is performed. The preprocessed relative motion state signal is input to a nonlinear interference observer. The observer estimates the target maneuver acceleration interference in real time through the dynamic evolution and feedback correction of its internal state variables.

[0110] Next, the optimal guidance law is calculated. The relative motion state signal and the disturbance estimation signal are input to the adaptive dynamic programming optimal guidance law module: the value function estimation network outputs the value function estimate based on the input state signal; the network weights are dynamically adjusted through a dedicated weight update law to make the value function estimate asymptotically approach the optimal value function; the optimal feedback control signal is derived based on the optimal value function. Figure 3 This is a graph showing the rate of change of the single network weights used in the adaptive dynamic programming algorithm of the aircraft trajectory optimization method provided by this invention, such as... Figure 3 As shown, this embodiment of the invention balances convergence speed and stability by designing a dedicated weight update law.

[0111] Next, composite guidance command generation and overload limiting are performed. The interference estimation signal is used as a feedforward compensation signal and linearly superimposed with the optimal feedback control signal to generate the initial composite guidance command. The initial command is input to the overload limiting module, and after amplitude limiting processing by a saturation function, the final composite guidance command is output, ensuring that the acceleration corresponding to the command does not exceed the maximum threshold of 40g.

[0112] Finally, trajectory optimization and interception are performed. The final composite guidance command is sent to the missile's thrusters and servos, driving the missile to adjust its flight trajectory. Figure 4 This is a trajectory optimization / guidance command variation diagram of the aircraft trajectory optimization method provided by the present invention, such as... Figure 4As shown, the method provided by this embodiment of the invention greatly improves interception accuracy. During flight, the rate of change of missile-target distance and line-of-sight angle is monitored in real time. When the missile-target distance is less than 30m, the interception is determined to be successful, and guidance control is stopped. Figure 5 This is a dynamic graph showing the convergence of the line-of-sight rotation rate of the aircraft trajectory optimization method provided by this invention. Figure 6 This is a motion trajectory diagram of an aircraft intercepting a maneuvering target using the aircraft trajectory optimization method provided by this invention, such as... Figure 5 and Figure 6 As shown, during the trajectory optimization process, the combined effect of feedforward compensation and feedback control enables the line-of-sight angle change rate to converge rapidly to near zero. When the distance between the missile and the target is less than the preset threshold, the interception is considered successful.

[0113] Compared with traditional methods, the method provided in this embodiment of the invention has higher interception accuracy, lower energy consumption, and stronger robustness, and has good prospects for engineering applications.

[0114] The aircraft trajectory optimization method provided in this embodiment of the invention takes into account the actual engineering constraints of the aircraft through overload limitation processing, effectively avoids overload saturation phenomenon, and improves the engineering practicality of the method.

[0115] The apparatus provided in the embodiments of the present invention will be described below. The apparatus described below can be referred to in correspondence with the method described above.

[0116] Figure 7 This is a schematic diagram of the structure of the aircraft trajectory optimization device provided by the present invention, as shown below. Figure 7 As shown, the device includes a data acquisition module 710, a feedforward module 720, a feedback module 730, and a control module 740 connected in sequence.

[0117] The acquisition module 710 is used to acquire relative motion state data between the aircraft and the tracked target; The feedforward module 720 is used to input the relative motion state data into a nonlinear interference observer to determine interference estimation data; the nonlinear interference observer is constructed based on the relative motion model between the aircraft and the tracked target; Feedback module 730 is used to construct a value function estimation network, inputting the relative motion state data into the value function estimation network to determine feedback control data; the value function estimation network adjusts the weights of the value function estimation network using a weight update law; The control module 740 is used to determine a composite guidance command based on the interference estimation data and the feedback control data, so as to control the aircraft based on the composite guidance command.

[0118] The aircraft trajectory optimization device provided in this invention accurately characterizes the strong coupling characteristics between the line-of-sight azimuth and pitch angles by constructing a three-dimensional nonlinear relative motion model, effectively avoiding the truncation error caused by traditional linearized models and laying a model foundation for high-precision control. Through a nonlinear disturbance observer, real-time and accurate estimation and feedforward compensation of target maneuver disturbances are achieved without requiring prior information on target maneuvers or measuring target acceleration, significantly enhancing the system's robustness against large target maneuvers. Furthermore, an innovative weight update mechanism is implemented, based on an adaptive dynamic programming method to design a feedback-optimal guidance law. The optimal guidance law ensures the rapid and stable convergence of the value function estimation network, meeting the millisecond-level real-time control requirements. Through the linear superposition of feedforward and feedback, and overload processing, the interference effect is eliminated at the source, the optimality of the flight trajectory is guaranteed, and the control commands are kept within the safe range of the aircraft's physical actuators, avoiding the risk of runaway due to overload saturation. Through the synergistic effect of multiple links such as model building, interference estimation, optimal guidance law design, command fusion and trajectory control, the interception accuracy and response speed of the aircraft against highly maneuverable targets in high-dynamic and highly interference environments are significantly improved, which has extremely high engineering application value.

[0119] Figure 8 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 8 As shown, the electronic device may include a processor 810, a communications interface 820, a memory 830, and a communications bus 840, wherein the processor 810, the communications interface 820, and the memory 830 communicate with each other via the communications bus 840. The processor 810 can call logical commands stored in the memory 830 to execute the methods described in the above embodiments, for example: The system acquires relative motion state data between the aircraft and the tracked target; inputs this relative motion state data into a nonlinear interference observer to determine interference estimation data; the nonlinear interference observer is constructed based on the relative motion model between the aircraft and the tracked target; a value function estimation network is constructed, and the relative motion state data is input into the value function estimation network to determine feedback control data; the value function estimation network adjusts its weights using a weight update law; based on the interference estimation data and the feedback control data, a composite guidance command is determined to control the aircraft based on the composite guidance command.

[0120] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0121] The processor in the electronic device provided in this embodiment of the invention can call logical instructions in the memory to implement the above method. Its specific implementation method is the same as the aforementioned method implementation method and can achieve the same beneficial effects, which will not be repeated here.

[0122] This invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the methods provided in the above embodiments.

[0123] The specific implementation method is the same as the aforementioned method implementation method and can achieve the same beneficial effects, so it will not be repeated here.

[0124] This invention provides a computer program product, including a computer program that, when executed by a processor, implements the method described above.

[0125] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0126] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0127] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for optimizing aircraft trajectory, characterized in that, include: Acquire relative motion data between the aircraft and the tracked target; The relative motion state data is input into the nonlinear disturbance observer to determine the disturbance estimation data; The nonlinear interference observer is constructed based on the relative motion model between the aircraft and the tracked target; A value function estimation network is constructed, and the relative motion state data is input into the value function estimation network to determine the feedback control data; The value function estimation network uses a weight update law to adjust the weights of the value function estimation network; Based on the interference estimation data and the feedback control data, a composite guidance command is determined to control the aircraft based on the composite guidance command.

2. The aircraft trajectory optimization method according to claim 1, characterized in that, The nonlinear disturbance observer includes internal state variables and an observer gain matrix; the structure and parameters of the observer gain matrix are determined based on the Lyapunov stability criterion.

3. The aircraft trajectory optimization method according to claim 2, characterized in that, The step of inputting the relative motion state data into the nonlinear disturbance observer to determine the disturbance estimation data includes: Based on the relative motion state data, update the internal state variables in the nonlinear disturbance observer; The disturbance estimation data is output based on a linear combination of the internal state variables and the relative motion state data.

4. The aircraft trajectory optimization method according to claim 1, characterized in that, The construction of the value function estimation network involves inputting the relative motion state data into the value function estimation network to determine the feedback control data, including: The value function estimation network is constructed based on a single hidden layer neural network; The relative motion state data is input into the value function estimation network to obtain the value function estimate; Based on the estimated value of the value function, combined with the weight matrix of the control input energy consumption term in the performance index function and the control input matrix of the relative motion model, the feedback control data is determined.

5. The aircraft trajectory optimization method according to claim 1, characterized in that, The weight update law includes an error feedback term; The error feedback term is determined by the performance index function and the error function determined by the value function estimation network; The error feedback term is used to dynamically adjust the update rate of the weights in the value function estimation network.

6. The aircraft trajectory optimization method according to claim 1, characterized in that, The step of determining the composite guidance command based on the interference estimation data and the feedback control data includes: The interference estimation data and the feedback control data are linearly superimposed to generate an initial composite guidance command; Based on the preset maximum acceleration threshold of the aircraft, the initial composite guidance command is subjected to amplitude limiting processing to obtain the composite guidance command.

7. An aircraft trajectory optimization device, characterized in that, include: The data acquisition module is used to acquire relative motion state data between the aircraft and the target being tracked. The feedforward module is used to input the relative motion state data into the nonlinear interference observer to determine the interference estimation data; the nonlinear interference observer is constructed based on the relative motion model between the aircraft and the tracked target; The feedback module is used to construct a value function estimation network, and inputs the relative motion state data into the value function estimation network to determine the feedback control data; The value function estimation network uses a weight update law to adjust the weights of the value function estimation network; The control module is used to determine composite guidance commands based on the interference estimation data and the feedback control data, so as to control the aircraft based on the composite guidance commands.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the aircraft trajectory optimization method as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the aircraft trajectory optimization method as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the aircraft trajectory optimization method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Aircraft formation cooperative control method based on self-learning

    CN120335486A

  • Stratospheric airship trajectory tracking method based on reinforcement learning optimal control

    WO2024216870A1