Aircraft trajectory optimization method and device, electronic equipment and storage medium
By constructing a three-dimensional relative motion model and an adaptive dynamic programming network combined with integral sliding mode control of a semi-positive definite obstacle function, the problem of insufficient accuracy and robustness in aircraft trajectory optimization is solved, and high-precision and low-bounce trajectory optimization is achieved in the scenario of intercepting maneuvering targets.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INST OF AUTOMATION CHINESE ACAD OF SCI
- Filing Date
- 2026-01-27
- Publication Date
- 2026-05-12
AI Technical Summary
Existing linear model-based aircraft trajectory optimization methods are insufficient to meet the accuracy and robustness requirements of aircraft in intercepting maneuvering targets. Traditional adaptive dynamic programming lacks robustness, while sliding mode control suffers from chattering and has high computational complexity.
An integral sliding mode control method combining an adaptive dynamic programming network and a semi-definite obstacle function is adopted. By constructing a three-dimensional relative motion model, the first and second control laws are determined, the flight control commands are optimized to achieve trajectory optimization, the optimal control of the nominal system is achieved by using an adaptive dynamic programming network, and the semi-definite obstacle function is combined to suppress unknown disturbances.
It improves the accuracy and robustness of the aircraft trajectory, meets the requirements of intercepting maneuvering targets, reduces chattering, lowers computational complexity, and optimizes energy consumption.
Smart Images

Figure CN122018524A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of trajectory optimization technology, and in particular to a method, apparatus, electronic device, and storage medium for optimizing aircraft trajectories. Background Technology
[0002] Aircraft trajectory optimization is a key technology for enabling aircraft to accurately complete missions in complex environments, especially in scenarios involving the interception of maneuvering targets, where high precision and robustness of trajectory optimization are required.
[0003] Currently, trajectory optimization for aircraft is typically achieved through control laws based on linear models. However, this method struggles to guarantee the accuracy and robustness of the aircraft trajectory, making it difficult to meet the requirements for aircraft intercepting maneuvering targets. Summary of the Invention
[0004] This invention provides a method, apparatus, electronic device, and storage medium for optimizing aircraft trajectories, which addresses the shortcomings of existing technologies where linear model-based control laws cannot guarantee the accuracy and robustness of aircraft trajectories, and thus cannot meet the requirements of aircraft in intercepting maneuvering targets.
[0005] This invention provides a method for optimizing aircraft trajectory, comprising the following steps: Construct a three-dimensional relative motion model between the aircraft and the maneuvering target; Based on the aforementioned three-dimensional relative motion model and adaptive dynamic programming network, the first control law is determined; Based on the three-dimensional relative motion model and the first control law, the integral sliding mode variable is determined; The second control law is determined based on a positive semi-definite barrier function, wherein the positive semi-definite barrier function uses the integral sliding mode variable as the independent variable. Flight control commands are determined based on the first control law and the second control law, and the trajectory of the aircraft is optimized based on the flight control commands to obtain an optimized trajectory.
[0006] According to the present invention, a method for optimizing the trajectory of an aircraft, wherein determining a first control law based on the three-dimensional relative motion model and an adaptive dynamic programming network includes: Based on the state variable penalty term and control variable penalty term of the three-dimensional relative motion model, the optimal control performance function is determined. The Hamiltonian function is determined based on the gradient information of the optimal control performance function with respect to the state variables, as well as the dynamic terms and control input terms of the three-dimensional relative motion model. The objective function is determined based on the Hamiltonian function; The weights of the adaptive dynamic programming network are updated based on the objective function and the auxiliary function to obtain the adaptive dynamic programming update network. The first control law is determined based on the output value of the adaptive dynamic programming update network, the control penalty term, and the control input term of the three-dimensional relative motion model.
[0007] According to a method for optimizing an aircraft trajectory provided by the present invention, the step of updating the weights of the adaptive dynamic programming network based on the objective function and the auxiliary function to obtain an adaptive dynamic programming update network includes: The objective function is subjected to normalized gradient descent to obtain the first weight update term; The second weight update term is determined based on the auxiliary function; The weights of the adaptive dynamic programming network are updated in real time based on the first weight update term and the second weight update term until the adaptive dynamic programming network converges, thus obtaining the adaptive dynamic programming updated network.
[0008] According to the present invention, a method for optimizing an aircraft trajectory, wherein determining the second control law based on a positive semi-definite obstacle function includes: Substitute the real-time value of the integral sliding mode variable into the positive semidefinite barrier function to obtain the barrier function output value; Based on the output value of the obstacle function and the control input matrix of the three-dimensional relative motion model, the sliding mode control gain is determined. The second control law is determined based on the sliding mode control gain and the sign function of the integral sliding mode variable.
[0009] According to the present invention, a method for optimizing the trajectory of an aircraft, wherein determining the integral sliding mode variable based on the three-dimensional relative motion model and the first control law includes: Obtain the real-time and initial values of the state variables in the three-dimensional relative motion model; The sum of the dynamic term in the three-dimensional relative motion model and the control term of the first control law is integrated over time to obtain the dynamic integral term; The integral sliding mode variable is determined based on the real-time value of the state variable, the initial value, and the dynamic integral term.
[0010] According to the present invention, a method for optimizing the trajectory of an aircraft includes constructing a three-dimensional relative motion model between the aircraft and a maneuvering target, comprising: A three-dimensional relative motion model is constructed based on state variables and flight time, wherein the state variables are the line-of-sight tilt rate and line-of-sight deflection rate of the aircraft relative to the maneuvering target; The three-dimensional relative motion model includes dynamic terms related to the state variables, control input terms describing the normal acceleration of the aircraft, and disturbance input terms describing the normal acceleration of the maneuvering target.
[0011] According to the present invention, an aircraft trajectory optimization method is provided, wherein the auxiliary function is a Lyapunov function, and the step of determining a second weight update term based on the auxiliary function includes: The second weight update term is generated based on the gradient information of the Lyapunov function with respect to the state variable.
[0012] The present invention also provides an aircraft trajectory optimization device, comprising the following modules: The relative motion modeling module is used to construct a three-dimensional relative motion model between the aircraft and the maneuvering target; The first control law determination module is used to determine the first control law based on the three-dimensional relative motion model and the adaptive dynamic programming network. The integral sliding mode construction module is used to determine the integral sliding mode variables based on the three-dimensional relative motion model and the first control law; The second control law determination module is used to determine the second control law based on a positive semi-definite barrier function, wherein the positive semi-definite barrier function uses the integral sliding mode variable as the independent variable. The trajectory optimization module is used to determine flight control commands based on the first control law and the second control law, and to optimize the trajectory of the aircraft based on the flight control commands to obtain an optimized trajectory.
[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the aircraft trajectory optimization method as described above.
[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the aircraft trajectory optimization method as described above.
[0015] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the aircraft trajectory optimization method as described above.
[0016] The aircraft trajectory optimization method, device, electronic equipment, and storage medium provided by this invention achieve optimal control of the nominal system by utilizing an adaptive dynamic programming network, and combine it with integral sliding mode control based on a semi-positive definite obstacle function to adaptively suppress unknown disturbances, thereby improving the accuracy and robustness of the aircraft trajectory and meeting the requirements of the aircraft in intercepting maneuvering targets. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating the aircraft trajectory optimization method provided by the present invention.
[0019] Figure 2 This is a schematic diagram of the process for determining the first control law provided by the present invention.
[0020] Figure 3 This is a schematic diagram of the process for determining the adaptive dynamic programming update network provided by the present invention.
[0021] Figure 4 This is a schematic diagram of the process for determining the second control law based on a semi-positive definite barrier function provided by the present invention.
[0022] Figure 5 This is a schematic diagram of the process for determining integral sliding mode variables based on a three-dimensional relative motion model and a first control law, provided by the present invention.
[0023] Figure 6 This is a schematic diagram illustrating the principle of the aircraft trajectory optimization method provided by the present invention.
[0024] Figure 7 This is a schematic diagram of the three-dimensional motion relationship between the aircraft and the maneuvering target provided by the present invention.
[0025] Figure 8 This is a schematic diagram of the trajectory of an aircraft intercepting a maneuvering target, provided by the present invention.
[0026] Figure 9 This is a graph showing the change in line-of-sight angular velocity provided by the present invention.
[0027] Figure 10 This is a curve showing the change of the combined trajectory optimization control law provided by the present invention.
[0028] Figure 11 This is a graph showing the change in the evaluation network weights provided by the present invention.
[0029] Figure 12 This is a schematic diagram of the aircraft trajectory optimization device provided by the present invention.
[0030] Figure 13 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0032] It should be noted that, in the description of this invention, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0033] The terms "first," "second," etc., used in this invention are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more.
[0034] To facilitate a full understanding of the technical solution of this application, the following content is hereby introduced: With the rapid development of modern aircraft technology, various aircraft are increasingly widely used in national defense, civil aviation, and other fields, placing increasingly higher demands on the accuracy, robustness, and real-time performance of trajectory optimization. Especially in typical scenarios such as intercepting maneuvering targets, aircraft need to face a variety of complex factors, including unknown disturbances caused by target maneuvers, external environmental interference, and model uncertainties. Traditional trajectory optimization methods struggle to simultaneously meet the dual requirements of optimality and robustness.
[0035] Adaptive dynamic programming, as an optimal control method based on reinforcement learning, can approximate the optimal control law of nonlinear systems through iterative learning and has received widespread attention in the field of aircraft trajectory optimization. However, trajectory optimization methods based solely on adaptive dynamic programming have significant shortcomings: poor robustness, and trajectory tracking accuracy will decrease significantly in the presence of unknown disturbances, making it difficult to meet the requirements of high-precision tasks.
[0036] Sliding mode control effectively suppresses the effects of disturbances and has unique advantages in the field of robust control. However, traditional sliding mode control has inherent technical defects: its high-frequency switching characteristics can lead to system chattering, and frequent excitation of actuators not only causes severe wear on the actuators but also increases the energy consumption of the aircraft. In addition, the gain of traditional sliding mode control usually needs to be preset according to the disturbance boundary. When the disturbance boundary is unknown, overestimation of the gain will exacerbate chattering, while underestimation of the gain will fail to effectively resist disturbances, making it difficult to balance the requirements of robustness and chattering suppression.
[0037] Meanwhile, classic adaptive dynamic programming algorithms often employ a dual-network framework of evaluator and executor, requiring the construction of two independent neural networks to implement evaluation and execution functions respectively. This results in high computational complexity and stringent requirements for the real-time processing capabilities of the aircraft, making it difficult to meet the ultra-real-time dynamic performance demands such as terminal trajectory optimization. Currently, no method can simultaneously address the problems of insufficient robustness of adaptive dynamic programming, significant chattering in sliding mode control, and high computational overhead. Therefore, there is an urgent need for an aircraft trajectory optimization method that balances optimality, robustness, real-time performance, and low chattering.
[0038] The following is combined Figures 1-13 This invention describes the aircraft trajectory optimization method, apparatus, electronic device, and storage medium provided by the present invention.
[0039] Figure 1 This is a flowchart illustrating the aircraft trajectory optimization method provided by the present invention, as shown below. Figure 1 As shown, the execution entity of the aircraft trajectory optimization method provided by this invention can be an onboard computer, guidance and control system, ground control station, server, or electronic device containing a processor. Unless otherwise specified, the onboard computer of the aircraft will be used as an example in the following embodiments.
[0040] As an optional embodiment, the aircraft trajectory optimization method mainly includes, but is not limited to, the following steps: Step 110: Construct a three-dimensional relative motion model between the aircraft and the maneuvering target.
[0041] An aircraft refers to a vehicle capable of flying in the air and having trajectory control capabilities. For example, an aircraft can be a missile, a drone, an interceptor, or other aerospace vehicle that requires high-precision trajectory optimization and guidance control.
[0042] A maneuvering target refers to an object that an aircraft needs to track, intercept, or escort, and that object usually has the ability to change its own state of motion. For example, a maneuvering target can be an enemy fighter jet, ballistic missile, cruise missile, or other aerial target with unknown maneuvering strategies.
[0043] A three-dimensional relative motion model refers to a mathematical model used to describe the relative position and relative motion state changes of an aircraft and a maneuvering target in three-dimensional space. For example, a three-dimensional relative motion model can be a set of nonlinear differential equations established in the line-of-sight coordinate system, used to characterize the dynamic laws of the evolution of the line-of-sight tilt angle, line-of-sight deflection angle and their rate of change over time.
[0044] A three-dimensional relative motion model can be constructed using the vector derivative method combined with coordinate system transformation. For example, first, a line-of-sight coordinate system is established with the aircraft's center of mass as the origin, and a reference inertial coordinate system is defined. By analyzing the relative position vector and relative velocity vector between the aircraft and the target, the derivative of the relative velocity with respect to time in the inertial coordinate system is calculated using the vector derivative method. At the same time, the rotational angular velocity of the line-of-sight coordinate system relative to the reference inertial coordinate system is considered. The acceleration term in the relative motion equation is decomposed into the normal acceleration component of the aircraft and the normal acceleration component of the target. Finally, a three-dimensional relative motion model is obtained with the line-of-sight angular rate as the state variable, the aircraft acceleration as the control input, and the target acceleration as the disturbance input.
[0045] Step 120: Determine the first control law based on the three-dimensional relative motion model and the adaptive dynamic programming network.
[0046] An adaptive dynamic programming network refers to a neural network structure built based on the principles of reinforcement learning. For example, an adaptive dynamic programming network can be a single network structure containing only an evaluation network. This evaluation network uses a single hidden layer neural network to approximate the optimal performance function, using the activation function of the state variables as the basis function, and adjusts the network weights in real time through a weight update law, thereby outputting an estimate of the system performance function.
[0047] The first control law refers to the optimal control strategy designed for a nominal three-dimensional relative motion model that does not consider external unknown disturbances. It aims to achieve system state convergence and optimize energy consumption. For example, the first control law can be a nominal optimal control law, which is constructed based on the output weights of an adaptive dynamic programming network, the gradient information of the activation function, and the control input matrix of the three-dimensional relative motion model. It is used to control the normal acceleration of the aircraft under undisturbed conditions, so that the line-of-sight angular rate converges to zero.
[0048] Step 130: Determine the integral sliding mode variables based on the three-dimensional relative motion model and the first control law.
[0049] Integral sliding mode variable refers to a composite variable constructed within the integral sliding mode control framework, comprising state variables and their integral terms. This ensures that the variable lies on the sliding surface from the initial moment, thereby eliminating the approaching process of traditional sliding mode control. For example, the integral sliding mode variable can be a vector calculated from the real-time values of the state variables, the initial values of the state variables, and the dynamic integral terms of the three-dimensional relative motion model. The design of this integral sliding mode variable guarantees that it strictly follows the optimal nominal trajectory determined by the first control law when unperturbed.
[0050] Step 140: Determine the second control law based on the positive semi-definite barrier function, which uses the integral sliding mode variable as the independent variable.
[0051] A positive semi-definite barrier function is a continuous even function that is strictly increasing within a preset interval and takes the value of zero at the origin. It is used to constrain system states or error variables within a specific range. For example, a positive semi-definite barrier function can be a function defined between zero and a preset boundary value (such as 0.15). Its expression is the ratio of the absolute value of the integral sliding mode variable to the difference between the preset boundary value and that absolute value. As the integral sliding mode variable approaches the preset boundary value, the value of the positive semi-definite barrier function approaches infinity, thus forming a constraint barrier on the state.
[0052] The second control law refers to an adaptive robust control strategy designed using the characteristics of the obstacle function, aiming to cope with unknown disturbances such as the acceleration of maneuvering targets and suppress chattering. For example, the second control law can be an adaptive sliding mode control law, which is constructed by multiplying the output value of the semi-positive definite obstacle function with the inverse matrix of the control input matrix in the three-dimensional relative motion model to obtain the adaptive control gain, and then multiplying this adaptive control gain with the sign function of the integral sliding mode variable. Since the semi-positive definite obstacle function is continuous and zero at the zero point, this second control law can adaptively adjust the gain according to the magnitude of the sliding mode variable and remain continuous near the sliding surface, thereby effectively avoiding high-frequency switching of the control quantity.
[0053] Step 150: Determine flight control commands based on the first and second control laws, and optimize the trajectory of the aircraft based on the flight control commands to obtain the optimized trajectory.
[0054] Flight control commands refer to the comprehensive control signals that are ultimately applied to the aircraft's actuators. They are composed of the optimal control components for the nominal system and the robust control components for disturbance suppression. For example, a flight control command can be the vector sum of the first control law and the second control law, specifically characterized as the normal acceleration component required by the aircraft in the line-of-sight coordinate system, used to drive the aircraft's autopilot or control surface deflection to change the aircraft's flight attitude.
[0055] An optimized trajectory refers to a spatial flight path generated by an aircraft under flight control commands that meets specific performance indicators. This spatial flight path takes into account both optimal energy consumption and robustness against disturbances. For example, an optimized trajectory can be a flight path within a boundary region where the line-of-sight angular rate between the aircraft and the maneuvering target converges to near zero within a finite time. This enables the aircraft to accurately intercept the maneuvering target in a quasi-parallel approach, while ensuring a smooth flight process without significant overload maneuvers.
[0056] The aircraft trajectory optimization method provided by this invention achieves optimal control of the nominal system by utilizing an adaptive dynamic programming network, and combines it with integral sliding mode control based on a semi-positive definite obstacle function to adaptively suppress unknown disturbances, thereby improving the accuracy and robustness of the aircraft trajectory and meeting the requirements of the aircraft in intercepting maneuvering targets.
[0057] Figure 2 This is a flowchart illustrating the process of determining the first control law provided by the present invention, as shown below. Figure 2 As shown, as another optional embodiment provided by the present invention, the first control law is determined based on a three-dimensional relative motion model and an adaptive dynamic programming network, including but not limited to the following steps: Step 210: Determine the optimal control performance function based on the state variable penalty term and control variable penalty term of the three-dimensional relative motion model.
[0058] Optionally, the three-dimensional relative motion model is as shown in formula (1): (1) in, Let be the line-of-sight angle vector of the aircraft relative to the maneuvering target. The line-of-sight angular velocity vector (i.e.) (first derivative) The rate of change of the line-of-sight angular velocity (i.e. (the second derivative) This refers to the nonlinear dynamic term in the three-dimensional relative motion model. To control the input matrix, The nominal control input vector. This is the perturbation input matrix, used to describe the weights of the influence of external perturbations on the rate of change of the line-of-sight angular rate. The disturbance input vector consists of the normal acceleration components of the maneuvering target in the line-of-sight coordinate system, representing the unknown external disturbances present in the relative motion model.
[0059] Alternatively, for the case where external unknown disturbances (i.e., nominal system) are not considered, the three-dimensional relative motion model is as shown in Equation (2): (2) in, Let be the line-of-sight angle vector of the aircraft relative to the maneuvering target. The line-of-sight angular velocity vector (i.e.) (first derivative) The rate of change of the line-of-sight angular velocity (i.e. (the second derivative) This refers to the nonlinear dynamic term in the three-dimensional relative motion model. To control the input matrix, The nominal control input vector is the nominal guidance law designed based on adaptive dynamic programming.
[0060] The state variable penalty term refers to the part used to constrain the magnitude of the system's state variables. Its purpose is to make the line-of-sight angular rate converge to zero quickly, so as to achieve a quasi-parallel approach between the aircraft and the target. For example, the state variable penalty term can be expressed as the product of the transpose of the state variable vector and the state penalty matrix, and then multiplied by the state variable vector. The state penalty matrix is a positive definite symmetric matrix, and the size of its diagonal elements reflects the weight of the convergence requirements of the line-of-sight tilt rate and the line-of-sight deflection rate.
[0061] The control penalty term refers to the part used to constrain the nominal control input amplitude. Its purpose is to prevent excessive control from causing excessive energy consumption of the aircraft or saturation of the actuators. For example, the control penalty term can be expressed as the product of the transpose of the nominal control input vector and the control penalty matrix, and then multiplied by the nominal control input vector. The control penalty matrix is a positive definite symmetric matrix used to achieve a balance between tracking accuracy and energy consumption.
[0062] The optimal control performance function refers to the integral of the utility function over the infinite time domain. It is used to evaluate the cumulative cost of the system evolving from its current state to its final equilibrium state. For example, the optimal control performance function is shown in Equation (3): (3) in, The optimal control performance function is... For the current moment, Let ∞ be the integral variable, and let ∞ represent the infinite time domain. for The line-of-sight angular velocity vector at time t. Let be the state penalty matrix. T Represents the transpose of a vector. for The nominal control input vector at time t. To control the penalty matrix, This is the penalty term for the state variable. This is a penalty item for controlling quantity.
[0063] Step 220: Determine the Hamiltonian function based on the gradient information of the optimal control performance function with respect to the state variables, as well as the dynamic terms and control input terms of the three-dimensional relative motion model.
[0064] Specifically, the Hamiltonian function is constructed based on optimal control theory. It integrates the gradient information of the optimal control performance function with respect to the state variables, and incorporates the dynamic evolution law in the three-dimensional relative motion model. For example, the Hamiltonian function is shown in formula (4): (4) in, For Hamiltonian functions, The line-of-sight angular velocity vector, Let be the state penalty matrix. The nominal control input vector. T Represents the transpose of a vector. To control the penalty matrix, For state variable penalty terms, To control the amount of penalty items, The optimal control performance function with respect to the state variables gradient vector, This refers to the nonlinear dynamic term in the three-dimensional relative motion model. This is the control input matrix in the three-dimensional relative motion model.
[0065] Step 230: Determine the objective function based on the Hamiltonian function.
[0066] Specifically, the objective function is an error metric function used to train the weights of the adaptive dynamic programming network. It is usually designed to approximate the square form of the Hamiltonian function, aiming to minimize the objective function to make the system approach the optimal control law. For example, the objective function is shown in Equation (5): (5) in, Let be the objective function. The approximate Hamiltonian function is obtained by substituting the estimated performance function and its gradient from the output of the adaptive dynamic programming network into the Hamiltonian function. The line-of-sight angular velocity vector, This is the current estimated weight vector of the adaptive dynamic programming network (i.e., the evaluation network).
[0067] Step 240: Update the weights of the adaptive dynamic programming network based on the objective function and auxiliary function to obtain the adaptive dynamic programming updated network.
[0068] An adaptive dynamic programming update network refers to a neural network that iteratively corrects its internal weight parameters based on a specific weight update law at the current time. Its purpose is to make the estimated performance function output by the adaptive dynamic programming network gradually approach the true optimal performance function.
[0069] Step 250: Based on adaptive dynamic programming, update the network's output value, control penalty term, and control input term of the three-dimensional relative motion model to determine the first control law.
[0070] Specifically, the first control law is derived based on the extremum condition of the Hamiltonian function, and it uses the gradient information of the optimal performance function approximated by the adaptive dynamic programming network to construct the nominal optimal control signal.
[0071] Optionally, the first control law is as shown in formula (6): (6) in, This is the first control law. This is the inverse of the control penalty matrix, which is derived from the control quantity penalty term; This is the transpose of the control input matrix in the three-dimensional relative motion model. This is the transpose of the gradient of the activation function with respect to the state variables in an adaptive dynamic programming network. Update the output weight vector of the network for adaptive dynamic programming.
[0072] The aircraft trajectory optimization method provided by this invention constructs an optimal control performance function based on state variable penalty terms and control variable penalty terms, and uses Hamiltonian function and auxiliary function to guide the weight update of adaptive dynamic programming network. Under the premise of ensuring the convergence of network weight update and system stability, it can achieve accurate approximation of the nominal optimal control law, thereby balancing trajectory tracking accuracy and flight energy consumption optimization under undisturbed conditions.
[0073] Figure 3 This is a schematic diagram of the process for determining the adaptive dynamic programming update network provided by the present invention, as shown below. Figure 3 As shown, as another optional embodiment provided by the present invention, the weights of the adaptive dynamic programming network are updated based on the objective function and the auxiliary function to obtain an adaptive dynamic programming update network, including but not limited to the following steps: Step 310: Perform normalized gradient descent processing on the objective function to obtain the first weight update term.
[0074] Normalized gradient descent processing refers to an improved gradient descent optimization algorithm that adjusts the step size of gradient updates by introducing a normalization factor to prevent instability or divergence in weight updates when the gradient magnitude is too large. For example, normalized gradient descent processing can calculate the gradient of the square of the approximate Hamiltonian function with respect to the network weights and divide this gradient by a normalization term related to the gradient norm, thereby obtaining the update direction and magnitude used to minimize the objective function.
[0075] The first weight update term refers to the main component in the weight update law that drives the network to approach the optimal performance function. Its purpose is to minimize the Bellman error through iteration. For example, the first weight update term can be obtained based on the learning rate, the normalized gradient vector, and the approximate Hamiltonian function. This first weight update term makes the network weights evolve in the direction of reducing the Hamiltonian function residual.
[0076] Step 320: Determine the second weight update term based on the auxiliary function.
[0077] Auxiliary functions are scalar functions used to assist in analyzing system stability or constructing stability constraints. They are usually designed as radially unbounded positive definite functions. For example, the partial derivatives of auxiliary functions with respect to state variables can reflect the changing trend of system energy and are used to provide stability guarantees during the weight learning process.
[0078] The second weight update term refers to the auxiliary component in the weight update law used to ensure the stability of the closed-loop system. Its role is to prevent state divergence during the exploration of the optimal solution. For example, the second weight update term can be a term constructed based on the gradient of the auxiliary function with respect to the state variable, the Jacobian matrix of the activation function, and the model correlation matrix, combined with a switching parameter. When the evolution trend of the system state may violate the stability condition, the second weight update term intervenes in the weight update process to keep the system stable.
[0079] Step 330: Update the weights of the adaptive dynamic programming network in real time based on the first weight update term and the second weight update term until the adaptive dynamic programming network converges, thus obtaining the adaptive dynamic programming updated network.
[0080] The convergence of an adaptive dynamic programming network refers to the fact that, as time progresses and the number of iterations increases, the weight parameters of the evaluation network no longer change significantly, or the weight estimation error converges to a predetermined bounded region. For example, the convergence of an adaptive dynamic programming network is manifested by the approximate value of the Hamiltonian function approaching zero and the norm of the weight vector remaining stable. At this point, the network is an adaptive dynamic programming update network, and its output can accurately reflect the optimal performance index under the current state.
[0081] The aircraft trajectory optimization method provided by this invention constructs a first weight update term to minimize the objective function using a normalized gradient descent algorithm, and introduces a second weight update term based on an auxiliary function as a stability constraint. This method can ensure the convergence of weights and the stability of the closed-loop system of the adaptive dynamic programming network during the online learning process while guaranteeing the stability of numerical calculations. It effectively avoids weight divergence and achieves a fast and stable approximation of the optimal control law.
[0082] Figure 4 This is a flowchart illustrating the process of determining the second control law based on a positive semi-definite barrier function, as provided by the present invention. Figure 4 As shown, as another optional embodiment provided by the present invention, the second control law is determined based on the positive semi-definite barrier function, including but not limited to the following steps: Step 410: Substitute the real-time value of the integral sliding mode variable into the positive semidefinite obstacle function to obtain the output value of the obstacle function.
[0083] Specifically, the semi-definite barrier function is a continuous even function defined on a preset interval. It takes the value of zero at its zero point, and its value approaches infinity as the input variable approaches the boundary of the preset interval. It is used to constrain the integral sliding mode variable within a safe range. For example, the output value of the barrier function is shown in formula (7): (7) in, The output value of the barrier function. This represents the real-time value of the integral sliding mode variable. For the modulus of the integral sliding mode variable, The value is a preset positive constant, representing the boundary value at which the integral sliding mode variable is allowed to fluctuate. This positive semi-definite barrier function ensures that the output is small when the sliding mode variable is small, and increases sharply when the sliding mode variable approaches the boundary.
[0084] Step 420: Determine the sliding mode control gain based on the obstacle function output value and the control input matrix of the three-dimensional relative motion model.
[0085] Specifically, the sliding mode control gain is an adaptive parameter used to counteract model nonlinearity and external disturbances. Its magnitude is automatically adjusted according to the change in the output value of the barrier function, without the need to pre-estimate the upper bound of the disturbance. For example, the sliding mode control gain is shown in formula (8): (8) in, For sliding mode control gain, This is the inverse of the control input matrix in the three-dimensional relative motion model. This is the output value of the barrier function.
[0086] Step 430: Determine the second control law based on the sliding mode control gain and the sign function of the integral sliding mode variable.
[0087] Specifically, the second control law is a robust control term that uses the sliding mode control gain to adjust the system state through feedback. Due to the zero-value characteristic of the barrier function at the origin, the control law remains continuous throughout the entire domain, thereby effectively suppressing the chattering phenomenon of traditional sliding mode control. For example, the second control law is shown in formula (9): (9) in, This is the second control law, namely the adaptive robust control law. For sliding mode control gain, The sign function for the integral sliding mode variable.
[0088] It should be noted that the second control law forces the integral sliding mode variable to remain within the preset range through a negative feedback mechanism, and eventually converges to near zero.
[0089] The aircraft trajectory optimization method provided by this invention dynamically calculates the sliding mode control gain by real-time input of the integral sliding mode variable into the positive semi-definite barrier function, and constructs a second control law by combining it with the sign function. It can adaptively adjust the control strength to resist unknown disturbances without knowing the disturbance boundary by utilizing the growth characteristics of the barrier function. At the same time, it uses the continuity of the barrier function at the zero point to ensure the smooth transition of the control law, thereby effectively suppressing system chattering and reducing wear on the actuator.
[0090] Figure 5 This is a flowchart illustrating the process of determining integral sliding mode variables based on a three-dimensional relative motion model and a first control law, as provided by the present invention. Figure 5 As shown, as another optional embodiment provided by the present invention, the integral sliding mode variable is determined based on the three-dimensional relative motion model and the first control law, including but not limited to the following steps: Step 510: Obtain the real-time and initial values of the state variables in the three-dimensional relative motion model.
[0091] The real-time value of the state variable refers to the line-of-sight angular rate vector of the aircraft relative to the maneuvering target at the current moment, while the initial value refers to the line-of-sight angular rate vector at the start of the trajectory optimization process. For example, the current line-of-sight angular rate can be measured and calculated in real time by the aircraft's seeker or inertial navigation system, and the initial line-of-sight angular rate recorded in the memory can be read.
[0092] Step 520: Integrate the sum of the dynamic term and the control term of the first control law in the three-dimensional relative motion model over time to obtain the dynamic integral term.
[0093] Specifically, the dynamic integral term characterizes the theoretical evolution trajectory of the system state under nominal control, which is obtained by accumulating the sum of the nonlinear dynamic function of the three-dimensional relative motion model and the control acceleration generated by the first control law in the time domain.
[0094] Step 530: Determine the integral sliding mode variable based on the real-time value, initial value, and dynamic integral term of the state variable.
[0095] Specifically, the integral sliding mode variable is a vector constructed by subtracting the initial value from the real-time value of the state variable and then subtracting the dynamic integral term. The design of this variable ensures that the initial state of the system naturally lies on the sliding surface, thereby eliminating the approaching process of traditional sliding mode control. For example, the integral sliding mode variable is shown in formula (10): (10) in, For integral sliding mode variables, This refers to the real-time value of the state variable. The initial values for the state variables. Represents the dynamic integral term. for The nonlinear dynamic term of the three-dimensional relative motion model at time t, for The control input matrix at each time step, for The first control law at time (i.e., the nominal optimal control law), this formula (10) ensures that .
[0096] The aircraft trajectory optimization method provided by this invention constructs an integral sliding mode variable that includes initial values of state variables and dynamic integral terms, which enables the initial state of the system to naturally lie on the sliding mode surface. This eliminates the approaching process present in traditional sliding mode control, ensuring that the aircraft has robustness against unknown disturbances from the initial moment, and achieving rapid response to maneuvering targets and high-precision tracking throughout the entire process.
[0097] In another embodiment of the present invention, constructing a three-dimensional relative motion model between the aircraft and the maneuvering target includes: constructing a three-dimensional relative motion model based on state variables and flight time, wherein the state variables are the line-of-sight tilt rate and line-of-sight deflection rate of the aircraft relative to the maneuvering target; the three-dimensional relative motion model includes dynamic terms related to the state variables, control input terms describing the normal acceleration of the aircraft, and disturbance input terms describing the normal acceleration of the maneuvering target.
[0098] Specifically, this three-dimensional relative motion model is derived based on the dynamic equations in the line-of-sight coordinate system. It calculates the second derivative of the relative position vector between the aircraft and the target, and introduces the rotational angular velocity of the line-of-sight coordinate system relative to the reference inertial coordinate system. This projects the vector equations of relative motion onto the axis of the line-of-sight coordinate system, thus establishing a nonlinear mapping relationship between the line-of-sight angular acceleration and the system state, control input, and external disturbances. For example, the specific component forms of the three-dimensional relative motion model are shown in equations (11) and (12): (11) (12) in, The rate of change of the line-of-sight tilt angle. The rate of change of the line-of-sight deflection angle. The relative distance between the aircraft and the target. The rate of change of relative distance. and These are the line of sight tilt angle and the line of sight deflection angle, respectively. and These are the line-of-sight tilt rate and the line-of-sight deflection rate, respectively. and These represent the normal acceleration components of the aircraft, corresponding to the control inputs. and These are the normal acceleration components of the maneuvering target, corresponding to the disturbance input terms.
[0099] The aircraft trajectory optimization method provided by this invention establishes a nonlinear three-dimensional relative motion model with line-of-sight tilt rate and line-of-sight deflection rate as state variables, and explicitly includes aircraft control input and maneuvering target disturbance input. This model can accurately describe the dynamic geometric relationship of the interception process in three-dimensional space and precisely map the unknown disturbances caused by target maneuvering into the system model. Thus, it provides a complete mathematical basis for achieving quasi-parallel approach guidance with converged line-of-sight angular rate and subsequent high-precision anti-disturbance control design.
[0100] In another embodiment provided by the present invention, the auxiliary function is a Lyapunov function, and determining the second weight update term based on the auxiliary function includes: generating the second weight update term according to the gradient information of the Lyapunov function with respect to the state variables.
[0101] The Lyapunov function is a radially unbounded positive definite scalar function used to prove the stability of a system. Its gradient information with respect to the state variables reflects the sensitivity of the system's energy to state changes. By introducing this gradient information into the weight update law, the weight evolution direction of the evaluation network can be forced to satisfy the Lyapunov stability condition.
[0102] For example, the second weight update term is shown in formula (13): (13) in, For the second weight update item, This is a stability switching indicator function. It takes a value of 1 when the system no longer meets the preset stability degradation condition, and 0 otherwise. For stability learning rate, The Jacobian matrix of the activation function of an adaptive dynamic programming network with respect to the state variables. This is a positive definite matrix related to the three-dimensional relative motion model and the control penalty matrix. Let be the gradient vector of the Lyapunov function with respect to the state variables.
[0103] The aircraft trajectory optimization method provided by this invention utilizes the Lyapunov function as an auxiliary function and constructs a second weight update term based on its gradient information with respect to the state variables. This introduces explicit stability constraints into the online learning process of the adaptive dynamic programming network, preventing the weight update direction from deviating from the stable region. Thus, while ensuring optimal approximation, it theoretically and rigorously guarantees the eventual uniformity and boundedness of the closed-loop control system.
[0104] Figure 6 This is a schematic diagram illustrating the principle of the aircraft trajectory optimization method provided by the present invention, as shown below. Figure 6 As shown, the control architecture employs a composite control strategy, primarily consisting of an upper integral sliding mode robust guidance law module and a lower adaptive optimal guidance law module. First, the system calculates the relative motion state based on the motion information of the target aircraft and the missile (i.e., the aircraft), and simultaneously inputs this state information into both modules. In the adaptive optimal guidance law module, an evaluation neural network is used as the core approximation performance function, and the approximation error is calculated based on the network output. Simultaneously, the network weights are updated and constrained in real time using Lyapunov stability conditions to calculate the approximate optimal control (i.e., the first control law). Meanwhile, in the integral sliding mode robust guidance law module, integral sliding mode variables are constructed based on the relative motion state and model information, and robust control components (i.e., the second control law) are generated to suppress disturbances. Finally, the system superimposes the near-optimal control output by the adaptive optimal guidance law module with the robust control component output by the integral sliding mode robust guidance law module to generate a total flight control command that is applied to the missile. The updated motion state of the missile is then differentially calculated with the state of the target aircraft through the feedback loop, thus forming a complete closed-loop control system to achieve robust trajectory optimization for maneuvering targets.
[0105] Figure 7 This is a schematic diagram of the three-dimensional motion relationship between the aircraft and the maneuvering target provided by the present invention, as shown in the figure. Figure 7 As shown, in order to accurately describe the relative motion state between the aircraft and the maneuvering target, the center of mass of the aircraft is first considered as the reference point. M Establish a reference inertial coordinate system with the origin, which is defined by the solid line axis. x , y , z It is configured to provide a static reference standard; and is based on the aircraft's center of mass. M Establish a line-of-sight coordinate system with the origin as the origin, which is defined by the dashed axis. x 4 , y 4 , z 4 Composition, in which x 4 The axis always points towards the target's center of mass. T The direction. Based on this, define the line-of-sight angle. q ϵ and line of sight angle q β To characterize the rotational relationship between the line-of-sight coordinate system and the reference inertial coordinate system, where q β This indicates that the line of sight is on the reference plane (e.g., M xz Projection on a plane and x The angle between the axes, q ϵ Indicates line of sight M x4 The angle between the projection and the object, along with these two angles and their rate of change, constitute the core state variables for constructing the three-dimensional relative motion model.
[0106] As an optional embodiment, to verify the effectiveness of the proposed aircraft trajectory optimization method based on integral sliding mode control and adaptive dynamic programming, the following simulation conditions and parameters were set: In the ground coordinate system, the initial position coordinates of the maneuvering target were set to [5500, 3500, 2500] meters, and the initial velocity was set to [300, 0, 0] meters / second; the initial position coordinates of the aircraft (missile) were set to [0, 0, 0] meters, and the initial velocity was set to [1500, 0, 0] meters / second. The initial values of the evaluation network weights were set to vectors with three components of 20 each; the parameters in the adaptive dynamic programming algorithm were set as follows: the state variable penalty matrix was selected as a positive definite symmetric matrix with diagonal elements of 10000 and 100; the nominal control variable penalty matrix... R Choose a two-dimensional identity matrix multiplied by 0.000001; the learning rate in the gradient descent weight update law. α 1 The learning rate for the stability auxiliary term is set to 25. α 2Set to 0.01. The preset interval parameter of the positive semi-definite obstacle function is set to 0.15 to ensure that the integral sliding mode variable is constrained within this interval. The maximum normal acceleration of the aircraft is limited to 40g (g is the acceleration due to gravity); the normal acceleration component of the maneuvering target is set to 2g, i.e., the target is maneuvering to escape. The simulation stops when the relative distance between the aircraft and the target is less than 30 meters.
[0107] Figure 8 This is a schematic diagram of the trajectory of an aircraft intercepting a maneuvering target, as provided by the present invention. Figure 8 As shown, in the case of X (m) , Y(m) and Z(m) In a three-dimensional coordinate system defined by axes, solid curves represent the target trajectory, and dashed curves represent the missile trajectory. The simulation sets the initial position of the missile (vehicle) at the origin (0, 0, 0), the initial position of the target at (5500, 3500, 2500), and the target has an unknown maneuvering acceleration. Figure 8 As can be seen, under the control of the trajectory optimization method proposed in this invention, the missile can quickly respond from its initial position and adjust its attitude. Its flight trajectory (dashed line) is smooth and continuous, without violent jitter or large-scale maneuvering. Finally, it precisely intersects with the trajectory of the maneuvering target (solid line) in three-dimensional space, achieving successful interception of the target. This intuitively verifies that the method of this invention has excellent robustness and guidance accuracy when facing target maneuvering interference, and that the generated optimized trajectory can ensure the stability of the flight process and effectively optimize energy consumption.
[0108] Figure 9 This is a graph showing the change in line-of-sight angular velocity provided by the present invention, such as... Figure 9 As shown, the horizontal axis represents simulation time in seconds (s); the vertical axis represents the line-of-sight angle rate of change in radians per second (rad / s). The solid curve corresponds to the line-of-sight tilt rate (…). q ϵ The trajectory of the change, the dashed curve corresponds to the rate of line of sight deflection ( ) q β The trajectory of change.
[0109] from Figure 9As can be seen, at the simulation start time (0s), both the line-of-sight tilt rate and the line-of-sight deflection rate are non-zero (approximately 0.06 rad / s and -0.08 rad / s, respectively), indicating a relative rotation between the aircraft and the target in the initial stage. Under the combined control law of integral sliding mode and adaptive dynamic programming proposed in this invention, both curves show a rapid convergence trend. Around 4 to 5 seconds into the simulation, the line-of-sight tilt rate (represented by the solid line) and the line-of-sight deflection rate (represented by the dashed line) have smoothly converged to near the zero mark, and remain stably in the zero-value region for the subsequent time (5s to 8s), without significant divergence or oscillation. This phenomenon of state variables converging to zero intuitively verifies that the method of this invention can effectively resist unknown disturbances caused by target maneuvers, successfully achieving the zeroing of the line-of-sight angular rate, that is, achieving a quasi-parallel approach between the aircraft and the maneuvering target, ensuring the guidance accuracy of the terminal trajectory.
[0110] Figure 10 This is a graph showing the variation of the combined trajectory optimization control law provided by this invention, such as... Figure 10 As shown, the horizontal axis represents the simulation time in seconds (s), and the vertical axis represents the magnitude of the guidance command in meters per second squared (m / s). 2 The solid line represents the normal acceleration component of the aircraft in the direction corresponding to the line-of-sight tilt angle. aMϵ The dashed line represents the normal acceleration component of the aircraft in the direction corresponding to the line-of-sight angle. aMβ These two components together constitute the total control input driving the aircraft's motion. As can be seen from the curve trend, the guidance commands in both directions exhibit smooth and continuous changes. Throughout the simulation, the high-frequency switching of control quantities, i.e., chattering, commonly found in traditional sliding mode control, is completely avoided. This is mainly due to the continuous characteristic of the second control law based on the semi-definite obstacle function at the zero point, which enables adaptive and smooth adjustment of the control gain.
[0111] Furthermore, observation of the amplitude of the control commands reveals that although both curves show an upward trend in the initial stage of interception in order to quickly eliminate the line-of-sight angular velocity, their peak values are strictly controlled at 300 m / s. 2 The following value is significantly lower than the maximum available overload of the aircraft (e.g., 40g, approximately 392m / s), indicating that no overload saturation occurred during the entire interception process. This strongly verifies that the first control law in this invention effectively optimizes and constrains the energy consumption of the aircraft by introducing a control quantity penalty term, and can ensure the smooth operation of the actuator while meeting the requirements of high-precision interception.
[0112] Figure 11 This is a graph showing the change in network weights provided by the present invention, such as... Figure 11As shown, the horizontal axis represents the simulation time in seconds (s), and the vertical axis represents the numerical value of the evaluation network weights. Three different curves (solid, dashed, and dotted lines) represent the dynamic evolution of different components in the evaluation network weight vector. It can be clearly seen that at the initial moment of the simulation (t=0s), all three curves start from a preset initial value of 20. As the simulation progresses, driven by the weight update law which includes a normalized gradient descent term and a Lyapunov stability auxiliary term, each weight component is rapidly adjusted online based on feedback information from the system state. Approximately between 3 and 4 seconds into the simulation, the dramatic changes in the three curves gradually subside, converging and stabilizing near different constant values (e.g., the dashed line converges to approximately 31, the solid line to approximately 24, and the dotted line to approximately 12), and remain horizontal straight lines in the subsequent time, without significant fluctuations. The characteristic that the weights converge to a steady state quickly within a finite time intuitively demonstrates that the single-network structure adopted in this invention has higher computational efficiency and faster online learning capability compared to the traditional dual-network architecture. At the same time, it verifies that the weight estimation error is eventually uniformly bounded, thereby ensuring that the stability and real-time requirements of the control system are met when the aircraft performs trajectory optimization.
[0113] Figure 12 This is a schematic diagram of the structure of the aircraft trajectory optimization device provided by the present invention, as shown below. Figure 12 As shown, it mainly includes, but is not limited to: The relative motion modeling module 1210 is used to construct a three-dimensional relative motion model between the aircraft and the maneuvering target.
[0114] The first control law determination module 1220 is used to determine the first control law based on the three-dimensional relative motion model and the adaptive dynamic programming network. The integral sliding mode construction module 1230 is used to determine the integral sliding mode variables based on the three-dimensional relative motion model and the first control law; The second control law determination module 1240 is used to determine the second control law based on a positive semi-definite barrier function, wherein the positive semi-definite barrier function uses an integral sliding mode variable as the independent variable. The trajectory optimization module 1250 is used to determine flight control commands based on the first control law and the second control law, and to optimize the trajectory of the aircraft based on the flight control commands to obtain an optimized trajectory.
[0115] It should be noted that the aircraft trajectory optimization device provided by the present invention can execute the aircraft trajectory optimization method described in any of the above embodiments during specific operation, which will not be elaborated in this embodiment.
[0116] The aircraft trajectory optimization device provided by this invention achieves optimal control of the nominal system by utilizing an adaptive dynamic programming network, and combines it with integral sliding mode control based on a semi-positive definite obstacle function to adaptively suppress unknown disturbances, thereby improving the accuracy and robustness of the aircraft trajectory and meeting the requirements of the aircraft in the scenario of intercepting maneuvering targets.
[0117] Figure 13 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 13 As shown, the electronic device may include a processor 1310, a communications interface 1320, a memory 1330, and a communication bus 1340. The processor 1310, communications interface 1320, and memory 1330 communicate with each other via the communication bus 1340. The processor 1310 can call logical instructions from the memory 1330 to execute an aircraft trajectory optimization method. This method includes: constructing a three-dimensional relative motion model between the aircraft and the maneuvering target; determining a first control law based on the three-dimensional relative motion model and an adaptive dynamic programming network; determining integral sliding mode variables based on the three-dimensional relative motion model and the first control law; determining a second control law based on a positive semi-definite obstacle function, where the positive semi-definite obstacle function uses the integral sliding mode variables as independent variables; determining flight control commands based on the first and second control laws; and optimizing the aircraft trajectory based on the flight control commands to obtain an optimized trajectory.
[0118] Furthermore, the logical instructions in the aforementioned memory 1330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0119] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the aircraft trajectory optimization method provided by the above methods. The method includes: constructing a three-dimensional relative motion model of the aircraft and the maneuvering target; determining a first control law based on the three-dimensional relative motion model and an adaptive dynamic programming network; determining an integral sliding mode variable based on the three-dimensional relative motion model and the first control law; determining a second control law based on a positive semi-definite obstacle function, wherein the positive semi-definite obstacle function uses the integral sliding mode variable as its independent variable; determining flight control commands based on the first and second control laws; and optimizing the trajectory of the aircraft based on the flight control commands to obtain an optimized trajectory.
[0120] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the aircraft trajectory optimization method provided by the above methods. The method includes: constructing a three-dimensional relative motion model of the aircraft and a maneuvering target; determining a first control law based on the three-dimensional relative motion model and an adaptive dynamic programming network; determining an integral sliding mode variable based on the three-dimensional relative motion model and the first control law; determining a second control law based on a positive semi-definite obstacle function, wherein the positive semi-definite obstacle function uses the integral sliding mode variable as its independent variable; determining flight control commands based on the first and second control laws; and optimizing the aircraft trajectory based on the flight control commands to obtain an optimized trajectory.
[0121] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0122] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for optimizing aircraft trajectory, characterized in that, include: Construct a three-dimensional relative motion model between the aircraft and the maneuvering target; Based on the aforementioned three-dimensional relative motion model and adaptive dynamic programming network, the first control law is determined; Based on the three-dimensional relative motion model and the first control law, the integral sliding mode variable is determined; The second control law is determined based on a positive semi-definite barrier function, wherein the positive semi-definite barrier function uses the integral sliding mode variable as the independent variable. Flight control commands are determined based on the first control law and the second control law, and the trajectory of the aircraft is optimized based on the flight control commands to obtain an optimized trajectory.
2. The aircraft trajectory optimization method according to claim 1, characterized in that, The determination of the first control law based on the three-dimensional relative motion model and the adaptive dynamic programming network includes: Based on the state variable penalty term and control variable penalty term of the three-dimensional relative motion model, the optimal control performance function is determined. The Hamiltonian function is determined based on the gradient information of the optimal control performance function with respect to the state variables, as well as the dynamic terms and control input terms of the three-dimensional relative motion model. The objective function is determined based on the Hamiltonian function; The weights of the adaptive dynamic programming network are updated based on the objective function and the auxiliary function to obtain the adaptive dynamic programming update network. The first control law is determined based on the output value of the adaptive dynamic programming update network, the control penalty term, and the control input term of the three-dimensional relative motion model.
3. The aircraft trajectory optimization method according to claim 2, characterized in that, The step of updating the weights of the adaptive dynamic programming network based on the objective function and auxiliary function to obtain an adaptive dynamic programming update network includes: The objective function is subjected to normalized gradient descent to obtain the first weight update term; The second weight update term is determined based on the auxiliary function; The weights of the adaptive dynamic programming network are updated in real time based on the first weight update term and the second weight update term until the adaptive dynamic programming network converges, thus obtaining the adaptive dynamic programming updated network.
4. The aircraft trajectory optimization method according to claim 1, characterized in that, The determination of the second control law based on the positive semi-definite barrier function includes: Substitute the real-time value of the integral sliding mode variable into the positive semidefinite barrier function to obtain the barrier function output value; Based on the output value of the obstacle function and the control input matrix of the three-dimensional relative motion model, the sliding mode control gain is determined. The second control law is determined based on the sliding mode control gain and the sign function of the integral sliding mode variable.
5. The method according to claim 1, characterized in that, The determination of integral sliding mode variables based on the three-dimensional relative motion model and the first control law includes: Obtain the real-time and initial values of the state variables in the three-dimensional relative motion model; The sum of the dynamic term in the three-dimensional relative motion model and the control term of the first control law is integrated over time to obtain the dynamic integral term; The integral sliding mode variable is determined based on the real-time value of the state variable, the initial value, and the dynamic integral term.
6. The method according to claim 1, characterized in that, The construction of the three-dimensional relative motion model between the aircraft and the maneuvering target includes: A three-dimensional relative motion model is constructed based on state variables and flight time, wherein the state variables are the line-of-sight tilt rate and line-of-sight deflection rate of the aircraft relative to the maneuvering target; The three-dimensional relative motion model includes dynamic terms related to the state variables, control input terms describing the normal acceleration of the aircraft, and disturbance input terms describing the normal acceleration of the maneuvering target.
7. The method according to claim 3, characterized in that, The auxiliary function is a Lyapunov function, and the step of determining the second weight update term based on the auxiliary function includes: The second weight update term is generated based on the gradient information of the Lyapunov function with respect to the state variable.
8. An aircraft trajectory optimization device, characterized in that, include: The relative motion modeling module is used to construct a three-dimensional relative motion model between the aircraft and the maneuvering target; The first control law determination module is used to determine the first control law based on the three-dimensional relative motion model and the adaptive dynamic programming network. The integral sliding mode construction module is used to determine the integral sliding mode variables based on the three-dimensional relative motion model and the first control law; The second control law determination module is used to determine the second control law based on a positive semi-definite barrier function, wherein the positive semi-definite barrier function uses the integral sliding mode variable as the independent variable. The trajectory optimization module is used to determine flight control commands based on the first control law and the second control law, and to optimize the trajectory of the aircraft based on the flight control commands to obtain an optimized trajectory.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the aircraft trajectory optimization method as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the aircraft trajectory optimization method as described in any one of claims 1 to 7.