Tethered flight vehicle maneuver planning method based on dynamic modeling of maneuverable region
Patent Information
- Application Number
- CN202610915364.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-24
- Publication Date
- 2026-08-18
- Estimated Expiration
- 2046-06-24
AI Technical Summary
[0004]本申请提供一种基于可机动区域动态建模的绳系拖曳飞行器机动规划方法,旨在解决绳系拖曳飞行器在复杂动态风场中难以自主规避高代价区域、安全机动规划困难的问题
[0004] This application provides a maneuver planning method for tethered aircraft based on dynamic modeling of maneuverable areas, aiming to solve the problems of tethered aircraft having difficulty autonomously avoiding high-cost areas and the difficulty in safe maneuver planning in complex dynamic wind fields.
Smart Images

Figure CN122431416B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of aircraft control technology, specifically to a maneuver planning method for tethered towed aircraft based on dynamic modeling of maneuverable areas. Background Technology
[0002] Tethered aircraft are a type of specialized aircraft that coordinates flight with a main aircraft via a towed cable. With the increasing complexity of aviation missions, tethered aircraft are being used in various complex mission scenarios, such as soft-air refueling, tethered decoy munitions, and towed airborne recovery. Taking autonomous soft-air refueling missions as an example, the controllable refueling drogue needs to actively maneuver to dock with the refueling receiver.
[0003] However, tethered aircraft face multiple complex constraints during actual autonomous maneuvering: the complex flow field disturbances generated by the wake of the main aircraft can cause uncertainties in the attitude of the tethered aircraft, disrupting its motion stability; the actuators of the tethered aircraft themselves have obvious capability boundaries, and their output torque and range of motion are limited by physical structure and performance, making it difficult to achieve precise tracking of arbitrary trajectories; at the same time, the dynamic stability of the tethered aircraft is easily affected by changes in its position and the coupling with the external environment, further increasing the difficulty of autonomous maneuvering control. The combination of these factors poses a severe challenge to the autonomous and safe maneuvering control of tethered aircraft. Currently, related technologies lack modeling of the safe maneuvering area of tethered aircraft in high-dynamic environments, as well as safe maneuvering planning for tethered aircraft in dynamic environments. Therefore, safe maneuvering trajectory planning for tethered aircraft is an urgent technical problem to be solved. Summary of the Invention
[0004] This application provides a maneuver planning method for tethered aircraft based on dynamic modeling of maneuverable areas, aiming to solve the problems of tethered aircraft having difficulty autonomously avoiding high-cost areas and the difficulty in safe maneuver planning in complex dynamic wind fields.
[0005] To achieve the above objectives, this application provides the following technical solution: In a first aspect, embodiments of this application provide a maneuver planning method for tethered towed aircraft based on dynamic modeling of maneuverable areas, including: Acquire the real-time position and speed of the tethered aircraft, as well as the gust-induced speed; The nominal cost map is dynamically corrected based on the gust-induced velocity to obtain the comprehensive cost map. The nominal cost map is a pre-constructed comprehensive impact of each location in space on the maneuverability and safety of the tethered aircraft without real-time wind field correction. The boundary of the high-cost area is determined based on the comprehensive cost map, and the current safe distance threshold is determined based on the real-time speed. A barrier function is then constructed based on the boundary of the high-cost area and the current safe distance threshold. Flight planning indicators are determined based on the tracking error between the real-time position and the target position, control energy consumption, the cost of the comprehensive cost map, and the obstacle function; among them, control energy consumption is used to characterize the energy consumption of the actuator output of the tethered aircraft. The control commands are solved with the goal of minimizing the flight planning index, resulting in target control commands. These target control commands are used to control the tethered aircraft to perform high-cost area avoidance and safe maneuver trajectory tracking.
[0006] In some embodiments of this application, the method further includes: The operational airspace of the tethered towed aircraft is discretized into a grid using the three-dimensional voxel grid method, resulting in airspace grid points. For each airspace grid point, determine the wake vortex-induced velocity field of the main aircraft, the boundary value of the control capability of the tethered aircraft, and the dynamic stability value. The nominal cost of each spatial grid point is obtained by fusing the wake vortex-induced velocity field, the boundary value of control capability, and the dynamic stability value to form a nominal cost map.
[0007] In some embodiments of this application, determining the wake vortex-induced velocity field of the main aircraft, the control capability boundary value and dynamic stability value of the tethered aircraft includes: The tail vortex induced velocity of the main aircraft is determined based on the weight, flight speed, air density, and fuselage length of the main aircraft, and the tail vortex induced velocity field is obtained. The boundary value of control capability is determined based on the actuator control output amplitude required for the tethered aircraft to maintain stability at the airspace grid points. The dynamic stability cost is determined based on the lateral deviation between the position of the tethered aircraft and the nominal equilibrium position of the same tether length.
[0008] In some embodiments of this application, the nominal cost map is dynamically corrected based on the gust-induced velocity to obtain a comprehensive cost map, including: The composite induced velocity is obtained by adding the wake vortex induced velocity field and the gust induced velocity. The dynamic flow field induced velocity is obtained by adding the synthetic induced velocity and the equivalent induced velocity of atmospheric turbulence; where the equivalent induced velocity of atmospheric turbulence is the product of a preset ratio and the synthetic induced velocity. The dynamic flow field induced velocity is truncated and normalized to obtain the dynamic flow field cost of each spatial grid point. The nominal cost of each spatial grid point is added to the cost of the dynamic flow field to obtain the comprehensive cost map.
[0009] In some embodiments of this application, determining the boundary of a high-cost region based on a comprehensive cost map includes: The spatial region formed by the first grid point in the comprehensive cost map whose cost exceeds a preset threshold is identified as a high-cost region. Within the lateral plane where the tethered aircraft is currently located, a local search algorithm is used to search and determine the boundary of the high-cost region.
[0010] In some embodiments of this application, a current safe distance threshold is determined based on real-time velocity to construct a barrier function based on the high-cost region boundary and the current safe distance threshold, including: The current safe distance threshold is determined based on the lateral velocity component in the real-time velocity; where the current safe distance threshold is positively correlated with the lateral velocity component. The minimum distance between the real-time location and the boundary is taken as the safe distance; The obstacle function is determined based on the safe distance and the current safe distance threshold.
[0011] In some embodiments of this application, flight planning indicators are determined based on the tracking error between the real-time location and the target location, control energy consumption, the cost of the comprehensive cost map, and the obstacle function, including: The weighting values are determined based on tracking error, control energy consumption, the cost of the comprehensive cost map, the obstacle function, and their respective weights. The weighted values are summed, and the sum is integrated over time to obtain the flight planning indicators.
[0012] In some embodiments of this application, the control commands are solved with the goal of minimizing flight planning indicators to obtain target control commands, including: With the minimum flight planning index as the optimization objective, the control commands are solved by a preset dynamic programming solver to obtain the target control commands; The preset dynamic programming solver includes a first network and a second network; the first network is used to evaluate the value estimate corresponding to the candidate control command output by the second network, and the second network is used to update the candidate control command based on the value estimate output by the first network.
[0013] In some embodiments of this application, the method further includes: The Hamiltonian function is calculated based on the tracking error, candidate control commands, the cost of the integrated cost map, and the obstacle function. The Hamiltonian Jacobi Bellman equation residuals are calculated based on the Hamiltonian function, and the loss function of the first network is determined based on the Hamiltonian Jacobi Bellman equation residuals, so as to update the first network using the loss function of the first network. The loss function of the second network is determined based on the residual of the candidate control command relative to the first derivative of the Hamiltonian function, so as to update the second network using the loss function of the second network.
[0014] Secondly, embodiments of this application provide an electronic device, including a processor and a memory storing processor-executable instructions; when the instructions are executed by the processor, a tethered towed aircraft maneuver planning method based on dynamic modeling of maneuverable areas is implemented. Attached Figure Description
[0015] To more intuitively illustrate the prior art and this application, several exemplary figures are provided below. It should be understood that the specific shapes and structures shown in the figures should not generally be regarded as limiting conditions for implementing this application; for example, based on the technical concept disclosed in this application and the exemplary figures, those skilled in the art are able to easily make conventional adjustments or further optimizations to the addition / reduction / classification, specific shapes, positional relationships, connection methods, size ratios, etc. of certain units (components).
[0016] Figure 1 A schematic diagram illustrating the implementation process of the tethered towed aircraft maneuver planning method based on dynamic modeling of maneuverable areas provided in this application embodiment; Figure 2 A schematic diagram illustrating the implementation framework of the tethered towed aircraft maneuver planning scheme provided in this application embodiment; Figure 3 A schematic diagram of Actor-Critic dual-network adaptive dynamic programming provided in an embodiment of this application; Figure 4 A schematic diagram showing the comparison of position errors during the maneuvering process provided in the embodiments of this application; Figure 5 A schematic diagram illustrating the comparison of maneuver process cost values provided in the embodiments of this application; Figure 6 This is a schematic diagram of the composition structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. Any combination of different embodiments is possible.
[0018] In the description of this application: unless otherwise stated, "a plurality of" means two or more. The terms "first," "second," "third," etc., in this application are intended to distinguish the objects referred to and do not have any special meaning in terms of technical connotation (e.g., they should not be construed as an emphasis on importance or order). Expressions such as "including," "comprising," and "having" also mean "not limited to" (certain units, components, materials, steps, etc.).
[0019] This application provides a method for planning the maneuverability of tethered aircraft based on dynamic modeling of the maneuverable region, aiming to solve the problem of safe maneuver trajectory planning for tethered aircraft. For example... Figure 1 As shown, the maneuver planning method for tethered aircraft based on dynamic modeling of maneuverable areas may include the following steps: Step 101: Obtain the real-time position and speed of the tethered aircraft, as well as the gust-induced speed.
[0020] In embodiments of this application, the electronic device can acquire the real-time position and speed of the tethered aircraft, as well as the gust-induced speed.
[0021] In the embodiments of this application, a tethered towed aircraft refers to an aircraft that is connected to a main aircraft by a tow rope and can move relative to it within a certain space, such as a controllable refueling drogue.
[0022] In the embodiments of this application, the main aircraft refers to the aircraft in front of the tethered aircraft, and its wake vortex field has a significant impact on the tethered aircraft.
[0023] In the embodiments of this application, the gust-induced velocity can be the lateral gust velocity component that is measured or estimated in real time by the main aircraft's onboard atmospheric data system.
[0024] In some embodiments of this application, the three-dimensional position and velocity of the tethered aircraft can be acquired in real time using a combination of airborne global navigation satellite system / inertial navigation system or other positioning sensors.
[0025] In some embodiments of this application, the lateral gust velocity component can be measured or estimated using the main aircraft's atmospheric data system, serving as the gust-induced velocity. This data is typically transmitted wirelessly to the tethered aircraft's onboard computer.
[0026] Step 102: Dynamically correct the nominal cost map based on the gust-induced velocity to obtain a comprehensive cost map; wherein, the nominal cost map is a pre-constructed comprehensive impact of each spatial location on the maneuverability and safety of the tethered aircraft without real-time wind field correction.
[0027] In the embodiments of this application, the electronic device can obtain the real-time position and speed of the tethered aircraft, as well as the gust-induced speed, and then dynamically correct the nominal cost map based on the gust-induced speed to obtain a comprehensive cost map. The nominal cost map is a pre-constructed comprehensive impact of each spatial position on the maneuverability and safety of the tethered aircraft without real-time wind field correction.
[0028] In the embodiments of this application, the nominal cost map can be understood as an offline pre-built three-dimensional voxel grid map, where each voxel (grid point) stores a nominal cost value, which reflects the impact of the spatial location on the maneuverability and safety of the tethered aircraft (including wake vortex disturbance, control capability margin, and dynamic stability) without real-time wind field correction.
[0029] In the embodiments of this application, the comprehensive cost map is an online real-time corrected cost map, that is, on the basis of the nominal cost map, the dynamic flow field induced velocity cost (obtained by conservative estimation from real-time gusts and atmospheric turbulence) is superimposed to reflect the actual cost distribution under the current environment.
[0030] Understandably, since the nominal cost map is built offline in advance, the nominal cost needs to be corrected during online flight due to changes in the actual flow field caused by gusts and atmospheric turbulence.
[0031] In some embodiments of this application, the electronic device can discretize the operating airspace of the tethered aircraft into a grid using a three-dimensional voxel grid method to obtain airspace grid points; for each airspace grid point, the wake vortex induced velocity field of the main aircraft, the control capability boundary value and dynamic stability value of the tethered aircraft are determined; the wake vortex induced velocity field, the control capability boundary value and the dynamic stability value are fused to obtain the nominal value of each airspace grid point to form a nominal cost map.
[0032] In some embodiments of this application, when discretizing the operating airspace of a tethered aircraft using a three-dimensional voxel grid method, a typical airspace below and behind the main aircraft can be selected (e.g., within 200m behind the main aircraft, within 50m below, and ±20m laterally). This three-dimensional space is uniformly divided into cubic voxels with side lengths of 0.5m or 1m, with each voxel's center serving as a spatial grid point.
[0033] In some embodiments of this application, when determining the vortex-induced velocity field of the main aircraft, the control capability boundary value and dynamic stability value of the tethered aircraft, the electronic device can determine the vortex-induced velocity of the main aircraft based on the weight, flight speed, air density and fuselage length of the main aircraft, thus obtaining the vortex-induced velocity field; determine the control capability boundary value based on the actuator control output amplitude required for the tethered aircraft to maintain stability at the airspace grid point; and determine the dynamic stability value based on the lateral deviation distance between the position of the tethered aircraft and the nominal equilibrium position with the same rope length.
[0034] In some embodiments of this application, the wake vortex induced velocity field can be calculated using the Hallock-Burnham wake vortex model, which calculates the tangential induced velocity (vector) generated by the two vortex cores on the left and right sides of the main aircraft at the grid point, and decomposes it into horizontal (y) and vertical (z) components.
[0035] For example, the Hallock-Burnham model is used to calculate the induced velocity generated by the wake vortex. The formula for calculating the tangential induced velocity of a single vortex is as follows: (1); in, For tangential induced velocity, Vortex ring quantity is used to characterize vortex intensity; The radial distance from a spatial point to the vortex core. Vortex core radius. To calculate the above variables, the weight of the main aircraft must be given. Flight speed air density and fuselage length The formula for calculating the vortex ring quantity is as follows: (2); Among them, the initial vortex ring quantity , Here is the dissipation coefficient. This is the reference time for vortex dissipation.
[0036] The induced velocity of the main aircraft's wake vortex includes the tangential induced velocity generated by the left and right vortex cores, i.e., the total induced velocity is: (3); Decompose it into horizontal (y) and vertical (z) directions, as shown in the following formulas: (4); (5); (6); (7); in, The horizontal component (y-axis direction, unit: m / s) of the tangential induced velocity generated by the left vortex nucleus at this spatial point. , ) are the coordinates of a point in space, ( , ) is the center coordinate of the left vortex core, ( , () represents the center coordinates of the right vortex core. The vertical component of the tangential induced velocity generated by the left vortex core at the target point is denoted as . The horizontal (y-direction) induced velocity component generated by the right vortex core. The vertical (z-direction) induced velocity component generated by the right vortex core.
[0037] In some embodiments of this application, the control capability boundary value is obtained by calculating the additional control force (total aerodynamic force minus static aerodynamic force) required to stably hover the tethered aircraft at that grid point. A larger force amplitude indicates a smaller remaining control margin and a higher cost. The control capability boundary value is obtained after truncation and normalization.
[0038] For example, the control capability boundary value obtained after truncation and normalization can be expressed as the following formula: (8); in, , The boundary value of the functional load control capability in the horizontal and vertical directions are respectively. The additional control force amplitude required for a tethered aircraft to maintain stable hovering in the horizontal direction (y-axis) can be obtained by subtracting the static aerodynamic force from the total aerodynamic force. The additional control force amplitude required to maintain stability in the vertical direction (z-axis); This represents the maximum permissible output amplitude of the actuator.
[0039] In some embodiments of this application, the dynamic stability cost is obtained by calculating the lateral deviation distance between the grid point and the nominal equilibrium position with the same rope length (the tethered aircraft is hovering directly below and behind the main aircraft and within the longitudinal symmetry plane). The larger the deviation, the worse the dynamic stability and the higher the cost. The dynamic stability cost is obtained after truncation and normalization.
[0040] For example, the lateral deviation distance between a grid point and its nominal equilibrium position with the same rope length can be expressed as: The dynamic stability cost, obtained by truncating and normalizing this term, is expressed by the following formula: (9); in, This indicates the horizontal (y-axis) deviation of the current grid point from the nominal equilibrium position. This indicates the vertical (z-axis) deviation of the current grid point from the nominal equilibrium position. This indicates the maximum permissible deviation distance (truncation threshold).
[0041] In some embodiments of this application, when the electronic device dynamically corrects the nominal cost map based on the gust-induced velocity to obtain a comprehensive cost map, it can add the wake vortex-induced velocity field and the gust-induced velocity to obtain the composite induced velocity; add the composite induced velocity and the equivalent induced velocity of atmospheric turbulence to obtain the dynamic flow field induced velocity; wherein, the equivalent induced velocity of atmospheric turbulence is the product of a preset ratio and the composite induced velocity; the dynamic flow field induced velocity is truncated and normalized to obtain the dynamic flow field cost value of each spatial grid point; and the nominal cost value of each spatial grid point is added to the dynamic flow field cost value to obtain the comprehensive cost map.
[0042] For example, the synthesis induction rate can be expressed as the following formula: (10); in, For the wake vortex induced velocity field, This refers to the gust-induced speed.
[0043] The induced velocity of the dynamic flow field can be expressed by the following formula: (11); in, This is a preset ratio.
[0044] After truncating and normalizing the induced velocity of the dynamic flow field, the cost of the dynamic flow field can be expressed by the following formula: (12); in, The cutoff threshold represents the maximum allowable dynamic flow field induced velocity amplitude (or maximum cost). This represents the truncation function. This represents the normalization function.
[0045] In the embodiments of this application, the specific value of the preset ratio is not limited, for example, it can be 20%.
[0046] Step 103: Determine the boundary of the high-cost area based on the comprehensive cost map, and determine the current safe distance threshold based on the real-time speed, so as to construct a barrier function based on the boundary of the high-cost area and the current safe distance threshold.
[0047] In the embodiments of this application, the electronic device can dynamically correct the nominal cost map according to the gust-induced speed to obtain a comprehensive cost map, determine the boundary of the high cost area according to the comprehensive cost map, and determine the current safe distance threshold based on the real-time speed, so as to construct a barrier function based on the boundary of the high cost area and the current safe distance threshold.
[0048] In some embodiments of this application, when determining the boundary of a high-cost region based on a comprehensive cost map, the electronic device can identify the spatial region formed by the first grid points in the comprehensive cost map whose cost is higher than a preset threshold as a high-cost region; and within the lateral plane where the tethered aircraft is currently located, a local search algorithm is used to search and determine the boundary of the high-cost region.
[0049] In the embodiments of this application, the specific value of the preset threshold is not limited, for example, it can be 0.5.
[0050] It is understandable that the boundary of the high-cost region is the dividing line between the high-cost region and the non-high-cost region.
[0051] For example, the boundary of a region with a cost value higher than 0.5 is found in the comprehensive cost map and is used as the boundary of a high-cost region.
[0052] In some embodiments of this application, when an electronic device determines a current safe distance threshold based on real-time speed and constructs an obstacle function based on the high-cost region boundary and the current safe distance threshold, it can determine the current safe distance threshold according to the lateral velocity component in the real-time speed; wherein, the current safe distance threshold is positively correlated with the lateral velocity component; the minimum distance between the real-time position and the boundary is taken as the safe distance; and the obstacle function is determined according to the safe distance and the current safe distance threshold.
[0053] In the embodiments of this application, the obstacle function tends to infinity when the safe distance approaches the current safe distance threshold, thereby creating a "hard constraint" effect in the optimization.
[0054] In the embodiments of this application, a reciprocal obstacle function is constructed using the safe distance and the current safe distance threshold, such that when the safe distance metric approaches the current safe distance threshold, the obstacle function value increases sharply, thereby "rejecting" the aircraft from approaching the high-cost region in subsequent optimization.
[0055] For example, the minimum distance between the current position of the tethered aircraft and the boundary of an area with a cost value greater than 0.5, as a safety distance, can be expressed by the following formula: (13); in,( , , ) represents the real-time three-dimensional position coordinates of the tethered aircraft. Vertically, In the horizontal direction, In the vertical direction. , , () represents a point located on the boundary of a high-cost region. Specifically, these points satisfy the following conditions: the cost value of the comprehensive cost map is exactly equal to 0.5 (a preset threshold), and their x-coordinate is perpendicular to the tethered towed aircraft. Same, meaning only in the current longitudinal position of the aircraft. Search for boundary points within the lateral plane (yz plane).
[0056] For example, the barrier function can be expressed as the following formula: (14); in, This represents the current safe distance threshold.
[0057] For example, the current safe distance threshold is a dynamic form that adjusts in real time with the lateral movement speed, and can be expressed as the following formula: (15); in, For the nominal safe distance, As a dynamic adjustment factor, and These represent the current velocity components of the tethered aircraft in the horizontal (y-axis) and vertical (z-axis) directions, respectively. This represents the magnitude of the lateral velocity.
[0058] Step 104: Determine flight planning indicators based on the tracking error between the real-time position and the target position, control energy consumption, the cost of the comprehensive cost map, and the obstacle function; among which, control energy consumption is used to characterize the energy consumption of the actuator output of the tethered aircraft.
[0059] In the embodiments of this application, the electronic device can determine the boundary of the high-cost area based on the comprehensive cost map and determine the current safe distance threshold based on the real-time speed. After constructing an obstacle function based on the boundary of the high-cost area and the current safe distance threshold, the flight planning index is determined based on the tracking error between the real-time position and the target position, the control energy consumption, the cost value of the comprehensive cost map, and the obstacle function. Among them, the control energy consumption is used to characterize the energy consumption of the actuator output of the tethered aircraft.
[0060] In the embodiments of this application, the tracking error can be a three-dimensional spatial position deviation vector between the real-time position of the tethered aircraft and the desired target position.
[0061] In the embodiments of this application, control energy consumption can be the sum of squares of control commands, used to characterize the energy consumption of the actuator output.
[0062] In the embodiments of this application, the flight planning index is a performance index that integrates the weighted sum of four factors—tracking error, control energy consumption, cost map, and obstacle function—over time, and serves as the optimization objective for the optimal control problem.
[0063] In some embodiments of this application, when an electronic device determines flight planning indicators based on the tracking error between the real-time position and the target position, control energy consumption, the cost value of the comprehensive cost map, and the obstacle function, it can determine weighted values based on the tracking error, control energy consumption, the cost value of the comprehensive cost map, the obstacle function, and their respective weights; sum the weighted values, and perform time integration on the summation result to obtain the flight planning indicators.
[0064] For example, performance metrics can be expressed as the following formula: (16); in, To track the error vector, which is typically a three-dimensional position deviation, its quadratic form is... This represents the sum of squares of the errors. Weights for tracking errors; To control the command vector (such as control surface deflection, thrust, etc.), its quadratic form This indicates the control of energy consumption. To control the weight of energy consumption items; To represent the cost of the comprehensive cost map, a square form is used here. To emphasize punishment in high-cost areas, The weight of the substitution item; The barrier function value is expressed in squared form. The penalty gradient can be further increased when approaching high-cost regions. The weights of the barrier function terms; The variable for integration is time.
[0065] Step 105: Solve for the control commands with the goal of minimizing the flight planning index to obtain the target control commands; among which, the target control commands are used to control the tethered aircraft to perform high-cost area avoidance and safe maneuver trajectory tracking.
[0066] In the embodiments of this application, the electronic device can determine the flight planning index based on the tracking error between the real-time position and the target position, the control energy consumption, the cost value of the comprehensive cost map, and the obstacle function. Then, it can solve for the control command with the goal of minimizing the flight planning index to obtain the target control command. The target control command is used to control the tethered aircraft to perform high-cost area avoidance and safe maneuver trajectory tracking.
[0067] In the embodiments of this application, solving for control commands with the goal of minimizing flight planning indicators is a continuous-time nonlinear optimal control problem. Due to the complexity of the system, it is difficult to solve the Hamiltonian Jacobi Bellman (HJB) equations analytically. Therefore, this application employs an adaptive dynamic programming method based on an actor-critic dual network for online approximate solution. The solved target control commands (e.g., control surface deflection angle, engine thrust, etc.) are sent to the actuators to drive the tethered aircraft, achieving high-cost area avoidance and safe maneuver trajectory tracking.
[0068] In some embodiments of this application, when an electronic device solves for control commands with the goal of minimizing flight planning indicators to obtain target control commands, it can use a preset dynamic programming solver to solve for control commands with the goal of minimizing flight planning indicators to obtain target control commands. The preset dynamic programming solver includes a first network and a second network. The first network is used to evaluate the value estimate corresponding to the candidate control commands output by the second network, and the second network is used to update the candidate control commands based on the value estimate output by the first network.
[0069] In the embodiments of this application, the preset dynamic programming solver is an adaptive dynamic programming dual-network structure; wherein, the first network is the critic network, used to approximate the optimal value function; the second network is the actor network, used to approximate the optimal control law. The two learn together to form an approximate optimal control strategy.
[0070] In some embodiments of this application, the electronic device may also calculate the Hamiltonian function based on the tracking error, candidate control commands, the cost value of the integrated cost map, and the obstacle function; calculate the Hamilton-Jacobi-Bellman equation residual based on the Hamiltonian function, and determine the loss function of the first network based on the Hamilton-Jacobi-Bellman equation residual, so as to update the first network using the loss function of the first network; and determine the loss function of the second network based on the residual of the candidate control command relative to the first derivative of the Hamiltonian function, so as to update the second network using the loss function of the second network.
[0071] In the embodiments of this application, the Hamiltonian function is a scalar function in optimal control theory, consisting of the integrand of the performance index, the system dynamic equation, and the adjoint variables, used to derive the optimality condition.
[0072] In the embodiments of this application, the Hamilton-Jacobi-Bellman equation (HJB equation) is a partial differential equation in optimal control theory, and its solution is the optimal value function. Due to the nonlinearity of actual systems, approximate solution methods are usually adopted.
[0073] For example, the Hamiltonian function can be expressed as the following formula: (17); in, The optimal value function is the function that starts from the current state. Starting from the minimum future cumulative performance index achievable according to the optimal control law. , for Regarding the system state (here referring to the tracking error vector) (gradient) It is a kind of Vectors of the same dimension; yes The first derivative with respect to time.
[0074] For example, the Hamilton-Jacobi-Bellman equation can be expressed as the following formula: (18).
[0075] Furthermore, based on the principle of minimum value, through calculation... The target control command can be obtained, expressed as the following formula: (19); in, The input matrix of the system, also known as the control distribution matrix, describes the control commands. How to affect system state The derivative of .
[0076] For example, the optimization objective of the first network is to minimize the residuals of the HJB equation, which can be expressed as the following formula: (20); in, The weight matrix for tracking error, The weight matrix is used to control the energy consumption term.
[0077] The loss function of the first network can be expressed as the following formula: (twenty one); Then, using the policy gradient method, the weight update law of the first network is obtained as follows: (twenty two); in, It is the learning rate of the first network. It is the weight vector of the first network. It is an intermediate variable. , It is the feature mapping function of the first network. Represents feature mapping Regarding state error The gradient (Jacobi matrix).
[0078] For example, the optimization objective of the second network is to make the control law satisfy the optimality condition; the residual of the candidate control command with respect to the first derivative of the Hamiltonian function (the optimality condition residual) can be expressed as follows: (twenty three); in, That is, the Hamiltonian function defined by the aforementioned formula (17), Indicates a candidate control command. That is, the first derivative (gradient) of the Hamiltonian function with respect to the candidate control command should be zero according to the minimum principle. It is the first network output Regarding system status gradient, This is the weight matrix of the second network. It is the feature mapping function of the second network. It is the feature mapping of the second network. Regarding the status The gradient transpose.
[0079] The loss function of the second network can be expressed as the following formula: (twenty four); Then, using the policy gradient method, the weight update law of the second network is obtained as follows: (25); in, Let be the weight vector of the second network. It is the learning rate of the second network.
[0080] This application provides a maneuver planning method for tethered aircraft based on dynamic modeling of maneuverable regions. Compared with existing technologies, this application has the following technical advantages: It offline constructs a nominal cost map that comprehensively considers the wake vortex of the main aircraft, the control capability boundary of the towed aircraft, and dynamic stability, and online corrects it in real time based on gusts and atmospheric turbulence to generate a dynamic comprehensive cost map. This "offline + online" flow field modeling method reduces online computation while accurately reflecting complex time-varying wind fields, making the division of maneuverable regions more scientific and real-time. By constructing a safe distance threshold based on lateral velocity, the aircraft automatically expands its safety margin during high-speed maneuvers and allows it to approach danger zones at low speeds, balancing maneuverability and safety. Addressing system nonlinearity and wind field uncertainty, this application does not rely on precise analytical model solutions but instead uses online learning of two neural networks to gradually approximate the optimal control law, overcoming the "curse of dimensionality" of traditional discretized ADP and achieving stable control under model uncertainty and time-varying disturbances. It is evident that this application solves the problem of safe maneuver planning for tethered aircraft in complex wind fields by organically combining dynamic modeling of maneuverable areas with adaptive dynamic programming, and has high practical value and promising prospects for promotion.
[0081] Based on the above embodiments, in another embodiment of this application, exemplarily, as follows: Figure 2As shown, the maneuver planning method for tethered aircraft based on dynamic modeling of maneuverable regions mainly includes six core steps: Step 1: Modeling the wake vortex-induced velocity cost of the main aircraft; Based on the wake vortex model of the main aircraft (such as the Hallock-Burnham model), an offline nominal cost map is constructed to reflect the flight cost caused by wake vortex-induced velocity at various spatial locations without real-time wind field correction. Step 2: Online correction of flow field-induced velocity cost; Combining real-time measured gust wind speeds, a strategy of "offline nominal + online gust correction + atmospheric turbulence conservative compensation" is adopted to calculate the dynamic flow field-induced velocity cost, dynamically updating the nominal cost map from Step 1 to obtain a comprehensive cost map. Step 3: Modeling the boundary cost of aircraft control capability; Based on the additional control output amplitude of the actuators required for the tethered aircraft to maintain stable hovering at different spatial locations, a boundary cost model of control capability is established. The greater the required control force, the smaller the remaining control margin, and the higher the cost. Step 4: Modeling the dynamic stability cost of the aircraft; The lateral deviation distance between the current spatial location and the nominal equilibrium position with the same rope length is used to measure dynamic stability. The greater the deviation distance, the worse the dynamic stability and the higher the cost. Step 5: Modeling dynamic threshold constraints for safe distance; based on the real-time lateral velocity of the tethered aircraft, dynamically adjust the safe distance threshold (the faster the speed, the greater the required safe distance), and construct a reciprocal obstacle function to transform the safe distance constraint into an optimization penalty term. Step 6: Solving using adaptive dynamic programming (Actor-Critic ADP) based on an actor-critic dual network; combine tracking error, control energy consumption, comprehensive cost map value, and the dynamic safe distance obstacle function into a comprehensive performance index. Use the Critic network to approximate the optimal value function and the Actor network to approximate the optimal control law. Through an iterative process of policy evaluation and improvement, solve for the optimal control command online to achieve high-cost area avoidance and safe maneuver trajectory planning.
[0082] An "offline + online" flow field induced velocity cost model was established; an offline main vehicle wake vortex induced velocity model was established based on the Hallock-Burnham model; the induced velocity model was corrected in real time based on the gust estimation information measured by the main vehicle; and a conservative strategy was adopted to further correct the cost model for the impact of atmospheric turbulence.
[0083] Establish a flight performance cost model for a tethered aircraft, including a control capability boundary cost model to characterize the maximum maneuverability of the tethered aircraft, and a dynamic stability cost model to characterize the dynamic stability of the tethered aircraft at different spatial locations.
[0084] A dynamic constraint model for a tethered aircraft is established, defining regions with a cost value exceeding a certain threshold as high-cost regions and other regions as maneuverable regions; a dynamic constraint model for safe distance is established using the distance between the tethered aircraft and the boundary of the high-cost region.
[0085] Based on the established dynamic cost model and safety distance dynamic constraint model, a comprehensive performance index is established, and an adaptive dynamic programming algorithm based on the Actor-Critic dual network is designed to realize the avoidance and safe maneuver trajectory planning of high-cost areas such as the wake vortex center wind field.
[0086] Through the above technical solutions, this application comprehensively considers various factors such as complex flow fields, control capabilities, and flight stability, and establishes a complete evaluation index for the flight cost of tethered aircraft. Compared with existing technologies, it has significant advantages in terms of computational efficiency and model completeness, enabling the maneuverable area model of tethered aircraft to balance real-time performance and accuracy. The above technical solutions of this application consider the safe distance threshold between the tethered aircraft and the boundary of high-cost areas during maneuvering, and use the lateral velocity of the aircraft and the threshold for dynamic correction. Compared with existing technologies, it further considers the synergistic constraints of maneuverability and safety during maneuvering, enabling the tethered aircraft to maintain a safe distance from high-cost areas as much as possible while ensuring the completion of the maneuvering mission. Furthermore, the above technical solutions of this application adopt an adaptive dynamic programming algorithm under the Actor-Critic network architecture, and learn the optimal control strategy online through a dual-network architecture. Compared with existing technologies, it is more adaptable to dynamically changing flight environments and effectively avoids high-risk areas such as wake vortices. This application solves the technical problems of unclear maneuverable area and difficulty in planning safe maneuver trajectories in the existing tethered towed aircraft through the above technical solution, and effectively avoids the large-scale swaying of tethered towed aircraft under interference such as complex flow fields.
[0087] In some embodiments of this application, the scheme for establishing a flow field-induced velocity cost model and correcting it in real time includes: To reduce the computational load, the flow field behind the main aircraft was characterized into three parts: a predictable, slow-time-varying wake vortex flow field; a large-scale, long-period, coarsely measurable gust wind field; and a small-scale, high-frequency changing atmospheric turbulence. Based on the characteristics of these three flow fields, the wake vortex-induced velocity of the main aircraft was first modeled offline. The induced velocity generated by the wake vortex was calculated using the Hallock-Burnham model. The formula for calculating the tangential induced velocity of a single vortex is shown in formula (1) above; the formula for calculating vortex circulation is shown in formula (2) above.
[0088] Generally speaking, the induced velocity of the tail vortex of the main aircraft includes the induced velocity generated by the left and right vortex cores, that is, the total induced velocity is as shown in the aforementioned formula (3), which is decomposed into the horizontal direction y and the vertical direction z, as shown in the aforementioned formulas (4) to (7).
[0089] Since the gusts are roughly the same in a certain airspace, the lateral gust induced velocity components measured in real time by the airborne system are superimposed on the nominal induced velocity field of the main aircraft tail vortex obtained by offline pre-calculation to complete the online preliminary correction of the flow field induced velocity, as shown in the aforementioned formula (10).
[0090] Unlike the low-frequency, large-scale disturbance characteristics of gusts, atmospheric turbulence is a high-frequency, small-scale, and highly random time-varying wind field. Due to limitations in the sampling frequency, spatial measurement range, and response bandwidth of airborne sensing systems, accurate real-time measurement across the entire frequency band and airspace is not possible. To balance the real-time performance of airborne online calculations with the robustness of control law design, a conservative processing strategy is adopted: the equivalent lateral induced velocity amplitude of atmospheric turbulence is set to 20% of the sum of the lateral induced velocities of the main aircraft wake vortex and the gusts at the same spatial location. The induced velocity of the flow field after dynamic correction by gusts and atmospheric turbulence is shown in the aforementioned formula (11).
[0091] After truncating and normalizing the total induced velocity amplitude, the dynamic flow field induced velocity cost is finally established as shown in the aforementioned formula (12).
[0092] In some embodiments of this application, the scheme for establishing a flight performance cost model for a tethered aircraft includes: Control Capability Boundary Cost Model: For a tethered aircraft to maintain position and attitude stability at different points in the three-dimensional airspace, the required actuator control output varies significantly. Furthermore, the physical output of the actuators has inherent saturation constraints, with its maximum permissible output amplitude being... For any target point in the airspace, the larger the actuator control output amplitude required for stable maintenance, the lower the residual control margin of the system, the weaker the robustness against unknown external disturbances such as sudden gusts, and the higher the risk of instability. Based on this, a boundary cost index for the control capability of a tethered aircraft is introduced to quantify the actuator control cost required for stable maintenance at any point in the airspace. The cost of this index is positively correlated with the actuator control output amplitude required for the point. The boundary cost model for the control capability of a tethered aircraft at any point in space is calculated according to the following procedure: 1. Given the three-dimensional coordinates of the spatial position of the tethered aircraft relative to the host aircraft; 2. Calculate the magnitude and direction of the cable tension required to maintain stability based on the catenary model; 3. Calculate the total aerodynamic force required for stability based on the equilibrium equations; 4. Calculate the static dynamics of the tethered aircraft based on the wake vortex induced velocity; 5. Subtract the total aerodynamic force and static aerodynamic force calculated in steps 3 and 4 to obtain the additional control force required to maintain stability at this point; 6. After calculating the required control force, truncation and normalization are performed to obtain the functional load control capability boundary cost calculation formula as shown in the aforementioned formula (8).
[0093] Dynamic Stability Cost Model for Tethered Aircraft: The nominal equilibrium position of a tethered aircraft lies within the vertically symmetrical plane of the suspending point. Given a fixed rope length, when the tethered aircraft deviates from this symmetrical plane and enters a non-equilibrium position, an additional steady-state control force must be generated through active control to maintain the target position. Under the same rope length constraint, the greater the lateral deviation of the tethered aircraft from its nominal equilibrium position, the higher the amplitude of the steady-state control force required to maintain the position. This control force is achieved by adjusting the aerodynamic configuration of the tethered aircraft; a large aerodynamic configuration deflection directly weakens the dynamic stability margin of the system, and the amplitude of the steady-state control force is significantly negatively correlated with the system's dynamic stability. Therefore, a dynamic stability cost index for tethered aircraft is introduced to quantify the lateral deviation between the current position of the tethered aircraft and the corresponding nominal equilibrium position with the same rope length. The cost of this indicator is positively correlated with the lateral deviation distance. Similarly, this term is truncated and normalized to obtain the formula for calculating the dynamic stability cost of the tethered aircraft as shown in the aforementioned formula (9).
[0094] In some embodiments of this application, the scheme for establishing the dynamic constraint model of the tethered towed aircraft includes: To ensure the safety of the tethered aircraft during maneuvering and prevent it from entering high-cost regions, the relative distance between the tethered aircraft and regions with cost exceeding the threshold is introduced as a hard constraint into the motion planning model of the tethered aircraft. To address the drawbacks of high computational redundancy and excessive ineffective consumption of computational resources in global traversal search of high-cost region boundaries, a local search algorithm is used to quickly solve for the boundary distance, thereby reducing computational overhead and improving the real-time performance of the planning algorithm. The minimum distance from the current position of the tethered aircraft to the boundary of the high-cost region within the same yz plane is used to approximate the global minimum distance between the tethered aircraft and the high-cost region in the three-dimensional space. The specific calculation method is as follows: find the boundary of the region with a cost greater than 0.5 within the yz plane where the tethered aircraft is located, and then calculate the minimum distance from the current position of the tethered aircraft to the above boundary, as shown in the aforementioned formula (13).
[0095] To further prevent tethered aircraft from entering high-cost areas, a dynamic threshold strategy is introduced. The reciprocal obstacle function is defined as shown in the aforementioned formula (14).
[0096] Furthermore, the safety distance threshold is designed as a dynamic form that adjusts in real time with the lateral movement speed, as shown in the aforementioned formula (15).
[0097] In some embodiments of this application, a scheme for establishing comprehensive performance indicators is developed based on the established dynamic cost model and the dynamic constraint model for safe distance, such as... Figure 3 As shown, it includes: Taking into account the expected position tracking error, the cost model established above, and the dynamic safety distance constraint model, a comprehensive performance index is established as shown in the aforementioned formula (16).
[0098] Define the Hamiltonian function as shown in the aforementioned formula (17).
[0099] The HJB equation is shown in the aforementioned formula (18).
[0100] Based on the principle of minimum value, through calculation The optimal control strategy can be obtained: .
[0101] An Actor-Critic network structure is adopted, and its approximate solution is obtained through adaptive learning. The Critic network is used to approximate the optimal value function, as shown in Equation (26); while the Actor network is used to approximate the optimal control law, as shown in Equation (27). (26); (27).
[0102] The optimization objective of the Critic network is to minimize the residual of the HJB equation, as shown in the aforementioned formula (20).
[0103] The loss function of the Critic network can be defined as shown in the aforementioned formula (21).
[0104] Using the policy gradient method, the weight update law of the Critic network is obtained as shown in the aforementioned formula (22).
[0105] The optimization objective of the Actor network is to make the control law satisfy the optimality condition, and the residual is shown in the aforementioned formula (23).
[0106] Its loss function can be defined as the aforementioned formula (24).
[0107] Using the policy gradient method, the weight update law of the Actor network is obtained as shown in the aforementioned formula (25).
[0108] In some embodiments of this application, polynomial basis functions (such as radial basis functions, Hermitian polynomials, etc.) can be used as activation functions or feature maps of hidden layers of neural networks to approximate value functions and control laws.
[0109] Furthermore, the effectiveness of the method described in this application is compared with that of a maneuver control method that does not consider avoidance of high-risk areas. The comparison of position error changes during the maneuver of the tethered towed aircraft is as follows: Figure 4 As shown, the comparison of changes in cost during the maneuver is as follows: Figure 5 As shown, it can be observed that the cost of the tethered aircraft using the method in this application is less than 0.5 throughout the maneuver, while the cost of the comparative algorithm can reach a maximum of 0.7. This means that the tethered aircraft enters a high-cost region during maneuver, leading to... Figure 4 Under the method described in this application, the maximum deviation of the z-direction error of the tethered towed aircraft is only 0.5m, while the maximum deviation of the z-direction error under the comparative algorithm is nearly 3m. The results prove that the method described in this application can achieve safe maneuver control.
[0110] Based on the above embodiments, another embodiment of this application provides an electronic device. For example... Figure 6 As shown, the electronic device 1 proposed in this application embodiment may include a processor 11 and a memory 12 storing instructions executable by the processor 11; further, the electronic device 1 may also include a communication interface 13 and a bus 14 for connecting the processor 11, the memory 12 and the communication interface 13.
[0111] In the embodiments of this application, the processor 11 can be at least one of the following: Application-Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Digital Signal Processing Device (DSPD), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), Central Processing Unit (CPU), controller, microcontroller, and microprocessor. It is understood that for different devices, the electronic device used to implement the above-mentioned processor function can also be other, and this application embodiment does not specifically limit it. The memory 12 can be connected to the processor 11, wherein the memory 12 is used to store executable program code, which includes computer operation instructions. The memory 12 may include high-speed RAM memory, and may also include non-volatile memory, such as at least two disk drives.
[0112] In embodiments of this application, bus 14 is used to connect communication interface 13, processor 11 and memory 12 to enable communication between these devices.
[0113] In embodiments of this application, memory 12 is used to store instructions and data.
[0114] In practical applications, the aforementioned memory 12 can be volatile memory, such as random-access memory (RAM), or non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); or a combination of the above types of memory, and provide instructions and data to the processor 11.
[0115] Furthermore, in this embodiment, the functional modules can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional module.
[0116] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the method of this embodiment.
[0117] Specifically, the program instructions corresponding to the tethered aircraft maneuver planning method based on dynamic modeling of maneuverable areas in this embodiment can be stored on storage media such as optical discs and hard disks. When the program instructions corresponding to the tethered aircraft maneuver planning method based on dynamic modeling of maneuverable areas in the storage media are read or executed by an electronic device, the following steps are included: Acquire the real-time position and speed of the tethered aircraft, as well as the gust-induced speed; The nominal cost map is dynamically corrected based on the gust-induced velocity to obtain the comprehensive cost map. The nominal cost map is a pre-constructed comprehensive impact of each location in space on the maneuverability and safety of the tethered aircraft without real-time wind field correction. The boundary of the high-cost area is determined based on the comprehensive cost map, and the current safe distance threshold is determined based on the real-time speed. A barrier function is then constructed based on the boundary of the high-cost area and the current safe distance threshold. Flight planning indicators are determined based on the tracking error between the real-time position and the target position, control energy consumption, the cost of the comprehensive cost map, and the obstacle function; among them, control energy consumption is used to characterize the energy consumption of the actuator output of the tethered aircraft. The control commands are solved with the goal of minimizing the flight planning index, resulting in target control commands. These target control commands are used to control the tethered aircraft to perform high-cost area avoidance and safe maneuver trajectory tracking.
[0118] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.
[0119] This application is described with reference to schematic and / or block diagrams of implementations of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each step and / or block in the schematic and / or block diagrams, as well as combinations thereof, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more steps of the schematic and / or one or more blocks of the block diagrams.
[0120] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in the implementation flow diagram. Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0121] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0122] The above embodiments are merely preferred embodiments provided to fully illustrate this application, and the scope of protection of this application is not limited thereto. Equivalent substitutions or modifications made by those skilled in the art based on this application are all within the scope of protection of this application.
Claims
1. A maneuver planning method for tethered towed aircraft based on dynamic modeling of maneuverable areas, characterized in that, The method includes: The operational airspace of the tethered towed aircraft is discretized into a grid using the three-dimensional voxel grid method, resulting in airspace grid points. For each of the aforementioned airspace grid points, the vortex-induced velocity of the main aircraft is determined based on its weight, flight speed, air density, and fuselage length, resulting in a vortex-induced velocity field. The control capability boundary value is determined based on the actuator control output amplitude required for the tethered aircraft to maintain stability at the airspace grid point. The dynamic stability value is determined based on the lateral deviation between the position of the tethered aircraft and its nominal equilibrium position with the same rope length. The vortex-induced velocity field, the control capability boundary value, and the dynamic stability value are fused to obtain the nominal value for each airspace grid point, forming a nominal cost map. This nominal cost map is a pre-constructed map used to characterize the comprehensive impact of spatial locations on the maneuverability and safety of the tethered aircraft without real-time wind field correction. Acquire the real-time position and speed of the tethered aircraft, as well as the gust-induced speed; The composite induced velocity is obtained by adding the wake vortex induced velocity field and the gust induced velocity. The dynamic flow field induced velocity is obtained by adding the synthetic induced velocity and the equivalent induced velocity of atmospheric turbulence; wherein the equivalent induced velocity of atmospheric turbulence is the product of a preset ratio and the synthetic induced velocity. The induced velocity of the dynamic flow field is truncated and normalized to obtain the dynamic flow field cost of each spatial grid point. The nominal cost value of each spatial grid point is added to the dynamic flow field cost value to obtain the comprehensive cost map; The boundary of the high-cost area is determined based on the comprehensive cost map, and the current safe distance threshold is determined based on the real-time speed, so as to construct a barrier function based on the boundary of the high-cost area and the current safe distance threshold; Flight planning indicators are determined based on the tracking error between the real-time position and the target position, control energy consumption, the cost value of the integrated cost map, and the obstacle function; wherein, the control energy consumption is used to characterize the energy consumption of the actuator output of the tethered aircraft. The control commands are solved with the goal of minimizing the flight planning index to obtain the target control commands; wherein, the target control commands are used to control the tethered towed aircraft to perform high-cost area avoidance and safe maneuver trajectory tracking.
2. The maneuver planning method for tethered towed aircraft based on dynamic modeling of maneuverable areas according to claim 1, characterized in that, Determining the boundary of the high-cost region based on the comprehensive cost map includes: The spatial region formed by the first grid point in the comprehensive cost map whose cost is higher than a preset threshold is defined as a high-cost region. Within the lateral plane where the tethered aircraft is currently located, a local search algorithm is used to search and determine the boundary of the high-cost region.
3. The maneuver planning method for tethered towed aircraft based on dynamic modeling of maneuverable areas according to claim 2, characterized in that, The step of determining the current safe distance threshold based on the real-time speed, and constructing a barrier function based on the high-cost region boundary and the current safe distance threshold, includes: The current safe distance threshold is determined based on the lateral velocity component in the real-time speed; wherein the current safe distance threshold is positively correlated with the lateral velocity component. The minimum distance between the real-time location and the boundary is taken as the safety distance; The obstacle function is determined based on the safe distance and the current safe distance threshold.
4. The maneuver planning method for tethered towed aircraft based on dynamic modeling of maneuverable areas according to claim 1, characterized in that, The process of determining flight planning indicators based on the tracking error between the real-time location and the target location, control energy consumption, the cost value of the integrated cost map, and the obstacle function includes: The weighting values are determined based on the tracking error, the control energy consumption, the cost value of the integrated cost map, the obstacle function, and their respective weights. The weighted values are summed, and the summation result is integrated over time to obtain the flight planning index.
5. The maneuver planning method for tethered towed aircraft based on dynamic modeling of maneuverable areas according to claim 1, characterized in that, The process of solving for control commands with the goal of minimizing the flight planning indicators to obtain target control commands includes: With the minimum flight planning index as the optimization objective, the control commands are solved using a preset dynamic programming solver to obtain the target control commands; The preset dynamic programming solver includes a first network and a second network; the first network is used to evaluate the value estimate corresponding to the candidate control command output by the second network, and the second network is used to update the candidate control command based on the value estimate output by the first network.
6. The maneuver planning method for tethered towed aircraft based on dynamic modeling of maneuverable areas according to claim 5, characterized in that, The method further includes: The Hamiltonian function is calculated based on the tracking error, the candidate control command, the cost value of the integrated cost map, and the obstacle function. The Hamiltonian Jacobi Bellman equation residuals are calculated based on the Hamiltonian function, and the loss function of the first network is determined based on the Hamiltonian Jacobi Bellman equation residuals, so as to update the first network using the loss function of the first network. The loss function of the second network is determined based on the residual of the candidate control command relative to the first derivative of the Hamiltonian function, so as to update the second network using the loss function of the second network.
7. An electronic device, characterized in that, It includes a processor and a memory storing executable instructions of the processor; when the instructions are executed by the processor, it implements the tethered aircraft maneuver planning method based on dynamic modeling of maneuverable areas as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Flight control unit for parafoil type unmanned plane
CN105404308A
Unmanned flight base station automatic cruise method and equipment based on communication blind spots
CN116540775A