Wing sweep angle deformation optimization method, electronic device and storage medium based on aircraft fuel economy
By using the DDPG algorithm to build an intelligent network, the aircraft's wing surface sweep angle is optimized, and the fuel economy and range problems caused by the fixed wing surface sweep angle in traditional design are solved, achieving more efficient fuel use and range improvement.
Patent Information
- Application Number
- CN202411937778.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-12-26
AI Technical Summary
In traditional aircraft design, the sweep angle of the wing surface is fixed, which cannot provide the best fuel economy and range under different flight conditions, and it is difficult for the existing technology to effectively use artificial intelligence technology to make intelligent decisions.
Deep deterministic policy gradient (DDPG) algorithm is used to construct an agent network to optimize the sweep angle of the wing surface, and intelligent decisions on the fuel economy of the aircraft are achieved by training the actor network and critic network.
Through intelligent decision-making, the wing surface sweep angle is optimized, which significantly improves the aircraft's range and fuel economy, adapts to complex and changing flight conditions, and achieves accurate and real-time adjustment of the wing surface sweep angle.
Smart Images

Figure CN119740324B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of aerospace engineering, and particularly relates to a method for optimizing the deformation of the wing sweep angle based on the fuel economy of an aircraft, an electronic device, and a storage medium. Background Art
[0002] The fuel economy and range of hypersonic aircraft have always been the focus of attention in the field of aerospace engineering. This is mainly determined by two factors: the design of the aircraft and the flight conditions. In aircraft design, the wing sweep angle is an important factor that directly affects the aerodynamic performance and fuel consumption of the aircraft.
[0003] In traditional aircraft design, the wing sweep angle is usually fixed. This design is simple, but to a certain extent, it limits the fuel economy and range of the aircraft under different flight conditions. For example, for hypersonic aircraft, changes in altitude and speed will significantly affect the aerodynamic performance and fuel consumption. In this case, a fixed wing sweep angle may not provide the best fuel economy.
[0004] To solve this problem, one solution is to improve the fuel economy and range of the aircraft by changing the wing sweep angle of the aircraft. However, how to determine the optimal wing sweep angle under specific flight conditions to achieve the best fuel economy and range is a challenge. Traditional methods usually make decisions based on experience or simple rules, and the effect of this method is not ideal, and it cannot adapt to complex and changing flight conditions.
[0005] In this context, artificial intelligence technology provides new possibilities for intelligent decision-making. The Deep Deterministic Policy Gradient (DDPG) algorithm is a method that combines deep learning and reinforcement learning. The DDPG algorithm can achieve intelligent decision-making in a continuous action space by training an agent. This provides a new idea for solving the problem of aircraft fuel economy decision-making.
[0006] However, in the existing technology, there has been no report on applying the DDPG algorithm to aircraft fuel economy decision-making, especially the decision-making of the optimal wing sweep angle of an aircraft with a variable sweep angle. This may be due to the complexity of the aircraft fuel economy decision-making problem and the implementation difficulty of the DDPG algorithm. Summary of the Invention
[0007] The problem to be solved by the present invention is to achieve intelligent decision-making of the optimal wing sweep angle to improve the range, and to propose a method for optimizing the deformation of the wing sweep angle based on the fuel economy of an aircraft, an electronic device, and a storage medium.
[0008] To achieve the above object, the present invention is realized by the following technical solutions:
[0009] A method for optimizing the sweep angle deformation of a wing surface based on the fuel economy of an aircraft, comprising the following steps:
[0010] S1. Network design: Construct an agent network, including an actor network and a critic network, where the input of the actor network is the state and the output is the action;
[0011] S2. Environment design: Set the initial environment of the hypersonic aircraft, define the state corresponding to the hypersonic aircraft including flight altitude, speed, range and mass, define the action corresponding to the hypersonic aircraft as the wing sweep angle, and define the reward corresponding to the hypersonic aircraft as the fuel consumption and its influence relationship;
[0012] S3. Input the environment designed in step S2 into the network designed in step S1 to obtain the actor network and critic network based on the hypersonic aircraft. After interactive training to collect experience values, then use the DDPG algorithm improved by the successful sample replay method, and use the collected experience to train the obtained actor network and critic network based on the hypersonic aircraft to obtain a wing sweep angle deformation optimization model based on the fuel economy of the aircraft;
[0013] S4. Install the wing sweep angle deformation optimization model based on the fuel economy of the aircraft obtained in step S3 on the hypersonic aircraft computer to make a wing sweep angle optimization decision under flight conditions.
[0014] Further, the specific implementation method of step S1 includes the following steps:
[0015] S1.1. Construct an actor network for learning and making decisions on actions from the environmental state, including 1 input layer, 2 hidden layers and 1 output layer. The expressions for the input and output of the actor network are:
[0016] a = π(s) (1)
[0017] where a is the action, s is the state, and π is a non-linear function relationship expressed by the actor network;
[0018] S1.2. Construct a critic network for scoring the actions of the actor network, including 1 input layer, 2 hidden layers and 1 output layer. The expressions for the input and output of the critic network are:
[0019]
[0020] where Q is the value of the optimal problem, is a non-linear function relationship expressed by the critic network;
[0021] S1.3. Set the functional relationships π and Q of the actor and critic networks, which are calculated by the forward calculation method of the neural network. In the actor network, the output of the l-th layer is denoted as a (l) , where l = {1, 2}, and the calculation is divided into two steps: linear transformation and non-linear activation function. The expressions are as follows:
[0022] z (l) = W (l) a (l-1) + b (l) (3)
[0023] a (l) = σ (l) (z (l) ) (4)
[0024] Among them, z (l) is the linear transformation of the l-th layer, W (l) is the weight matrix of the l-th layer, b (l) is the bias vector of the l-th layer, and σ (l) is the activation function of the l-th layer. σ (l) (x) is the non-linear activation function of the l-th layer, which is set to tanh(x); θ is the weight of the neural network;
[0025] The output of the entire network is the activation of the last layer, so a = a (2) ;
[0026] For the critic network, it is calculated by the forward calculation method of the neural network, and the output is:
[0027] Q = a (2) (5)
[0028] Furthermore, the input layer of the actor network in step S1.1 is set with 4 input nodes corresponding to 4 state variables, the hidden layer includes 256 nodes, and the output layer includes 1 node;
[0029] The input layer of the critic network in step S1.2 is set with 4 input nodes corresponding to 3 state variables and 1 action, the hidden layer includes 64 nodes, and the output layer includes 1 node to score the action of the actor network.
[0030] Furthermore, the specific implementation method of step S2 includes the following steps:
[0031] S2.1. Set the initial environment of the hypersonic vehicle, including the flight altitude range of 15 - 30 km, the flight speed range of 1200 - 2400 m / s, the initial total fuel m fuel of 500 kg, and the net weight m dry of 3000 t, and obtain the initial weight m0 = m fuel+m dry ;
[0032] S2.2. Define the states corresponding to the hypersonic vehicle, including flight altitude h, speed v, range R, and current mass m. Then, the expression for the states corresponding to the hypersonic vehicle is:
[0033] s = [h, v, R, m] (6)
[0034] Define the action corresponding to the hypersonic vehicle as the wing sweep angle Λ, and the expression is:
[0035] a = Λ (7)
[0036] where the range of Λ is from 0 to 30 degrees;
[0037] S2.3. Define that the fuel consumption is mainly affected by the following factors: h affects the atmospheric density and aerodynamic drag, v affects the power demand and aerodynamic drag, and Λ affects the aerodynamic characteristics of the vehicle, including lift and drag. Then, the expression for the fuel consumption E(h, v, Λ) is:
[0038] E(h, v, Λ) = k1·D(h, v, Λ) + k2·P(h, v, Λ) (8)
[0039] where D(h, v, Λ) is the aerodynamic drag, P(h, v, Λ) is the thrust, and k1 and k2 are the first positive constant and the second positive constant;
[0040] Design the reward corresponding to the hypersonic vehicle to be related to the fuel consumption and its influence. The expression is:
[0041] r(s, a) = w1ΔE + w2ΔR (9)
[0042] where r(s, a) is the reward obtained by taking action a in state s, ΔE is the improvement in fuel efficiency compared with the previous decision, and ΔE = E(h', v', Λ') - E(h, v, Λ);
[0043] where h', v', Λ' and R' are the parameters corresponding to the target state s', and w1 and w2 are the first weight parameter and the second weight parameter, which are used to adjust the current optimization weight value of the fuel efficiency;
[0044] S2.4. Through the target state, calculate the state transition F to give the recurrence relation expression of s' = F(s, a) as:
[0045]
[0046] where ζ is the flight path angle.
[0047] Furthermore, the specific implementation method of step S3 includes the following steps:
[0048] S3.1. Collect experience values: Execute actions in the environment, and store each step of the transition (s, a, r, s′) in the experience replay buffer. The expression is:
[0049]
[0050] where is the O-U noise, which is obtained from the stochastic differential equation ω is the rate parameter, which determines the speed at which the system returns to the mean μ; μ is the long-term mean, and the process oscillates around this mean; υ is the intensity of the noise, which affects the amplitude of the noise fluctuation; dW t is the increment of Brownian motion, representing the random perturbation over time;
[0051] S3.2. Set the improved DDPG algorithm with the successful sample replay method:
[0052] In the replay buffer, assume that the agent reaches the state s* before the fuel is about to run out in a certain attempt. Even if it may cause the fuel mass of the aircraft to be negative, set the state s* as the new target g' = s*, recalculate the reward, and the new target is to reach the state s*. Set the new reward r HER as the negative exponent of the distance to the target state. The expression is:
[0053]
[0054] Input g' and s' into the network and recalculate the value of the critic
[0055] The goal of the improved DDPG algorithm with the successful sample replay method is to release the constraints in the case of task failure, so that the agent can extract the trend of reducing the fuel consumption rate in this case;
[0056] S3.3. Sample from the experience replay buffer, randomly extract samples for training the network, and obtain the training dataset S. The expression is:
[0057] S = {(s i , g i , a i , r i , s i ′, g i ′) | i = 1, 2,..., M} (13)
[0058] where M is the fixed sampling sample, and M is set to 128;
[0059] S3.4. Train the critic network using the training dataset obtained in step S3.3. For each (s, g, a, r, s′, g′) sampled from S, calculate the value of the current critic network using the target critic network and the target actor network to calculate the value Q′ of the target optimal problem, with the expression:
[0060]
[0061] Then calculate the target value y using the value of the target optimal problem adjusted by HER, with the expression:
[0062] y = r + γQ′ (15)
[0063] where γ is the discount factor, generally taken as γ = 0.99;
[0064] Define the expression for the metric L of the critic network as:
[0065]
[0066] For L with respect to the weights of the critic network calculate the gradient
[0067] Set the gradient to zero to obtain the weights of the critic network that minimize L S3.5. Using the updated critic network, obtain π Define the metric for the actor network as J = ∑r. For J with respect to the weights θ
[0068]
[0069] calculate the gradient, obtaining the expression: Set the gradient π' to zero to obtain the weights θ of the actor network that maximize the metric J
[0070] S3.6. Update the weights of the target actor and critic networks through a soft update method, respectively:
[0071]
[0072] where τ is the update weight, set as τ = 0.1.
[0073] Furthermore, in step S3, define the target g = m - 3000 > 0, that is, the fuel must not be exhausted.
[0074] An electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method for optimizing the deformation of the wing sweep angle based on the fuel economy of the aircraft are implemented.
[0075] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the method for optimizing the deformation of the wing sweep angle based on the fuel economy of the aircraft is implemented.
[0076] Advantages of the present invention:
[0077] For the method for optimizing the deformation of the wing sweep angle based on the fuel economy of the aircraft according to the present invention, an agent is trained using the DDPG algorithm. The state of the agent consists of the flight altitude and speed of the aircraft, and the action is to adjust the wing sweep angle of the aircraft. Through interaction with the environment, the agent learns how to decide the optimal wing sweep angle according to the current flight conditions (i.e., flight altitude and speed) to achieve optimal fuel economy, thereby increasing the range.
[0078] For the method for optimizing the deformation of the wing sweep angle based on the fuel economy of the aircraft according to the present invention, the trained agent is deployed into the aircraft. During flight, the agent decides the optimal wing sweep angle in real time according to the current flight altitude and speed, and guides the aircraft to adjust the wing sweep angle.
[0079] For the method for optimizing the deformation of the wing sweep angle based on the fuel economy of the aircraft according to the present invention, by combining intelligent decision-making and the design of an aircraft with a variable sweep angle, precise and real-time adjustment of the wing sweep angle can be achieved, thereby significantly increasing the range of the aircraft. Compared with traditional decision-making methods based on experience or simple rules, the method of the present invention is more intelligent and flexible, and can adapt to complex and changing flight conditions. In addition, due to the use of the DDPG algorithm, the method of the present invention can also handle continuous action spaces, which is very important for the precise adjustment of the sweep angle. The method of the present invention has broad application prospects in the field of optimizing the range of deformable aircraft. Description of the drawings
[0080] Figure 1 is a flowchart of the method for optimizing the deformation of the wing sweep angle based on the fuel economy of the aircraft according to the present invention;
[0081] Figure 2 is a schematic diagram of the training process of the agent of the present invention;
[0082] Figure 3 is a schematic diagram of the reward design of the present invention;
[0083] Figure 4Schematic diagram of exploration - utilization process of the present invention, where (a) is the influence curve of height offset on sweep angle change under on - line use conditions, and (b) is the influence curve of height offset on sweep angle change under the condition of adding exploration noise during off - line training;
[0084] Figure 5 Variation curve of the loss function during the training of the present invention;
[0085] Figure 6 Variation curve of the reward function during the training of the present invention;
[0086] Figure 7 Variation curve of the aircraft weight during the invocation of the present invention;
[0087] Figure 8 Variation curve of the aircraft sweep during the invocation of the present invention;
[0088] Figure 9 Variation curve of the aircraft height during the invocation of the present invention;
[0089] Figure 10 Variation curve of the aircraft speed during the invocation of the present invention. Detailed implementation manners
[0090] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. It should be understood that the specific implementation manners described herein are only used to explain the present invention and are not used to limit the present invention, that is, the specific implementation manners described are only a part of the implementation manners of the present invention, rather than all of the specific implementation manners. Usually, the components of the specific implementation manners of the present invention described and shown in the accompanying drawings herein can be arranged and designed in various different configurations, and the present invention can also have other implementation manners.
[0091] Therefore, the detailed description of the specific implementation manners of the present invention provided in the accompanying drawings below is not intended to limit the scope of the claimed present invention, but only represents the selected specific implementation manners of the present invention. All other specific implementation manners obtained by those skilled in the art based on the specific implementation manners of the present invention without creative efforts belong to the scope of protection of the present invention.
[0092] To further understand the content, features and effects of the present invention, the following specific implementation manners are exemplified and are accompanied by Figure 1 - Accompanying Figure 10 The details are as follows:
[0093] Example 1:
[0094] A method for optimizing the deformation of the wing sweep angle based on the fuel economy of an aircraft, comprising the following steps:
[0095] S1. Network Design: Construct an agent network, including an actor network and a critic network. The input of the actor network is the state, and the output is the action;
[0096] Furthermore, the specific implementation method of step S1 includes the following steps:
[0097] S1.1. Construct an actor network for learning and making decisions on actions from the environmental state, including 1 input layer, 2 hidden layers, and 1 output layer. The expressions for the input and output of the actor network are:
[0098] a = π(s) (1)
[0099] where a is the action, s is the state, and π is a non-linear functional relationship expressed by the actor network;
[0100] Furthermore, the input layer of the actor network in step S1.1 is set with 4 input nodes corresponding to 4 state variables, the hidden layer includes 256 nodes, and the output layer includes 1 node;
[0101] The input layer of the critic network in step S1.2 is set with 4 input nodes corresponding to 3 state variables and 1 action, the hidden layer includes 64 nodes, and the output layer includes 1 node to score the action of the actor network;
[0102] S1.2. Construct a critic network for scoring the action of the actor network, including 1 input layer, 2 hidden layers, and 1 output layer. The expressions for the input and output of the critic network are:
[0103]
[0104] where Q is the value of the optimal problem, is a non-linear functional relationship expressed by the critic network;
[0105] S1.3. Set the functional relationships π and Q of the actor and critic networks to be calculated by the forward calculation method of the neural network. Let the output of the l-th layer in the actor network be denoted as a (l) , l = {1, 2}. The calculation is divided into two steps, linear transformation and non-linear activation function, and the expressions are:
[0106] z (l) = W (l) a (l-1) + b (l) (3)
[0107] a (l) = σ (l) (z (l) ) (4)
[0108] where z(l) is the linear transformation of the l-th layer, W (l) is the weight matrix of the l-th layer, b (l) is the bias vector of the l-th layer, σ (l) is the activation function of the l-th layer σ (l) (x) is the non-linear activation function of the l-th layer, set to tanh(x); θ is the weight of the neural network;
[0109] The output of the whole network is the activation of the last layer, so a = a (2) ;
[0110] For the critic network, it is calculated by the forward calculation method of the neural network, and the output is:
[0111] Q = a (2) (5)
[0112] S2. Environment design: Set the initial environment of the hypersonic vehicle, define the states corresponding to the hypersonic vehicle including flight altitude, speed, range and mass, define the action corresponding to the hypersonic vehicle as the wing sweep angle, and define the reward corresponding to the hypersonic vehicle as fuel consumption and its influence relationship;
[0113] Furthermore, the specific implementation method of step S2 includes the following steps:
[0114] S2.1. Set the initial environment of the hypersonic vehicle, including the flight altitude range of 15 - 30 km, the flight speed range of 1200 - 2400 m / s, the initial total fuel m fuel of 500 kg, the net weight m dry of 3000 t, and obtain the initial weight m0 = m fuel + m dry ;
[0115] S2.2. Define the states corresponding to the hypersonic vehicle including flight altitude h, speed v, range R and current mass m, then the expression of the states corresponding to the hypersonic vehicle is:
[0116] s = [h, v, R, m] (6)
[0117] Define the action corresponding to the hypersonic vehicle as the wing sweep angle Λ, and the expression is:
[0118] a = Λ (7)
[0119] where, the range of Λ is from 0 to 30 degrees;
[0120] S2.3. It is defined that the fuel consumption is mainly affected by the following factors: h affects the atmospheric density and aerodynamic drag, v affects the power demand and aerodynamic drag, and Λ affects the aerodynamic characteristics of the aircraft, including lift and drag. Thus, the expression of fuel consumption E(h, v, Λ) is:
[0121] E(h, v, Λ) = k1·D(h, v, Λ) + k2·P(h, v, Λ) (8)
[0122] Among them, D(h, v, Λ) is the aerodynamic drag, P(h, v, Λ) is the thrust, and k1 and k2 are the first positive constant and the second positive constant;
[0123] The reward corresponding to the design of a hypersonic aircraft is related to the fuel consumption and its influence. The expression is:
[0124] r(s, a) = w1ΔE + w2ΔR (9)
[0125] Among them, r(s, a) is the reward obtained by taking action a in state s, ΔE is the improvement in fuel efficiency compared with the previous decision, and ΔE = E(h', v', Λ') - E(h, v, Λ);
[0126] Among them, h', v', Λ' and R' are the parameters corresponding to the target state s', and w1 and w2 are the first weight parameter and the second weight parameter, which are used to adjust the current optimization weight value of fuel efficiency;
[0127] S2.4. Through the target state, calculate the state transition F to give the recurrence relation expression of s' = F(s, a) as:
[0128]
[0129] Among them, ζ is the flight path angle;
[0130] S3. Input the environment designed in step S2 into the network designed in step S1 to obtain the actor network and critic network based on the hypersonic aircraft. After interactive training to collect experience values, then use the DDPG algorithm improved by the successful sample replay method, and use the collected experience to train the obtained actor network and critic network based on the hypersonic aircraft to obtain the wing sweep angle deformation optimization model based on the fuel economy of the aircraft;
[0131] Furthermore, the specific implementation method of step S3 includes the following steps:
[0132] S3.1. Collect experience values: Execute actions in the environment and store each step of the transition (s, a, r, s′) in the experience replay buffer. The expression is:
[0133]
[0134] Among them, is the O-U noise, which is obtained from the stochastic differential equation , where ω is the rate parameter that determines the speed at which the system returns to the mean μ; μ is the long-term mean around which the process oscillates; υ is the intensity of the noise that affects the amplitude of the noise fluctuation; dW t is the increment of Brownian motion, representing the random perturbation over time;
[0135] S3.2. Set the improved DDPG algorithm with the successful sample replay method:
[0136] In the replay buffer, assume that the agent reaches the state s* before the fuel is about to run out in a certain attempt. Even if it may cause the fuel mass of the aircraft to be negative, set the state s* as the new target g' = s*, recalculate the reward, and the new target is to reach the state s*. Set the new reward r HER as the negative exponent of the distance to the target state, and the expression is:
[0137]
[0138] Input g' and s' into the network and recalculate the value of the critic
[0139] The goal of the improved DDPG algorithm with the successful sample replay method is to release the constraints in the case of task failure, so that the agent can extract the trend of reducing the fuel consumption rate in this case;
[0140] S3.3. Sample from the experience replay buffer, randomly extract samples for training the network, and obtain the training dataset S, and the expression is:
[0141] S = {(s i , g i , a i , r i , s i ′, g i ′) | i = 1, 2,..., M} (13)
[0142] Among them, M is the fixed sampling sample, and M is set to 128;
[0143] S3.4. Train the critic network based on the training dataset obtained in step S3.3. Use each (s, g, a, r, s′, g′) in the sampled S to calculate the current critic network Use the target critic network and the target actor network to calculate the value Q′ of the target optimal problem, and the expression is:
[0144]
[0145] Then, use the value of the target optimal problem after HER adjustment to calculate the target value y, and the expression is:
[0146] y = r + γQ′ (15)
[0147] Where γ is the discount factor, and generally γ = 0.99;
[0148] Define the expression of the metric L of the critic network as:
[0149]
[0150] For L with respect to the weights of the critic network Find the gradient Set the gradient to zero to obtain the weights of the critic network that minimizes L
[0151] S3.5. Using the updated critic network, obtain Define the metric of the actor network as J = ∑r, and for J with respect to the weights θ of the actor network π Find the gradient, and the obtained expression is:
[0152]
[0153] Set the gradient To zero to obtain the weights θ of the actor network that maximizes the metric J π' ;
[0154] S3.6. Update the weights of the target actor and critic networks through the soft update method, which are respectively:
[0155]
[0156] Where τ is the update weight, and τ = 0.1 is set.
[0157] Furthermore, in step S3, define the target g = m - 3000 > 0, that is, the fuel must not be exhausted.
[0158] S4. Install the wing sweep angle deformation optimization model based on the fuel economy of the aircraft obtained in step S3 on the hypersonic aircraft computer to perform wing sweep angle optimization decision-making under flight conditions.
[0159] An optimization method for wing sweep angle deformation based on the fuel economy of an aircraft described in this embodiment:
[0160] 1. State definition and action space:
[0161] The state consists of the flight altitude h and speed v of the aircraft, i.e., the state s = (h, v); the action space is defined as the adjustment range of the wing sweep angle, i.e., the action a ∈ [Λ min , Λ max , where Λ min and Λ max represent the minimum and maximum values of the wing sweep angle respectively.
[0162] 2. Agent Training:
[0163] The agent is trained using the DDPG algorithm. The training of the DDPG algorithm includes two types of networks: the actor and the critic. The actor network makes online decisions, and the critic network scores the effect of the online decisions, and further optimizes the actor network based on the score
[0164] 3. Real-time Decision-making:
[0165] During the flight, the agent makes a real-time decision on the best wing sweep angle according to the current flight altitude and speed through the actor network based on the flight state.
[0166] 4. Reward Design:
[0167] To train the agent to achieve the optimization of fuel economy, a mechanism with the negative value of fuel consumption as the reward is designed.
[0168] 5. Method of Replaying Successful Samples:
[0169] The improvement of the present invention compared with the conventional DDPG algorithm lies in adding the replay of successful samples during training. In the guidance link of the hypersonic morphing aircraft, the wing sweep angle has a significant impact on the cruise conditions, so there are large differences in the training samples, which in turn causes the conventional algorithm to converge slowly. The replay of successful samples speeds up the learning process by preferentially selecting those experiences that are most helpful for learning.
[0170] Through the above features, this embodiment realizes an intelligent and real-time decision-making method for aircraft fuel economy, which is applicable to aircraft with variable sweep angles and has good theoretical basis and practical value.
[0171] A method for optimizing the wing sweep angle deformation based on aircraft fuel economy described in this embodiment can effectively handle the continuous action space, enabling the aircraft to intelligently adjust the wing sweep angle during flight, thereby achieving the goal of improving fuel economy. The training results of the agent are as follows Figure 5 Figure 6 shown, Figure 5 represents the curve of the loss function changing with the number of training episodes. The fact that the loss function tends to 0 indicates that the training converges. In this example, the training terminates at 5000 episodes, Figure 6It shows the change curve of the reward function with the number of training rounds. An increase in the reward function indicates an improvement in the performance of the agent. In this example, the reward converges at around 3000 rounds, indicating that the agent has been fully trained. In addition, for a case where the altitude decreases from 27 km to 18 km and the speed decreases from 2210 m / s to 1860 m / s, the simulation results are as Figures 7 to 10 shown, Figure 7 Table of fuel change curve, Figure 8 Table of sweep angle change curve, Figure 9 It shows the altitude change curve of the aircraft, Figure 10 It shows the speed change curve of the aircraft. It can be seen that at around 100 seconds, the aircraft has completed the predetermined deformation, altitude reduction, and speed reduction flight mission with less fuel consumption.
[0172] Embodiment 2:
[0173] An electronic device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, it implements the steps of the method for optimizing the deformation of the wing sweep angle based on the fuel economy of the aircraft described in Embodiment 1.
[0174] The computer device of the present invention may be a device including a processor and a memory, such as a single-chip microcomputer including a central processing unit. And, when the processor is used to execute the computer program stored in the memory, it implements the steps of the recommended method for the method for optimizing the deformation of the wing sweep angle based on the fuel economy of the aircraft described above.
[0175] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays
[0176] Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0177] The memory may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area may store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0178] Embodiment 3:
[0179] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, it implements the method for optimizing the deformation of the wing sweep angle based on the fuel economy of an aircraft described in Embodiment 1.
[0180] The computer-readable storage medium of the present invention can be any form of storage medium readable by the processor of a computer device, including but not limited to non-volatile memory, volatile memory, ferroelectric memory, etc. A computer program is stored on the computer-readable storage medium. When the processor of the computer device reads and executes the computer program stored in the memory, the steps of the above-mentioned method for optimizing the deformation of the wing sweep angle based on the fuel economy of an aircraft can be implemented.
[0181] The computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0182] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0183] Although the present application has been described above with reference to specific embodiments, various improvements can be made thereto and components thereof can be replaced with equivalents without departing from the scope of the present application. In particular, as long as there is no structural conflict, the features in the specific embodiments disclosed in the present application can be combined with each other in any way, and the exhaustive description of the combinations is not given in this specification only for the sake of saving space and resources. Therefore, the present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. A method for optimizing wing sweep angle deformation based on aircraft fuel economy, characterized in that: The steps include: S1. Network design: Build an agent network, including an actor network and a critic network, where the input of the actor network is the state and the output is the action; S2. Environment design: Set the initial environment of the hypersonic aircraft, define the states of the hypersonic aircraft including flight altitude, speed, range and mass, define the action of the hypersonic aircraft as the wing sweep angle, and define the reward of the hypersonic aircraft as fuel consumption and its influence relationship; The specific implementation method of step S2 includes the following steps: S2.
1. Set the initial environment of the hypersonic vehicle; S2.
2. Define the state corresponding to the hypersonic vehicle; S2.
3. Definition: Fuel consumption is mainly affected by the following factors: flight altitude h affects atmospheric density and aerodynamic drag, speed v affects power demand and aerodynamic drag, and wing sweep angle Λ affects the aerodynamic characteristics of the aircraft, including lift and drag; S2.
4. Through the target state, calculate the state transition F to give the recursive relationship of the target state s'=F(s,a), where a is the action and s is the state; S3. Input the environment designed in step S2 into the network designed in step S1 to obtain an actor network and a critic network based on a hypersonic aircraft, collect experience values through interactive training, and then use the DDPG algorithm improved by the successful sample playback method to train the actor network and critic network based on the hypersonic aircraft using the collected experience to obtain a wing sweep angle deformation optimization model based on aircraft fuel economy; The specific implementation method of step S3 includes the following steps: S3.
1. Collect experience: Execute actions in the environment and store the transformation (s, a, r, s′) of each step in the experience replay buffer; S3.
2. Set the improved DDPG algorithm for successful sample playback method: In the replay buffer, assume that the agent reaches state s* before the fuel is about to run out in a certain attempt, set state s* as the new goal g'=s*, recalculate the reward, and the new goal is to reach state s*, set the new reward r HER is the negative index of the distance to the target state; S3.
3. Sampling from the experience replay buffer, randomly extracting samples for training the network, and obtaining a training data set; S3.
4. Train the critic network based on the training data set obtained in step S3.3, use each (s, g, a, r, s′, g′) in the sampled S to calculate the current critic network C(s, a), and use the target critic network C′ and the target actor network to calculate the value Q′ of the target optimal problem; S3.
5. Using the updated critic network, we get the value of the optimal problem Q = C(s, a), define the index of the actor network as J = ∑r, and the weight θ of J with respect to the actor network π Find the gradient; S3.
6. Update the weights of the target actor and critic network by soft updating method; S4. The wing sweep angle deformation optimization model based on aircraft fuel economy obtained in step S3 is installed on the hypersonic aircraft computer to make an optimization decision on the wing sweep angle under flight conditions.
2. The method for optimizing wing sweep angle deformation based on aircraft fuel economy according to claim 1, characterized in that: The specific implementation method of step S1 includes the following steps: S1.
1. Construct an actor network to learn and decide actions from the environment state, including 1 input layer, 2 hidden layers and 1 output layer. The input and output expressions of the actor network are: a=π(s) (1) Among them, a is the action, s is the state, and π is the nonlinear function relationship expressed by the actor network; S1.
2. Construct a critic network to score the actions of the actor network, including 1 input layer, 2 hidden layers and 1 output layer. The input and output expressions of the critic network are: Q=C(s,a) (2) Among them, Q is the value of the optimal problem, and C is the nonlinear function relationship expressed by the critic network; S1.
3. Set the functional relationship π and Q of the actor and critic networks to be calculated by the forward calculation method of the neural network. In the actor network, the output of the lth layer is represented as a (l) , l={1,2}, the calculation is divided into two steps, linear transformation and nonlinear activation function, the expression is: z (l) =W (l) a (l-1) +b (l) (3) a (l) =σ (l) (z (l) ) (4) Among them, z (l) is the linear transformation of the lth layer, W (l) is the weight matrix of the lth layer, b (l) is the bias vector of the lth layer, σ (l) is the activation function σ of the lth layer (l) (x) is the nonlinear activation function of the lth layer, set to tanh(x); θ is the weight of the neural network; The output of the entire network is the activation of the last layer, so a = a (2) ; For the critic network, the output is calculated by the forward calculation method of the neural network: Q=a (2) (5)。 3. The method for optimizing wing sweep angle deformation based on aircraft fuel economy according to claim 2, characterized in that: The input layer of the actor network in step S1.1 is set with 4 input nodes corresponding to 4 state variables, the hidden layer includes 256 nodes, and the output layer includes 1 node; The input layer of the critic network in step S1.2 is set with 4 input nodes corresponding to 3 state variables and 1 action, the hidden layer includes 64 nodes, and the output layer includes 1 node to score the actor network action.
4. The method for optimizing wing sweep angle deformation based on aircraft fuel economy according to claim 3, characterized in that: The specific implementation method of step S2 includes the following steps: S2.
1. Set the initial environment of the hypersonic aircraft, including the flight altitude range of 15-30km, the flight speed range of 1200-2400m / s, and the initial total fuel m fuel 500kg, net weight m dry is 3000t, and the initial weight is m0=m fuel +m dry ; S2.
2. Define the state corresponding to the hypersonic aircraft to include flight altitude h, speed v, range R and current mass m. The expression of the state corresponding to the hypersonic aircraft is: s=[h,v,R,m] (6) The action corresponding to the hypersonic aircraft is defined as the wing sweep angle Λ, which is expressed as: a=Λ (7) Among them, the range of Λ is 0 to 30 degrees; S2.
3. Definition: Fuel consumption is mainly affected by the following factors: h affects atmospheric density and aerodynamic drag, v affects power demand and aerodynamic drag, and Λ affects the aerodynamic characteristics of the aircraft, including lift and drag. The expression of fuel consumption E(h, v, Λ) is: E(h,v,Λ)=k1·D(h,v,Λ)+k2·P(h,v,Λ) (8) Where D(h,v,Λ) is the aerodynamic drag, P(h,v,Λ) is the thrust, k1 and k2 are the first and second normal constants; The reward for designing a hypersonic vehicle is related to fuel consumption and its impact, and the expression is: r(s,a)=w1△E+w2△R (9) Where r(s,a) is the reward for taking action a in state s, △E is the improvement in fuel efficiency compared with the previous decision, △E=E(h',v',Λ')-E(h,v,Λ); Among them, h', v', Λ' and R' are parameters corresponding to the target state s', w1 and w2 are the first weight parameter and the second weight parameter, which are used to adjust the current optimization weight of fuel efficiency; S2.
4. By calculating the state transition F through the target state, the recursive relation expression of s'=F(s,a) is given as: where ζ is the flight path angle.
5. The method for optimizing wing sweep angle deformation based on aircraft fuel economy according to claim 4, characterized in that: The specific implementation method of step S3 includes the following steps: S3.
1. Collecting experience: Execute actions in the environment and store the transformation (s, a, r, s′) of each step into the experience replay buffer, expressed as: Where O is OU noise, which is given by the stochastic differential equation dO t =ω(μ-O t )dt+υdW t We can get: ω is the rate parameter, which determines the speed at which the system returns to the mean μ; μ is the long-term mean, around which the process will oscillate; υ is the intensity of the noise, which affects the amplitude of the noise fluctuation; dW t is the Brownian motion increment, which represents the random disturbance over time; S3.
2. Set the improved DDPG algorithm for successful sample playback method: In the replay buffer, assume that the agent reaches state s* before the fuel is about to run out in a certain attempt, set state s* as the new goal g'=s*, recalculate the reward, and the new goal is to reach state s*, set the new reward r HER is the negative index of the distance to the target state, and the expression is: Input g' and s' into the network and recalculate the critic's value C(s,g′,a); The goal of the DDPG algorithm improved by the successful sample replay method is to release the constraints in the case of task failure, so that the agent can extract the trend that reduces the fuel consumption rate in this case; S3.
3. Sampling from the experience replay buffer, randomly extracting samples for training the network, and obtaining the training data set S, expressed as: S={(s i ,g i ,a i ,r i ,s i ′,g i ′)∣i=1,2,…,M} (13) Among them, M is a fixed sampling sample, and M is set to 128; S3.
4. Train the critic network based on the training data set obtained in step S3.3, use each (s, g, a, r, s′, g′) in the sampled S, calculate the current critic network C(s, a), and use the target critic network C′ and the target actor network to calculate the value Q′ of the target optimal problem, expressed as: Q′=C′(s′,π(s′)) (14) Then use the value of the target optimal problem adjusted by HER to calculate the target value y, expressed as: y=r+γQ′ (15) Among them, γ is the discount coefficient, and γ=0.99; The expression for the indicator L of the critic network is defined as: The weight θ of L on the critic network C Finding the Gradient Let the gradient be zero and get the critic network weight θ that minimizes L C ; S3.
5. Using the updated critic network, we get Q = C(s, a), define the actor network index as J = ∑r, and give J the weight θ with respect to the actor network π Find the gradient and get the expression: Let the gradient is zero, and the actor network weight θ that maximizes the index J is obtained π' ; S3.
6. Update the weights of the target actor and critic networks through the soft update method, respectively: Among them, τ is the update weight, and τ=0.1 is set.
6. The method for optimizing wing sweep angle deformation based on aircraft fuel economy according to claim 5, characterized in that: In step S3, the target g=m-3000>0 is defined, that is, the fuel must not be exhausted.
7. An electronic device, characterized in that: The invention comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of a method for optimizing wing sweep angle deformation based on aircraft fuel economy as described in any one of claims 1 to 6 when executing the computer program.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for optimizing wing sweep angle deformation based on aircraft fuel economy as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Heterogeneous wireless network vertical switching method based on reinforcement learning TD3 algorithm
CN113784410A
Target hierarchical architecture-based deformable aircraft intelligent planning method and system
CN117268391A