Aircraft high lift device two-dimensional to three-dimensional optimization method based on deep reinforcement learning

By decomposing the optimization problem of the three-dimensional aircraft lifting device into multiple two-dimensional profiles, using deep reinforcement learning and flow field localization, the problems of medium and high cost and high-dimensional action space of three-dimensional optimization are solved, and efficient and fast three-dimensional optimization is achieved, with versatility and adaptability.

CN120470904APending Publication Date: 2025-08-12BEIHANG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510550317.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

In the optimization of deep reinforcement learning, there are problems of inefficiency in strategy exploration caused by high-cost three-dimensional flow field calculation and high-dimensional action space in the optimization of three-dimensional aircraft lifting devices.

Method used

The three-dimensional optimization problem is simplified into multiple two-dimensional profile optimization problems, and the agent is trained using two-dimensional computing fluid mechanics environment. Through the flow field locality and strategic translation invariance, a two-dimensional to three-dimensional optimization method for aircraft lifting devices based on deep reinforcement learning is established, and the two-dimensional to three-dimensional sections are decomposed into multiple two-dimensional profiles independently optimized.

Benefits of technology

It significantly reduces the training cost of three-dimensional optimization problems, improves the efficiency of strategy exploration, and achieves efficient and fast three-dimensional optimization, which is versatile and adaptable, and can cope with variable design conditions and enhance the appearance of the device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120470904A_ABST
    Figure CN120470904A_ABST
Patent Text Reader

Abstract

The invention discloses an aircraft high lift device two-dimensional to three-dimensional optimization method based on deep reinforcement learning, and belongs to the technical field of aircrafts. According to the method, in an environment based on two-dimensional computational fluid mechanics, an efficient and universal aircraft high lift device optimization strategy suitable for a target three-dimensional optimization problem is trained, the training cost of the three-dimensional optimization problem is remarkably reduced, and high calculation overhead of traditional three-dimensional flow field simulation is avoided; the flow field locality and strategy translation invariance are utilized to decompose a three-dimensional problem into a plurality of two-dimensional profiles for independent optimization, so that an intelligent agent only needs to process a low-dimensional action space, and the problem of curse of dimensionality caused by a high-dimensional action space is effectively solved; when the trained intelligent agent is used for solving the similar optimization problem, the optimization efficiency is high, the convergence speed is high, and the method has high universality for variable design working conditions and the appearance of the high lift device by means of the previous experience for solving the similar problem.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of aircraft technology, and in particular to a two-dimensional to three-dimensional optimization method for an aircraft high-lift device based on deep reinforcement learning. Background Art

[0002] Deep reinforcement learning is an efficient and versatile aerodynamic shape optimizer. It has been successfully applied in the optimization of aerodynamic shapes such as airfoils, wings, lift-enhancing devices, and propellers. It has very bright development prospects and high research value.

[0003] The advantage of deep reinforcement learning-based aerodynamic shape optimization lies in the versatility of the optimization strategies learned from examples across different design conditions and initial geometries, as well as their high optimization efficiency. This means that once the agent is trained, it can serve as an efficient and universal optimizer for similar shape optimization problems.

[0004] Although the trained agent can still complete the optimization task efficiently at a low computational cost, the data-driven nature of aerodynamic shape optimization based on deep reinforcement learning makes it difficult to perform high-cost aerodynamic analysis and high-dimensional design space, and the training cost of the agent strategy for three-dimensional problems is very high.

[0005] Transfer learning is a commonly used method to reduce the cost of aerodynamic analysis. A neural network is pre-trained in a two-dimensional environment, and then its parameters are transferred to the three-dimensional environment of the target optimization problem for further training. Although transfer learning can significantly reduce training costs, the high cost of three-dimensional flow field calculations is still unavoidable.

[0006] In deep reinforcement learning, agents must fully explore their environment, and the action space can have more than a dozen dimensions. Exploring and learning strategies in such a high-dimensional action space requires massive sampling. In active flow control based on deep reinforcement learning, the locality and translational invariance of both the controlled system and the agent are exploited to address the "curse of dimensionality" in the action space. High-dimensional actions are split among multiple parallel agents, each of which only needs to explore and learn within a small action subspace, significantly improving sampling efficiency. Scholars have pointed out that locality and translational invariance are key to addressing the "curse of dimensionality" in deep reinforcement learning for three-dimensional problems.

[0007] Based on this, a two-dimensional to three-dimensional optimization method for aircraft high-lift devices based on deep reinforcement learning is proposed. Summary of the Invention

[0008] The purpose of the present invention is to provide a two-dimensional to three-dimensional optimization method for aircraft high-lift devices based on deep reinforcement learning to solve the problems in the background technology.

[0009] To achieve the above objectives, the present invention provides a two-dimensional to three-dimensional optimization method for aircraft high-lift devices based on deep reinforcement learning, comprising the following steps:

[0010] S1. Modeling the target three-dimensional optimization problem for the aircraft high-lift device, simplifying the target three-dimensional optimization problem into a corresponding two-dimensional optimization problem and modeling it; both the target three-dimensional optimization problem and the two-dimensional optimization problem include definitions of objective functions, constraints, and design variables;

[0011] S2. Establish a two-dimensional reinforcement learning model based on two-dimensional computational fluid dynamics for aerodynamic optimization and obtain an optimization strategy for a simplified two-dimensional optimization problem;

[0012] S3. Testing the optimization strategy under different design conditions or different lift-enhancing device shapes to further obtain the optimal optimization strategy;

[0013] S4. Based on the two-dimensional reinforcement learning model and optimal optimization strategy, a three-dimensional reinforcement learning model based on three-dimensional computational fluid dynamics is established to solve the target three-dimensional optimization problem. After a small number of iterations, the optimal design variables for the target three-dimensional optimization problem are obtained.

[0014] S5. Extend the optimal optimization strategy to three-dimensional optimization problems with the same objective function, constraints, and design variables under multiple design conditions and different high-lift device shapes, thus obtaining an efficient and universal aircraft three-dimensional high-lift device optimizer.

[0015] Preferably, in S1:

[0016] 1) For a three-dimensional optimization problem: the objective function is to maximize the maximum lift coefficient at the stall angle of attack, and the constraint is that the lift coefficient at medium angles of attack is not less than the initial value;

[0017] The geometric shape of the three-dimensional flap in the aircraft high-lift device is decomposed into multiple spanwise sections. Design variables are independently defined for each section. Depending on the complexity of the flap, the design variables of a flap include one or two sets of dimensionless slot parameters. Each set of slot parameters includes the overlap amount and slot width of the local spanwise section.

[0018] 2) For the two-dimensional optimization problem: the objective function and constraints are the same as those for the three-dimensional optimization problem. The design variables are a set of seam parameters, including the overlap amount and seam width of the local spanwise section.

[0019] Preferably, in S2, the specific steps of using the two-dimensional reinforcement learning model for aerodynamic optimization are:

[0020] S21. Uniformly and densely discretize the design space corresponding to the design variables in the two-dimensional optimization problem. Then, through geometric modeling, mesh deformation, and numerical calculation, obtain aerodynamic data for each sample design point, and establish a sample library consisting of the design variables and their corresponding aerodynamic data; the aerodynamic data includes the objective function, constraints, and states corresponding to the values of each design variable.

[0021] S22. Establish a two-dimensional reinforcement learning model based on two-dimensional computational fluid dynamics. The two-dimensional reinforcement learning model includes the definition of the agent, environment, action, state, and reward function, and establishes the interaction process between the agent and the two-dimensional environment.

[0022] S23. Based on the sample library, the value-based reinforcement learning algorithm is used to perform offline training on the interaction process between the agent and the environment to obtain the optimization strategy for the simplified two-dimensional optimization problem.

[0023] Preferably, the geometric modeling update, mesh deformation and numerical calculation in S21 are specifically as follows: geometric modeling of the aircraft high-lift device is performed, and the geometric model is updated according to the values of the design variables, the initial mesh of the geometric model is deformed, and the RANS equation and the kω-sst turbulence model are used to solve the flow field information including the lift coefficient and the velocity field at a given angle of attack.

[0024] Preferably, in S22, the intelligent agent is an algorithm for optimizing strategy learning and execution, the environment in which the intelligent agent interacts is a flow field simulation based on two-dimensional computational fluid dynamics, the action is a change in a discretized design variable, and the state is a velocity field near the flap at a stall angle of attack and a medium angle of attack;

[0025] The reward function is related to the objective function and constraints of the two-dimensional optimization problem. When the action causes the value of the design variable to exceed the design space or the constraint is not satisfied, the reward is negative. In other cases, the reward function is the increment of the objective function at the current time step.

[0026] The initial design variables of the interaction process between the agent and the environment are randomly selected from the sample library. The specific interaction process is: at a certain time step in a certain round, the agent inputs the environment state of the current time step and outputs the action to be taken in the current time step. The environment uses the action as input and feeds back to the agent the reward function value of the current time step and the state of the next time step.

[0027] Preferably, in S23, the value estimation function of the value-based reinforcement learning algorithm adopts a Q network, and during the offline training process, the intelligent agent takes the action corresponding to the highest value with a gradually increasing probability.

[0028] Preferably, in S3, the specific process of testing the optimization strategy is: testing the performance of the optimization strategy learned by the intelligent agent when the design working conditions of the sample library or the shape of the lift-enhancing device are different from those in S23; if the optimal design variables cannot be efficiently searched, return to S23, and modify the design working conditions or the shape of the lift-enhancing device and retest until the conditions are met and the optimal optimization strategy is obtained.

[0029] Preferably, the three-dimensional reinforcement learning model in S4 includes definitions of the agent, environment, action, state, and reward function, and establishes a real-time interaction process between the agent and the three-dimensional environment;

[0030] Among them, the intelligent agent adopts the optimal optimization strategy obtained by S3, and the environment in which the intelligent agent interacts is a flow field simulation based on three-dimensional computational fluid dynamics. The action is the change in the discretized design variables of each spanwise section, and the state is the stall angle of attack of each spanwise section and the velocity field near the flap at medium angle of attack.

[0031] Preferably, during the real-time interaction between the intelligent agent and the three-dimensional environment, each time step includes the geometric modeling update, mesh deformation and numerical calculation of the three-dimensional lift-enhancing device. The intelligent agent synchronously processes the states from different span-wise sections respectively and outputs the actions of the local sections. The environment receives these actions, synchronously modifies the design variables of the different span-wise sections respectively, and uses the three-dimensional computational fluid dynamics method to obtain the state of each span-wise section in the next time step and feeds it back to the intelligent agent.

[0032] Preferably, in said S5, when the optimal optimization strategy is applied to a design operating condition or a high-lift device shape other than that used in the S3 training stage, no additional training process is required.

[0033] Therefore, the present invention provides a two-dimensional to three-dimensional optimization method for aircraft high-lift devices based on deep reinforcement learning, which has the following beneficial effects:

[0034] (1) By training the optimization strategy completely based on a two-dimensional computational fluid dynamics environment, the high-cost three-dimensional flow field simulation calculations in traditional methods are avoided, significantly reducing the training cost of three-dimensional optimization problems.

[0035] (2) By utilizing the locality of flow field information and the translation invariance of the strategy, the three-dimensional problem is decomposed into multiple two-dimensional sections for independent optimization. Each agent only needs to process the low-dimensional action space, which overcomes the problem of low strategy exploration efficiency caused by the high-dimensional action space in traditional deep reinforcement learning and effectively addresses the "curse of dimensionality".

[0036] (3) When solving similar optimization problems, the trained intelligent agent can draw on the experience of solving similar problems in the past, with high optimization efficiency and fast convergence speed. It is highly versatile for various design conditions and lift-enhancing device shapes.

[0037] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is a flow chart of an embodiment of the present invention;

[0039] Figure 2 Schematic diagram of various cross-sectional design variables of a high-lift device according to an embodiment of the present invention, wherein a represents a flap and b represents a main wing;

[0040] Figure 3 A schematic diagram of a calculation grid for a two-dimensional lift-enhancing device according to an embodiment of the present invention;

[0041] Figure 4 Schematic diagram of the interaction process between an agent and an environment according to an embodiment of the present invention;

[0042] Figure 5 Schematic diagram of a dual-depth Q network algorithm according to an embodiment of the present invention;

[0043] Figure 6 A schematic diagram of a Q network according to an embodiment of the present invention;

[0044] Figure 7 A schematic diagram of a calculation grid for a three-dimensional lift-enhancing device according to an embodiment of the present invention;

[0045] Figure 8 This is a schematic diagram of the application of the agent strategy migration to a three-dimensional problem according to an embodiment of the present invention;

[0046] Figure 9 Schematic diagram of the process of solving the optimization problem from two dimensions to three dimensions according to an embodiment of the present invention. DETAILED DESCRIPTION

[0047] The technical solution of the present invention is further described below with reference to the accompanying drawings and embodiments.

[0048] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments.

[0049] DLR-F11 is the implementation object of the performance test of the high-lift device optimization method commonly used in the field of aircraft design. Its configuration is a wing-body combination. The high-lift device of the wing includes leading edge slats, main wing and flaps. The configuration of DLR-F11 is as follows Figure 2 shown.

[0050] Example

[0051] Such as Figure 1As shown, the present invention provides a two-dimensional to three-dimensional optimization method for aircraft high-lift devices based on deep reinforcement learning. The method is applied to the slot parameter optimization of the DLR-F11 flap, comprising the following steps:

[0052] S1. Modeling the target three-dimensional optimization problem of the aircraft high-lift device, simplifying the target three-dimensional optimization problem into a corresponding two-dimensional optimization problem and modeling it. The simplified two-dimensional optimization problem is actually the optimization of the two-dimensional high-lift device;

[0053] Among them, both the target three-dimensional optimization problem and the two-dimensional optimization problem include the definition of objective function, constraints, and design variables:

[0054] 1) For a three-dimensional optimization problem: the objective function is to maximize the maximum lift coefficient at the stall angle of attack, with the constraint that the lift coefficient at intermediate angles of attack is no less than the initial value; the design variables are the dimensionless slot parameters of the flap, including the overlap and slot width;

[0055] The geometric shape of the three-dimensional flap in the aircraft high-lift device is decomposed into multiple spanwise sections. Design variables are independently defined for each section. Depending on the complexity of the flap, the design variables for a flap include one or two sets of slot parameters. Each set of slot parameters includes the overlap amount and slot width of the local spanwise section.

[0056] In this embodiment, the design variables are the slot parameters of the DLR-F11 flap, including the overlaps O1 and O2 of the inner and outer spanwise sections and the slot widths G1 and G2;

[0057] The ranges of the design variables are:

[0058] -0.2%≤O1≤2%,0.5%≤G1≤2.5%;

[0059] -0.2%≤O2≤2%,0.5%≤G2≤2.5%;

[0060] The optimization goal is to maximize the maximum lift coefficient at an angle of attack of 20°, and the constraint is that the lift coefficient at an angle of attack of 10° is not less than 2.0.

[0061] 2) For the 2D optimization problem: the objective function and constraints are the same as those for the 3D optimization problem, and the design variables are a set of seam parameters, including the overlap and seam width of the local spanwise section.

[0062] In this embodiment, the object of the two-dimensional optimization problem is a DLR-F11 multi-section airfoil with a chord length of 1 m, and the design variables are the flap slot parameters of the multi-section airfoil, such as Figure 2 As shown, it includes overlap O and seam width G.

[0063] The ranges of the design variables are:

[0064] -0.2%≤O≤2%,0.5%≤G≤2.5%;

[0065] The optimization goal is to maximize the maximum lift coefficient at an angle of attack of 23°, and the constraint is that the lift coefficient at an angle of attack of 10° is not less than 3.6.

[0066] S2. Based on the simplified two-dimensional optimization problem, a two-dimensional reinforcement learning model based on two-dimensional computational fluid dynamics is established for aerodynamic optimization, and the optimization strategy for the simplified two-dimensional optimization problem is obtained, specifically:

[0067] S21. The design space corresponding to the design variables in the two-dimensional optimization problem is discretized uniformly and densely. The discrete design points constitute a discrete sample library. Then, through geometric modeling, mesh deformation, and numerical calculation, aerodynamic data for each sample design point is obtained, and a sample library consisting of the design variables and their corresponding aerodynamic data is established. The aerodynamic data includes the objective function, constraints, and states corresponding to the values of each design variable.

[0068] The discretization of the design space should be sufficiently uniform and dense so that all design variables that may be accessed during the interaction between the agent and the environment can be searched in the sample library. To this end, the initial design variables of the interaction process should be randomly selected from the sample library.

[0069] In this embodiment, uniform sampling is performed in the design space, and the discrete step lengths of the design variables O and G are both 0.1%, so the sample size is 23×21=483, and the sample library is established. When generating the sample library, the design conditions are defined as Mach number and Reynolds number, which are Ma=0.2 and Re=3×10 6 The computational grid uses a structured grid, such as Figure 3 As shown, the total number of grids is about 46,000.

[0070] The geometric modeling update, mesh deformation and numerical calculation are specifically as follows: geometric modeling of the aircraft high-lift device is carried out, and the geometric model is updated according to the values of the design variables. The initial mesh of the geometric model is deformed to adapt to the updated geometric model. The RANS equation and kω-sst turbulence model are used to solve the flow field information including lift coefficient and velocity field under a given angle of attack.

[0071] S22. Establish a two-dimensional reinforcement learning model based on two-dimensional computational fluid dynamics. The two-dimensional reinforcement learning model includes the definition of the agent, environment, action, state, and reward function, and establishes the interaction process between the agent and the two-dimensional environment.

[0072] In this embodiment, the intelligent agent is an algorithm for optimizing strategy learning and execution, and the reinforcement learning algorithm selected is a dual-depth Q-network algorithm; the environment in which the intelligent agent interacts is a flow field simulation based on two-dimensional computational fluid dynamics; the state is the velocity field near the flap at a stall angle of attack of 23° and a medium angle of attack of 10°, and the velocity field extraction range is x / c∈0.83,1.14],y / c∈[-0.165,0.075]; x and y are coordinate system parameters, and c is the local airfoil chord length in meters.

[0073] The action is the change of the discretized design variables ΔO and ΔG, which contains 4 sets of discrete values, namely:

[0074]

[0075] The reward function is related to the objective function and constraints of the two-dimensional optimization problem. When the action causes the value of the design variable to exceed the design space or the constraint is not satisfied, the reward is negative. In other cases, the reward function is the increment of the objective function at the current time step.

[0076] The reward function of this embodiment is defined as follows:

[0077]

[0078] Where k is the scaling factor, (C L,α=23° ) t+1 is the objective function for the next time step, (C L,α=23° ) t is the objective function of the current time step, (C L,α=10° ) t+1 is the constraint for the next time step.

[0079] The initial design variables of the interaction process between the agent and the environment are randomly selected from the sample library. The specific interaction process is: at a certain time step in a certain round, the agent inputs the environmental state of the current time step and outputs the action to be taken in the current time step. The environment takes the action as input and searches the sample library for the design variables and its objective function, constraints and state for the next time step, and feeds back to the agent the reward function value of the current time step and the state of the next time step.

[0080] In this embodiment, the interaction process between the agent and the environment is as follows: Figure 4 As shown in the figure, the training consists of 30,000 rounds, with a maximum of 20 interaction steps per round. If the action causes the design variables to exceed the design space, the aerodynamic data of the next time step will not be searched, the next state will be recorded as the terminal state, the reward will be returned, and the round will be terminated.

[0081] S23. Based on the sample library, the value-based reinforcement learning algorithm is used to conduct offline training on the interaction process between the agent and the environment, and the optimization strategy for the simplified two-dimensional optimization problem is obtained;

[0082] The value estimation function of the value-based reinforcement learning algorithm uses a Q-network, a convolutional neural network whose input data is a three-dimensional tensor whose shape is consistent with the state, and whose output data is a vector whose dimension is consistent with the number of possible actions. During offline training, the agent takes the action corresponding to the highest value with gradually increasing probability.

[0083] In this embodiment, the reinforcement learning algorithm is selected as the dual-depth Q network algorithm, and its general process is as follows: Figure 5 As shown, the pseudo code of the dual-depth Q network algorithm is as follows:

[0084]

[0085]

[0086] The structure of the Q network is as follows Figure 6 As shown, the input is an H×W×2 state, where H is the number of rows in the velocity field matrix and W is the number of columns. The tensor depth is 2, representing a state containing a velocity field with two angles of attack. The neural network output is a 4-dimensional vector, whose elements correspond one to one to the value functions of the four discrete actions in the input state. The hidden layer consists of two convolutional layers, two pooling layers, and one fully connected layer. The final output layer uses a linear activation function, while the activation functions of all other layers are Reinforced Luminance (ReLU). The pooling operation in the pooling layer is max pooling with a window size of 2×2. The stride of each convolutional layer is 1, and the stride of each pooling layer is 2. The neural network is trained using Adam with a learning rate of 0.001.

[0087] S3. Test the optimization strategy under different design conditions or different lift-enhancing device shapes to further obtain the optimal optimization strategy, specifically:

[0088] The aerodynamic data in the sample library supports the interaction process between the intelligent agent and the environment. Using a suitable reinforcement learning algorithm, after several rounds of offline training with multiple time steps in each round, the optimal optimization strategy for the simplified two-dimensional optimization problem is obtained.

[0089] The specific process of testing the optimization strategy is as follows: the performance of the optimization strategy learned by the intelligent agent is tested when the design working conditions or the shape of the lift-boosting device in the sample library are different from those in S23. If the optimal design variables cannot be efficiently searched, the intelligent agent returns to S23 and modifies the design working conditions or the shape of the lift-boosting device to test again until the conditions are met (the optimal design variables can be efficiently searched), thereby obtaining the optimal optimization strategy.

[0090] In this embodiment, the variation range of the Mach number and Reynolds number in the design operating conditions during the inspection phase is:

[0091] 0.15≤Ma≤0.25;

[0092] 9×10 5 ≤Re≤9×10 6 ;

[0093] In addition, the shape of the lift-enhancing device is not limited to the multi-section airfoil of S1, and other single-slotted flap airfoils with the same deflection angle can be selected.

[0094] S4. Based on the two-dimensional reinforcement learning model and optimal optimization strategy, a three-dimensional reinforcement learning model based on three-dimensional computational fluid dynamics is established to solve the target three-dimensional optimization problem. After a small number of iterations, the optimal design variables for the target three-dimensional optimization problem are obtained.

[0095] The three-dimensional reinforcement learning model, in which the definitions of the agent, environment, action, state, and reward function correspond to those of the two-dimensional reinforcement learning model, establishes a real-time interaction process between the agent and the three-dimensional environment. The locality and translation invariance of the flow field information and the agent's strategy ensure the effectiveness of the learned strategy based on the two-dimensional environment in the three-dimensional computational fluid dynamics environment.

[0096] Among them, the intelligent agent adopts the optimal optimization strategy and strategy execution algorithm obtained by S3. The environment in which the intelligent agent interacts is a flow field simulation based on three-dimensional computational fluid dynamics. The state is the velocity field near the flap at the stall angle of attack of 20° and the medium angle of attack of 10° in each spanwise section of the inner and outer sections. The extraction range of the velocity field is x / c∈[0.83,1.14],y / c∈-0.165,0.075]. At this time, c contains the local wing chord length of the inner and outer sections, and the unit is m.

[0097] Action is the change in the discretized design variables of each spanwise section, which contains 4 groups of discrete values, namely:

[0098]

[0099] Where i = 1, 2, representing the inner and outer profiles respectively. Since the strategy is no longer updated, the reward function does not need to be defined.

[0100] During the real-time interaction between the intelligent agent and the three-dimensional environment, each time step includes the geometric modeling update, mesh deformation and numerical calculation of the three-dimensional lift-enhancing device (corresponding to the geometric modeling update, mesh deformation and numerical calculation in S21). The intelligent agent synchronously processes the states from different span-wise sections and outputs the actions of the local sections. The environment receives these actions, synchronously modifies the design variables of different span-wise sections, and uses three-dimensional computational fluid dynamics methods to obtain the states of each span-wise section in the next time step and feeds them back to the intelligent agent.

[0101] In this embodiment, the interaction process between the agent and the environment occurs separately, synchronously, and in real time on the inner and outer sections. As described in S1 and S4, the optimization problem and reinforcement learning modeling have the same definitions of design variables, actions, and states for the inner and outer sections. The agent trained in the two-dimensional environment is applicable to the environments of the inner and outer sections respectively. The flow field information required for each interaction is based on the three-dimensional computational fluid dynamics simulation of the entire aircraft. The numerical calculation adopts the following method: Figure 7 The structured grid shown in Figure 1 has a total of approximately 4 million grid cells. For each spanwise section, the agent receives the state of the local velocity field and determines the action for the next time step. This action acts on the design variables of the local section. After the design variables of all sections are updated, the environment feeds back the state of the local section for the next time step to the agent through geometric modeling, mesh deformation, and numerical calculation, and the cycle continues. The agents for all spanwise sections use the same optimization strategy, such as Figure 8 shown.

[0102] S5. Extend the optimal optimization strategy to three-dimensional optimization problems with the same objective function, constraints, and design variables for multiple design conditions and different high-lift device shapes within a certain range. Thus, an efficient and universal aircraft three-dimensional high-lift device optimizer is obtained.

[0103] Because this method takes advantage of the locality of flow field information and the translational invariance of the strategy, no additional training process is required when the optimal optimization strategy is applied to design conditions or lift-enhancing device shapes that are not used in the S3 training phase. During the training phase, the optimization strategy learned by the agent through the two-dimensional computational fluid dynamics environment can capture the general laws of aerodynamic optimization of lift-enhancing devices, and these laws remain applicable when different design conditions or shapes change. In addition, the translational invariance of the strategy enables the agent to directly transfer experience in the two-dimensional environment to the three-dimensional environment. Even when faced with new design conditions or shapes, the agent can efficiently adjust the design variables based on existing experience without the need for retraining. This feature significantly improves the versatility and adaptability of the method, enabling it to quickly respond to diverse optimization needs.

[0104] In this embodiment, the variation range of Mach number and Reynolds number is:

[0105] 0.15≤Ma≤0.25;

[0106] 9×10 5 ≤Re≤9×10 6 ;

[0107] The high-lift device's shape is not limited to the DLR-F11; other single-slotted flap airfoils with the same deflection angle can be selected. The number of high-lift devices deployed along the span is arbitrary, with each flap typically having a maximum of two design profiles. In the target three-dimensional optimization problem, the agent can simultaneously optimize the design variables of multiple spanwise profiles of multiple flaps during each interaction with the environment.

[0108] In summary, this method trains an efficient and general aircraft high-lift device optimization strategy for the target three-dimensional optimization problem in a two-dimensional computational fluid dynamics environment. The concise flow chart of the method is shown in the attached figure. Figure 9 shown.

[0109] Therefore, the present invention provides a two-dimensional to three-dimensional optimization method for aircraft lift-enhancing devices based on deep reinforcement learning. By training the optimization strategy in an environment completely based on two-dimensional computational fluidics, the training cost of three-dimensional optimization problems is significantly reduced, and the high computational overhead of traditional three-dimensional flow field simulation is avoided. At the same time, by utilizing the locality of the flow field and the translation invariance of the strategy, the three-dimensional problem is decomposed into multiple two-dimensional sections for independent optimization, so that the intelligent agent only needs to process the low-dimensional action space, effectively solving the "curse of dimensionality" problem caused by the high-dimensional action space. In addition, when solving similar optimization problems, the trained intelligent agent can draw on the experience of solving similar problems in the past, with high optimization efficiency and fast convergence speed, and has strong versatility for variable design conditions and lift-enhancing device shapes.

[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solutions of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A two-dimensional to three-dimensional optimization method for aircraft high-lift devices based on deep reinforcement learning, characterized in that: The following steps are involved: S1. Modeling the target three-dimensional optimization problem of the aircraft high-lift device, simplifying the target three-dimensional optimization problem into a corresponding two-dimensional optimization problem and modeling it; Both the target three-dimensional optimization problem and the two-dimensional optimization problem include the definition of the objective function, constraints, and design variables; S2. Establish a two-dimensional reinforcement learning model based on two-dimensional computational fluid dynamics for aerodynamic optimization and obtain an optimization strategy for a simplified two-dimensional optimization problem; S3. Testing the optimization strategy under different design conditions or different lift-enhancing device shapes to further obtain the optimal optimization strategy; S4. Based on the two-dimensional reinforcement learning model and optimal optimization strategy, a three-dimensional reinforcement learning model based on three-dimensional computational fluid dynamics is established. After a small number of iterations, the optimal design variables for the target three-dimensional optimization problem are obtained. S5. Extend the optimal optimization strategy to similar three-dimensional optimization problems with multiple design conditions and different high-lift device shapes, thus obtaining an efficient and universal aircraft three-dimensional high-lift device optimizer.

2. The method for optimizing an aircraft high-lift device from two dimensions to three dimensions based on deep reinforcement learning according to claim 1, characterized in that: In S1: 1) For a three-dimensional optimization problem: the objective function is to maximize the maximum lift coefficient at the stall angle of attack, and the constraint is that the lift coefficient at medium angles of attack is not less than the initial value; The geometric shape of the three-dimensional flap in the aircraft high-lift device is decomposed into multiple spanwise sections. Design variables are independently defined for each section. Depending on the complexity of the flap, the design variables of a flap include one or two sets of dimensionless slot parameters. Each set of slot parameters includes the overlap amount and slot width of the local spanwise section. 2) For the two-dimensional optimization problem: the objective function and constraints are the same as those for the three-dimensional optimization problem. The design variables are a set of seam parameters, including the overlap amount and seam width of the local spanwise section.

3. The method for optimizing an aircraft high-lift device from two dimensions to three dimensions based on deep reinforcement learning according to claim 1, characterized in that: In S2, the specific steps of using the two-dimensional reinforcement learning model for aerodynamic optimization are: S21. Uniformly and densely discretize the design space corresponding to the design variables in the two-dimensional optimization problem. Then, through geometric modeling, mesh deformation, and numerical calculation, obtain aerodynamic data for each sample design point, and establish a sample library consisting of the design variables and their corresponding aerodynamic data; the aerodynamic data includes the objective function, constraints, and states corresponding to the values of each design variable. S22. Establish a two-dimensional reinforcement learning model based on two-dimensional computational fluid dynamics. The two-dimensional reinforcement learning model includes the definition of the agent, environment, action, state, and reward function, and establishes the interaction process between the agent and the two-dimensional environment. S23. Based on the sample library, the value-based reinforcement learning algorithm is used to perform offline training on the interaction process between the agent and the environment to obtain the optimization strategy for the simplified two-dimensional optimization problem.

4. The method for optimizing an aircraft high-lift device from two dimensions to three dimensions based on deep reinforcement learning according to claim 3, characterized in that: The geometric modeling update, mesh deformation and numerical calculation in S21 are specifically as follows: geometric modeling of the aircraft high-lift device is performed, and the geometric model is updated according to the values of the design variables, the initial mesh of the geometric model is deformed, and the RANS equation and the kω-sst turbulence model are used to solve the flow field information including the lift coefficient and the velocity field at a given angle of attack.

5. The method for optimizing an aircraft high-lift device from two dimensions to three dimensions based on deep reinforcement learning according to claim 4, characterized in that: In S22, the intelligent agent is an algorithm for optimizing strategy learning and execution, the environment in which the intelligent agent interacts is a flow field simulation based on two-dimensional computational fluid dynamics, the action is the change in the discretized design variable, and the state is the velocity field near the flap at stall angle of attack and medium angle of attack; The reward function is related to the objective function and constraints of the two-dimensional optimization problem. When the action causes the value of the design variable to exceed the design space or the constraint is not satisfied, the reward is negative. In other cases, the reward function is the increment of the objective function at the current time step. The initial design variables of the interaction process between the agent and the environment are randomly selected from the sample library. The specific interaction process is: at a certain time step in a certain round, the agent inputs the environment state of the current time step and outputs the action to be taken in the current time step. The environment uses the action as input and feeds back to the agent the reward function value of the current time step and the state of the next time step.

6. The method for optimizing an aircraft high-lift device from two dimensions to three dimensions based on deep reinforcement learning according to claim 5, characterized in that: In S23, the value estimation function of the value-based reinforcement learning algorithm adopts the Q network, and during the offline training process, the intelligent agent takes the action corresponding to the highest value with a gradually increasing probability.

7. The method for optimizing an aircraft high-lift device from two dimensions to three dimensions based on deep reinforcement learning according to claim 6, characterized in that: In S3, the specific process of testing the optimization strategy is: testing the performance of the optimization strategy learned by the intelligent agent when the design working conditions or the shape of the lift-enhancing device in the sample library are different from those in S23. If the optimal design variables cannot be efficiently searched, return to S23, modify the design working conditions or the shape of the lift-enhancing device, and retest until the conditions are met to obtain the optimal optimization strategy.

8. The method for optimizing an aircraft high-lift device from two dimensions to three dimensions based on deep reinforcement learning according to claim 1, characterized in that: The three-dimensional reinforcement learning model in S4 includes the definition of the agent, environment, action, state and reward function, and establishes the real-time interaction process between the agent and the three-dimensional environment; Among them, the intelligent agent adopts the optimal optimization strategy obtained by S3, and the environment in which the intelligent agent interacts is a flow field simulation based on three-dimensional computational fluid dynamics. The action is the change in the discretized design variables of each spanwise section, and the state is the stall angle of attack of each spanwise section and the velocity field near the flap at medium angle of attack.

9. The method for optimizing an aircraft high-lift device from two dimensions to three dimensions based on deep reinforcement learning according to claim 8, characterized in that: During the real-time interaction between the intelligent agent and the three-dimensional environment, each time step includes the geometric modeling update, mesh deformation and numerical calculation of the three-dimensional lift-enhancing device. The intelligent agent synchronously processes the states from different span-wise sections and outputs the actions of the local sections. The environment receives these actions, synchronously modifies the design variables of the different span-wise sections, and uses three-dimensional computational fluid dynamics methods to obtain the states of each span-wise section in the next time step and feeds them back to the intelligent agent.

10. The method for optimizing an aircraft high-lift device from two dimensions to three dimensions based on deep reinforcement learning according to claim 1, characterized in that: In S5, when the optimal optimization strategy is applied to design conditions or high-lift device shapes other than those used in the S3 training phase, no additional training process is required.

Citation Information

Cited By

  • Three-dimensional building grid generation method based on human feedback reinforcement learning

    CN120874210A