Two-dimensional principal component analysis method for reinforcement learning of multi-section airfoil optimization strategy

Through two-dimensional principal component analysis of dimensional flow field information and combined with deep Q network training agents, the generality and high cost problems of reinforcement learning aerodynamic appearance optimization methods under different operating conditions are solved, and efficient optimization strategy learning across operating conditions and across configurations is achieved.

CN120470905APending Publication Date: 2025-08-12BEIHANG UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510550746.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

In the prior art, reinforcement learning aerodynamic appearance optimization method with geometric parameters as the state is insufficient in different design conditions, and the state dimension is too large when using flow field information as the state, resulting in high cost of training of neural networks.

Method used

The flow field information is reduced by dimensionality using the two-dimensional principal component analysis method, and the feature field after dimensionality reduction is used as the state input for reinforcement learning. Combined with the deep Q network algorithm, the agent is trained to learn cross-work conditions and cross-configuration optimization strategies.

Benefits of technology

It realizes efficient optimization of the agent in different design scenarios and multi-stage airfoil shapes, reduces the number of parameters and calculation costs of neural network training, and has the ability to directly migrate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120470905A_ABST
    Figure CN120470905A_ABST
Patent Text Reader

Abstract

The invention discloses a two-dimensional principal component analysis method for reinforcement learning of a multi-section airfoil optimization strategy, and belongs to the technical field of aircrafts, and the method comprises the following steps: S1, defining a multi-section airfoil optimization problem; s2, establishing a pneumatic data sample library by using Latin hypercube sampling; s3, performing dimension reduction on the velocity field matrix by adopting a two-dimensional principal component analysis method; s4, establishing a reinforcement learning model of the optimization strategy; and S5, training the reinforcement learning model to obtain an optimal optimization strategy with direct migration capability. According to the method, while the effectiveness of the reinforcement learning agent in observing the environment state is ensured, the dimensionality of the state is effectively reduced, the number of layers of the neural network and the number of parameters required by optimization strategy learning are reduced, and the trained optimization strategy is suitable for various design working conditions and multi-section airfoil profiles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of aircraft technology, and in particular to a two-dimensional principal component analysis method for reinforcement learning of a multi-segment airfoil optimization strategy. Background Art

[0002] Reinforcement learning, as an efficient and versatile aerodynamic shape optimization method, has demonstrated its unique advantages. The optimization strategies derived through sample learning not only exhibit excellent adaptability to operating conditions and geometric generalization, but also enable efficient parameter optimization. This property enables the trained agent to be directly used as a general optimizer, effectively solving similar aerodynamic shape optimization problems. This method has demonstrated significant advantages in aerodynamic optimization applications such as airfoils, wings, lift-enhancing devices, and propellers.

[0003] State is a fundamental concept in reinforcement learning. It defines the agent's observation of the environment in a specific control or optimization problem. This observation is usually incomplete. From the perspective of the state space, aerodynamic shape optimization methods based on reinforcement learning can be divided into two categories: value function estimation using geometric parameters as state and value function estimation using flow field information as state.

[0004] In aerodynamic shape optimization, the state definition is mostly based on geometry, with only a few based on physical flow field characteristics such as velocity and pressure. However, aerodynamic shape optimization based on reinforcement learning, which uses geometric parameters as states, is more like an agent's reinforced memorization of the optimal path from the initial design to the optimal design. The generalizability of the trained agent to other design conditions (such as different Mach and Reynolds numbers) is questionable. Given that the flow field is more directly related to the objective function than the geometry, flow field information can more effectively reflect the value of the current state of the environment and more universally guide action decisions.

[0005] Despite this, using the current design variables and updated design variables as states and actions, respectively, has been a common choice in research over the past five years. The few studies that use flow field information as state have a very small state dimension, which cannot fully represent the flow field characteristics. Using the entire flow field as state inevitably leads to an excessively large state dimension, which increases the training cost of the neural network.

[0006] Therefore, how to provide an aerodynamic shape optimization method that takes flow field information as input and introduces two-dimensional principal component analysis technology to reduce the dimension of flow field information is a problem that technical personnel in this field urgently need to solve. Summary of the Invention

[0007] The purpose of the present invention is to provide a two-dimensional principal component analysis method for reinforcement learning of multi-segment airfoil optimization strategy to solve the problems in the background technology.

[0008] To achieve the above objectives, the present invention provides a two-dimensional principal component analysis method for reinforcement learning of multi-segment airfoil optimization strategy, comprising the following steps:

[0009] S1. Define the optimization objective, constraints, and design variables for the multi-segment airfoil optimization problem.

[0010] S2. Use Latin hypercube sampling to collect samples in the design space composed of design variables, perform geometric modeling, mesh deformation, and numerical calculation on the samples, and obtain the velocity field matrix near the flap corresponding to each design variable value. All velocity field matrices near the flap constitute an aerodynamic data sample library;

[0011] S3, using the two-dimensional principal component analysis method to learn the projection matrix, using the projection matrix to reduce the velocity field matrix into a feature field, retaining key flow field information;

[0012] S4. Build a reinforcement learning model for the optimization strategy based on S1. The reinforcement learning model includes the definition of the agent, the environment, the agent's actions, the state of the environment's feedback to the agent, and the reward function.

[0013] S5. Based on computational fluid dynamics simulation, the interaction process between the intelligent agent and the environment is established to train the reinforcement learning model and obtain the optimal optimization strategy with direct transfer capability.

[0014] Preferably, in S1, the optimization objective is to maximize the maximum lift coefficient of the multi-section airfoil; the design variables are the slot parameters of the flap, including the overlap amount and the slot width; and the constraint is that the maximum lift coefficient is not less than the initial value.

[0015] Preferably, the geometric modeling, mesh deformation and numerical calculation in S2 are specifically as follows:

[0016] A geometric model of a multi-segment airfoil is constructed based on the sample, the grid topology of the geometric model is changed, and the numerical calculation uses a solver based on the RANS equation and SA turbulence model to obtain the velocity field information at the maximum lift coefficient.

[0017] Preferably, in S3, the dimension reduction of the velocity field matrix is performed in both row and column directions, and the row dimension reduction and the column dimension reduction correspond to different projection matrices respectively.

[0018] Preferably, the specific process of learning the projection matrix using the two-dimensional principal component analysis method is:

[0019] 1) For the velocity field data set, solve the projection vector X of row dimension reduction, satisfying

[0020] in, It means finding the projection vector X that maximizes the objective function J(X). The objective function is expressed as:

[0021] J(X)=X T G t X;

[0022] Where G t is the covariance matrix in the row direction, represents the mean operation, A is the velocity field dataset, is the mean matrix of the velocity field;

[0023] 2) Calculate the eigenvalues and eigenvectors of the row covariance matrix, arrange the eigenvectors in descending order of eigenvalue, select the first d eigenvectors in turn, and construct the projection matrix U for row dimension reduction, U = [X1,…,X d ];

[0024] 3) According to steps 1) and 2), the projection matrix of column dimension reduction is calculated. The column dimension reduction projection matrix is expressed as:

[0025] V=[Z1,…,Z q ];

[0026] Where Z i is the column-wise covariance matrix G t The eigenvectors corresponding to the first q largest eigenvalues of ′;

[0027] 4) Based on the results of step 2) and step 3), the velocity field matrix is reduced in both directions at the same time. The feature field after dimension reduction is expressed as: C = V T AU.

[0028] Preferably, in said S4, the intelligent agent is a reinforcement learning algorithm for optimizing strategy training and execution;

[0029] The environment is a flow field simulation based on computational fluid dynamics;

[0030] The agent's action is the change in the design variables, which are randomly initialized in the design space at the beginning of each round of optimization strategy learning.

[0031] Preferably, in said S4, the state fed back to the intelligent agent by the environment is the dimensionality-reduced feature field of the velocity field matrix near the flap at the maximum lift coefficient;

[0032] The reward function fed back to the agent by the environment is related to the objective function of the multi-segment airfoil optimization problem. The reward function rule is: when the action causes the value of the design variable to exceed the design space or the computational fluid dynamics simulation does not converge, the reward is negative; otherwise, the reward is the increment of the maximum lift coefficient at the current time step.

[0033] Preferably, the training in S5 is performed over multiple rounds, each round including multiple time steps;

[0034] In each time step, the agent takes the state of the current environment as input and the optimal action of the current state as output. The environment takes the action as input and updates the design variables. After the same geometric modeling, mesh deformation and numerical calculation as S2, the reward function value of the current time step and the state of the next time step are output. This cycle repeats until the current round ends.

[0035] Preferably, the condition for terminating the round is one of reaching a preset time step, the action causing the design variable to exceed the design space, and the computational fluid dynamics simulation not converging.

[0036] Preferably, the training in S5 adopts a deep Q-network algorithm, and a convolutional neural network is used as the value estimation function of the deep Q-network algorithm.

[0037] Therefore, the two-dimensional principal component analysis method for reinforcement learning of multi-segment airfoil optimization strategy of the present invention has the following beneficial effects:

[0038] (1) By adopting the dimension-reduced feature field based on flow field information as the state input of reinforcement learning, the physical characteristics of the flow field are retained, enabling the intelligent agent to learn the general flow characteristic laws that are decoupled from specific geometric parameters, thereby having the ability to directly migrate across working conditions and configurations, and is suitable for different design scenarios and multi-segment airfoil shapes.

[0039] (2) Two-dimensional principal component analysis is used to reduce the dimensionality of the velocity field, which significantly reduces the dimensionality of the state space and the number of parameters and computational cost required for neural network training, enabling the intelligent agent to learn the optimal strategy more quickly.

[0040] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is a flow chart of an embodiment of the present invention;

[0042] Figure 2 A schematic diagram of a numerical calculation grid for a multi-segment airfoil according to an embodiment of the present invention;

[0043] Figure 3 Schematic diagram of multi-section airfoil slot parameters according to an embodiment of the present invention, where a represents the main wing and b represents the flap;

[0044] Figure 4 A schematic diagram of a two-dimensional principal component analysis according to an embodiment of the present invention;

[0045] Figure 5 Schematic diagram of the interaction process between the intelligent agent and the environment of the reinforcement learning model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0046] The technical solution of the present invention is further described below with reference to the accompanying drawings and embodiments.

[0047] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments.

[0048] The 30P30N is a three-section airfoil consisting of a leading edge slat, main wing, and flaps.

[0049] Example

[0050] like Figure 1 As shown, the present invention provides a two-dimensional principal component analysis method for reinforcement learning of multi-segment airfoil optimization strategy. The method is applied to the optimization of the 30P30N flap, the flap angle is set to 30°, the reference chord length is 1m, and the coordinate origin is located at the leading edge of the airfoil when the slats are retracted.

[0051] Here are the steps:

[0052] S1. Define the optimization objective, constraints, and design variables for the multi-segment airfoil optimization problem.

[0053] The optimization goal is to maximize the maximum lift coefficient of the multi-section airfoil; the constraint is that the maximum lift coefficient is not less than the initial value; in this embodiment, the design variables are the slot parameters of the 30P30N flap, such as Figure 3 As shown, it includes overlap O and seam width G.

[0054] S2. Use Latin hypercube sampling to collect sufficient samples in the design space composed of design variables, perform geometric modeling, mesh deformation, and numerical calculation on the samples, and obtain the velocity field matrix near the flap corresponding to each design variable value. All velocity field matrices near the flap constitute an aerodynamic data sample library;

[0055] Among them, geometric modeling, mesh deformation and numerical calculation are specifically as follows:

[0056] A geometric model of a multi-segment airfoil is constructed based on the sample, and the grid topology of the geometric model is changed to adapt to the updated geometric model. The numerical calculation uses a solver based on the RANS equation and the SA turbulence model to obtain the velocity field information at the maximum lift coefficient.

[0057] In this embodiment, the design space is:

[0058] -0.2%≤O≤2%,0.5%≤G≤2.5%;

[0059] Latin hypercube sampling was used to randomly obtain 1000 sets of design variables in the design space. The fixed flow parameters and design conditions used in generating the aerodynamic data sample library were set to: Ma = 0.2, Re = 5×10 6 ; In the formula, Ma represents the Mach number, Re represents the Reynolds number; the computational grid uses a structured grid, such as Figure 2 As shown, the total number of grids is about 200,000. The extraction range of the velocity field near the flap is x∈[0.83,1.14],y∈-0.165,0.075], where x and y are coordinate system parameters in meters. The velocity field matrix of each sample constitutes the aerodynamic data sample library

[0060] S3, using a two-dimensional principal component analysis method to learn the projection matrix of velocity field dimensionality reduction in the aerodynamic data sample library, and using the projection matrix to reduce the velocity field matrix into a low-dimensional feature field, which represents the key compression information of the velocity field;

[0061] Among them, the dimensionality reduction of the velocity field matrix is performed along both the row and column directions (i.e., the width and height directions of the matrix), and the row dimensionality reduction and column dimensionality reduction correspond to different projection matrices respectively.

[0062] In this embodiment, the shape of the velocity field matrix is H=150 rows and W=200 columns, and the shape of the feature field is h=80 rows and w=100 columns. Figure 4 As shown in Figure 2, the bidirectional two-dimensional principal component analysis method is as follows:

[0063] 1) For velocity field dataset Solve the projection vector X of row dimension reduction, satisfying That is, find X that maximizes the objective function J(X);

[0064] in,

[0065] Where J(X) is the objective function, which is used to measure the variance of the projected data, G t is the covariance matrix in the row direction, is the mean matrix of the velocity field, Represents the mean operation.

[0066] Therefore, X is G t The eigenvector corresponding to the maximum eigenvalue of G. t The eigenvalues and eigenvectors of , and the eigenvectors are arranged in descending order according to the eigenvalues, and the first d eigenvectors are selected in turn to form the row compression projection matrix U = [X1,…,X d ].

[0067] Similarly, the column-compressed projection matrix V = [Z1,…,Z q ]; Zi It's G t The eigenvectors corresponding to the first q largest eigenvalues of ′, G t ′ is the covariance matrix in the column direction;

[0068] Then the feature field after dimensionality reduction C=V T AU=C q×d .

[0069] S4. Based on the optimization objectives, constraints, and design variables in S1, a reinforcement learning model for the optimization strategy is established. The reinforcement learning model includes the definition of the agent, the environment, the agent's actions, the state of the environment's feedback to the agent, and the reward function. Specifically:

[0070] 1) The agent is a reinforcement learning algorithm that optimizes policy training and execution;

[0071] 2) The environment is a flow field simulation based on computational fluid dynamics;

[0072] 3) The agent's action is the change in the design variable, corresponding to an increase or decrease in the overlap amount or seam width, which is a number of small discrete values. At the beginning of each round of optimization strategy learning, the design variable is randomly initialized in the design space;

[0073] In this embodiment, the action is the change of the design variables ΔO and ΔG, which includes 8 sets of discrete values, namely:

[0074]

[0075] 4) The state fed back to the agent by the environment is the reduced-dimensional feature field of the velocity field matrix near the flap at the maximum lift coefficient;

[0076] 5) The reward function fed back to the agent by the environment is related to the objective function of the multi-segment airfoil optimization problem. The reward function rule is: when the action causes the value of the design variable to exceed the design space or the computational fluid dynamics simulation does not converge, the reward is negative; otherwise, the reward is the increment of the maximum lift coefficient at the current time step;

[0077] The reward function of this embodiment is defined as follows:

[0078]

[0079] Where k is the scaling factor, (C Lmax ) t+1 is the optimization objective for the next time step,

[0080] C Lmax ) t is the optimization objective at the current time step.

[0081] S5. Based on computational fluid dynamics simulation, the interaction process between the intelligent agent and the environment is established to train the reinforcement learning model, and the optimal optimization strategy with direct transfer capability is obtained (so that the intelligent agent learns the optimal optimization strategy).

[0082] Among them, the training goes through multiple rounds, and each round includes multiple time steps;

[0083] In each time step, the agent takes the state of the current environment as input and the optimal action of the current state as output. The environment takes the action as input and updates the design variables. After the same geometric modeling, mesh deformation and numerical calculation as S2, the reward function value of the current time step and the state of the next time step are output. This cycle repeats until the current round ends. The round termination conditions are: reaching the preset time step, the action causes the design variables to exceed the design space, or the computational fluid dynamics simulation does not converge.

[0084] In this embodiment, the interaction process between the agent and the environment is as follows: Figure 5 As shown in Figure 2, the training consists of 10,000 rounds, with a maximum of 10 interaction steps per round. If an action causes the design variables to exceed the design space or the computational fluid dynamics simulation does not converge, the aerodynamic data for the next time step will not be calculated, the next state will be recorded as the terminal state, the reward will be returned, and the round will be terminated. The design conditions for the training phase are defined as: Ma = 0.2, Re = 5×10 6 The computational grid is as follows: Figure 2 The structured grid shown.

[0085] The training adopts the deep Q network algorithm, and the convolutional neural network is used as the value estimation function of the deep Q network algorithm. The input of the convolutional neural network is the state, which is an 80×100 feature field in this embodiment. The output is a vector (8-dimensional vector) containing the value estimation of each action, and its dimensions correspond one to one to the value function of 8 discrete actions under the input state. The hidden layer of the convolutional neural network contains 3 groups of convolutional layers and pooling layers, and 1 fully connected layer. The step size of the convolutional layer and each pooling layer is 1 and 2, respectively. The convolution kernel size of the convolutional layer is 3×3, and the number of channels output by each convolutional layer is 16, 32, and 64, respectively. The pooling operation of the pooling layer is maximum pooling, and the window size is 2×2. The activation function of the hidden layer is ReLu, and the output layer uses a linear activation function.

[0086] By modifying the design conditions and multi-segment airfoil shape of the multi-segment airfoil optimization problem, the efficiency and versatility of the optimal optimization strategy are verified;

[0087] The results show that in this embodiment, the variation range of Mach number and Reynolds number is:

[0088] 0.15≤Ma≤0.25;

[0089] 1×106 ≤Re≤9×10 6 ;

[0090] It can be seen that this strategy maintains a significant optimization efficiency advantage. This efficiency stems from the reinforcement learning model's ability to quickly reason about the reduced-dimensional feature field and the general optimization experience accumulated by the agent during the training phase. At the same time, the optimal optimization strategy can be directly applied to design conditions or multi-segment airfoil shapes that are different from those in the training phase of step S5. In other words, the optimal optimization strategy has the ability to directly transfer across conditions and configurations. The principle is as follows:

[0091] A universal state characterization method based on flow field characteristics is employed: Two-dimensional principal component analysis is used to reduce the velocity field's dimensionality. The reinforcement learning model then learns universal flow characteristics that are decoupled from specific geometric parameters. This data-driven optimization strategy, independent of specific geometry parameters or operating conditions, captures the essence of aerodynamic characteristics that are universal across operating conditions. This enables excellent transferability and adaptability, enabling direct application to new design scenarios without retraining.

[0092] In addition, the type of multi-segment airfoil is not limited to 30P30N, and multi-segment airfoils of various shapes can be applied.

[0093] Therefore, the present invention provides a two-dimensional principal component analysis method for reinforcement learning of multi-segment airfoil optimization strategies. By adopting a dimension reduction feature field based on flow field information as the state input for reinforcement learning, the physical characteristics of the flow field are retained, so that the intelligent agent can learn the general flow characteristic laws decoupled from specific geometric parameters, thereby having the ability to directly migrate across working conditions and configurations, and is suitable for different design scenarios and multi-segment airfoil shapes; the two-dimensional principal component analysis is used to reduce the dimension of the velocity field, which significantly reduces the dimension of the state space, reduces the number of parameters and computational cost required for neural network training, and enables the intelligent agent to learn the optimal strategy faster.

[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solutions of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A two-dimensional principal component analysis method for reinforcement learning of multi-segment airfoil optimization strategy, characterized in that: The following steps are involved: S1. Define the optimization objective, constraints, and design variables for the multi-segment airfoil optimization problem. S2. Use Latin hypercube sampling to collect samples in the design space composed of design variables, perform geometric modeling, mesh deformation, and numerical calculation on the samples, and obtain the velocity field matrix near the flap corresponding to each design variable value. All velocity field matrices near the flap constitute an aerodynamic data sample library; S3, using the two-dimensional principal component analysis method to learn the projection matrix, using the projection matrix to reduce the velocity field matrix into a feature field, retaining key flow field information; S4. Build a reinforcement learning model for the optimization strategy based on S1. The reinforcement learning model includes the definition of the agent, the environment, the agent's actions, the state of the environment's feedback to the agent, and the reward function. S5. Based on computational fluid dynamics simulation, the interaction process between the intelligent agent and the environment is established to train the reinforcement learning model and obtain the optimal optimization strategy with direct transfer capability.

2. A two-dimensional principal component analysis method for reinforcement learning of multi-segment airfoil optimization strategy according to claim 1, characterized in that: In S1, the optimization objective is to maximize the maximum lift coefficient of the multi-section airfoil; the design variables are the slot parameters of the flap, including the overlap amount and the slot width; and the constraint is that the maximum lift coefficient is not less than the initial value.

3. A two-dimensional principal component analysis method for reinforcement learning of multi-segment airfoil optimization strategy according to claim 1, characterized in that: The geometric modeling, mesh deformation and numerical calculation in S2 are specifically as follows: constructing a geometric model of a multi-segment airfoil according to the sample, changing the mesh topology of the geometric model, and using a solver based on the RANS equation and the SA turbulence model for numerical calculation to obtain the velocity field information at the maximum lift coefficient.

4. The two-dimensional principal component analysis method for reinforcement learning of multi-segment airfoil optimization strategy according to claim 1, characterized in that: In S3, the dimension reduction of the velocity field matrix is performed along both row and column directions, and the row dimension reduction and the column dimension reduction correspond to different projection matrices respectively.

5. The two-dimensional principal component analysis method for reinforcement learning of multi-segment airfoil optimization strategy according to claim 1, characterized in that: In S4, the agent is a reinforcement learning algorithm that optimizes policy training and execution; The environment is a flow field simulation based on computational fluid dynamics; The agent's action is the change in the design variables, which are randomly initialized in the design space at the beginning of each round of optimization strategy learning.

6. The two-dimensional principal component analysis method for reinforcement learning of multi-segment airfoil optimization strategy according to claim 1, characterized in that: In S4, the state fed back to the agent by the environment is the reduced-dimensional feature field of the velocity field matrix near the flap when the lift coefficient is maximum; The reward function fed back to the agent by the environment is related to the objective function of the multi-segment airfoil optimization problem. The reward function rule is: when the action causes the value of the design variable to exceed the design space or the computational fluid dynamics simulation does not converge, the reward is negative. Otherwise, the reward is the increment of the maximum lift coefficient at the current time step.

7. The two-dimensional principal component analysis method for reinforcement learning of multi-segment airfoil optimization strategy according to claim 1, characterized in that: The training in S5 is carried out over multiple rounds, each round including multiple time steps; In each time step, the agent takes the state of the current environment as input and the optimal action of the current state as output. The environment takes the action as input and updates the design variables. After the same geometric modeling, mesh deformation and numerical calculation as S2, the reward function value of the current time step and the state of the next time step are output. This cycle repeats until the current round ends.

8. The two-dimensional principal component analysis method for reinforcement learning of multi-segment airfoil optimization strategy according to claim 7, characterized in that: The condition for terminating the round is one of reaching a preset time step, the action causing the design variable to exceed the design space, and the computational fluid dynamics simulation not converging.

9. The two-dimensional principal component analysis method for reinforcement learning of multi-segment airfoil optimization strategy according to claim 1, characterized in that: The training in S5 adopts a deep Q-network algorithm, and a convolutional neural network is used as a value estimation function of the deep Q-network algorithm.

Citation Information

Cited By

  • Aerodynamic characteristic intelligent prediction method based on flow field relevance

    CN121257329A

  • An intelligent prediction method for aerodynamic characteristics based on flow field correlation

    CN121257329B

  • Deep learning neural network wing design method based on manifold dimension reduction

    CN121278863A

  • A Deep Learning Neural Network-Based Wing Design Method Based on Manifold Dimensionality Reduction

    CN121278863B