Unsteady flow control method and system based on reinforcement learning
Through reinforcement learning methods, the problem of non-stable flow in aircraft engine compressors and turbine machinery is solved, and efficient and accurate flow stability control is achieved, suitable for different flow scenarios such as compressors and turbine blades.
Patent Information
- Application Number
- CN202510831682.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to effectively solve the complex phenomenon of non-static flow in aircraft engine compressors and turbine machinery. Traditional methods rely on experimental data or empirical formulas and are difficult to adapt to the strong dynamic characteristics of non-static flow and the high-dimensional non-linear coupling relationship.
The non-stable flow control method based on reinforcement learning is adopted, and the flow control parameters are optimized by defining the set of flow control parameters, action space and reward function, combined with dynamic modal decomposition, and the flow control parameters are optimized, and the PPO and GAE algorithms are used to improve the algorithm stability and training efficiency.
It has achieved the improvement of multi-objective optimization efficiency in complex flow scenarios, the accuracy of flow stability control is improved, the adaptability and engineering applicability are enhanced, and the calculation load is reduced, ensuring that the control strategy is easy to implement in actual projects.
Smart Images

Figure CN120337830A_ABST
Abstract
Description
Technical Field
[0001] The present invention proposes an unsteady flow control method and system based on reinforcement learning, belonging to the field of flow control. Background Art
[0002] Unsteady flow control is a key technology in fields such as aeroengine compressors and turbomachinery. Its goal is to suppress complex unsteady phenomena such as flow separation and rotational instability through active or passive control means, thereby improving aerodynamic performance and operating safety. Traditional flow control methods can be divided into two categories: passive control methods such as casing treatment and geometric profile optimization, which change the flow field characteristics through fixed geometric structures or preset parameters. Such methods require parametric design based on a large amount of experimental data or empirical formulas. For example, the parameter combinations such as the channel width and inclination angle of the casing treatment are determined through a trial-and-error method. Active control methods such as tip jet and boundary layer suction, which achieve dynamic control by adjusting the actuator parameters (such as jet velocity and frequency) in real time. Their design relies on expert experience or simplified theoretical models. For example, jet parameters are selected based on linear stability theory, or the mapping relationship between control parameters and flow field response is calibrated through experiments.
[0003] In addition, traditional optimization algorithms (such as gradient descent method and genetic algorithm) are often used in the prior art to optimize control parameters. For example, the parameters are adjusted based on the gradient information of the objective function (such as pressure ratio and efficiency), or a heuristic search algorithm is used to find the local optimal solution that satisfies the constraints in the discrete parameter space. However, such methods are usually based on static or quasi-steady assumptions and are difficult to adapt to the strong dynamic characteristics and high-dimensional nonlinear coupling relationships of unsteady flows. Summary of the Invention
[0004] The present invention provides an unsteady flow control method and system based on reinforcement learning to solve the above-mentioned problems: An unsteady flow control method based on reinforcement learning proposed by the present invention, the method includes: S101. Obtain the research object, select the control means according to the research object, define the flow control parameter set based on the control means, set the value range of the parameters in the flow control parameter set, and define the action space according to the value range of the parameters ; S102. Gradually adjust through the action space to obtain the flow control parameters; S103. Input the flow control parameters into the CFD model to predict the unsteady flow field, obtain the unsteady flow field data based on the predicted unsteady flow field, obtain the performance parameters according to the unsteady flow field data , obtain the unsteady flow field statistics by decomposing the unsteady flow field data through dynamic mode decomposition , and S103 is integrated in the environment module; S104. Output the state space , where the state space includes current flow control parameters , performance parameters and unsteady flow field statistics . Define a reward function and perform flow control optimization based on the state space, action space, and reward function.
[0005] Furthermore, define the reward function, including:[[]]
[0006] where represents the reward function represents the penalty term represents the preset target threshold of the performance metric.
[0007] Furthermore, calculate the generalized advantage estimate; S305. Optimize the policy network based on the calculated generalized advantage estimate; S306. Determine whether the optimized policy network meets the training termination conditions. If it meets, the training terminates. If it does not meet, return to step S302 to continue training. The training termination conditions include that the unsteady statistics reach the preset target value and the number of training steps exceeds the preset threshold.
[0008] Furthermore, initialize the policy network and value network, including:[[]] Define the policy network as , where represents the policy represents the trainable parameters of the policy network represents the specific action instruction generated by the agent according to the policy network. The policy network inputs the state and outputs the probability distribution parameters of the action . The probability distribution parameters include the mean and variance of the Gaussian distribution; Define the value network as , where represents the trainable parameters of the value network. The value network inputs the state and outputs the state value estimate; The structures of the policy network and value network adopt fully connected neural networks.
[0009] Furthermore, generate interaction data for optimizing the policy network through interactive sampling and trajectory collection. The interaction data is stored in the trajectory buffer pool, including:[[]] Randomly initialize the flow control parameters , and obtain and That is, the initial state is obtained ; Select an action according to the current policy and generate new flow control parameters to implement action decision-making and state update; Call the environment module to obtain and and calculate the reward to update the state ; Record the single-interaction data into the trajectory buffer pool to implement trajectory storage.
[0010] Furthermore, based on the trajectory buffer pool and the value network, calculate the value network loss, and optimize the value network based on the value network loss, including: Calculate the value network loss:
[0011] wherein, represents the value network loss function, represents the expectation operator, represents the value network function, represents the Monte Carlo return;
[0012] wherein, represents the immediate reward obtained at time step k, γ represents the discount factor, t represents the starting time step for calculating the Monte Carlo return, and T represents the ending time step for calculating the Monte Carlo return; Minimize the mean square error of the value function, and optimize the value network based on the calculated minimized mean square error of the value function.
[0013] Furthermore, after optimizing the value network, calculate the generalized advantage estimation based on the trajectory buffer pool and the optimized value network, including: Calculate the generalized advantage estimation based on the trajectory data in the trajectory buffer pool and the optimized value network to balance the immediate reward and the long-term benefit:
[0014] wherein, represents the generalized advantage estimation function, represents the discount factor, and λ represents the GAE hyperparameter, .
[0015] Furthermore, optimize the policy network based on the calculated generalized advantage estimation, including: Calculate the probability ratio of the new and old policy actions: , where represents the probability that the new policy, i.e., with parameter θ, selects action under state ; represents the probability that the old policy, i.e., with parameter , selects action under the same state ; represents the probability ratio of the new and old policy actions; Update the policy parameters through the clipped objective function:
[0016] where represents the clipping threshold to prevent policy mutation, represents the expectation operator, represents clipping to a specified interval, represents the clipped objective function.
[0017] Furthermore, record the single-interaction data to the trajectory buffer to achieve trajectory storage, including: Measure the similarity weight of the new data and the existing data in the buffer; Obtain the information gain by quantifying the information amount of the new data through the information entropy of the value function of the policy network; Dynamically adjust the storage probability by combining the similarity weight and the information gain.
[0018] A non-steady flow control system based on reinforcement learning proposed by the present invention calculates the advantage function based on trajectory data, including: Define an action space module for obtaining the research object, selecting control means according to the research object, defining a set of flow control parameters based on the control means, setting the value range of the parameters in the set of flow control parameters, and defining the action space according to the value range of the parameters ; Obtain a flow control parameter module for gradually adjusting and obtaining flow control parameters through the action space; An environment module for inputting the flow control parameters into a CFD model, predicting an unsteady flow field, obtaining unsteady flow field data based on the predicted unsteady flow field, obtaining performance parameters from the unsteady flow field data, and obtaining unsteady flow field statistics by decomposing the unsteady flow field data through dynamic mode decomposition ; A control optimization module for outputting a state space that contains the current flow control parameter , Performance parameters and unsteady flow field statistics , define the reward function, and perform flow control optimization based on the state space, action space, and reward function.
[0019] Advantages of the present invention: (1) High efficiency of multi-objective collaborative optimization. Through the design of the reward function and the multi-dimensional fusion of the state space, the present invention realizes the dynamic balance between the performance goal and unsteady suppression. The reward function takes the unsteady statistic as the core optimization index, and at the same time introduces a penalty term to forcefully satisfy performance constraints such as efficiency and pressure ratio, ensuring that the optimization process always advances within the feasible region. The state space integrates control parameters, performance indicators, and unsteady flow field characteristics, enabling the policy network to simultaneously perceive the system performance and flow dynamic characteristics, thus taking into account both stability and efficiency. Compared with the traditional method that relies on manual trial and error for parameter tuning, this framework autonomously optimizes through data-driven methods, significantly improving the multi-objective optimization efficiency in complex flow scenarios. (2) Improvement of algorithm stability and training efficiency. Aiming at the challenges of easy divergence and difficult convergence of reinforcement learning in flow field control, the present invention combines proximal policy optimization (PPO) and generalized advantage estimation (GAE) to enhance the robustness of the algorithm. PPO limits the policy update amplitude through probability ratio clipping to avoid oscillations caused by sudden policy changes during training; GAE balances immediate rewards and long-term benefits to reduce local optimal traps. The policy network and value network adopt a lightweight fully connected structure, reducing the computational load while ensuring the ability to extract flow field characteristics. This design enables the algorithm to converge stably in the high-dimensional parameter space, providing a reliable guarantee for global optimization.
[0020] (3) Strong adaptability and scalability of the control method. The parameterized control framework of the present invention combines flexibility and engineering applicability. The control parameter set is compatible with geometric parameters (such as channel size and position) and dynamic parameters (such as jet velocity), supports the joint optimization of multiple control means, and can be adapted to different flow scenarios such as compressor and turbine blades. The action space is defined as the incremental adjustment of parameters, avoiding the destruction of physical feasibility caused by sudden parameter changes through progressive optimization, ensuring that the control strategy is easy to implement in actual engineering. For example, in the tip region, the optimal solution can be gradually approached by the coordinated fine-tuning of the channel width and jet velocity, reducing the cost of experimental trial and error. (4) Precise suppression of unsteady flow. The present invention innovatively integrates dynamic mode decomposition (DMD) deeply into the reinforcement learning framework, directly optimizing for the physical essence of flow instability. DMD extracts unsteady statistics from transient flow field data, quantifies dynamic characteristics such as vortex shedding and flow separation, and uses them as the core optimization objectives of the reward function. Compared with traditional methods that rely on empirical parameters or local flow field indicators, this scheme realizes more precise flow stability control by capturing the spectral characteristics of broadband unsteady perturbations, providing technical support for the safe operation of high-load fluid machinery. Description of the Drawings
[0021] Figure 1 Schematic diagram of a method for unsteady flow control based on reinforcement learning according to the present invention; Figure 2 Schematic diagram of the control of tip jet flow by a method for unsteady flow control according to the present invention; Figure 3 Top view of tip jet. Specific implementation manners
[0022] In order to more clearly understand the above objects, features and advantages of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other.
[0023] In the following description, many specific details are set forth in order to fully understand the present invention. The described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments, and are not intended to limit the present invention.
[0025] An embodiment of the present invention, a method for unsteady flow control based on reinforcement learning, the method includes: S101. Obtain a research object, select a control means according to the research object, define a set of flow control parameters based on the control means, set the value range of the parameters in the set of flow control parameters, and define an action space according to the value range of the parameters ; S102. Gradually adjust and obtain flow control parameters through the action space; S103. Input the flow control parameters into a CFD model to predict an unsteady flow field, obtain unsteady flow field data based on the predicted unsteady flow field, obtain performance parameters according to the unsteady flow field data , obtain unsteady flow field statistics by decomposing the unsteady flow field data by dynamic mode decomposition , and S103 is integrated in an environment module; S104. Output a state space , the state space includes current flow control parameters , performance parameters and unsteady flow field statistics , define a reward function and perform flow control optimization based on the state space, action space, and reward function.
[0026] The working principle and effects of the above technical solution are as follows: First, obtain a specific research object. According to the characteristics and requirements of this research object, select appropriate control means. For example, according to the research object (such as the compressor tip region), select control means (such as casing treatment channels, tip jet holes). Based on the selected control means, define a flow control parameter set containing multiple relevant parameters. The flow control parameter set includes geometric parameters (position, size, shape) or dynamic parameters (jet velocity, frequency). Then set the value ranges of these parameters (such as the channel width ∈[0.5mm, 2mm], the jet velocity ∈[10m / s, 50m / s]. Finally, define the action space according to the value ranges of the parameters , represents the adjustment amount of the i-th flow control parameter . Through the action space, clarify how the parameters can be changed. Parameter adjustment (S102): Based on the defined action space, gradually adjust the flow control parameters, that is, change each flow control parameter according to the adjustment amount allowed by the action space , so as to explore the situations under different parameter combinations; Flow field analysis and parameter acquisition (S103): Input the adjusted flow control parameters into the computational fluid dynamics (CFD) model. This model will predict the unsteady flow field (i.e., the flow field that changes over time) based on the input parameters. Based on the predicted unsteady flow field, relevant unsteady flow field data are obtained, such as the flow velocity and pressure at different positions. Then, performance parameters such as efficiency and pressure ratio are calculated according to these flow field data. At the same time, the unsteady flow field data are processed by a mathematical method called dynamic mode decomposition (DMD) to obtain unsteady flow field statistics (such as the mean square value of the DMD spectrum of dynamic mode decomposition). Define a reward function, which is used to evaluate the advantages and disadvantages of different state spaces in order to suppress the unsteadiness of the flow field. Then, based on the state space, action space, and reward function, by continuously adjusting the parameters in the action space (i.e., adjusting the flow control parameters), observe the changes in the state space, and judge whether the current parameter setting is better according to the reward function, so as to realize the optimization process of flow control and find the optimal combination of flow control parameters. Precise flow control: By defining the set of flow control parameters and the action space, the flow characteristics of the research object can be precisely adjusted and controlled in a targeted manner, improving the accuracy and flexibility of flow control; In-depth flow field analysis: Using the CFD model to predict the unsteady flow field and obtaining unsteady flow field statistics through dynamic mode decomposition can deeply analyze the characteristics and variation laws of the flow field, providing more comprehensive and accurate data support for flow control; Performance optimization: Based on the state space and reward function for flow control optimization, the optimal combination of flow control parameters can be automatically searched, thereby optimizing the performance of the research object and improving the overall performance and efficiency of the system.
[0027] In an embodiment of the present invention, the defined reward function includes:
[0028] Among them, represents the reward function, represents the penalty term, represents the preset target threshold of the performance index.
[0029] The working principle and effect of the above technical solution are as follows: Two-stage optimization strategy, the first stage: Performance compliance first. When the performance parameter does not reach the preset target threshold target, apply a fixed penalty of -K to force the agent to give priority to improving the basic performance. This is equivalent to setting a hard constraint to ensure that the control strategy does not waste time on secondary goals until the key performance is met; The second stage: Fine flow field optimization. Once meets the requirements, the reward function switches to , guiding the agent to further optimize the dynamic characteristics of the flow field. At this time, the reward is directly related to the flow field statistics, realizing the progressive optimization of multiple physical objectives; due to the strong guiding nature of K, the penalty term K is usually set as a positive number much larger than the typical value (e.g., K = 1e6), ensuring that the reward obtained by the agent when is significantly lower than the reward interval after meeting the standard, forming a strong behavior guidance and avoiding the policy lingering in the invalid area. Decoupling multi-objective conflicts with clear priorities, taking the performance parameter as a rigid prerequisite to avoid direct joint optimization of and and the possible objective conflicts that may occur; focusing in stages, first ensuring that the system meets the basic performance requirements, and then finely adjusting the dynamic characteristics of the flow field, which conforms to the design logic of "ensuring safety first and then seeking optimization" in engineering practice; improving training efficiency and alleviating sparse rewards. Traditional RL is easily troubled by the sparse reward problem in complex flow control (such as binary rewards relying only on the final performance). Here, a more intensive gradient signal is provided through a two-stage design: when not meeting the standard, -K provides clear negative feedback; after meeting the standard, varies continuously with the flow field dynamics, guiding the policy to fine-tune; exploring-exploiting balance. In the initial stage, the agent focuses on exploring how to make meet the standard (exploiting large-range parameter adjustments), and in the later stage, it focuses on optimization (local fine search) to avoid ineffective exploration; enhancing policy robustness and anti-interference ability. If environmental disturbances cause to be temporarily lower than the target, the penalty term -K will drive the policy to quickly recover to the safe interval and avoid falling into an unstable state; dynamic adaptability. When the flow field conditions change and the original policy fails, the reward function will automatically trigger the first-stage optimization to re-find the new control parameters that satisfy ≥ target.
[0030] An embodiment of the present invention performs flow control optimization based on the state space, action space, and reward function, including: S301. Initialize the policy network and value network; S302. Generate interaction data for optimizing the policy network through interactive sampling and trajectory collection, and store the interaction data in the trajectory buffer pool; S303. Calculate the value network loss based on the trajectory buffer pool and the value network, and optimize the value network based on the value network loss; S304. After optimizing the value network, calculate the generalized advantage estimation based on the trajectory buffer pool and the optimized value network; S305. Optimize the policy network based on the calculated generalized advantage estimation; S306. Determine whether the optimized policy network meets the training termination condition. If it meets, the training terminates; if not, return to step S302 to continue training.
[0031] The working principle and effect of the above technical solution are as follows: First, initialize the policy network and the value network. The policy network is used to learn how to select appropriate actions according to the current state, and the value network is used to evaluate the value (expected return) of taking different actions in a certain state. Initialize the parameters of the network, such as weights and biases, to lay the foundation for the subsequent learning process. By allowing the agent to interact with the environment for sampling, that is, the agent takes actions in the environment and observes the feedback of the environment (including new states, rewards, etc.), collect the trajectory information in these interaction processes to form interaction data for optimizing the policy network. These data record the actions taken by the agent in different states and the rewards obtained, etc., and store them in the trajectory buffer pool. The purpose of doing this is to accumulate enough experience data for subsequent learning and optimization. Based on the interaction data in the trajectory buffer pool and the current value network, calculate the loss function of the value network. The loss function measures the gap between the value predicted by the value network and the actual obtained value. By minimizing this loss function, adjust the parameters of the value network so that it can more accurately evaluate the value of the state. Generalized Advantage Estimation (S304). After optimizing the value network, use the data in the trajectory buffer pool and the optimized value network to calculate the Generalized Advantage Estimation (GAE). Generalized Advantage Estimation is a method for evaluating the advantage of actions. It comprehensively considers the immediate reward of taking a certain action in the current state and the rewards that may be obtained in the future, and measures the relative advantages and disadvantages of actions by introducing a discount factor and an advantage function. Calculating GAE can help the policy network better understand the value differences of taking different actions in different states. Optimize the policy network according to the calculated Generalized Advantage Estimation. The policy network adjusts its parameters according to the Generalized Advantage Estimation, making the agent more inclined to choose those actions with higher advantages, thereby improving the performance and return of the agent in the environment. Determine whether the optimized policy network meets the preset training termination condition, such as reaching a certain performance index, the number of iterations reaching the upper limit, etc. If the condition is met, stop the training. At this time, it is considered that the policy network has learned a better policy. If the condition is not met, return to step S302 to continue interactive sampling and trajectory collection, and further train the policy network until the termination condition is met.
[0032] An embodiment of the present invention initializes the policy network and the value network, including: Define the policy network as , where represents the policy, represents the trainable parameters of the policy network. represents the specific action instruction generated by the agent according to the policy network, and the policy network inputs the state , and outputs the action probability distribution parameters, and the probability distribution parameters include the mean and variance of the Gaussian distribution; Define the value network as , where represents the trainable parameters of the value network, and the value network inputs the state , and outputs the state value estimation; The structures of the policy network and the value network adopt fully connected neural networks.
[0033] In one embodiment of the present invention, interactive data for policy optimization is generated through interactive sampling and trajectory collection, and interactive data for optimizing the policy network is generated through interactive sampling and trajectory collection. The interactive data is stored in the trajectory buffer pool, including: Randomly initialize the flow control parameter , and obtain and through the environment module, that is, the initial state is obtained; Select the action according to the current policy , and generate a new flow control parameter to implement action decision-making and state update; Call the environment module to obtain and , calculate the reward , and update the state ; Record the single-interaction data to the trajectory buffer pool to implement trajectory storage.
[0034] The working principle and effect of the above technical solution are as follows: Initial state acquisition, first randomly initialize the flow control parameter The flow control parameter represents the initial control setting of the research object. Then, through the environment module (which integrates relevant calculation functions such as CFD models), the performance parameter and the unsteady flow field statistic are calculated according to the initial flow control parameter. The data of these three parts together constitute the initial state of the agent. At this time, the agent has a basic understanding of the initial situation of the environment; Action decision-making and state update, the agent selects an action from the action space according to the current policy network . The action It is expressed as the adjustment amount of the flow control parameter , by adding the adjustment amount to the current flow control parameter , a new flow control parameter is generated, thus realizing the update of the system state by the action decision. This process reflects the behavior of the agent changing the environmental state by executing actions. New state evaluation and reward calculation: Call the environment module, and recalculate according to the updated flow control parameter to obtain new performance parameters and unsteady flow field statistics . Based on these new data, combined with the pre-set reward mechanism (such as according to the optimization of performance parameters, etc.), the reward is calculated. At the same time, a new state is composed of the new flow control parameter, performance parameter and unsteady flow field statistics, and the agent obtains the feedback information of the environment after executing the action in this way. Trajectory storage: Record the current state , the executed action , the obtained reward and the updated state obtained during a single interaction, form a quadruple , and store it in the trajectory buffer pool . With continuous interaction, the trajectory buffer pool will accumulate a large amount of such experience data, providing rich learning materials for subsequent policy optimization.
[0035] In one embodiment of the present invention, the value network loss is calculated based on the trajectory buffer pool and the value network, and the value network is optimized based on the value network loss, including: Calculating the value network loss:
[0036] Among them, represents the value network loss function, represents the expectation operator, represents the value network function, represents the Monte Carlo return;
[0037] Among them, represents the immediate reward obtained at time step k, γ represents the discount factor, t represents the starting time step for calculating the Monte Carlo return, and T represents the ending time step for calculating the Monte Carlo return; Minimize the mean square error of the value function, and optimize the value network based on the calculated minimized mean square error of the value function.
[0038] The working principle and effect of the above technical solution are as follows: the value network is used to estimate the long-term expected return (i.e., state value), and the Monte Carlo return represents the actual cumulative discounted reward from the current time step to the termination time step . The loss function makes the value network more accurately approximate the true state value by minimizing the mean square error (MSE) between the predicted value and the true return .
[0039] In one embodiment of the present invention, after optimizing the value network, the generalized advantage estimation is calculated based on the trajectory buffer pool and the optimized value network, including: Calculating the generalized advantage estimation based on the trajectory data in the trajectory buffer pool and the optimized value network to balance the immediate reward and the long-term benefit:
[0040] Among them, represents the generalized advantage estimation function, represents the discount factor, and λ represents the GAE hyperparameter, .
[0041] The working principle and effect of the above technical solution are as follows: , represents the one-step TD error, and represents the current reward plus the difference between the next state value and the current state value . represents the discount factor, weighing the importance of the current reward and future rewards, represents the GAE hyperparameter , controlling the weight decay rate of the TD errors of different steps; by performing exponential decay weighting on the TD errors of each step , when , the weight is 1, corresponding to the one-step TD error ; when , the weight is , corresponding to the two-step TD error ; as increases, the weight decays exponentially according to . Introducing to smooth and interpolate the multi-step TD error, and trade off between bias and variance by adjusting represents being close to the Monte Carlo method, with high variance and low bias; Indicates close to one-step TD, with low variance and high bias. Exponential weighting Ensures that the contribution of TD errors in the more distant future gradually weakens, avoiding long-term dependence on noise; one-step TD error Combines the immediate reward and the prediction of the value network , retaining both actual experience and making use of the learned value function. Provides a smooth estimate of future returns, reducing dependence on complete trajectories; through the optimized Improves the accuracy of TD errors, thereby improving advantage estimation.
[0042] An embodiment of the present invention, an optimized policy network based on the calculated generalized advantage estimation, includes: Calculate the probability ratio of the new and old policy actions: , where Represents the probability that the new policy, i.e., with parameter θ, selects action under state , Represents the old policy, i.e., with parameter selects action under the same state , Represents the probability ratio of the new and old policy actions; Update the policy parameters by clipping the objective function:
[0043] where Represents the clipping threshold to prevent sudden policy changes, Represents the expectation operator, Represents clipping to the specified interval, Represents the clipped objective function.
[0044] An embodiment of the present invention records single-interaction data Trajectory buffer pool to implement trajectory storage, including: Measure the similarity weight between the new data and the existing data in the buffer pool; Obtain the information gain by quantifying the information content of the new data through the information entropy of the value function of the policy network; Dynamically adjust the storage probability by combining the similarity weight and the information gain.
[0045] Define the similarity weight of the new data and the buffer pool :
[0046] where Denote the state encoder (such as the variational autoencoder VAE), which maps the state to a low-dimensional embedding space. Denote the kernel bandwidth parameter, which controls the similarity sensitivity. Denote the current number of samples in the buffer pool.
[0047] Calculate the new state through the above function. The average similarity with all states in the buffer pool. The larger the value, the higher the redundancy.
[0048] Based on the value network Define the information gain based on the prediction uncertainty:
[0049] Where Denote the current interaction sample The information gain weight of Denote that through multiple forward propagations of Monte Carlo Dropout, calculate the value distribution probability of the state (Discretized into bins of classes), Denote the value network For the state prediction result The information entropy, where K represents the number of discretized bins, that is, the probability of each interval after discretizing the continuous distribution of the predicted value into K intervals.
[0050] Define the role of the function of information gain. The higher the information entropy, the more uncertain the value estimation of the current state by the current policy, and it needs to be stored preferentially.
[0051] The adaptive storage probability combines the similarity weight and the information gain to define the storage probability:
[0052] Denote the Sigmoid function, which maps the value to [0, 1]. Denote the adjustment coefficient, which controls the storage tendency. Is a very small constant (such as ) to prevent division by zero errors.
[0053] When the information gain of the new data is high and the similarity with the data in the pool is low, the storage probability approaches 1; conversely, if the data is redundant (high similarity) and the information content is low, the storage probability approaches 0. If The second preset threshold, then store , otherwise discard.
[0054] The technical effects of this application are as follows: Through the dynamic balance of similarity weights and information gain, the adaptive storage of empirical data is achieved, and the problems of redundant data accumulation and loss of high-value samples in the experience replay pool (Replay Buffer) in reinforcement learning are solved. It is specifically divided into the following three steps: Step 1: Measure redundancy through similarity weights, state encoding: Map high-dimensional states to a low-dimensional embedding space through a variational autoencoder (VAE) to capture the semantic similarity between states; Kernel density estimation: Based on the RBF kernel function, calculate the similarity mean of the new state and all states in the buffer pool; The larger the value, the more similar it is to the existing data in the pool, and the higher the redundancy; Step 2: Measure value uncertainty through information gain, Monte Carlo Dropout: Calculate the value distribution probability of the state through multiple forward propagations (discretized into classes), Information entropy: Quantify the prediction uncertainty of the value network for the state; The larger the value, the more uncertain the value estimation of the current policy for this state, and it needs to be stored preferentially; Step 3: Adaptive storage probability, through the adjustment coefficient, control the storage tendency; The Sigmoid function maps the value to [0, 1] to prevent division-by-zero errors; High information gain + low similarity: , store preferentially; Low information gain + high similarity: , discard. Achieve efficient storage, reduce redundant data, and enhance the diversity of the experience replay pool; Store high-value samples preferentially to accelerate model convergence; Robustness, adapt to high-dimensional state spaces through state encoding and kernel density estimation; Information entropy quantifies uncertainty to avoid overfitting; Applicable to offline reinforcement learning (Offline RL) and online learning scenarios; Through the dynamic trade-off between redundancy and information volume, intelligent management of the experience replay pool is achieved. Its core value lies in: improving sample efficiency, reducing redundancy, and focusing on high-value samples; enhancing model generalization and improving policy robustness through diverse samples; adapting to complex tasks and achieving efficient learning in dynamic environments; combining reinforcement learning optimization algorithms to further reduce computational overhead and improve real-time performance. Map to the low-dimensional embedding space, capturing the semantic similarity between states; Kernel density estimation: Based on the RBF kernel function, calculate the similarity mean of the new state and all states in the buffer pool; The larger the value, the more similar it is to the existing data in the pool, and the higher the redundancy; Step 2: Measure value uncertainty through information gain, Monte Carlo Dropout: Calculate the value distribution probability of the state through multiple forward propagations, (discretized into classes), Information entropy: Quantify the prediction uncertainty of the value network for the state; The larger the value, the more uncertain the value estimation of the current policy for this state, and it needs to be stored preferentially; Step 3: Adaptive storage probability, through the adjustment coefficient , control the storage tendency; The Sigmoid function maps the value to [0, 1], preventing division-by-zero errors; High information gain + low similarity: , store preferentially; Low information gain + high similarity: , discard. Achieve efficient storage, reduce redundant data, and enhance the diversity of the experience replay pool; Store high-value samples preferentially to accelerate model convergence; Robustness, through state encoding and kernel density estimation, adapt to high-dimensional state spaces; Information entropy quantifies uncertainty to avoid overfitting; Applicable to offline reinforcement learning (Offline RL) and online learning scenarios; Through the dynamic trade-off between redundancy and information volume, intelligent management of the experience replay pool is achieved. Its core value lies in: improving sample efficiency, reducing redundancy, and focusing on high-value samples; enhancing model generalization and improving policy robustness through diverse samples; adapting to complex tasks and achieving efficient learning in dynamic environments; combining reinforcement learning optimization algorithms to further reduce computational overhead and improve real-time performance.
[0055] This application automatically finds and optimizes the best parameter combination for flow control through high-efficiency reinforcement learning, making the flow performance better, such as making aircraft engines more fuel-efficient, quieter, or more efficient. Taking the flow control in the compressor tip region as an example, the performance of the method of this invention is compared with that of traditional reinforcement learning algorithms.
[0056] In a single-channel tip region of a low-pressure compressor of an aeroengine, aiming at the unsteady fluctuations caused by tip flow, a method of injecting air at the tip is proposed to suppress them.
[0057] Set of flow control parameters :
[0058] The number of jets is 1 - 5, the jet position is the range coordinates across the tip, the jet flow rate is 0 - 1 kg / s, the jet is from the casing into the channel, and the angle between the direction and the radial direction is in the range of -80° to 80°.
[0059] Action space: , and the number of jet holes, jet position, jet flow rate, and jet direction step size are 1, 0.01 m, 0.01 kg / s, and 1° respectively.
[0060] The performance parameters such as the efficiency and pressure ratio of the compressor are obtained through CFD and defined as . It is required that the performance parameters be greater than a certain value. (For example, the efficiency is greater than 85% and the pressure ratio is greater than 4); By performing dynamic mode decomposition (DMD) on the flow field (such as tip pressure snapshots), unsteady statistics (such as the mean square value of the DMD spectrum) can be obtained and defined as .
[0061] Define the state space . It covers the control parameters and the corresponding performance parameters and unsteady statistics. The action space represents the change of the flow control variable .
[0062] CFD solver and reinforcement learning environment, CFD model: ANSYS Fluent is used for transient LES simulation, and the time interval for flow field sampling , and the physical time for a single simulation ; RL environment and interface: Python calls Fluent + PyTorch.
[0063] Reward function (two stages of this application):
[0064] Take , the target efficiency includes an efficiency of 0.85 and a pressure ratio of 4.
[0065] The present invention: PPO + GAE (Proximal Policy Optimization + Generalized Advantage Estimation, with the same reward); Comparison 1: DDPG (Deep Deterministic Policy Gradient, and the reward function takes , no stage segmentation); Comparison 2: Conventional PPO (without GAE, reward definition is the same as DDPG); Performance indicators: Convergence steps / round: The number of steps required for the reward to steadily increase to the target and reach the optimal control parameters.
[0066] Strategy stability: 10 rounds of testing after training, standard deviation of performance indicators.
[0067] System efficiency: the total number of steps required for CFD simulation, the efficiency improvement value at the optimal time, and the unsteadiness (DMD index).
[0068] 2. Experimental process 1. Randomly initialize parameters ( ); 2. RL agent samples actions and adjusts ( ), input CFD simulation; 3. Extract flow field efficiency, pressure ratio and DMD statistics, calculate rewards, and the environment returns to a new state; 4. Optimize the value network and policy network; 5. When the termination condition is met (the system reaches the specified efficiency and the reward fluctuation is <1%), stop training.
[0069] 3. Experimental results (convergence steps of each algorithm and optimal efficiency improvement)
[0070] Average convergence Epi: the average number of rounds to reach the target in each experiment; Stability of reaching the target: the standard deviation of the optimal efficiency of multiple samplings in the test phase; Based on the above reinforcement learning framework, the corresponding flow control parameters are output (For example, the number of jets is 1, the jet position is 0.1 times the chord length downstream of the leading edge of the blade, the jet volume is 0.1 kg / s, the jet direction is at an angle of -30° to the radial direction, and at an angle of 20° to the meridian plane), and finally a flow control method suitable for the current research object and working conditions is obtained (such as Figure 2 as shown).
[0071] 4. Analysis and explanation 1. Improved training efficiency: The method of the present invention converges within 38 episodes on average, which is 57.7% less than the traditional DDPG and 29.6% less than the conventional PPO. From the early to mid-stage of the experiment, the reward signal is significantly more dense, and the experience reuse efficiency is high; 2. Improved algorithm stability: The final standard deviation of the PPO+GAE group in multiple rounds of testing was only 0.003, which was much lower than DDPG (0.012) and conventional PPO (0.007). In addition, when the environment was disturbed (perturbed CFD inlet boundary ±5%), the strategy of the present invention could automatically recover to the efficient parameters, showing strong robustness. 3. Optimization of flow control performance: Under the control of PPO+GAE, the mean square value of the DMD spectrum decreases the most, and the suppression effect of unsteadiness is significantly better than that of the control group (-63% vs -36% / -45%); 4. Contribution of reward design: The two-stage reward strategy guides the agent to preferentially improve , avoid ineffective exploration of sub-optimal targets in the early stage, greatly reduce the sparse reward dilemma, and make the most critical working conditions easier to achieve; 5. Effect of experience pool management: An adaptive experience pool based on information gain and similarity weight is introduced. The diversity of replay experience is improved, and the repeated samples are reduced to 22% of the total experience pool, accelerating the model convergence and without performance oscillation.
[0072] An embodiment of the present invention, an unsteady flow control system based on reinforcement learning, calculates the advantage function based on trajectory data, including: Define an action space module, which is used to obtain the research object, select control means according to the research object, define a set of flow control parameters based on the control means, set the value range of the parameters in the set of flow control parameters, and define the action space according to the value range of the parameters ; Obtain the flow control parameter module, which is used to gradually adjust and obtain the flow control parameters through the action space; The environment module is used to input the flow control parameters into the CFD model, predict the unsteady flow field, obtain the unsteady flow field data based on the predicted unsteady flow field, obtain the performance parameters according to the unsteady flow field data , and obtain the unsteady flow field statistics by decomposing the unsteady flow field data through dynamic mode decomposition ; The control optimization module is used to output the state space , the state space includes the current flow control parameters , performance parameters and unsteady flow field statistics , define the reward function, and perform flow control optimization based on the state space, action space and reward function.
[0073] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A method for unsteady flow control based on reinforcement learning, characterized in that, The method includes: S101. Obtain the research object, select control means according to the research object, define a set of flow control parameters based on the control means, set the value range of the parameters in the set of flow control parameters, and define the action space according to the value range of the parameters ; S102. Obtain flow control parameters by gradually adjusting the action space; S103. Input the flow control parameters into the CFD model to predict the unsteady flow field, obtain unsteady flow field data based on the predicted unsteady flow field, and acquire performance parameters according to the unsteady flow field data , and obtain unsteady flow field statistics by performing dynamic mode decomposition on the unsteady flow field data , and S103 is integrated within the environment module; S104. Output the state space , where the state space includes current flow control parameters , performance parameters and unsteady flow field statistics . Define a reward function and perform flow control optimization based on the state space, action space, and reward function.
2. The unsteady flow control method based on reinforcement learning according to claim 1, wherein Define a reward function, including: Among them, represents the reward function, represents the penalty term, represents the preset target threshold of the performance metric.
3. The unsteady flow control method based on reinforcement learning according to claim 1, characterized in that, Perform flow control optimization based on the state space, action space, and reward function, including: S301. Initialize the policy network and the value network; S302. Generate interaction data for optimizing the policy network through interactive sampling and trajectory collection, and store the interaction data in a trajectory buffer pool; S303. Calculate the value network loss based on the trajectory buffer pool and the value network, and optimize the value network based on the value network loss; S304. After optimizing the value network, calculate the generalized advantage estimation based on the trajectory buffer pool and the optimized value network; S305. Optimize the policy network based on the calculated generalized advantage estimation; S306. Determine whether the optimized policy network meets the training termination condition. If it meets, the training terminates. If it does not meet, return to step S302 to continue training. The training termination conditions include that the non-steady statistic reaches a preset target value and the number of training steps exceeds a preset threshold.
4. The unsteady flow control method based on reinforcement learning according to claim 3, characterized in that Initialize the policy network and the value network, including: Define the policy network as , the policy network takes the input state and outputs the probability distribution parameters of the action . The probability distribution parameters include the mean and variance of the Gaussian distribution; Define the value network as , where represents the trainable parameters of the value network, and the value network takes an input state and outputs an estimated state value; The structures of the policy network and the value network adopt fully connected neural networks.
5. The unsteady flow control method based on reinforcement learning according to claim 3, characterized in that Generate interaction data for policy optimization through interactive sampling and trajectory collection, and store the interaction data in the trajectory buffer pool. Storing the interaction data in the trajectory buffer pool includes: Randomly initialize the flow control parameters , and obtain them through the environment module and , that is, the initial state is obtained ; According to the current strategy Select an action , and generate new flow control parameters to implement action decision-making and state update; Call the environment module to obtain and , calculate the reward , update the status ; Record single interaction data To the trajectory buffer pool Realize trajectory storage.
6. The unsteady flow control method based on reinforcement learning according to claim 3, wherein Calculate the value network loss based on the trajectory buffer pool and the value network, and optimize the value network based on the value network loss, including: Calculate the value network loss: Among them, represents the value network loss function, represents the expectation operator, represents the value network function, represents the Monte Carlo return; wherein, represents the immediate reward obtained at time step k, γ represents the discount factor, t represents the starting time step for calculating the Monte Carlo return, and T represents the ending time step for calculating the Monte Carlo return; Minimize the mean square error of the value function, and optimize the value network based on the calculated minimized mean square error of the value function.
7. The unsteady flow control method based on reinforcement learning according to claim 3, characterized in that After optimizing the value network, calculate the generalized advantage estimation based on the trajectory buffer pool and the optimized value network, including: Calculate the Generalized Advantage Estimation (GAE) based on the trajectory data in the trajectory buffer and the optimized value network , balancing immediate rewards and long-term benefits: Among them, represents the generalized advantage estimation function, represents the discount factor, and λ represents the GAE hyperparameter. represents the difference between the immediate reward and the long-term value at time step t. represents the difference between the immediate reward and the long-term value at time step t + k. , represents that the state is when the value estimate is represents that the state is when the value estimate is represents the immediate reward obtained at time step t.
8. The unsteady flow control method based on reinforcement learning according to claim 3, wherein Optimize the policy network based on the calculated generalized advantage estimation, including: Calculate the probability ratio of the new and old policy actions: , where represents the probability of the new policy, i.e., with parameter θ, selecting action in state . represents the probability of the old policy, i.e., with parameter selecting action in the same state . represents the probability ratio of the new and old policy actions; Update the policy parameters by clipping the objective function: Among them, represents a clipping threshold to prevent policy mutation, represents an expectation operator, represents to clip to a specified interval, represents a clipped objective function.
9. The unsteady flow control method based on reinforcement learning according to claim 5, characterized in that Record single interaction data To the trajectory buffer pool Implement trajectory storage, including: Measure the similarity weight of the new data and the existing data in the buffer pool; Quantify the information amount of the new data through the information entropy of the value function of the policy network to obtain the information gain; Dynamically adjust the storage probability by combining the similarity weight and the information gain.
10. A non-steady flow control system based on reinforcement learning, characterized in that, Calculate the advantage function based on the trajectory data, including: Define the action space module, which is used to obtain the research object, select control means according to the research object, define a set of flow control parameters based on the control means, set the value range of the parameters in the set of flow control parameters, and define the action space according to the value range of the parameters ; A flow control parameter acquisition module for obtaining flow control parameters by gradually adjusting the action space; An environment module is used to input the flow control parameters into a CFD model, predict an unsteady flow field, obtain unsteady flow field data based on the predicted unsteady flow field, and obtain performance parameters according to the unsteady flow field data , and obtain unsteady flow field statistics by decomposing the unsteady flow field data through dynamic mode decomposition ; A control optimization module for outputting a state space , where the state space includes current flow control parameters , performance parameters and unsteady flow field statistics , define a reward function, and perform flow control optimization based on the state space, action space, and reward function.
Citation Information
Patent Citations
Traffic signal control method based on deep reinforcement learning
CN112216124A
Data-driven reduced-order model method for predicting turbine rotor-stator interference unsteady flow
CN117688858A
Square column vortex-induced vibration active flow control method based on deep reinforcement learning
CN117993286A
Autonomous underwater robot three-dimensional dynamic trajectory planning method and system based on PPO-IIFDS
CN120103861A