Micro-channel flow control method based on DRL
By applying a DRL-based method in the microchannel, the flow characteristics research and the active control of synthetic jets are solved, and the problems of low flow efficiency and safety hazards in the microchannel are achieved, and efficient and stable flow control is achieved.
Patent Information
- Application Number
- CN202411939312.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-13
AI Technical Summary
The flow in the microchannel is due to the constant laminar flow state and low mass transfer efficiency, making it difficult to meet the needs of efficient transportation. At the same time, the channel vortex vibration caused by the reaction rate increases the equipment strength and manufacturing cost, and there is a risk of blockage, especially in the preparation of energy-containing materials, which may cause safety hazards.
The method based on deep reinforcement learning (DRL) is adopted, combined with the idea of deep neural network and reinforcement learning, and the flow characteristics of microchannels are studied. Through the active control of the synthetic jet, the optimal control strategy is learned to achieve flow field oscillation suppression and reduce flow resistance and lift fluctuations.
Accurate control of flow in microchannels is achieved, flow resistance and lift fluctuations are reduced, mass transfer efficiency is improved, the needs of efficient transportation are met, equipment strength and manufacturing costs are reduced, and safety hazards are reduced.
Smart Images

Figure CN119987194A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of flow stability and active flow control, and in particular to a flow control method in a microchannel based on deep reinforcement learning (DRL). Background Art
[0002] As a chemical reactor with a characteristic size of millimeters or less, microreactors have important application potential in the preparation of nanomaterials, especially energetic materials, due to their advantages of miniaturization, integration, high yield, strong product uniformity and precise controllable reaction conditions. However, due to the limitations of channel characteristics, the flow in the microchannel is usually in a steady laminar state, with low mass transfer efficiency, which is difficult to meet the needs of efficient transportation. At the same time, the channel vortex-induced vibration caused by the accelerated reaction rate increases the flow resistance loss and the inner wall pressure, which increases the equipment strength and manufacturing cost, and there is a risk of blockage, which may cause safety hazards especially in the preparation of energetic materials. Traditional microchannel flow control methods are mostly passive control, mainly focusing on size and structure design, and have limited multi-dimensional information optimization capabilities in the face of dynamic flow fields. The DRL method combines reinforcement learning and deep learning, and uses the feature extraction capabilities of deep neural networks to improve the learning efficiency of traditional reinforcement learning. On the basis of data-driven, the microchannel flow mass transfer is actively controlled. After 500 rounds of training interactions, the jet action can be accurately controlled to achieve oscillation suppression of the flow field and reduce the flow resistance and lift fluctuations on the channel structure. Summary of the invention
[0003] The purpose of the present invention is to provide a microchannel flow control method based on DRL, which uses deep neural networks combined with the idea of reinforcement learning to study the flow characteristics of microchannels, train neural networks to use synthetic jets to actively control the flow, and learn optimal control strategies to achieve stable flow control in microchannels designed with reward functions, reduce structural resistance, and reduce lift fluctuations to meet the needs of microreactor preparation.
[0004] The technical solution to achieve the purpose of the present invention is: a method for controlling flow in a microchannel based on DRL, comprising the following steps:
[0005] Step 1: Geometry construction and meshing of the computational domain of the microchannel:
[0006] First, use Gmsh to draw the geometric model of the microchannel. Its size consists of three parts. The first is a rectangular subdomain with flow lengths of x1 and x3, with (n1+n3)×2n y nodes; the second is a square array cylindrical subdomain of length x2, with (2n2+4n y )×n b nodes; the third is a subdomain that uses an O-type mesh to densify the area around each cylinder, with (3n y+n2)×n c nodes. Then the spatial discretization grid is divided by the spectral element method. The velocity field is calculated by the N=7th order Legendre polynomial interpolation points through the Gauss-Lobato-Legendre integration points in each spectral unit. Finally, numerical simulation is performed, and the two-step backward difference formula method is used for time integration to simulate the unsteady flow in the microchannel. The time step is set to meet the numerical stability criterion, CFL≤1; the simulation results are calculated by the high-precision solver Nek5000. Go to step 2.
[0007] Step 2: Linearized NS equations and flow stability analysis:
[0008] First, the established microchannel flow field is represented by the set (U, P) of velocity field U and pressure field P, and decomposed into the steady state (U b ,P b ) and small disturbances The sum of σ is substituted into the NS equation, and the high-order small terms are ignored to obtain the linearized NS equation. Then the control equation is organized into a matrix form and converted into a generalized eigenvalue problem. The matrix-free numerical method IRAM based on global linear stability is used to simulate the eigenvalue σ that affects the flow stability in the microchannel calculation domain. The flow stability is analyzed based on this result. The steady-state term (U b ), using the selective frequency damping method, by adding a low-pass filter on the right side of the NS equation to suppress the unsteady oscillation of the flow field, the steady-state solution of the nonlinear flow system is obtained as the reference mode for convergence in the subsequent control objective. Go to step 3.
[0009] Step 3: Calculate the location and characteristics of the wave-making area:
[0010] The adjoint method based on the linearized NS equations is used to perform sensitivity analysis on the microchannel flow field. and pressure disturbance The accompanying eigenvector of and The NS equations are linearized and the Arnoldi method is used to transform the large-scale eigenvalue problem into the Hessenberg form. The Krylov subspace K is generated by constructing an implicit Jacobian matrix. m , and use the Gram-Schmidt orthogonalization process to create an orthogonal basis to solve the eigenvectors, and finally calculate the direct eigenvectors and the accompanying eigenvector Will and Overlap to obtain the wave-making area that affects flow stability; go to step 4.
[0011] Step 4: Build a DRL flow control framework, train it, and obtain a control model:
[0012] Based on the analysis of flow stability in microchannels and the calculation of wave-making zones, a DRL flow control framework consisting of three parts, agent, environment, and reward, was constructed. During training, the agent part used a multilayer perceptron with two hidden layers. The goal of the reward function was designed to minimize the energy value of vortex shedding in the environment. By introducing the penalty term P e and the dynamic historical reward normalization function as feedback to the agent on the changes in the flow environment after the agent executes the control action; through the mechanism of maximizing the expected cumulative reward, the agent will cycle through the action execution-reward feedback interaction based on the real-time feedback of the flow environment and the reward information obtained. By repeating this process and performing a large number of interactions, the agent can adjust its strategy based on the new state and reward information to obtain the control model. Go to step 5.
[0013] Step 5: Based on the control model trained in step 4, a control test is performed on the flow in the microchannel, and the control effect of the DRL control method is evaluated by the flow field oscillation amplitude and vortex shedding energy.
[0014] Compared with the prior art, the advantages of the present invention are:
[0015] 1) Based on the DRL method, the present invention combines the advantages of reinforcement learning and deep learning to achieve active control of the flow in the microchannel in a data-driven manner. It can accurately learn the vortex structure characteristics in the flow field, effectively extract the core dynamic information in the complex flow, and achieve accurate regulation of vortex shedding and flow patterns, breaking through the limitations of traditional passive control methods that use structural dimensions to optimize the flow characteristics of microchannels.
[0016] 2) Using a dynamic normalization mechanism based on the historical reward function, by adjusting the reward value in real time and balancing short-term and long-term benefits, it can significantly improve the generalization ability of the control strategy and adapt to the changing flow field environment, and have higher robustness, which has been verified by simulation.
[0017] 3) Probes are arranged in the core sensitive areas that affect flow stability to avoid using a large amount of physical information of flow field states with high similarity and repetition, reduce data redundancy, provide agents with higher-value state differences, improve the convergence speed of deep strategy updates, and significantly reduce the difficulty and time of agents exploring and learning optimal control strategies in high-dimensional state space. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 The figure is a flow chart of a method for controlling flow in a microchannel based on DRL according to the present invention.
[0019] Figure 2 This is the DRL control framework diagram.
[0020] Figure 3 Schematic diagram of the flow field structure. DETAILED DESCRIPTION
[0021] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention are described in further detail below.
[0022] Combination Figure 1 to Figure 3 The microchannel flow control method based on DRL of the present invention comprises the following steps:
[0023] Step 1: Geometry construction and meshing of the computational domain of the microchannel:
[0024] First, use Gmsh to draw the geometric model of the microchannel. Its size consists of three parts. The first is a rectangular subdomain with flow lengths of x1 and x3, with (n1+n3)×2n y nodes; the second is a square array cylindrical subdomain of length x2, with (2n2+4n y )×n b nodes; the third is a subdomain that uses an O-type mesh to densify the area around each cylinder, with (3n y +n2)×n c nodes.
[0025] Then the space is discretely divided into grids by the spectral element method, and the velocity field is calculated by interpolating the N=7th order Legendre polynomials through the Gauss-Lobato-Legendre integration points in each spectral unit.
[0026] Finally, numerical simulation was carried out, and the two-step backward difference formula method was used for time integration to simulate the unsteady flow in the microchannel. The time step was set to meet the numerical stability criterion, CFL≤1; the simulation results were calculated by the high-precision solver Nek5000.
[0027] Numerical simulations were performed to obtain the independency of the simulation results due to different geometric dimensions and mesh densities in the microchannel.
[0028] Step 2: Linearize the NS equations and flow stability analysis, as follows:
[0029] First, the established microchannel flow field is represented by the set (U, P) of velocity field U and pressure field P, and decomposed into the steady state (U b ,P b ) and small disturbances Substitute it into the NS equation and ignore the high-order small terms to get the linearized NS equation;
[0030] Then the control equation is transformed into a matrix form to solve a generalized eigenvalue problem. The matrix-free numerical method IRAM based on global linear stability is used to simulate the eigenvalue σ that affects the flow stability in the microchannel calculation domain. The flow stability is analyzed based on this result. The steady-state term (U b ), using the selective frequency damping method, a low-pass filter is added to the right side of the NS equation to suppress the unsteady oscillation of the flow field and obtain the steady-state solution of the nonlinear flow system as the reference mode for convergence in the subsequent control objective.
[0031] Assuming that the disturbance flow field is in the form of a regular mode, its general solution is in the form of a time exponential function:
[0032]
[0033] In the formula, is a small disturbance of the velocity field and pressure field, is the solution of small perturbation, (x, y) is the position coordinate, t is the simulation time step, T is the transposition symbol, and σ represents the eigenvalue;
[0034] Substitute equation (1) into the NS equation:
[0035]
[0036] represents the Laplace operator, Re is the Reynolds number; the linearized ordinary differential equation is converted into a matrix form:
[0037]
[0038] Where B is the dependent ground state (U b ,P b )’s Jacobian matrix is obtained, and the eigenvalue σ=λ+iω is obtained, where (real part) λ represents the growth rate of the disturbance mode, and (imaginary part) ω represents the oscillation frequency of the disturbance mode. This is used to describe the characteristic modes of the main flow behaviors obtained by the linear stability analysis of the microchannel. i represents the imaginary part of the eigenvalue. When λ>0, the disturbance will grow with time and the flow will be unstable. When λ<0, the disturbance will decay and the flow will be stable. When λ=0, the disturbance remains constant and the flow is in a critical stable state. According to the results of the linear stability analysis, the unstable relationship between the blockage ratio and the Reynolds number in the microchannel is obtained, and the structural design and flow field parameters of the array cylinder are determined.
[0039] Step 3: Calculate the location and characteristics of the wave-making area:
[0040] The adjoint method based on the linearized NS equations is used to perform sensitivity analysis on the microchannel flow field. and pressure disturbance The accompanying eigenvector of and The NS equations are linearized and the Arnoldi method is used to transform the large-scale eigenvalue problem into the Hessenberg form. The Krylov subspace K is generated by constructing an implicit Jacobian matrix. m , and use the Gram-Schmidt orthogonalization process to create an orthogonal basis to solve the eigenvectors, and finally calculate the direct eigenvectors and the accompanying eigenvector Will and The overlap results in a wave-making zone that affects flow stability.
[0041] Using the adjoint method based on the linearized NS equations in step 2, the core region in the microchannel that is most sensitive to flow instability, namely the wave-making region, is identified as follows:
[0042] Step 3-1: Use velocity perturbation and pressure disturbance The accompanying eigenvector of and The linearized NS equation has the following adjoint form:
[0043]
[0044] Among them, σ * is the complex conjugate of the corresponding eigenvalue, and the solution to the equation represents the sensitivity of the perturbed growth rate to changes in the ground state or control parameters.
[0045] The boundary conditions for solving the adjoint equation are:
[0046]
[0047] Where n represents the boundary normal vector. The solution method is the same as that of linear stability analysis. The adjoint equation is converted into matrix form. The implicit restart Arnoldi algorithm is used to construct the implicit Jacobian matrix to solve the eigenvalues and eigenvectors.
[0048] Step 3-2: The direct eigenvector obtained from the linearization analysis and the adjoint eigenvector obtained in the adjoint mode Superposition is performed to obtain the position and intensity of the wave-making area:
[0049]
[0050] Among them, η represents the wave-making area; | | represents the vector modulus; < > represents the vector inner product.
[0051] The adjoint method can identify the size and location of the wave-making area that affects the flow stability in the microchannel, and arrange probes in this area to obtain the flow field DNS information.
[0052] Step 4: Based on the flow stability analysis and wave-making zone information in the microchannel, a DRL flow control framework is constructed and trained to obtain a control model.
[0053] Based on the analysis of flow stability in microchannels and the calculation of wave-making zones, a DRL flow control framework consisting of three parts, agent, environment, and reward, was constructed. During training, the agent part used a multilayer perceptron with two hidden layers. The goal of the reward function was designed to minimize the energy value of vortex shedding in the environment. By introducing the penalty term P e and the dynamic historical reward normalization function as feedback to the changes in the flow environment after the agent executes the control action; through the mechanism of maximizing the expected cumulative reward, the agent will cycle through the action execution-reward feedback interaction based on the real-time feedback of the flow environment and the reward information obtained. By repeating this process and performing a large number of interactions, the agent can adjust its strategy according to the new state and reward information to obtain the control model, as follows:
[0054] Step 4-1: The agent uses a feedforward neural network to process information, including two hidden layers of 512 neurons, and an MLP composed of an input layer and an output layer; its goal is to process the state information of the microchannel flow field environment s t , generate control action a t , and according to the reward feedback R t Optimization control strategy π(a t |s t ), and obtain the action strategy π that can maximize the subsequent reward expectation * Right now:
[0055]
[0056] Where γ represents the discount rate, r t Represents the immediate reward for executing the action; DNS information of the flow field obtained by the probe arranged in the core sensitive area affecting the flow instability in the microchannel, including position (x, y), velocity u, v, pressure p, and flow rate Q of the jet hole as the agent to execute the action jet , recorded as the flow field state:
[0057] s t = {u,v,p,Q jet} (9)
[0058] The vectors of multiple state dimensions As the input layer of the MLP.
[0059] Step 4-2: The state vector obtained from the input layer is used as the output of the 0th hidden layer, and is linked to the input of the first hidden layer through a fully connected method, and is calculated through a recursive formula:
[0060] h Ι =f(W Ι h Ι-1 Ι+b Ι ),Ι=1,2 (10)
[0061] Where h Ι represents the hidden layer, W Ι is the weight matrix of the connections between hidden layers, b Ι is the bias vector that corrects the linear transformation result.
[0062] The high-dimensional features are obtained after two layers of 512 neurons are processed by nonlinear transformation of ReLU activation function.
[0063] Step 4-3: The output layer of the agent converts the calculation results of the neural network into the synthetic jet flow action size a that needs to be actually executed. t , the calculation process is:
[0064] a t =σ z (W3h2+b3) (11)
[0065] Among them, σ z is the activation function Softmax for discrete action space, W3 and b3 are the weight matrix and bias vector connecting the hidden layer to the output layer. This item is used to generate the probability distribution of executing actions and select specific actions by sampling in different states:
[0066]
[0067] The output value of the neural network is o i Converted into probability p a (a t [i′]), ensuring that different flow field environments can correspond to different action selection probability distributions.
[0068] Step 4-4: In order to avoid the jet action update amplitude being too large during the DRL control process, which may cause the velocity and pressure in the environment to jump, causing the DNS to diverge and fail to converge, thus interrupting the DRL control training process, the learning rate α is set to 0.1, and the update strategy of the jet flow rate of the jet hole is set to:
[0069]
[0070] in, represents the jet flow rate at the previous time step, represents the jet flow value at the current time step, Q actionIt is an initial execution action value randomly selected by the agent in the action value interval [-0.20, 0.20] according to the initial environment state s0. The subsequent value is determined and updated by the agent every 50 time steps.
[0071] Step 4-5: When designing reward feedback in the DRL framework, a dynamic normalization method based on historical rewards is introduced to optimize the stability and convergence effect of the deep reinforcement learning training process, as follows:
[0072] The first step is to define the reward function Rw as minimizing the vortex shedding energy on the cylinder:
[0073]
[0074] Rw=P e -s e e g (15)
[0075] Among them, n node is the total number of grid nodes in the computational domain; (u xb ,u yb ) is the ground state of the flow field obtained by solving the SFD method; s e is the sum of the kinetic energy of the vortex shedding at all nodes in the computational domain; g is the maximum flow growth rate obtained by DMD evaluation. For a steady flow, g < 0, i.e., e g <1.0, indicating that the corresponding control strategy is effective and the reward increases; for neutral flow, g = 0, and the reward value is the original vortex shedding energy value; for unstable flow, g>0, that is, e g >1.0, the reward value of the corresponding action decreases, indicating that the control strategy is not effective; to adjust the magnitude of the reward value, a penalty term P is introduced e , which is used to ensure the stable propagation of gradients, avoid the problem of gradient update instability caused by too small reward values leading to slow learning process or too large reward values, and prevent gradient disappearance or explosion; the size of the penalty term is affected by factors such as spacing ratio, blocking ratio and Reynolds number; secondly, in order to avoid the reward value being biased towards positive or negative due to the empirically set initial kinetic energy difference, the historical reward dynamics are further standardized:
[0076]
[0077] Among them, d′ is the number of actions currently executed, Rw k is the reward value corresponding to the number of actions executed from d′-m+1 to d, μ s is the mean of historical rewards, standard deviation σ s Measures the discreteness of rewards, Rw s The reward value Rw obtained at the current time step d′It is centered by the mean and scaled by the standard deviation. The current reward value is adjusted in real time at each training step by considering the mean and standard deviation of the most recent m = 80 historical rewards. The normalization process smoothes the rewards to avoid problems with over-concentration or excessive fluctuations in the reward values.
[0078] Step 4-6: The state s obtained by the agent observing the environment at time t t , using the policy network to generate action a t The probability distribution of , and random sampling is performed according to the probability distribution to determine action a t ; The agent performs action a t After that, the environment feedbacks the new state s t+1 and reward r t ; The agent is based on the new state s t+1 Estimate the expected reward V(s) for the next action t+1 ), thereby approximately calculating the expected reward at time t and providing a basis for the next iteration; the above process is a Markov decision process, the goal is to maximize the expected discounted return:
[0079]
[0080] Among them, γ is the discount factor, which is used to balance the relationship between current rewards and future rewards, and ξ represents the number of rewards. The proximal policy optimization algorithm based on the Actor-Critic architecture is used. The action-value function is expressed in The expected return of taking an action from the current state under policy π; state-value function represents the expected return from the current state under strategy π; where the Actor network approximates the strategy function π based on the neural network parameters θ θ (a t |s t ), used to generate action probability distribution; the Critic network approximates the state value function V based on the parameter φ π (s t ), which is used to evaluate the expected return under different actions; PPO updates the policy network parameters θ by maximizing a clipped approximate objective function, which is:
[0081]
[0082] The strategy ratio is The clipping term limits the ratio to the range of [1-ε,1+ε], where ε is a hyperparameter with a value of 0.2, so that the difference between the current strategy and the old strategy is limited to a fixed range to prevent excessive policy parameter updates from causing training instability or overfitting; advantage function Execute action a under the current strategy tThe difference between the reward and the predicted reward measures the relative merits of the action; by using the Monte Carlo approximate gradient method, the gradient ascent of the parameter θ in the Actor network is maximized L CLIP (θ) to train the Actor network by accumulating discounted reward values:
[0083]
[0084] θ and θ * They are the policy network parameters before and after the gradient update; the error function is minimized by gradient descent on the Critic network parameter φ:
[0085]
[0086] Where R t is the cumulative discount, V(s t ) is the predicted value; the network parameter φ is updated based on this value evaluation; the neural network parameters are continuously updated, the action and environment interact to train the control strategy, and the control model is obtained.
[0087] Step 5: Use the two indicators of flow field oscillation amplitude and vortex shedding energy to evaluate the advantages and disadvantages of the DRL microchannel flow control method, as follows:
[0088] In the test results, the agent obtains the DNS flow field information of the corresponding time step after executing each action to calculate the lift coefficient C on the array cylindrical structure in the microchannel l :
[0089]
[0090] Where C is the cylindrical surface, dC is the cylindrical surface differential, and d yz Represents the diameter of a single cylinder, (n x , n y ) are the normal vectors in the streamwise and vertical directions, U τ represents the tangential velocity component at the integration point on the surface of the object, ρ represents the density of the fluid in the microchannel, and U max represents the maximum velocity of the fluid at the microchannel inlet, is the differential symbol.
[0091] After a certain number of control actions, the change of lift coefficient is observed; at the same time, the value s of vortex shedding energy of microchannel flow field after corresponding control actions is recorded during the interaction of DRL flow control framework. e , observe the change of vortex shedding energy with the control action.
[0092] Example 1
[0093] The DRL-based microchannel flow control method of the present invention comprises the following steps:
[0094] Step 1: Test the grid result accuracy of the total number of spectral flow field units of 1060, 1464, 2212, 3290, and 4836. The error between the densest and sparsest grids is within 0.001%. Test the distance between the inlet and outlet of the cylinder (x 1, x3)=(5,10), (8,16) and (10,20). When x1 and x3 exceed 5 and 10, the computational domain size no longer affects the calculation results. The values of x1 and x3 used in the present invention are 5 and 10, the number of nodes of n1 and n3 are 15 and 40, x2 has the same value as the watershed distance of 2, n2 is set to 21, and the total number of spectral units is 2212.
[0095] Step 2: Linearized NS equations and flow stability analysis process, as follows:
[0096] Assuming that the disturbance flow field is in the form of a regular mode, its general solution is in the form of a time exponential function:
[0097]
[0098] In the formula, is a small disturbance of the velocity field and pressure field, is the solution of small perturbation, (x, y) is the position coordinate, t is the simulation time step, T is the transposition symbol, and σ represents the eigenvalue;
[0099] Substitute equation (1) into the NS equation:
[0100]
[0101] represents the Laplace operator, Re is the Reynolds number; the linearized ordinary differential equation is converted into a matrix form:
[0102]
[0103] Where B is the dependent ground state (U b ,P b ), and the eigenvalue σ=λ+iω is obtained. The real part λ represents the growth rate of the disturbance mode, and the imaginary part ω represents the oscillation frequency of the disturbance mode. This is used to describe the characteristic mode of the main flow behavior obtained by the linear stability analysis of the microchannel. i represents the imaginary part of the eigenvalue. When λ>0, the disturbance will grow with time and the flow will be unstable; when λ<0, the disturbance will decay and the flow will be stable; and when λ=0, the disturbance remains constant and the flow is in a critical stable state. According to the results of the linear stability analysis, the unstable relationship between the blockage ratio and the Reynolds number in the microchannel is obtained, and the blockage ratio of the array cylinder is determined to be 0.2, the Reynolds number is 106, and the spacing ratio is 2.0.
[0104] Step 3: Using the adjoint method based on the linearized NS equations in step 2, the core area in the microchannel that is most sensitive to flow instability, i.e., the wave-making area, is identified as follows:
[0105] Step 3-1: Use velocity perturbation and pressure disturbance The accompanying eigenvector of and The linearized NS equation has the following adjoint form:
[0106]
[0107] Among them, σ * is the complex conjugate of the corresponding eigenvalue, and the solution of the equation represents the sensitivity of the perturbed growth rate to changes in the ground state or control parameters;
[0108] The boundary conditions for solving the adjoint equation are:
[0109]
[0110]
[0111] In the formula, n represents the boundary normal vector. The solution method is the same as that of linear stability analysis. The adjoint equation is converted into a matrix form. The implicit restart Arnoldi algorithm is used to construct an implicit Jacobian matrix to solve the eigenvalues and eigenvectors.
[0112] Step 3-2: The direct eigenvector obtained from the linearization analysis and the adjoint eigenvector obtained in the adjoint mode Superposition is performed to obtain the position and intensity of the wave-making area:
[0113]
[0114] Wherein, η represents the wave-making area; | | represents the vector modulus; < > represents the vector inner product;
[0115] The adjoint method can identify the size and location of the wave-making area that affects the flow stability in the microchannel, and arrange probes in this area to obtain the flow field DNS information.
[0116] Step 4: Based on the flow stability analysis and wave-making zone information in the microchannel, a DRL flow control framework is constructed and trained to obtain a control model, as follows:
[0117] Step 4-1: The agent uses a feedforward neural network to process information, including two hidden layers of 512 neurons, and an MLP composed of an input layer and an output layer; its goal is to process the state information of the microchannel flow field environment s t , generate control action at , and according to the reward feedback R t Optimization control strategy π(a t |s t ), and obtain the action strategy π that can maximize the subsequent reward expectation * Right now:
[0118]
[0119] Where γ represents the discount rate, r t Represents the immediate reward for executing the action; DNS information of the flow field obtained by the probe arranged in the core sensitive area affecting the flow instability in the microchannel, including position (x, y), velocity u, v, pressure p, and flow rate Q of the jet hole as the agent to execute the action jet , recorded as the flow field state:
[0120] s t = {u,v,p,Q jet} (9)
[0121] The vectors of multiple state dimensions As the input layer of MLP;
[0122] Step 4-2: The state vector obtained from the input layer is used as the output of the 0th hidden layer, and is linked to the input of the first hidden layer through a fully connected method, and is calculated through a recursive formula:
[0123] h Ι =f(W Ι h Ι-1 Ι+b Ι ),Ι=1,2 (10)
[0124] Where h Ι represents the hidden layer, W Ι is the weight matrix of the connections between hidden layers, b Ι is the bias vector, correcting the linear transformation result;
[0125] The high-dimensional features are obtained after two layers of 512 neurons are processed by the nonlinear transformation of the ReLU activation function;
[0126] Step 4-3: The output layer of the agent converts the calculation results of the neural network into the synthetic jet flow action size a that needs to be actually executed. t , the calculation process is:
[0127] a t =σ z (W3h2+b3) (11)
[0128] Among them, σ zis the activation function Softmax for discrete action space, W3 and b3 are the weight matrix and bias vector connecting the hidden layer to the output layer. This item is used to generate the probability distribution of executing actions and select specific actions by sampling in different states:
[0129]
[0130] The output value of the neural network is o i Converted into probability p a (a t [i′]), ensuring that different flow field environments can correspond to different action selection probability distributions;
[0131] Step 4-4: In order to avoid the jet action update amplitude being too large during the DRL control process, which may cause the velocity and pressure in the environment to jump, causing the DNS to diverge and fail to converge, thus interrupting the DRL control training process, the learning rate α is set to 0.1, and the update strategy of the jet flow rate of the jet hole is set to:
[0132]
[0133] in, represents the jet flow rate at the previous time step, represents the jet flow value at the current time step, Q action It is an initial execution action value randomly selected by the agent in the action value interval [-0.20, 0.20] according to the initial environment state s0. The subsequent value is determined and updated by the agent every 50 time steps.
[0134] Step 4-5: When designing reward feedback in the DRL framework, a dynamic normalization method based on historical rewards is introduced to optimize the stability and convergence effect of the deep reinforcement learning training process; the specific description is as follows:
[0135] The first step is to define the reward function Rw as minimizing the vortex shedding energy on the cylinder:
[0136]
[0137] Rw=P e -s e e g (15)
[0138] Among them, n node is the total number of grid nodes in the computational domain; (u xb ,u yb ) is the ground state of the flow field obtained by solving the SFD method; s e is the sum of the kinetic energy of the vortex shedding at all nodes in the computational domain; g is the maximum flow growth rate obtained by DMD evaluation. For a steady flow, g < 0, i.e., e g<1.0, indicating that the corresponding control strategy is effective and the reward increases; for neutral flow, g = 0, and the reward value is the original vortex shedding energy value; for unstable flow, g>0, that is, e g >1.0, the reward value of the corresponding action decreases, indicating that the control strategy is not effective; to adjust the magnitude of the reward value, a penalty term P is introduced e , which is used to ensure the stable propagation of gradients, avoid the problem of gradient update instability caused by too small reward values leading to slow learning process or too large reward values, and prevent gradient disappearance or explosion; the size of the penalty term is affected by factors such as spacing ratio, blocking ratio and Reynolds number; secondly, in order to avoid the reward value being biased towards positive or negative due to the empirically set initial kinetic energy difference, the historical reward dynamics are further standardized:
[0139]
[0140] Among them, d′ is the number of actions currently executed, Rw k is the reward value corresponding to the number of actions executed from d′-m+1 to d, μ s is the mean of historical rewards, standard deviation σ s Measures the discreteness of rewards, Rw s The reward value Rw obtained at the current time step d′ It is centered by the mean and scaled by the standard deviation. The current reward value is adjusted in real time at each training step by considering the mean and standard deviation of the most recent m = 80 historical rewards. The normalization process smoothes the rewards to avoid problems with over-concentration or excessive fluctuations in the reward values.
[0141] Step 4-6: The state s obtained by the agent observing the environment at time t t , using the policy network to generate action a t The probability distribution of , and random sampling is performed according to the probability distribution to determine action a t ; The agent performs action a t After that, the environment feedbacks the new state s t+1 and reward r t ; The agent is based on the new state s t+1 Estimate the expected reward V(s) for the next action t+1 ), thereby approximately calculating the expected reward at time t and providing a basis for the next iteration; the above process is a Markov decision process, the goal is to maximize the expected discounted return:
[0142]
[0143] Among them, γ is the discount factor, which is used to balance the relationship between current rewards and future rewards, and ξ represents the number of rewards. The proximal policy optimization algorithm based on the Actor-Critic architecture is used. The action-value function is expressed in The expected return of taking an action from the current state under policy π; state-value function represents the expected return from the current state under strategy π; where the Actor network approximates the strategy function π based on the neural network parameters θ θ (a t |s t ), used to generate action probability distribution; the Critic network approximates the state value function V based on the parameter φ π (s t ), which is used to evaluate the expected return under different actions; PPO updates the policy network parameters θ by maximizing a clipped approximate objective function, which is:
[0144]
[0145]
[0146] The strategy ratio is The clipping term limits the ratio to the range of [1-ε,1+ε], where ε is a hyperparameter with a value of 0.2, so that the difference between the current strategy and the old strategy is limited to a fixed range to prevent excessive policy parameter updates from causing training instability or overfitting; advantage function Execute action a under the current strategy t The difference between the reward and the predicted reward measures the relative merits of the action; by using the Monte Carlo approximate gradient method, the gradient ascent of the parameter θ in the Actor network is maximized L CLIP (θ) to train the Actor network by accumulating discounted reward values:
[0147]
[0148] θ and θ * They are the policy network parameters before and after the gradient update; the error function is minimized by gradient descent on the Critic network parameter φ:
[0149]
[0150] Where R t is the cumulative discount, V(s t ) is the predicted value; the network parameter φ is updated based on this value evaluation; the neural network parameters are continuously updated, the action and environment interact to train the control strategy, and the control model is obtained.
[0151] Step 5: Use the two indicators of flow field oscillation amplitude and vortex shedding energy to evaluate the advantages and disadvantages of the DRL microchannel flow control method, as follows:
[0152] In the test results, the agent obtains the DNS flow field information of the corresponding time step after executing each action to calculate the lift coefficient C on the array cylindrical structure in the microchannel l :
[0153]
[0154] Where C is the cylindrical surface, dC is the cylindrical surface differential, and d yz Represents the diameter of a single cylinder, (n x , n y ) are the normal vectors in the streamwise and vertical directions, U τ represents the tangential velocity component at the integration point on the surface of the object, ρ represents the density of the fluid in the microchannel, and U max represents the maximum velocity of the fluid at the microchannel inlet, is the differential symbol;
[0155] After a certain number of control actions, the change of lift coefficient is observed; at the same time, the value s of vortex shedding energy of microchannel flow field after corresponding control actions is recorded during the interaction of DRL flow control framework. e , observe the change of vortex shedding energy with the control action.
[0156] The method proposed in the present invention is simulated on a high-performance computing host configured with AMD Ryzen Threadripper 3990X (64-core, 128-thread) and Nvidia 1080Ti GPU; the software environment uses the Ubuntu18.04 LTS operating system and auxiliary libraries: NumPy and SciPy; the model training uses Python to build a deep reinforcement learning neural network framework, and is implemented based on Tensorflow. The reinforcement learning algorithm uses PPO, and the agent network structure uses MLP. In order to speed up the network training speed, the optimizer uses a multi-step Adam algorithm with a learning rate of 1e10 -3 ; In order to make the trained control strategy tend to select actions evenly, rather than select certain actions too surely, leading to premature convergence to an optimal solution and insufficient exploration, entropy regularization is used, and the weight is set to 1e10 -2 ; To focus on the value change of long-term rewards, the decay factor in the generalized advantage estimation is set to λ = 0.97; the DNS time step of the simulation environment is set to 5e10 -4 , the jet flow execution action value is updated every 200 steps.
[0157] Table 1 Comparative test
[0158] Testing time <![CDATA[Uncontrolled C l > <![CDATA[Controlled C l > <![CDATA[Uncontrolled s e > <![CDATA[Controlled s e > t=0 0.0254 0.0280 1015.08 1017.01 t=8 -0.0408 -0.0012 987.23 463.43 t=16 0.0116 -0.0051 1006.85 199.46 t=24 0.0302 -0.0063 1009.75 75.62 t=32 -0.0391 -0.0064 988.28 20.96 t=40 0.0052 -0.0063 1025.21 8.46 t=48 0.0344 -0.0064 1004.75 7.35 t=56 -0.0364 -0.0063 991.65 7.21 t=64 0.0237 -0.0063 1016.80 7.19
[0159] From the comparison of simulation results of the time series in Table 1, it can be seen that the superiority of the flow control effect of the present invention in the microchannel is that compared with the lift fluctuation amplitude of the uncontrolled flow field, the lift coefficient fluctuation of the controlled flow field is suppressed within 2%; the vortex shedding energy of the controlled flow at t=64 is reduced by 99.29% compared with the uncontrolled initial state. This reflects the significant advantages of the microchannel flow control strategy based on DRL in terms of controllability and efficiency.
Claims
1. A microchannel flow control method based on DRL, characterized in that: The following steps are involved: Step 1: Geometry construction and meshing of the computational domain of the microchannel: First, use Gmsh to draw the geometric model of the microchannel. Its size consists of three parts. The first is a rectangular subdomain with flow lengths of x1 and x3, with (n1+n3)×2n y nodes; the second is a square array cylindrical subdomain of length x2, with (2n2+4n y )×n b nodes; the third is a subdomain that uses an O-type mesh to densify the area around each cylinder, with (3n y +n2)×n c nodes; Then the space is discretized and meshed by the spectral element method, and the velocity field is calculated by interpolating the N = 7th order Legendre polynomials through the Gauss-Lobato-Legendre integration points in each spectral unit; Finally, numerical simulation was performed, using the two-step backward difference formula method for time integration to simulate the unsteady flow in the microchannel. The time step was set to meet the numerical stability criterion, CFL≤1; The simulation results are calculated by the high-precision solver Nek5000; Go to step 2; Step 2: Linearized NS equations and flow stability analysis: First, the established microchannel flow field is represented by the set (U, P) of velocity field U and pressure field P, and decomposed into the steady state (U b ,P b ) and small disturbances Substitute it into the NS equation and ignore the high-order small terms to get the linearized NS equation; Then the control equation is transformed into a matrix form to solve a generalized eigenvalue problem. The matrix-free numerical method IRAM based on global linear stability is used to simulate the eigenvalue σ that affects the flow stability in the microchannel calculation domain. The flow stability is analyzed based on this result. The steady state term (U b ), using the selective frequency damping method, by adding a low-pass filter on the right side of the NS equation to suppress the unsteady oscillation of the flow field, the steady-state solution of the nonlinear flow system is obtained as the convergent reference mode in the subsequent control objective; Go to step 3; Step 3: Calculate the location and characteristics of the wave-making area: The adjoint method based on the linearized NS equations is used to perform sensitivity analysis on the microchannel flow field. and pressure disturbance The accompanying eigenvector of and The NS equations are linearized and the Arnoldi method is used to transform the large-scale eigenvalue problem into the Hessenberg form. The Krylov subspace K is generated by constructing an implicit Jacobian matrix. m , and use the Gram-Schmidt orthogonalization process to create an orthogonal basis to solve the eigenvectors, and finally calculate the direct eigenvectors and the accompanying eigenvector Will and The overlap results in a wave-making zone that affects flow stability; Go to step 4; Step 4: Build a DRL flow control framework, train it, and obtain a control model: Based on the analysis of flow stability in microchannels and the calculation of wave-making zones, a DRL flow control framework consisting of three parts, agent, environment, and reward, was constructed. During training, the agent part used a multilayer perceptron with two hidden layers. The goal of the reward function was designed to minimize the energy value of vortex shedding in the environment. By introducing the penalty term P e and the dynamic historical reward normalization function as feedback to the changes in the flow environment after the agent executes the control action; through the mechanism of maximizing the expected cumulative reward, the agent will cycle through the action execution-reward feedback interaction based on the real-time feedback of the flow environment and the reward information obtained. By repeating this process and performing a large number of interactions, the agent can adjust its strategy according to the new state and reward information to obtain the control model; Go to step 5; Step 5: Conduct a control test on the flow in the microchannel based on the control model, and evaluate the control effect of the DRL control method by using the flow field oscillation amplitude and vortex shedding energy.
2. A DRL-based microchannel flow control method according to claim 1, characterized in that: In step 1, numerical simulation is performed to obtain the independence of the simulation results for different geometric dimensions and mesh densities in the microchannel.
3. A DRL-based microchannel flow control method according to claim 2, characterized in that: In step 2, the linearized NS equations and flow stability analysis process are as follows: Assuming that the disturbance flow field is in the form of a regular mode, its general solution is in the form of a time exponential function: In the formula, is a small disturbance of the velocity field and pressure field, is the solution of small perturbation, (x, y) is the position coordinate, t is the simulation time step, T is the transposition symbol, and σ represents the eigenvalue; Substitute equation (1) into the NS equation: represents the Laplace operator, Re is the Reynolds number; the linearized ordinary differential equation is converted into a matrix form: Where B is the dependent ground state (U b ,P b ) is obtained by the Jacobian matrix of the microchannel, and the eigenvalue σ=λ+iω is obtained, where λ represents the growth rate of the disturbance mode and ω represents the oscillation frequency of the disturbance mode. This is used to describe the characteristic modes of the main flow behaviors obtained by the linear stability analysis of the microchannel, and i represents the imaginary part of the eigenvalue. When λ>0, the disturbance will grow with time and the flow will be unstable; when λ<0, the disturbance will decay and the flow will be stable; and when λ=0, the disturbance remains constant and the flow is in a critical stable state. According to the results of the linear stability analysis, the unstable relationship between the blockage ratio and the Reynolds number in the microchannel is obtained, and the structural design and flow field parameters of the array cylinder are determined.
4. A DRL-based microchannel flow control method according to claim 3, characterized in that: In step 3, the core area in the microchannel that is most sensitive to flow instability, i.e., the wave-making area, is identified using the adjoint method based on the linearized NS equations in step 2, as follows: Step 3-1: Use velocity perturbation and pressure disturbance The accompanying eigenvector of and The linearized NS equation has the following adjoint form: Among them, σ * is the complex conjugate of the corresponding eigenvalue, and the solution of the equation represents the sensitivity of the perturbed growth rate to changes in the ground state or control parameters; The boundary conditions for solving the adjoint equation are: In the formula, n represents the boundary normal vector. The solution method is the same as that of linear stability analysis. The adjoint equation is converted into a matrix form. The implicit restart Arnoldi algorithm is used to construct an implicit Jacobian matrix to solve the eigenvalues and eigenvectors. Step 3-2: The direct eigenvector obtained from the linearization analysis and the adjoint eigenvector obtained in the adjoint mode Superposition is performed to obtain the position and intensity of the wave-making area: Wherein, η represents the wave-making area; | | represents the vector modulus; < > represents the vector inner product; The adjoint method can identify the size and location of the wave-making area that affects the flow stability in the microchannel, and arrange probes in this area to obtain the flow field DNS information.
5. A DRL-based microchannel flow control method according to claim 4, characterized in that: In step 4, based on the flow stability analysis and wave-making zone information in the microchannel, a DRL flow control framework is constructed and trained to obtain a control model, as follows: Step 4-1, the agent uses a feedforward neural network to process information, including two hidden layers of 512 neurons, and an MLP composed of an input layer and an output layer; Its goal is to process the state information of the microchannel flow environment t , generate control action a t , and according to the reward feedback R t Optimization control strategy π(a t |s t ), and obtain the action strategy π that can maximize the subsequent reward expectation * Right now: Where γ represents the discount rate, r t Represents the immediate reward for executing the action; DNS information of the flow field obtained by the probe arranged in the core sensitive area affecting the flow instability in the microchannel, including position (x, y), velocity u, v, pressure p, and flow rate Q of the jet hole as the agent to execute the action jet , recorded as the flow field state: s t ={u,v,p,Q jet } (9) The vectors of multiple state dimensions As the input layer of MLP; Step 4-2: The state vector obtained from the input layer is used as the output of the 0th hidden layer, and is linked to the input of the first hidden layer through a fully connected method, and is calculated through a recursive formula: h Ι =f(W Ι h Ι-1 I+b Ι ),I=1.2 (10) Where h Ι represents the hidden layer, W Ι is the weight matrix of the connections between hidden layers, b Ι is the bias vector, correcting the linear transformation result; The high-dimensional features are obtained after two layers of 512 neurons are processed by the nonlinear transformation of the ReLU activation function; Step 4-3: The output layer of the agent converts the calculation results of the neural network into the synthetic jet flow action size a that needs to be actually executed. t , the calculation process is: a t =s z (W3h2+b3) (11) Among them, σ z is the activation function Softmax for discrete action space, W3 and b3 are the weight matrix and bias vector connecting the hidden layer to the output layer. This item is used to generate the probability distribution of executing actions and select specific actions by sampling in different states: The output value of the neural network is o i Converted into probability p a (a t [i′]), ensuring that different flow field environments can correspond to different action selection probability distributions; Step 4-4: In order to avoid the jet action update amplitude being too large during the DRL control process, which may cause the velocity and pressure in the environment to jump, causing the DNS to diverge and fail to converge, thus interrupting the DRL control training process, the learning rate α is set to 0.1, and the update strategy of the jet flow rate of the jet hole is set to: in, represents the jet flow rate at the previous time step, represents the jet flow value at the current time step, Q action It is an initial execution action value randomly selected by the agent in the action value interval [-0.20, 0.20] according to the initial environment state s0. The subsequent value is determined and updated by the agent every 50 time steps; Step 4-5: When designing reward feedback in the DRL framework, a dynamic normalization method based on historical rewards is introduced to optimize the stability and convergence effect of the deep reinforcement learning training process; Step 4-6: The state s obtained by the agent observing the environment at time t t , using the policy network to generate action a t The probability distribution of , and random sampling is performed according to the probability distribution to determine action a t ; The agent performs action a t After that, the environment feedbacks the new state s t+1 and reward r t ; The agent is based on the new state s t+1 Estimate the expected reward V(s) for the next action t+1 ), thereby approximately calculating the expected reward at time t and providing a basis for the next iteration; the above process is a Markov decision process, the goal is to maximize the expected discounted return: Among them, γ is the discount factor, which is used to balance the relationship between current rewards and future rewards, and ξ represents the number of rewards. The proximal policy optimization algorithm based on the Actor-Critic architecture is used. The action-value function is expressed in The expected return of taking an action from the current state under policy π; state-value function represents the expected return from the current state under strategy π; where the Actor network approximates the strategy function π based on the neural network parameters θ θ (a t |s t ), used to generate action probability distribution; the Critic network approximates the state value function V based on the parameter φ π (s t ), which is used to evaluate the expected return under different actions; PPO updates the policy network parameters θ by maximizing a clipped approximate objective function, which is: The strategy ratio is The clipping term limits the ratio to the range of [1-ε,1+ε], where ε is a hyperparameter with a value of 0.2, so that the difference between the current strategy and the old strategy is limited to a fixed range to prevent excessive policy parameter updates from causing training instability or overfitting; advantage function Execute action a under the current strategy t The difference between the reward and the predicted reward measures the relative merits of the action; by using the Monte Carlo approximate gradient method, the gradient ascent of the parameter θ in the Actor network is maximized L CLIP (θ) to train the Actor network by accumulating discounted reward values: θ and θ * They are the policy network parameters before and after the gradient update; the error function is minimized by gradient descent on the Critic network parameter φ: Where R t is the cumulative discount, V(s t ) is the predicted value; the network parameter φ is updated based on this value evaluation; the neural network parameters are continuously updated, the action and environment interact to train the control strategy, and the control model is obtained.
6. A DRL-based microchannel flow control method according to claim 5, characterized in that: In step 4-5, when designing reward feedback in the DRL framework, a dynamic normalization method based on historical rewards is introduced to optimize the stability and convergence effect of the deep reinforcement learning training process; the specific description is as follows: The first step is to define the reward function Rw as minimizing the vortex shedding energy on the cylinder: Rw=P e -s e e g (15) Among them, n node is the total number of grid nodes in the computational domain; (u xb ,u yb ) is the ground state of the flow field obtained by solving the SFD method; s e is the sum of the kinetic energy of the vortex shedding at all nodes in the computational domain; g is the maximum flow growth rate obtained by DMD evaluation. For a steady flow, g < 0, i.e., e g <1.0, indicating that the corresponding control strategy is effective and the reward increases; for neutral flow, g = 0, and the reward value is the original vortex shedding energy value; for unstable flow, g>0, that is, e g >1.0, the reward value of the corresponding action decreases, indicating that the control strategy is not effective; to adjust the magnitude of the reward value, a penalty term P is introduced e , which is used to ensure the stable propagation of gradients, avoid the problem of gradient update instability caused by too small reward values leading to slow learning process or too large reward values, and prevent gradient disappearance or explosion; the size of the penalty term is affected by factors such as spacing ratio, blocking ratio and Reynolds number; secondly, in order to avoid the reward value being biased towards positive or negative due to the empirically set initial kinetic energy difference, the historical reward dynamics are further standardized: Among them, d′ is the number of actions currently executed, Rw k is the reward value corresponding to the number of actions executed from d′-m+1 to d, μ s is the mean of historical rewards, standard deviation σ s Measures the discreteness of rewards, Rw s The reward value Rw obtained at the current time step d′ It is centered by the mean and scaled by the standard deviation. The current reward value is adjusted in real time at each training step by considering the mean and standard deviation of the most recent m = 80 historical rewards. The normalization process smoothes the rewards to avoid problems with over-concentration or excessive fluctuations in the reward values.
7. The DRL-based microchannel flow control method according to claim 5, characterized in that: In step 5, two indicators, flow field oscillation amplitude and vortex shedding energy, are used to evaluate the quality of the DRL microchannel flow control method, as follows: In the test results, the agent obtains the DNS flow field information of the corresponding time step after executing each action to calculate the lift coefficient C on the array cylindrical structure in the microchannel l : Where C is the cylindrical surface, dC is the cylindrical surface differential, and d yz Represents the diameter of a single cylinder, (n x , n y ) are the normal vectors in the streamwise and vertical directions, U τ represents the tangential velocity component at the integration point on the surface of the object, ρ represents the density of the fluid in the microchannel, and U max represents the maximum velocity of the fluid at the microchannel inlet, is the differential symbol; After a certain number of control actions, the change of lift coefficient is observed; at the same time, the value s of vortex shedding energy of microchannel flow field after corresponding control actions is recorded during the interaction of DRL flow control framework. e , observe the change of vortex shedding energy with the control action.
Citation Information
Cited By
Intelligent optimization design method for shape follow-up runner of mold based on deep learning
CN120724506A
Double-speed electric actuating mechanism control system
CN120742656A
A dual speed electric actuator control system
CN120742656B