A Smart Plasma Turbulent Friction Reduction Method Based on Deep Reinforcement Learning

The intelligent plasma turbulent friction drag reduction control system, which utilizes deep reinforcement learning and optimizes the perturbation strategy of the turbulent boundary layer using DDPG and LSTM networks, solves the problems of insufficient adaptability and efficiency in turbulent control and achieves a highly efficient turbulent friction drag reduction effect.

CN119439747BActive Publication Date: 2025-11-14AIR FORCE UNIV PLA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411603329.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-11-14
Estimated Expiration
2044-11-11

AI Technical Summary

Technical Problem

The chaotic nature of turbulence and its multi-scale, high-dimensional, and highly nonlinear behavior make it difficult for existing control methods to achieve efficient and adaptable turbulent friction drag reduction control, especially in complex flow environments where it is difficult to establish accurate theoretical models.

Method used

An intelligent plasma turbulent friction drag reduction control system based on deep reinforcement learning is adopted. It combines a high-voltage power supply, a plasma exciter, a state sensor, and an evaluation sensor. Real-time data learning and control are performed through a deep deterministic policy gradient network (DDPG) and a long short-term memory network (LSTM) to optimize the perturbation strategy of the turbulent boundary layer.

Benefits of technology

It achieves real-time closed-loop control of the turbulent boundary layer, improves adaptability to complex incoming flows and drag reduction efficiency, and reduces R&D costs and manpower and time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119439747B_ABST
    Figure CN119439747B_ABST
Patent Text Reader

Abstract

This paper discloses an intelligent plasma turbulent friction drag reduction control system based on deep reinforcement learning, including an environment setup, a host computer, and a controller. The environment setup includes a high-voltage power supply, a plasma exciter, state sensors, evaluation sensors, and a signal acquisition board. It also provides an intelligent plasma turbulent friction drag reduction method based on deep reinforcement learning. This invention selects hot wires and hot films as sensors and arrayed dielectric barrier discharge plasma with fast response as the exciter. Through a deep reinforcement learning algorithm deployed in a computer and FPGA controller, it continuously interacts at high speed with the turbulent boundary layer, learning and optimizing the control strategy of the intelligent agent, ultimately achieving real-time closed-loop control for turbulent friction drag reduction. Compared with traditional open-loop control, this method has advantages such as better adaptability to complex inflows and higher drag reduction efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of active flow control, and in particular to an intelligent plasma turbulent friction drag reduction method based on deep reinforcement learning. Background Technology

[0002] Turbulence is ubiquitous in engineering applications, with typical scenarios including oil transportation in pipelines, airflow on aircraft surfaces, and ship navigation in water. To improve transportation efficiency, it is often necessary to regulate turbulent flow to reduce wall friction resistance. Typical turbulent flow control methods can be divided into two main categories: passive control and active control. Groove surfaces are a typical example of passive control; although simple in structure, their drag reduction efficiency is weak under off-design conditions and may even increase drag. In contrast, active control methods can adjust the excitation intensity according to changes in external flow conditions, thus having broader application prospects. Based on whether there is sensor feedback signal in the control system, active flow control technology can be further divided into open-loop control and closed-loop control. Open-loop control uses pre-set excitation parameters, while closed-loop control can adaptively adjust the exciter output using instantaneous flow field signals acquired in real time by sensors, potentially maintaining high drag reduction efficiency. However, the chaotic nature of turbulence causes it to exhibit multi-scale, high-dimensional, and highly nonlinear behavior even on simple geometries such as smooth walls. The uncertainty of the flow state and the multiple degrees of freedom of the control input pose a significant challenge to achieving optimal closed-loop control in turbulent environments. Existing control methods typically rely on reduced-order models derived from simplifications of the physical problem, belonging to the category of physical model-based approaches. However, for complex flows, establishing an accurate theoretical model to describe turbulent behavior is nearly impossible. Model-free control, represented by deep reinforcement learning (DRL), does not depend on any underlying model description from input to output. It can achieve environmental cognition and optimization of control strategies through continuous trial and error, demonstrating the potential to undertake complex control tasks. Combining DRL with turbulent friction reduction is expected to improve profitability and adaptability, driving innovative development in this field. Summary of the Invention

[0003] This invention proposes an intelligent plasma turbulent friction reduction control system based on deep reinforcement learning, comprising environmental components, a host computer, and a controller, wherein...

[0004] Environmental components include a high-voltage power supply, a plasma actuator, status sensors, evaluation sensors, and signal acquisition boards;

[0005] A high-voltage power supply, used to drive the plasma exciter to discharge;

[0006] A plasma exciter is installed at the controlled position of the object to be controlled in order to generate effective disturbance to the turbulent boundary layer;

[0007] The state sensor, located upstream of the plasma actuator, is used to sense the state of the flow field and output the signal acquisition board to the host computer and controller.

[0008] An evaluation sensor, located at or downstream of the plasma exciter, assesses the changes in turbulent frictional resistance in the flow field after plasma excitation and outputs the results to the host computer via a signal acquisition board.

[0009] The signal acquisition board collects the real-time flow field data output by the status sensor and the evaluation sensor, and converts it from voltage signal to digital signal for output to the host computer.

[0010] The host computer receives digital signals output from the signal acquisition board, combines, processes, and analyzes the acquired data, and performs network learning and training through the DDPG algorithm deployed in the host computer.

[0011] The controller receives flow field status information from the status sensors, outputs action commands to the high-voltage power supply through the internally deployed Actor network, and transmits the current action data to the host computer.

[0012] In one specific embodiment of the present invention, the status sensor is a hot-wire sensor, and the evaluation sensor is a hot-film sensor.

[0013] In one embodiment of the present invention, the hot-wire anemometer converts the velocity and shear stress signals collected by the hot-wire sensor and the hot-film sensor into voltage signals, and then converts them into digital signals through a signal acquisition card and transmits them to the host computer; a hot-wire anemometer must be used in conjunction with the hot-wire or hot-film sensor.

[0014] In another embodiment of the present invention, the control system includes: a wind tunnel flat plate 1, a dual-wire wall hot wire 2, a plasma exciter 3, a hot film sensor 4, a high-voltage power supply 5, an IGBT high-voltage switch 6, a control computer, an FPGA controller 7, and a CTA anemometer 8.

[0015] The wind tunnel flat plate is set in the middle of the wind tunnel to generate a flat plate turbulent boundary layer, with the incoming flow direction from left to right; from upstream to downstream, a dual-wire wall hot wire 2, a plasma exciter 3, and a hot film sensor 4 are arranged in sequence; a high-voltage power supply 5, an IGBT high-voltage switch 6, a control computer and FPGA controller 7, and a CTA anemometer 8 are arranged on the outside of the wind tunnel and are connected to various devices in the wind tunnel through wires, BNC wires, and RS232 serial cables;

[0016] The dual-wire wall-mounted hot wires 2 are distributed along the spanwise array to capture the spanwise distribution of high- and low-velocity stripes in the turbulent boundary layer. The higher the experimental wind speed, the smaller the sensor spacing should be to match the scale of the strip structure. The dual-wire wall-mounted hot wires 2 are used to measure the time series u(t) and v(t) of the flow velocity in the direction (u) and normal (v) directions, respectively, and return the flow field state s. t =(u(t),v(t)), where "(u(t),v(t))" represents a combination of u(t) and v(t);

[0017] The positive and negative terminals of the high-voltage power supply 5 are connected to the exposed electrode and the grounding electrode of the plasma exciter 3 respectively through high-voltage resistant wires. An IGBT high-voltage switch 6 for controlling the high-frequency switching of the circuit is also connected in series between the plasma exciter 3 and the positive terminal of the high-voltage power supply 5.

[0018] A hot-film sensor 4 is attached downstream of the plasma exciter 3, arranged along the spanwise array. The spanwise measurement point is aligned with the upstream dual-wire wall hot wire 2 to assess the change in surface shear stress τ after convection, corresponding to the gain r. t =1-τ actuation / τ baseline , where τ baseline and τ actuation These correspond to the shear stress under reference and excitation conditions, respectively;

[0019] The dual-wire wall heating wire 2 and the thermal film sensor 4 collect flow field data and transmit the signals to the control computer and FPGA controller 7.

[0020] In yet another embodiment of the invention,

[0021] The dual-wire wall heating wires 2 are distributed in a spanwise array with an array spacing ranging from 5 to 20 mm; the normal position measured by the dual-wire wall heating wires 2 needs to be controlled within the viscous wall region y. + <Within 50, y + This indicates the dimensionless normal wall height; the height range of the dual-wire wall heating wire 2 is 10. <y + <50;

[0022] The plasma actuator 3 is a comb-shaped array, with the exposed electrode on the left and the grounded electrode on the right.

[0023] High-voltage power supply 5 is a high-voltage sinusoidal AC power supply, nanosecond pulse power supply, or pulse-DC power supply;

[0024] Both the dual-wire wall heating wire 2 and the hot film sensor 4 use a constant-temperature CTA anemometer 8 to collect flow field data.

[0025] In another specific embodiment of the present invention, the spacing L between the dual-wire wall heating wires 2 is... s =10mm; y+ =15;

[0026] The array spacing of plasma exciter 3 is the spanwise spacing L of the twin-wire wall-mounted hot wire 2. s Half of them, intersecting each other;

[0027] High-voltage power supply 5 is a high-voltage sinusoidal AC power supply, nanosecond pulse power supply, or pulse-DC power supply.

[0028] In another embodiment of the present invention, the dual-wire wall-mounted hot wire 2 comprises a hot wire base 21, a conical tip support 22, a solder joint 23, and a tungsten wire 24.

[0029] On the upper surface of the hot wire base 21, two conical pointed supports 22 are arranged along a diameter direction. The supports are conical in shape and form a tall, slender cone with the solder joint 23 arranged at the top of the supports. The two conical pointed supports 22 are arranged symmetrically about the center of the upper surface of the hot wire base 21.

[0030] On the upper surface of the hot wire base 21, two low cones are arranged along another diameter direction, and the two low cones are arranged symmetrically about the center of the upper surface of the hot wire base 21.

[0031] Each double-wire wall heating wire 2 has two tungsten wires 24 arranged on it. Each tungsten wire 24 is connected to a high and a low conical tip support 22 through a solder joint 23. The two tungsten wires 24 are X-shaped along the flow direction and cross each other.

[0032] A method for intelligent plasma turbulent friction drag reduction based on deep reinforcement learning is also provided. This method is based on the aforementioned intelligent plasma turbulent friction drag reduction control system based on deep reinforcement learning. The method comprises two parts: a control loop and a learning loop. The control loop interacts with the real-world scenario, while the learning loop learns from the experience accumulated in the control loop, as detailed below:

[0033] The control loop process is as follows;

[0034] Step 1: Real-time monitoring of the state parameters s of the turbulent boundary layer using a hot-wire sensor. t =[(u(t),v(t)),(u(t-τ),v(t-τ)),(u(t-2τ),v(t-2τ))], and set the state parameter s t The data is sent to the Actor network in the FPGA controller, where u(t) and v(t) represent the flow velocity and normal velocity at time t, respectively.

[0035] Step 2: The Actor network will convert the current state s t As input to the neural network, it performs the corresponding action a. t As the output of the network;

[0036] Step 3: The Deep Deterministic Policy Gradient Network (DDPG) adds a Gaussian noise function ε to the action policy and then sends the corresponding control command, i.e., action a, to the high-voltage power supply. t =φ(U,f m ,DC)+ε, where φ(.,.,.) represents U, f m A combination of DC, here it is an abstract combination of expression parameters, that is, action a is related to U, f m DC related; U, f m DC represents voltage, modulation frequency, and duty cycle, respectively.

[0037] Step 4: The high-voltage power supply, according to the instructions of the parameter combination transmitted by the Actor network, enables the plasma exciter to apply disturbance to the turbulent boundary layer with the specified excitation form, excitation intensity and frequency control.

[0038] Step 5: The turbulent boundary layer is disturbed, and the flow field state changes further. After the plasma actuator finishes driving, the hot-wire sensor monitors the new state parameter s again. ’ t And use it as input for the next round of the Actor network;

[0039] Step 6: The action of the hot-film sensor on the plasma actuator a t An evaluation was conducted to obtain the benefit r. t ;

[0040] Step 7: The FPGA controller processes the data obtained in this round [s] t ,a t ,r t ,s ’ t The package is sent to the host computer for use in the learning loop;

[0041] The specific learning cycle process is as follows;

[0042] The host computer obtains data information from multiple interactions between the FPGA controller and the flow field [s] t ,a t ,r t ,s ’ t ] Stored in the experience pool, when the amount of data in the memory pool reaches a threshold, multiple sets of data information are retrieved from the experience pool [s t ,a t ,r t ,s ’ t] Randomly select quantitative data as learning experience, calculate the loss function according to formula (1) and update the parameters of the Actor network and Critic network; The DDPG algorithm introduces two target networks on the basis of the current network architecture, and the data in the experience pool is used to update the parameters of the four networks; The learning optimization goal of the Actor network is to make the action-state value function Q output by the Critic network more efficient. μ (s t ,a t To maximize [-Q], the optimal strategy is obtained, i.e., gradient descent is used to make [-Q]... μ (s t ,a t The goal of the Critic network is to minimize the error between the current network and the target Critic network; the loss functions of the two networks are as follows:

[0043]

[0044] Among them, L Actor and L Critic Let θ and μ represent the two loss functions of the Actor network and Critic network, respectively; N represents the number of empirical groups selected; θ and μ represent the parameters of the Actor network and Critic network, respectively, corresponding to Q. μ (s,a|θ) is the Q-value of the Critic network with parameter μ and parameter θ, where s and a represent the current state s. t and current action a t r(s,a) represents the reward r corresponding to the current state and action. t Q μ (s,a) represents the Q-value of the Critic network output with parameter μ in the current state. γ∈[0,1] is the discount factor; when γ is close to 0, it means the agent values ​​short-term rewards more; when γ is close to 1, the agent values ​​long-term cumulative rewards more. Q μ (s ’ ,a ’ The target Critic network uses the next state s from the experience base. ’ t The output value, s ’ a ’ These represent the next state s respectively. ’ t and the next action a ’ t After the network parameters θ and μ of the Actor network and Critic network are updated, the current network parameters θ and μ are passed to the corresponding target networks at regular time steps. The moving average update method is used to update the network parameters θ of the two target networks. ’ and μ’ Through continuous experience playback and learning, the control strategy of the Actor network is optimized, enabling the agent trained by the FPGA controller to make optimal actions in the face of complex and ever-changing turbulent boundary layers, thereby reducing frictional resistance.

[0045] In one embodiment of the present invention, the algorithm training and learning process is as follows:

[0046] The first step is to establish a learning scenario. On the hardware level, the wind tunnel is turned on to establish a turbulent boundary layer, and the dual-wire hot-wire sensor, hot-film sensor, and plasma exciter are activated to put the acquisition equipment in real-time monitoring mode and the control equipment in standby mode. On the software level, the neural network structure required for the DDPG algorithm and the hyperparameters corresponding to the learning control are defined. Gaussian random noise is used to initialize the network parameters and distribute them to the four neural networks.

[0047] The second step is to accumulate learning experience; a dual-wire hot-wire sensor acquires the current turbulent boundary layer flow field state s. t =[(u(t),v(t)),(u(t-τ),v(t-τ)),(u(t-2τ),v(t-2τ))], sent to the Actor network, outputting the action a in the current state. t =φ(U,f m The plasma actuator is jointly controlled by a high-voltage power supply and an IGBT high-voltage switch, according to action a. t By changing the discharge waveform, excitations of different intensities and frequencies are achieved, thereby modulating the turbulent boundary layer; after the excitation ends, a dual-wire hot-wire sensor collects the next state s'. t The benefits r generated by the evaluation control of the hot film sensor t =1-τ actuation / τ baseline The FPGA controller will use the data information obtained from this round of interaction [s] t ,a t ,r t ,s ’ t After the package is sent to the control computer, the next round of interaction begins.

[0048] The third step is to replay the learning experience and update the network parameters. The control computer records the learning experience sent by the FPGA controller into the experience pool and judges in real time whether the data in the experience pool has reached the set threshold. If the experience pool is full, the program automatically samples randomly from the experience pool, calculates the corresponding loss function, and uses it to update the network parameters. Otherwise, it continues to accumulate learning experience.

[0049] The fourth step is the end of the learning process. The control computer determines whether the drag reduction rate has converged or whether the number of steps has reached the set number based on the data changes from the hot film sensor. If the conditions are met, the learning process ends, and the FPGA controller controls the turbulent boundary layer using the currently learned strategy. If the conditions are not met, the aforementioned learning process continues.

[0050] This invention selects hot wires and hot films as sensors and arrayed dielectric barrier discharge plasma with fast response as exciter. Through a deep reinforcement learning algorithm deployed in a computer and FPGA controller, it continuously interacts at high speed with the turbulent boundary layer, learning and optimizing the control strategy of the intelligent agent, ultimately achieving real-time closed-loop control for turbulent friction drag reduction. Compared to traditional open-loop control, this method has advantages such as better adaptability to complex inflows and higher drag reduction efficiency.

[0051] The advantages of this invention are as follows:

[0052] 1. Improved adaptability to external flow: In traditional open-loop control, the exciter operates independently of the flow field state, lacking adjustment capability and exhibiting poor adaptability to complex flow conditions. The control method provided in this invention, combined with the DDPG-LSTM algorithm, can perceive the state of the flow field in the turbulent boundary layer in real time and adjust the learning strategy according to changes in the flow field, thereby significantly improving adaptability.

[0053] 2. Higher drag reduction efficiency: The control method of this invention can achieve condition-based excitation according to the flow field state, eliminating unnecessary power consumption, especially when no excitation is required. It achieves higher control efficiency while improving drag reduction capability.

[0054] 3. Data-driven and cost-saving in R&D: Compared with model-based control methods, the control method of this invention does not rely on high-precision modeling. It only needs to be deployed in a real environment for interaction to learn the corresponding control strategy (model). In addition, this method does not require manual intervention during operation, significantly saving manpower and time and reducing construction costs. Attached Figure Description

[0055] Figure 1 The framework of an intelligent plasma turbulent friction drag reduction control system based on deep reinforcement learning is shown.

[0056] Figure 2 A schematic diagram of an intelligent closed-loop turbulent friction reduction scheme based on deep reinforcement learning is shown.

[0057] Figure 3 A schematic diagram of a dual-wire wall-mounted heating wire structure is shown.

[0058] Figure 4 A flowchart of an intelligent closed-loop turbulent friction reduction scheme based on deep reinforcement learning is shown.

[0059] Figure label:

[0060] 1. Wind tunnel flat plate 2. Dual-wire wall-mounted hot wire 2.1 Hot wire base 2.2 Conical tip support 2.3 Solder joint 2.4 Tungsten wire 3. Plasma exciter 4. Hot film sensor 5. High-voltage power supply 6. IGBT high-voltage switch 7. Control computer and FPGA controller 8. CTA anemometer / hot wire meter Detailed Implementation

[0061] The present invention will now be described in detail with reference to the accompanying drawings.

[0062] Figure 1 This invention presents a framework for an intelligent plasma turbulent friction reduction control system based on deep reinforcement learning. The control system framework consists of three parts: environment setup, host computer, and controller.

[0063] The environment setup mainly includes a high-voltage power supply, a plasma exciter, a status sensor (hot-wire sensor), an evaluation sensor (hot-film sensor), a hot-wire anemometer, and a signal acquisition board. The object to be controlled is the turbulent boundary layer.

[0064] The high-voltage power supply primarily provides the driving power for the plasma actuator, driving it to discharge and generate plasma with specified excitation parameters (voltage, modulation frequency, and duty cycle). The high-voltage power supply can be a high-voltage sinusoidal AC power supply, a nanosecond pulse power supply, or a pulse-DC power supply, etc., as long as it meets the requirement of driving the plasma actuator to discharge; no specific restrictions apply.

[0065] Plasma actuators come in various types (e.g., surface arc plasma actuators, synthetic jet plasma actuators, dielectric barrier discharge plasma actuators, etc.), with dielectric barrier discharge plasma actuators being preferred due to their simple structure and rapid response. Plasma actuators suitable for turbulent drag reduction have diverse structural forms, which can be broadly classified into spanwise jets, normal jets, and flow-oriented jets based on the different forms of disturbance they generate in the flow field. This invention only requires the plasma actuator to effectively disturb the turbulent boundary layer; its structural form is not limited. The plasma actuator is installed at the controlled location (turbulent boundary layer).

[0066] State sensors, as a crucial component of the control system, are used to sense the state of the flow field. They monitor parameters such as velocity, temperature, and shear stress, which directly reflect the flow field state, and are typically located upstream of the plasma exciter. Hot-wire sensors with high frequency response (above 10kHz) and low disturbance to the flow field are preferred, aiming to improve the accuracy of flow field sensing while minimizing the introduction of additional disturbances.

[0067] The evaluation sensor is primarily used to assess changes in turbulent frictional resistance (surface shear stress). Suitable sensors include force balances, Preston tubes, oil film interferometers, hot-film sensors, and wall-mounted hot wires for measuring wall shear stress. The sensor can be located at or downstream of the plasma exciter. Consistent with the requirements for the aforementioned condition monitoring sensors, a hot-film sensor with a high frequency response is preferred to avoid delays in control time due to a low sensor frequency response, thereby improving the control frequency of the overall closed-loop control system.

[0068] Hot-wire anemometers convert the velocity and shear stress signals collected by state sensors (e.g., hot-wire sensors) and evaluation sensors (e.g., hot-film sensors) into voltage signals, which are then converted into digital signals by a signal acquisition card and transmitted to a host computer. Note that a hot-wire anemometer is generally required when using hot-wire or hot-film sensors; however, it is usually not necessary when using other sensors.

[0069] The present invention does not impose specific requirements on the signal acquisition board, as long as it can meet the requirements for high-speed acquisition of sensor data signals.

[0070] The host computer and controller are responsible for the control hardware and software algorithm parts in the control system framework.

[0071] The host computer can be a mature computer, and the controller can be a computer, a microcontroller, an FPGA (Field Programmable Gate Array) controller, or other control devices that can operate stably and at high speed. This invention does not limit the specific implementation method and device of the algorithm. It is preferred to use an FPGA controller to interact with the environment. Compared with microcontroller controllers and traditional CPU control computers, FPGA controllers can reduce the control delay time to the microsecond level, which can fully meet the high-speed intelligent closed-loop control requirements of this invention for turbulence drag reduction control.

[0072] This invention primarily employs an algorithm combining Deep Deterministic Policy Gradient (DDPG) [Lillicrap TP, Hunt JJ, Pritzel A, et al. Continuous control with deep reinforcement learning[J]. arXiv preprintarXiv:1509.02971,2015.] with Long Short-Term Memory (LSTM) networks [Hochreiter S. Long Short-Term Memory[J]. Neural Computation MIT-Press,1997.] to explore and optimize the turbulence drag reduction control law. The DDPG algorithm is a powerful and flexible reinforcement learning algorithm. By combining the ideas of deep learning and policy gradient, compared to deep Q-network algorithms, the DDPG algorithm can achieve efficient policy learning and optimization in a continuous action space, and enhances its understanding of the environment by introducing a noise exploration mechanism.

[0073] The intelligent plasma turbulent friction drag reduction method based on deep reinforcement learning of the present invention is as follows:

[0074] The DDPG algorithm provides the learned experience gained from interacting with the environment to the simulated agent for training. The agent continuously updates its control strategy according to changes in the environment, ultimately training a suitable agent. Specifically, the DDPG algorithm is based on an actor-critic framework, constructing two network structures: actor and critic (the construction method is well known to those skilled in the art). The actor network is based on the current state s... t Perform the corresponding action a t The Critic network, on the other hand, considers the current state s. t The action taken by the Actor network, a t The quality of the action is judged (subsequently using the "value of the action-state value function Q"), that is, the action-state value function Q is output. μ (s t ,a t (This function is well known to those skilled in the art). Due to its unique memory and gate structures, the LSTM network architecture can effectively mine and learn highly correlated features in time series data. Compared to the fully connected neural network structure commonly used in the DDPG algorithm, the LSTM network has greater advantages in handling time series prediction and long-term dependencies, and is expected to be well-suited to the convection characteristics and correlations of strip structures in turbulent boundary layers. Therefore, all neural networks in this invention adopt the LSTM architecture.

[0075] This invention provides an intelligent plasma turbulent friction reduction method based on deep reinforcement learning. The method comprises two parts: a control loop and a learning loop. The control loop interacts with the real-world scenario to generate empirical data, while the learning loop learns from the experience accumulated in the control loop and updates the parameters of the algorithm network. The two loops are executed in parallel, as detailed below:

[0076] The control loop process is as follows.

[0077] Step 1: Real-time monitoring of the state parameters s of the turbulent boundary layer using a hot-wire sensor. t =[(u(t),v(t)),(u(t-τ),v(t-τ)),(u(t-2τ),v(t-2τ))] (tracing back two time steps τ), where u(t) and v(t) represent the flow velocity and normal velocity at time t, respectively, and the state parameter s is... t Send to the Actor network in the FPGA controller.

[0078] Step 2: The Actor network will convert the current state s t As input to the neural network, it performs the corresponding action a. t As the output of the network.

[0079] Step 3: To enhance the agent's exploration of the environment, DDPG adds a Gaussian noise function ε to the action policy and then sends the corresponding control command, i.e., action a, to the high-voltage power supply. t =φ(U,f m ,DC)+ε, where φ(.,.,.) represents U, f m A combination of DC, here it is an abstract combination of excitation parameters, that is, action a is related to U, f m For DC-related parameters, the user can define them according to their needs or the specific high-voltage power supply; no specific restrictions are imposed. U, f m DC represents voltage, modulation frequency, and duty cycle, respectively.

[0080] Step 4: Based on the parameter combination (voltage, frequency, and duty cycle) transmitted by the Actor network, the high-voltage power supply causes the plasma actuator to apply perturbation to the turbulent boundary layer with specific excitation form, excitation intensity, and frequency control. This method is well known to those skilled in the art and will not be described further.

[0081] Step 5: The turbulent boundary layer is disturbed, and the flow field state changes further. After the plasma actuator finishes driving, the hot-wire sensor monitors the new state parameter s again. ’ t And use it as input for the next round of the Actor network.

[0082] Step 6: The action of the hot-film sensor on the plasma actuator a t An evaluation was conducted (as described below) to obtain the benefit r. t .

[0083] Step 7: The FPGA controller processes the data obtained in this round [s] t ,a t ,r t ,s ’ t The package is sent to the host computer for learning loops.

[0084] The specific learning cycle process is as follows.

[0085] The host computer obtains data information from multiple interactions between the FPGA controller and the flow field [s] t ,a t ,r t ,s ’ t ] Stored in the experience pool, when the amount of data in the memory pool reaches a threshold, multiple sets of data information are retrieved from the experience pool [s t ,a t ,r t ,s ’ t A quantitative amount of data is randomly selected from the pool as learning experience (the specific amount is determined according to the experimental needs and is not specifically limited), and the loss function is calculated according to formula (1) to update the parameters of the Actor network and the Critic network. The DDPG algorithm introduces two target networks (target Actor network and target Critic network) on the basis of the current network architecture to improve the stability of the learning process. The data in the experience pool needs to be used to update the parameters of the four networks. Specifically, the learning optimization goal of the Actor network is to make the action-state value function Q output by the Critic network more efficient. μ (s t ,a t To maximize [-Q], the optimal strategy is obtained, i.e., gradient descent is used to make [-Q]... μ (s t ,a t The goal of the Critic network is to minimize the error between the current network and the target Critic network ("gradient descent" is well known to those skilled in the art); while the goal of the Critic network is to minimize the error between the current network and the target Critic network. The loss functions of the two networks are as follows:

[0086]

[0087] Among them, L Actor and L CriticLet θ and μ represent the two loss functions of the Actor network and Critic network, respectively; N represents the number of empirical groups selected; θ and μ represent the parameters of the Actor network and Critic network, respectively, corresponding to Q. μ (s,a|θ) is the Q-value of the Critic network with parameter μ and parameter θ, where s and a represent the current state s. t and current action a t r(s,a) represents the reward r corresponding to the current state and action. t Q μ (s,a) represents the Q-value of the Critic network output with parameter μ in the current state. γ∈[0,1] is the discount factor; when γ is close to 0, it means the agent values ​​short-term rewards more; when γ is close to 1, the agent values ​​long-term cumulative rewards more. Q μ (s ’ ,a ’ The target Critic network uses the next state s from the experience base. ’ t The output value, s ’ a ’ These represent the next state s respectively. ’ t and the next action a ’ t Furthermore, after the network parameters θ and μ of the Actor network and Critic network are updated (these two network parameters are known to those skilled in the art), the current network parameters θ and μ are passed to the corresponding target networks (i.e., the target Actor network and the target Critic network) at regular time steps, and the network parameters θ of the two target networks are updated using a moving average update method (this method is known to those skilled in the art). ’ and μ ’ Ultimately, through continuous experience playback and learning, the control strategy of the Actor network is optimized, enabling the agent trained by the FPGA controller to make optimal actions in the face of complex and ever-changing turbulent boundary layers, thereby reducing frictional resistance.

[0088] Specific implementation examples:

[0089] Figure 2 This is a schematic diagram of an intelligent closed-loop turbulent friction reduction scheme based on deep reinforcement learning. It mainly includes: a wind tunnel flat plate 1, a dual-wire wall-mounted hot wire 2, a plasma exciter 3, a hot-film sensor 4, a high-voltage power supply 5, an IGBT high-voltage switch 6, a control computer and FPGA controller 7, and a CTA anemometer 8. It should be noted that the dual-wire wall-mounted hot wire and the hot-film sensor correspond to the state sensor and evaluation sensor in the aforementioned control system, respectively. These devices are only illustrative here and are not strictly limiting.

[0090] The implementation case was conducted in a wind tunnel laboratory. The wind tunnel flat plate was placed in the middle of the wind tunnel to generate a flat plate turbulent boundary layer, with the incoming flow direction from left to right. From upstream to downstream, a dual-wire wall hot wire 2, a plasma exciter 3, and a hot film sensor 4 were arranged sequentially. A high-voltage power supply 5, an IGBT high-voltage switch 6, a control computer and FPGA controller 7, and a CTA anemometer 8 were arranged on the experimental table outside the wind tunnel and connected to the various devices inside the wind tunnel via wires, BNC cables, and RS232 serial cables.

[0091] The dual-wire wall-mounted heating wires 2 are arrayed along the spanwise (z-axis) to capture the distribution of high- and low-velocity stripes along the spanwise direction in the turbulent boundary layer. The array spacing ranges from 5 to 20 mm. The larger the experimental wind speed, the smaller the sensor spacing should be to match the scale of the strip structure. L is preferred. s =10mm. The structure of the dual-wire wall heating wire 2 is as follows: Figure 3 As shown, the system consists of a hot wire base 21, conical tip supports 22, solder joints 23, and tungsten wires 24. The shape of the hot wire base 21 is determined according to the specific experimental model installation location; it is shown as a cylinder in the figure. On the upper surface of the hot wire base 21, two conical tip supports 22 are arranged along one diameter direction. These supports are conical in shape and, together with the solder joints 23 at their apexes, form a tall, slender cone. The two conical tip supports 22 are arranged symmetrically about the center of the upper surface of the hot wire base 21. Similarly, two similar low cones are arranged along the other diameter direction, also symmetrically about the center of the upper surface of the hot wire base 21. Each sensor has two tungsten wires 24, each connected to one tall and one short conical tip support 22 via solder joints 23. The two tungsten wires 24 are arranged in an X-shape along the flow direction, intersecting each other, and jointly measure the time series u(t) and v(t) of the flow velocity's directional (u) and normal (v) components, respectively, returning the flow field state s. t =(u(t),v(t)), where "(u(t),v(t))" represents a combination of u(t) and v(t), which can be defined by the user according to their needs without specific restrictions. Considering that the wall-mounted hotwire needs to monitor the stripe structure distribution in the turbulent boundary layer in real time, the normal position measured by the hotwire needs to be controlled within the viscous wall region (y + <50), y + This represents the dimensionless normal wall height (well known to those skilled in the art), but near the wall, the hot wire inevitably suffers from measurement errors due to wall thermal effects. Therefore, the hot wire height range is 10. <y + <50, preferred y + =15, where the turbulence intensity of the flow velocity fluctuations reaches its peak.

[0092] In this case, the plasma actuator 3 has a common comb-shaped array structure, with exposed electrodes on the left and grounded electrodes on the right. The array spacing is not limited, but preferably has a spanwise spacing L between the twin-wire wall-mounted hot wires 2. s Half of them, intersecting each other.

[0093] The plasma exciter 3 is driven by a high-voltage power supply 5, which can be selected from: a high-voltage sinusoidal AC power supply, a nanosecond pulse power supply, a pulse-DC power supply, etc. The positive and negative terminals of the high-voltage power supply 5 are connected to the exposed electrode and the ground electrode of the plasma exciter 3 respectively through high-voltage resistant wires. An IGBT high-voltage switch 6 for controlling the high-frequency switching of the control circuit is also connected in series between the plasma exciter 3 and the positive terminal of the high-voltage power supply 5.

[0094] The hot-film sensor 4 is attached downstream of the plasma actuator 3, also at a spacing L. s Along the spanwise array, the spanwise measuring points are aligned with the upstream twin-wire wall hot wire 2 to assess the change in surface shear stress τ after convection, corresponding to the gain r. t =1-τ actuation / τ baseline , where τ baseline and τ actuation These correspond to the shear stress under reference and excitation conditions, respectively.

[0095] Both the dual-wire wall hot wire 2 and the hot film sensor 4 use a constant-temperature CTA anemometer 8 to collect flow field data and transmit the signals to the control computer and FPGA controller 7. The above measurement methods are common measurement methods in the field of flow measurement, and will not be described in detail here.

[0096] The following section introduces the algorithm training and learning process for a specific implementation case.

[0097] The first step is to establish the learning scenario. On the hardware side: The wind tunnel is activated to create a turbulent boundary layer. The dual-wire hot-wire sensor, hot-film sensor, and plasma exciter are started, putting the data acquisition equipment in real-time monitoring mode and the control equipment in standby mode. On the software side: The neural network structure required for the DDPG algorithm and the corresponding hyperparameters for learning control are defined. Gaussian random noise is used to initialize the network parameters, which are then distributed to the four neural networks.

[0098] The second step is to accumulate learning experience. A dual-wire hot-wire sensor acquires the current turbulent boundary layer flow field state s. t =[(u(t),v(t)),(u(t-τ),v(t-τ)),(u(t-2τ),v(t-2τ))], sent to the Actor network, outputting the action a in the current state. t =φ(U,f m The plasma actuator is jointly controlled by a high-voltage power supply and an IGBT high-voltage switch, according to action a.t By altering the discharge waveform, excitation at different intensities and frequencies is achieved, thereby modulating the turbulent boundary layer. After excitation ends, a dual-wire hot-wire sensor acquires the next state s'. t The benefits r generated by the evaluation control of the hot film sensor t =1-τ actuation / τ baseline The FPGA controller will process the data information obtained from this round of interaction [s] t ,a t ,r t ,s ’ t After the package is sent to the control computer, the next round of interaction begins.

[0099] The third step is experience playback and network parameter updates. The control computer records the learning experience sent by the FPGA controller into the experience pool and checks in real time whether the data in the experience pool has reached the set threshold. If the experience pool is full, the program automatically samples randomly from the experience pool, calculates the corresponding loss function, and uses it to update the network parameters; otherwise, it continues to accumulate learning experience.

[0100] The fourth step is the end of the learning process. The control computer, based on the data changes from the hot-film sensor, determines whether the drag reduction rate has converged or whether the set number of steps has been reached. If the conditions are met, the learning process ends, and the FPGA controller controls the turbulent boundary layer using the currently learned strategy. If the conditions are not met, the learning process continues.

Claims

1. A smart plasma turbulent friction drag reduction control system based on deep reinforcement learning, characterized in that, The system includes environmental components, a host computer, and a controller. in, Environmental components include a high-voltage power supply, a plasma actuator, status sensors, evaluation sensors, and signal acquisition boards; A high-voltage power supply, used to drive the plasma exciter to discharge; A plasma exciter is installed at the controlled position of the object to be controlled in order to generate effective disturbance to the turbulent boundary layer; The state sensor, located upstream of the plasma actuator, is used to sense the state of the flow field and output the signal acquisition board to the host computer and controller. An evaluation sensor, located at or downstream of the plasma exciter, assesses the changes in turbulent frictional resistance in the flow field after plasma excitation and outputs the results to the host computer via a signal acquisition board. The signal acquisition board collects the real-time flow field data output by the status sensor and the evaluation sensor, and converts it from voltage signal to digital signal for output to the host computer. The host computer receives digital signals output from the signal acquisition board, combines, processes, and analyzes the acquired data, and performs network learning and training through the DDPG algorithm deployed in the host computer. The controller receives flow field status information from the status sensors, outputs action commands to the high-voltage power supply through the internally deployed Actor network, and transmits the current action data to the host computer.

2. The intelligent plasma turbulent friction drag reduction control system based on deep reinforcement learning as described in claim 1, characterized in that, The status sensor uses a hot-wire sensor, and the evaluation sensor uses a hot-film sensor.

3. The intelligent plasma turbulent friction reduction control system based on deep reinforcement learning as described in claim 2, characterized in that, The hot-wire anemometer converts the velocity and shear stress signals collected by the hot-wire sensor and the hot-film sensor into voltage signals, which are then converted into digital signals by the signal acquisition card and transmitted to the host computer. A hot-wire anemometer must be used in conjunction with the hot-wire or hot-film sensor.

4. The intelligent plasma turbulent friction drag reduction control system based on deep reinforcement learning as described in claim 1, characterized in that, The control system includes: wind tunnel flat plate (1), dual-wire wall hot wire (2), plasma exciter (3), hot film sensor (4), high voltage power supply (5), IGBT high voltage switch (6), control computer, FPGA controller (7), and CTA anemometer (8). The wind tunnel flat plate is set in the middle of the wind tunnel to generate a flat plate turbulent boundary layer, with the incoming flow direction from left to right; from upstream to downstream, a double-wire wall hot wire (2), a plasma exciter (3) and a hot film sensor (4) are arranged in sequence; a high-voltage power supply (5), an IGBT high-voltage switch (6), a control computer and FPGA controller (7) and a CTA anemometer (8) are arranged on the outside of the wind tunnel and are connected to the various devices in the wind tunnel through wires, BNC wires and RS232 serial cables; The dual-wire wall-mounted hot wires (2) are distributed along the spanwise array to capture the distribution of high- and low-velocity strips along the spanwise in the turbulent boundary layer. The higher the experimental wind speed, the smaller the sensor spacing should be to match the scale of the strip structure. The dual-wire wall-mounted hot wires (2) are used to measure the time series u(t) and v(t) of the flow direction (u) and normal (v) components of the incoming flow velocity, and return the flow field state s. t =(u(t),v(t)), where "(u(t),v(t))" represents a combination of u(t) and v(t); The positive and negative terminals of the high voltage power supply (5) are connected to the exposed electrode and grounding electrode of the plasma exciter (3) respectively through high voltage resistant wires. An IGBT high voltage switch (6) for controlling the high frequency switching of the circuit is also connected in series between the plasma exciter (3) and the positive terminal of the high voltage power supply (5). A hot-film sensor (4) is attached downstream of the plasma exciter (3) and arranged along the spanwise array. The spanwise measurement point is aligned with the upstream dual-wire wall hot wire (2) to assess the change in surface shear stress τ after convection and the corresponding gain r. t =1-τ actuation / τ baseline , where τ baseline and τ actuation These correspond to the shear stress under reference and excitation conditions, respectively; The dual-wire wall heating wire (2) and the thermal film sensor (4) collect flow field data and transmit the signals to the control computer and FPGA controller (7).

5. The intelligent plasma turbulent friction drag reduction control system based on deep reinforcement learning as described in claim 4, characterized in that, The twin-wire wall heating wires (2) are distributed in a spanwise array with an array spacing of 5-20 mm; the normal position measured by the twin-wire wall heating wires (2) needs to be controlled within the viscous wall region y. + <Within 50, y + This indicates the dimensionless normal wall height; the height range of the double-wire wall heating wire (2) is 10. <y + <50; The plasma exciter (3) is a comb-shaped array, with the exposed electrode on the left and the grounded electrode on the right. The high-voltage power supply (5) is a high-voltage sinusoidal AC power supply, a nanosecond pulse power supply, or a pulse-DC power supply; Both the dual-wire wall heating wire (2) and the thermal film sensor (4) use a constant-temperature CTA anemometer (8) to collect flow field data.

6. The intelligent plasma turbulent friction drag reduction control system based on deep reinforcement learning as described in claim 5, characterized in that, Dual-wire wall heating wire (2) spacing L s =10mm; y + =15; The plasma exciter (3) array spacing is the spanwise spacing L of the double-wire wall-mounted hot wire (2). s Half of them, intersecting each other; The high-voltage power supply (5) is a high-voltage sinusoidal AC power supply, nanosecond pulse power supply or pulse-DC power supply.

7. The intelligent plasma turbulent friction drag reduction control system based on deep reinforcement learning as described in claim 4, characterized in that, The dual-wire wall-mounted heating wire (2) consists of a heating wire base (21), a conical tip support (22), solder joints (23), and a tungsten wire (24); On the upper surface of the hot wire base (21), two conical tip supports (22) are arranged along a diameter direction. The whole is conical and forms a tall and slender cone with the solder joint (23) arranged at its top. The two conical tip supports (22) are arranged symmetrically about the center of the upper surface of the hot wire base (21). On the upper surface of the hot wire base (21), two low cones are arranged along another diameter direction. The two low cones are arranged symmetrically about the center of the upper surface of the hot wire base (21). Each double-wire wall heating wire (2) has two tungsten wires (24) arranged on it. Each tungsten wire (24) is connected to a high and low conical tip support (22) through a solder joint (23). The two tungsten wires (24) are X-shaped along the flow direction and cross each other.

8. A method for intelligent plasma turbulent friction drag reduction based on deep reinforcement learning, which is based on the intelligent plasma turbulent friction drag reduction control system based on deep reinforcement learning as described in any one of claims 1 to 7, characterized in that, This method consists of two parts: a control loop and a learning loop. The control loop is used to interact with the real-world scenario, while the learning loop learns based on the experience accumulated in the control loop, as detailed below: The control loop process is as follows; Step 1: Real-time monitoring of the state parameters s of the turbulent boundary layer using a hot-wire sensor. t =[(u(t),v(t)),(u(t-τ),v(t-τ)),(u(t-2τ),v(t-2τ))], and set the state parameter s t The data is sent to the Actor network in the FPGA controller, where u(t) and v(t) represent the flow velocity and normal velocity at time t, respectively. Step 2: The Actor network will convert the current state s t As input to the neural network, it performs the corresponding action a. t As the output of the network; Step 3: The Deep Deterministic Policy Gradient Network (DDPG) adds a Gaussian noise function ε to the action policy and then sends the corresponding control command, i.e., the action, to the high-voltage power supply. in Representing U, f m A combination of DC, here it is an abstract combination of expression parameters, that is, action a is related to U, f m DC related; U, f m DC represents voltage, modulation frequency, and duty cycle, respectively. Step 4: The high-voltage power supply, according to the instructions of the parameter combination transmitted by the Actor network, enables the plasma exciter to apply disturbance to the turbulent boundary layer with the specified excitation form, excitation intensity and frequency control. Step 5: The turbulent boundary layer is disturbed, and the flow field state changes further. After the plasma actuator finishes driving, the hot-wire sensor monitors the new state parameter s' again. t And use it as input for the next round of the Actor network; Step 6: The action of the hot-film sensor on the plasma actuator a t An evaluation was conducted to obtain the benefit r. t ; Step 7: The FPGA controller processes the data obtained in this round [s] t ,a t ,r t ,s' t The package is sent to the host computer for use in the learning loop; The specific learning cycle process is as follows; The host computer obtains data information from multiple interactions between the FPGA controller and the flow field [s] t ,a t ,r t ,s' t ] Stored in the experience pool, when the amount of data in the memory pool reaches a threshold, multiple sets of data information are retrieved from the experience pool [s t ,a t ,r t ,s' t ] Randomly select quantitative data as learning experience, calculate the loss function according to formula (1) and update the parameters of Actor network and Critic network; The DDPG algorithm introduces two target networks into the current network architecture, and the data in the experience pool is used to update the parameters of the four networks; the learning and optimization objective of the Actor network is to make the action-state value function Q output by the Critic network more efficient. μ (s t ,a t To maximize [-Q], the optimal strategy is obtained, i.e., gradient descent is used to make [-Q]... μ (s t ,a t The goal of the Critic network is to minimize the error between the current network and the target Critic network; the loss functions of the two networks are as follows: Among them, L Actor and L Critic Let θ and μ represent the two loss functions of the Actor network and Critic network, respectively; N represents the number of empirical groups selected; θ and μ represent the parameters of the Actor network and Critic network, respectively, corresponding to Q. μ (s,a|θ) is the Q-value of the Critic network with parameter μ and parameter θ, where s and a represent the current state s. t and current action a t r(s,a) represents the reward r corresponding to the current state and action. t Q μ (s,a) represents the Q-value of the Critic network with parameter μ in the current state; γ∈[0,1] is the discount factor, when γ is close to 0, it means the agent values ​​short-term rewards more; when γ is close to 1, the agent values ​​long-term cumulative rewards more; Q μ (s',a') represents the target Critic network's next state s' based on the experience base. t The output values ​​s' and a' represent the next state s', respectively. t and the next action a' t ; After the network parameters θ and μ of the Actor network and Critic network are updated, the current network parameters θ and μ are passed to the corresponding target networks at regular time steps. The network parameters θ' and μ' of the two target networks are updated using the moving average update method. Through continuous experience playback and learning, the control strategy of the Actor network is optimized, enabling the agent trained by the FPGA controller to make optimal actions in the face of complex and ever-changing turbulent boundary layers, thereby reducing frictional resistance.

9. The intelligent plasma turbulent friction reduction method based on deep reinforcement learning as described in claim 8, characterized in that, The specific training and learning process of the algorithm is as follows: The first step is to establish a learning scenario. On the hardware level, the wind tunnel is turned on to establish a turbulent boundary layer, and the dual-wire hot-wire sensor, hot-film sensor, and plasma exciter are activated to put the acquisition equipment in real-time monitoring mode and the control equipment in standby mode. On the software level, the neural network structure required for the DDPG algorithm and the hyperparameters corresponding to the learning control are defined. Gaussian random noise is used to initialize the network parameters and distribute them to the four neural networks. The second step is to accumulate learning experience; a dual-wire hot-wire sensor acquires the current turbulent boundary layer flow field state s. t =[(u(t),v(t)),(u(t-τ),v(t-τ)),(u(t-2τ),v(t-2τ))], sent to the Actor network, outputting the action in the current state. The plasma actuator is jointly controlled by a high-voltage power supply and an IGBT high-voltage switch, according to action a. t By changing the discharge waveform, excitations of different intensities and frequencies are achieved, thereby modulating the turbulent boundary layer; after the excitation ends, a dual-wire hot-wire sensor collects the next state s'. t The benefits r generated by the evaluation control of the hot film sensor t =1-τ actuation / τ baseline The FPGA controller will use the data information obtained from this round of interaction [s] t ,a t ,r t ,s' t After the package is sent to the control computer, the next round of interaction begins. The third step is to replay the learning experience and update the network parameters. The control computer records the learning experience sent by the FPGA controller into the experience pool and judges in real time whether the data in the experience pool has reached the set threshold. If the experience pool is full, the program automatically samples randomly from the experience pool, calculates the corresponding loss function, and uses it to update the network parameters. Otherwise, it continues to accumulate learning experience. The fourth step is the end of the learning process. The control computer determines whether the drag reduction rate has converged or whether the number of steps has reached the set number based on the data changes from the hot film sensor. If the conditions are met, the learning process ends, and the FPGA controller controls the turbulent boundary layer using the currently learned strategy. If the conditions are not met, the aforementioned learning process continues.

Citation Information

Patent Citations

  • High-speed closed-loop flow control method based on reinforcement learning and FPGA neural network

    CN118011936A

  • Turbulence boundary layer plasma drag reduction system

    CN211930949U