Knowledge distillation-based high-speed flow interpretable artificial intelligence control method

Through knowledge distillation, the neural network model is transformed into a lightweight control model of symbolic expressions, which solves the interpretability problem of traditional deep reinforcement learning in the aerospace field, and achieves the efficiency and transparency of high-speed flow control.

CN120370720AInactive Publication Date: 2025-07-25AIR FORCE UNIV PLA
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510865077.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-07-25
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The traditional deep reinforcement learning flow control method has poor interpretability in the aerospace field, which leads to the inability to intuitively establish the connection between state and control instructions and cannot reflect the physical mechanism behind the perturbation strategy.

Method used

Using a knowledge distillation-based method, the pre-trained high-speed flow control model is converted into a lightweight control model in the form of symbolic expressions. Reinforcement learning is carried out through field programmable gate arrays (FPGAs) and image processing units (GPUs), and symbolic expressions are optimized using genetic programming to improve interpretability.

Benefits of technology

The high transparency and understanding of the high-speed flow control model are realized, which can directly reflect the functional relationship between the input variables and the output results, and meet the real-time requirements of supersonic flow control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120370720A_ABST
    Figure CN120370720A_ABST
Patent Text Reader

Abstract

The invention provides a high-speed flow interpretable artificial intelligence control method based on knowledge distillation. Wherein the electronic equipment obtains a pre-trained high-speed flow control model, and the high-speed flow control model is a model based on a neural network and can generate a disturbance strategy according to flow field state information; and distilling a lightweight control model in a symbol expression form by using a high-speed flow control model. Therefore, knowledge contained in the high-speed flow control model is extracted and converted into the lightweight control model in the symbol expression form, and the symbol expression is an explicit mathematical expression form and can directly reflect a function relationship between an input variable and an output result, so that compared with an implicit neural network model, the knowledge contained in the high-speed flow control model is extracted and converted into the lightweight control model in the symbol expression form; and the method has higher transparency and understandability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine learning, and more particularly, to an interpretable artificial intelligence control method for high-speed flows based on knowledge distillation. Background Art

[0002] Active flow control technology is a cutting-edge technology in the aerospace field. By introducing local controllable perturbations into the flow field, the macroscopic flow characteristics around an object can be changed, thereby achieving purposes such as lift increase and drag reduction, noise reduction and thrust increase. Moreover, after more than a century of development, active flow control has evolved from the open-loop steady control stage in the Prandtl era to the closed-loop adaptive control stage in the 21st century. Its core task is to find a closed-loop control law that can maximize the control benefit.

[0003] In recent years, data-driven control methods represented by deep reinforcement learning (DRL) have shown broad application prospects in the field of active flow control. By integrating the perception ability of deep learning and the decision-making ability of reinforcement learning, experience data is accumulated and the parameters of the deep neural network are updated during the continuous interaction and trial-and-error process with the controlled object, thereby enhancing its ability to solve complex non-linear decision-making problems. Research has shown that applying deep reinforcement learning technology to the low-speed flow around a cylinder and airfoil flow can automatically find a closed-loop control law that significantly reduces the drag of the cylinder. However, another problem faced by traditional deep reinforcement learning for flow control is poor interpretability, resulting in the control law obtained through training being a black-box deep neural network, unable to intuitively establish the connection between the state and the control command, and unable to reflect the physical mechanism behind the perturbation strategy. Summary of the Invention

[0004] To overcome at least one deficiency in the prior art, this application provides an interpretable artificial intelligence control method for high-speed flows based on knowledge distillation, specifically including: In a first aspect, this application provides an interpretable artificial intelligence control method for high-speed flows based on knowledge distillation, the method including: Obtain a pre-trained high-speed flow control model, where the high-speed flow control model is a neural network-based model that can generate a perturbation strategy according to the flow field state information; Use the high-speed flow control model to distill a lightweight control model in the form of a symbolic expression.

[0005] Compared with the prior art, this application has the following beneficial effects: This application provides an interpretable artificial intelligence control method for high-speed flow based on knowledge distillation. Among them, an electronic device obtains a pre-trained high-speed flow control model, where the high-speed flow control model is a neural network-based model that can generate a perturbation strategy according to the flow field state information; a lightweight control model in the form of a symbolic expression is distilled using the high-speed flow control model. In this way, by extracting and transforming the knowledge contained in the above high-speed flow control model into a lightweight control model in the form of a symbolic expression, since the symbolic expression is an explicit mathematical expression form that can directly reflect the functional relationship between the input variables and the output results, it has higher transparency and understandability compared to the implicit neural network model. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0007] Figure 1 One of the flow diagrams of the interpretable artificial intelligence control method for high-speed flow based on knowledge distillation provided by the embodiments of this application; Figure 2 The structural diagram of the test model provided by the embodiments of this application; Figure 3 The schematic diagram of the training principle provided by the embodiments of this application; Figure 4 Another flow diagram of the interpretable artificial intelligence control method for high-speed flow based on knowledge distillation provided by the embodiments of this application; Figure 5 The schematic diagram of the distillation and optimization principle provided by the embodiments of this application; Figure 6 The structural diagram of the interpretable control device for high-speed flow based on knowledge distillation provided by the embodiments of this application; Figure 7 The structural diagram of the electronic device provided by the embodiments of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0008] To make the objectives, technical solutions, and advantages of the embodiments of this application (hereinafter simply referred to as this embodiment) clearer, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. Usually, the components of the embodiments of this application described and shown in the drawings here can be arranged and designed in various different configurations.

[0009] Accordingly, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but merely represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts fall within the scope of protection of the present application.

[0010] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0011] In the description of the present application, it should be noted that the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance. In addition, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element.

[0012] Based on the above statements, as introduced in the background art, although it has been studied that applying deep reinforcement learning techniques to low-speed circular cylinder flow and airfoil flow can automatically find a closed-loop control law that significantly reduces the drag of the cylinder, another problem faced by traditional deep reinforcement learning flow control is poor interpretability, resulting in the control law obtained by training being a black-box deep neural network, unable to intuitively establish the connection between the state and the control command, and unable to reflect the physical mechanism behind the perturbation strategy.

[0013] Based on the discovery of the above technical problems, the following technical solutions are creatively proposed to solve or improve the above problems. It should be noted that the defects existing in the above prior art solutions are all the results obtained through practice and careful research. Therefore, the process of discovering the above problems and the solutions proposed by the embodiments of the present application below for the above problems should be regarded as contributions made to the present application in the process of invention and creation, and should not be understood as technical content well-known to those skilled in the art.

[0014] In view of this, the present embodiment provides an interpretable artificial intelligence control method for high-speed flow based on knowledge distillation as Figure 1 shown, and the method includes: S1. Obtain a pre-trained high-speed flow control model.

[0015] Among them, the high-speed flow control model is a neural network-based model that can generate disturbance strategies according to the flow field state information.

[0016] S2. Use the high-speed flow control model to distill a lightweight control model in the form of a symbolic expression.

[0017] In this way, by extracting and transforming the knowledge contained in the above high-speed flow control model into a lightweight control model in the form of a symbolic expression, since the symbolic expression is an explicit mathematical expression form that can directly reflect the functional relationship between the input variables and the output results, it has higher transparency and understandability compared with the implicit neural network model.

[0018] It should be understood that for the high-speed flow interpretable artificial intelligence control method based on knowledge distillation provided in this embodiment, the electronic device implementing this method can be, but is not limited to, a tablet computer, a laptop computer, a desktop computer, and an embedded control device.

[0019] To make the solution provided in this embodiment clearer, the following takes an embedded control device (hereinafter referred to as the control device) as the electronic device implementing this method to Figure 1 elaborate on each step of the method shown in detail. However, it should be understood that the operations in the flowchart can be implemented out of order, and the steps without logical context relationships can be reversed or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of this application. Continuing to refer to Figure 1 , the method includes: S1. Obtain a pre-trained high-speed flow control model.

[0020] Among them, the high-speed flow control model is a neural network-based model that can generate disturbance strategies according to the flow field state information. Therefore, it can be understood that the core function of the high-speed flow control model is to calculate and output a set of optimal disturbance strategies based on the input flow field state information (for example, physical quantities such as pressure, velocity, and temperature). These disturbance strategies are specifically manifested as control instructions for the flow control actuator, which are used to adjust the flow field characteristics to achieve the expected control objectives, such as drag reduction, noise reduction, or thrust enhancement. In other words, the model generates appropriate control actions by real-time analysis of the flow field state to change the macroscopic flow-around characteristics of the flow field.

[0021] Specifically, when the model receives the flow field state information collected by the sensor, it takes this as the input vector and sends it to the input layer of the neural network. Subsequently, the data undergoes weighted summation and non-linear transformation through the neurons in the hidden layer, and finally a set of perturbation strategies are generated in the output layer. These strategies usually appear as discrete or continuous numerical signals, directly corresponding to the action instructions of the actuator. For example, in the supersonic cavity flow control scenario, the model may calculate whether to activate a certain pulsed arc actuator and its specific drive parameters based on the instantaneous pressure value measured by the dynamic pressure sensor.

[0022] The research finds that the characteristic frequency of supersonic flow usually reaches above 10 kHz, which means that the control system needs to complete a series of operations such as state perception, decision-making calculation, and action execution within each cycle, and the time of the entire control cycle must be less than 100 μs. Such a stringent real-time requirement poses extremely high challenges to both the hardware performance and algorithm efficiency of the controller. However, traditional reinforcement learning methods mostly use fully connected neural networks or convolutional neural networks (CNNs) to represent control strategies. Although these deep neural networks have powerful non-linear modeling capabilities, their inference time is relatively long, with a typical value reaching above 10 ms. This results in the maximum control frequency of the control system based on such networks not exceeding 100 Hz, far from meeting the requirements of supersonic flow control.

[0023] After analysis, it is found that the root cause of the above problem lies in the high computational complexity of traditional deep neural networks. Since fully connected neural networks or CNNs usually contain multiple hidden layers and a large number of parameters, a large number of matrix operations need to be performed during forward propagation calculation. Even when running on high-performance computing devices, these operations still consume a lot of time, thus limiting the response speed of the control system.

[0024] In view of this, this embodiment provides an innovative solution, that is, to perform reinforcement learning through the cooperation of a Field-Programmable Gate Array (FPGA) and a Graphics Processing Unit (GPU) to obtain a pre-trained high-speed flow control model. Based on this concept, the following optional implementation manners of step S1 are provided in this embodiment: S1-1, run the current policy model to be trained through the field programmable gate array FPGA to obtain a sample perturbation strategy.

[0025] In this embodiment, the current to-be-trained policy model is run on an FPGA to process multiple consecutive historical flow field state information, and a sample perturbation policy is obtained. In this process, the FPGA is used to load and run the current version of the to-be-trained policy model. The model receives a series of consecutive historical flow field state information as input data (for example, time series information of physical quantities such as pressure, velocity, or temperature collected by sensors). Through the calculation and analysis of these input data by the to-be-trained policy model, the model outputs a set of specific control instructions, namely the so-called sample perturbation policy. These policies usually manifest as operation instructions for flow control actuators, such as whether to turn on a certain actuator and its specific drive parameter settings. In addition, it should be understood that during the training process, the model parameters will be continuously updated and adjusted. Therefore, in this embodiment, the current to-be-trained model specifically refers to the model version after the most recent parameter update.

[0026] S1-2. Obtain an empirical sample including the sample perturbation policy.

[0027] S1-3. According to the empirical sample, use a graphics processing unit (GPU) to update the current to-be-trained policy model to obtain a new to-be-trained policy model.

[0028] S1-4. If the new to-be-trained policy model does not meet the conditions for being a high-speed flow control model, deploy the new to-be-trained policy model to the FPGA and return to continue executing S1-1 until a high-speed flow control model is obtained.

[0029] Exemplarily, next, in combination with Figure 2 the shown test model and Figure 3 the schematic diagram of the training principle shown, a more intuitive description of the training process of the high-speed flow control model is given.

[0030] Continue to refer to Figure 2 , the test model is of a flat plate configuration and is made of ceramic, and can be used to simulate supersonic cavity flow under the action of a supersonic oncoming flow. The test model includes a flat plate 11, and the leading edge of the flat plate 11 is a downward-tilted wedge shape to ensure that the flow on the upper surface is a boundary layer flow not disturbed by shock waves. A cavity 12 is provided downstream of the flat plate 11, and when the mainstream flows through its upper part, huge pressure pulsations will be formed due to resonance. The goal of flow control is to reduce this pressure pulsation, thereby reducing the wall structure load and aerodynamic noise.

[0031] To sense the flow field, three dynamic pressure sensors 13 are installed at the bottom of the cavity 12, and the measured instantaneous pressure signals are used as the flow field state information. In addition, four pulsed arc exciters 14 are installed upstream of the leading edge of the cavity 12 and are used as controllers, and their control commands are discrete numerical signals (0 represents off, 1 represents on). Therefore, different perturbation strategies can be generated by adopting different control strategies for the four pulsed arc exciters 14.

[0032] In addition, each pulsed arc exciter 14 consists of an anode and a cathode. Among them, the anode and the cathode are the same in structure, and tungsten needles are selected as the electrode materials and inserted into the flat plate from bottom to top. The bottom of the tungsten needle is wrapped with an insulating layer to ensure that arcing does not occur at the bottom. The typical frequency responses of both the pulsed arc exciter 14 and the dynamic pressure sensor 13 can reach above 50 kHz, far exceeding the resonance frequency of the cavity 12 itself (the typical value is 1 - 10 kHz). Therefore, Figure 2 the shown test model can meet the requirements of closed-loop control.

[0033] Based on Figure 2 the shown test model, continue to refer to Figure 3 , the entire training process of the high-speed flow control model includes two closed-loop cycles. One is a high-speed real-time control cycle, and the other is a low-speed network training cycle. Among them, the real-time control cycle runs in the FPGA and completes the interaction with the high-speed flow field environment through high-frequency response exciters and high-frequency response sensors. The network training cycle runs in the GPU and is responsible for updating the parameters of the policy model to be trained. Considering that the GPU and the FPGA in the control device cannot communicate directly, therefore, in this example, the central processing unit (CPU) is used to transfer and store data between the two. Since the policy model to be trained is trained by means of reinforcement learning, therefore, in this example, the policy model to be trained is called the RBF model. The RBF model can only contain one hidden layer, so its parameter scale is smaller, and its running and training speeds are far superior to those of other fully connected deep neural networks.

[0034] In the real-time control cycle, the control device senses the real-time state of the high-speed flow through the high-frequency response sensor, and the generated analog voltage signal is used as the flow field state information and is temporarily stored in a first-in-first-out queue named I-FIFO after analog-to-digital conversion (ADC). The control device extracts multiple consecutive historical flow field state information from the I-FIFO and inputs it into the RBF model, and calculates the values corresponding to different perturbation strategies according to the corresponding network coefficients in the RBF model , and the perturbation strategy corresponding to the maximum value is used as the sample perturbation strategy at this time .

[0035] The control device uses a digital-to-analog converter (DAC) to convert the sample disturbance strategy into a trigger signal for the actuator driving power supply, and applies it to the actuator to complete the action adjustment in one control cycle. At the same time, the FPGA uses "current gas field state, current disturbance strategy, and gas field state after disturbance" as a control experience, expressed as Write to a first-in, first-out queue called S-FIFO.

[0036] It should be understood that in the traditional CPU control framework, data acquisition devices are usually independent peripheral devices that interact with the controller through high-speed cache and batch transmission mechanisms. After testing this mode, it was found that there is a control delay of several milliseconds (ms) in the entire control link. In the framework proposed in this example, the ADC and DAC modules are directly connected to the FPGA chip, and only a delay of two clock cycles is required to realize the reading of signals and the output of control instructions.

[0037] When the queue accumulates to a certain extent, the CPU of the control device obtains the accumulated control experience from the FPGA through the communication port. The protocol adopted by the communication port can be a serial port protocol (232, 485), a network communication protocol (TCP / IP, UDP), or a high-speed PCIE protocol, as long as it can meet the data transmission bandwidth requirements of the corresponding control task. In this example, UDP is used as the communication protocol based on the compromise between the communication distance and the transmission rate.

[0038] After obtaining the accumulated control experience, the control device will calculate the immediate reward corresponding to the disturbance strategy in each control experience according to the corresponding reward setting, which is expressed as , will control the experience and instant rewards Put together, it constitutes an experience sample In this example, the instant reward Set to the negative of the average value of the instantaneous pressure pulsation of multiple pressure sensors. The control device adds these experience samples to the experience library specially opened in the memory on the one hand, and writes them to the local disk on the other hand, recording these experience samples in the form of files.

[0039] The GPU and CPU in the control device communicate through PCIE to share the experience samples to the GPU. The GPU then uses these experience samples to calculate the model loss of the current policy model to be trained, and then updates its model parameters using the backpropagation algorithm. Taking the classic algorithm Deep Q-Network (DQN) in reinforcement learning as an example, the loss function is equal to the root mean square of the expected value calculated from the target Q-network and the current reward and the actual value predicted by the current Q-network. After 5-10 rounds of training, the GPU sends the new policy model to be trained to the CPU of the control device. The CPU of the control device further downloads the new policy model to be trained to the FPGA to achieve online update of the closed-loop control law. In addition, the control device can also write the new policy model to be trained to the local disk to facilitate users to trace the historical process of control.

[0040] For the training framework of the deep neural network in this example, this example does not make specific limitations, and it can be TensorFlow, PyTorch, etc. Similarly, for the reinforcement learning algorithm in this embodiment, no specific limitations are made either. Those skilled in the art can choose a value method (Q-learning, DQN, etc.) suitable for discrete control tasks according to the requirements of the control task, or can also choose a policy method (DDPG, PPO, etc.) suitable for continuous tasks.

[0041] In this way, compared with the traditional deep reinforcement learning training method of "CPU + data acquisition card", both the inference and training processes of the neural network are completed by the CPU. The training method of "CPU + FPGA + GPU" provided in this embodiment has the training process responsible for the GPU, and the disturbance control is completed by the FPGA suitable for parallel high-speed computing, greatly improving the operation speed and training speed.

[0042] Based on the description of the training method of the high-speed flow control model in the above embodiment, the following continues to Figure 1 describe step S2 in S2. Use the high-speed flow control model to distill a lightweight control model in the form of a symbolic expression.

[0043] It is found that although the number of model parameters has been limited when designing the high-speed flow control model, it still belongs to a gray-box model with poor interpretability. In order to reveal the physical mechanism behind the disturbance strategy, this embodiment uses the method of "knowledge distillation" to transform the gray-box neural network control law into an explicit symbolic control law. Therefore, based on this inventive concept, the following implementation manners of step S2 are provided in this embodiment: S2-1. Generate multiple symbolic expressions.

[0044] In this implementation, during the stage of generating multiple symbolic expressions, a series of possible explicit symbolic expressions are created as candidate models through specific algorithms or heuristic search methods. These symbolic expressions can be represented in the form of mathematical functions, with the input being the flow field state information (e.g., the instantaneous signal of a dynamic pressure sensor) and the output being the actuator control instruction (e.g., whether to turn on the plasma actuator). In a specific implementation, the structure and parameter range of these symbolic expressions can be preset according to the requirements of the actual application scenario to ensure that the generated expressions can cover possible control strategies.

[0045] S2-2. Evaluate multiple preferred expressions from the current multiple symbolic expressions using the tags generated by the high-speed flow control model.

[0046] It should be understood that the lightweight control model needs to be iteratively optimized multiple times to be finally obtained. This process can be regarded as a dynamic evolution process of gradually approaching the target. In each iteration, the system generates a new set of symbolic expressions based on the results of the previous round. Therefore, in this embodiment, the current multiple symbolic expressions refer to those generated in the latest round of iteration.

[0047] As Figure 5 shown, in the specific implementation process, for the above-generated symbolic expressions, the high-speed flow control model is used as the teacher model to evaluate the performance of each symbolic expression. Among them, the evaluation tags are generated by the teacher model based on the knowledge it has learned, and are used to supervise the learning process of the student model. It should be understood that the tags in this embodiment are divided into soft and hard, and the degree of "soft" and "hard" is determined by the distillation temperature parameter. A higher temperature parameter can make the target probability distribution smoother, thus preventing the student model from falling into a local optimal solution.

[0048] Therefore, during the evaluation process, the teacher model will score according to the control effects (such as the degree of reduction in pressure pulsation) produced by each symbolic expression in the simulation or experimental environment, and select multiple symbolic expressions with better performance as the preferred expressions.

[0049] S2-3. If there is no lightweight control model that meets the performance requirements among the multiple preferred expressions, generate multiple new symbolic expressions based on the multiple preferred expressions, and return to step S2-2 for execution until a lightweight control model that meets the performance requirements is obtained.

[0050] In this embodiment, the control device can perform crossover and mutation processing on multiple preferred expressions to obtain multiple new symbolic expressions. It can be understood that the core of the iterative optimization process lies in performing crossover and mutation processing on the preferred expressions through genetic programming to generate new symbolic expressions. Among them, genetic programming is actually a heuristic search algorithm based on the theory of biological evolution. By randomly combining or finely tuning the structure or parameters of the preferred expressions, new candidate expressions are generated. The newly generated symbolic expressions are again evaluated by the teacher model's labels to screen out a more optimal set of expressions. This process is repeated continuously until a lightweight control model that can meet the predetermined performance requirements is found.

[0051] Therefore, as a teacher model, the high-speed flow control model not only provides initial control strategy knowledge but also guides the learning direction of the student model (i.e., the symbolic expression) through the label generation mechanism. Since the structure of the symbolic expression is fixed and has a clear physical meaning, the finally obtained lightweight control model has stronger interpretability and can intuitively reflect the relationship between the flow field state and the control instructions.

[0052] In the practical process, it is found that although heuristic search methods such as genetic programming can optimize the structure and parameters of symbolic expressions to a certain extent, its exploration space is still much smaller than the weight adjustment range of the neural network model. Therefore, in the face of highly nonlinear or high-frequency dynamically changing flow field states, the symbolic expression may not be able to fully reproduce the fine control effect of the neural network model. In addition, the label data is generated by the teacher model based on the knowledge it has learned, but when passed to the student model, the amount of information has been compressed and simplified to a certain extent. Therefore, although the soft label can provide more probability distribution information than the hard label, it still cannot fully reflect the complex interaction relationships between the neurons inside the teacher model. In addition, the choice of the distillation temperature parameter will also affect the smoothness of the label, and too high or too low temperature settings may lead to the loss of some control details.

[0053] The above reasons lead to the inevitable loss of some complex or subtle control rules in the process of transferring the knowledge of the teacher model to the student model. To address this problem, as Figure 4 shown, the high-speed flow interpretable artificial intelligence control method based on knowledge distillation provided in this embodiment further includes: S3, optimizing the lightweight control model in the actual scenario to obtain an optimized lightweight control model.

[0054] In the specific implementation process, the control device optimizes the coefficients to be optimized through a reinforcement learning algorithm according to the perturbation strategy of the lightweight control model in the actual scenario to obtain an optimized lightweight control model.

[0055] Continuing to refer to 5, this embodiment implants a lightweight control model in the form of symbolic expression into the FPGA, and the optimization process is carried out based on the perturbation strategy of the lightweight control model in the actual scene. In the specific implementation process, the perturbation strategy may refer to the actuator control instructions generated according to the current flow field state information, for example, whether to turn on the plasma actuator and the specific excitation intensity or frequency. These control instructions directly act on the flow system in the experimental environment, and feedback the control effect through sensing devices such as dynamic pressure sensors. By accumulating a large amount of control experience data, an experience library containing the current state, action, next state and immediate reward can be constructed. These experience data provide necessary supervision information for the subsequent optimization process. When the new experience data accumulates to a certain extent, the parameters in the symbolic expression can be fine-tuned by the error back propagation method. This process is similar to the training mechanism of a neural network, but because the lightweight control model is a differentiable expression, it is aimed at the coefficients to be optimized in the symbolic expression.

[0056] In addition, considering that the output of the lightweight control model in symbolic expression is a continuous value, in order to realize discrete control instructions, the output can be first converted into a range of 0-1 through a Sigmoid activation function, and then the nearest integer method is used to obtain a control instruction output of 0 or 1 for driving the actuator.

[0057] Since the optimization process is carried out in a real experimental environment, it can fully consider the complex characteristics and dynamic change laws of the actual flow system. Therefore, through the iterative optimization of the reinforcement learning algorithm, the parameters in the symbolic expression gradually approach the optimal value, thereby achieving further improvement of the initial control law.

[0058] Based on the same inventive concept as the high-speed flow interpretable artificial intelligence control method based on knowledge distillation provided in this embodiment, this embodiment also provides a high-speed flow interpretable control device based on knowledge distillation, which includes at least one software function module that can be stored in a memory or solidified in an electronic device in the form of software. The processor in the electronic device is used to execute the executable module stored in the memory. For example, the software function module and computer program included in the device. Please refer to Figure 6 , functionally speaking, the device may include: A model generation module 21 is used to obtain a pre-trained high-speed flow control model, wherein the high-speed flow control model is a neural network-based model capable of generating a disturbance strategy according to flow field state information; The model distillation module 22 is used to distill a lightweight control model in the form of a symbolic expression using a high-speed flow control model.

[0059] In this embodiment, the model generation module 21 is used to implement Figure 1In step S1, the model distillation module 22 is used to implement Figure 1 In step S2. Therefore, for the specific implementation manners of the above modules, reference may be made to the specific implementation manners in the corresponding steps.

[0060] Since it has the same inventive concept as the high-speed flow interpretable artificial intelligence control method provided in this embodiment, the device also implements other steps or sub-steps of the method through the above modules.

[0061] Optionally, the model distillation module 22 is further used to: Optimize the lightweight control model in the actual scenario to obtain an optimized lightweight control model.

[0062] Optionally, the model distillation module 22 is specifically used to: According to the perturbation strategy of the lightweight control model in the actual scenario, optimize the coefficient to be optimized through a reinforcement learning algorithm to obtain an optimized lightweight control model.

[0063] Optionally, the model distillation module 22 is specifically used to: Generate a plurality of symbolic expressions; Evaluate a plurality of preferred expressions from the current plurality of symbolic expressions by using the labels generated by the high-speed flow control model; If there is no lightweight control model that meets the performance requirements among the plurality of preferred expressions, generate a plurality of new symbolic expressions according to the plurality of preferred expressions, and return to evaluating a plurality of preferred expressions from the current plurality of symbolic expressions by using the labels generated by the high-speed flow control model until a lightweight control model that meets the performance requirements is obtained.

[0064] Optionally, the model distillation module 22 is specifically used to: Perform crossover mutation processing on the plurality of preferred expressions to obtain a plurality of new symbolic expressions.

[0065] Optionally, the model generation module 21 is further used to: Run the current to-be-trained policy model through a field programmable gate array (FPGA) to obtain a sample perturbation strategy; Obtain an experience sample including the sample perturbation strategy; Update the current to-be-trained policy model by using a graphics processing unit (GPU) according to the experience sample to obtain a new to-be-trained policy model; If the new to-be-trained policy model does not meet the conditions for being a high-speed flow control model, deploy the new to-be-trained policy model to the FPGA, and return to running the current to-be-trained policy model through the field programmable gate array (FPGA) to obtain a sample perturbation strategy until a high-speed flow control model is obtained.

[0066] Optionally, the model generation module 21 is further used for: The current strategy model to be trained is run through FPGA to process multiple continuous historical flow field state information to obtain a sample disturbance strategy.

[0067] In addition, the functional modules in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.

[0068] It should also be understood that if the above implementation is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application.

[0069] Therefore, this embodiment also provides a storage medium, which is a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by the processor, the high-speed flow interpretable artificial intelligence control method based on knowledge distillation provided in this embodiment is implemented. Among them, the storage medium can be a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a disk or an optical disk, etc., which can store program codes.

[0070] This embodiment provides an electronic device that implements a high-speed flow interpretable artificial intelligence control method based on knowledge distillation. Figure 7 As shown, the electronic device includes a processor 32 and a memory 31. In addition, the memory 31 stores a computer program, and the processor implements the high-speed flow interpretable artificial intelligence control method based on knowledge distillation provided in this embodiment by reading and executing the computer program corresponding to the above implementation mode in the memory 31.

[0071] Continue to see Figure 7 The electronic device further includes a communication unit 33. The memory 31, the processor 32 and the communication unit 33 are electrically connected to each other directly or indirectly through a system bus 34 to achieve data transmission or interaction.

[0072] Among them, the memory 31 can be an information recording device based on any electronic, magnetic, optical or other physical principles, and is used to record execution instructions, data, etc. In some embodiments, the memory 31 can be, but is not limited to, a volatile memory, a non-volatile memory, a storage drive, etc.

[0073] In some embodiments, the volatile memory can be a Random Access Memory (RAM); in some embodiments, the non-volatile memory can be a Read Only Memory (ROM), a Programmable Read-Only Memory (PROM), an Erasable Programmable Read-Only Memory (EPROM), an Electric Erasable Programmable Read-Only Memory (EEPROM), a flash memory, etc.; in some embodiments, the storage drive can be a disk drive, a solid state drive, any type of storage disk (such as an optical disk, a DVD, etc.), or a similar storage medium, or a combination thereof, etc.

[0074] The communication unit 33 is used to transmit and receive data through a network. In some embodiments, the network can include a wired network, a wireless network, an optical fiber network, a telecommunication network, an intranet, the Internet, a Local Area Network (LAN), a Wide Area Network (WAN), a Wireless Local Area Networks (WLAN), a Metropolitan Area Network (MAN), a Wide Area Network (WAN), a Public Switched Telephone Network (PSTN), a Bluetooth network, a ZigBee network, or a Near Field Communication (NFC) network, etc., or any combination thereof. In some embodiments, the network can include one or more network access points. For example, the network can include a wired or wireless network access point, such as a base station and / or a network switching node, and one or more components of the service request processing system can be connected to the network through the access point to exchange data and / or information.

[0075] The processor 32 may be an integrated circuit chip with signal processing capabilities, and the processor may include one or more processing cores (e.g., a single-core processor or a multi-core processor). By way of example only, the above-mentioned processor may include a Central Processing Unit (CPU), an Application Specific Integrated Circuit (ASIC), an Application Specific Instruction-set Processor (ASIP), a Graphics Processing Unit (GPU), a Physics Processing Unit (PPU), a Digital Signal Processor (DSP), a Field Programmable Gate Array (FPGA), a Programmable Logic Device (PLD), a controller, a microcontroller unit, a Reduced Instruction Set Computing (RISC), or a microprocessor, etc., or any combination thereof.

[0076] It can be understood that Figure 7 The structure shown is only schematic. The electronic device may also have more or fewer components than Figure 7 shown, or have a configuration different from that Figure 7 shown. Figure 7 Each of the components shown may be implemented in hardware, software, or a combination thereof.

[0077] It should be understood that the devices and methods disclosed in the above embodiments can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0078] As described above, these are only various embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A high-speed flow interpretable artificial intelligence control method based on knowledge distillation, characterized in that, The method includes: Obtain a pre-trained high-speed flow control model, where the high-speed flow control model is a neural network-based model that can generate a perturbation strategy according to the flow field state information; Use the high-speed flow control model to distill a lightweight control model in the form of a symbolic expression.

2. The method for high-speed flow interpretable artificial intelligence control based on knowledge distillation according to claim 1, wherein The method further includes: Optimize the lightweight control model in an actual scenario to obtain an optimized lightweight control model.

3. The method for high-speed flow interpretable artificial intelligence control based on knowledge distillation according to claim 2, wherein The lightweight control model includes coefficients to be optimized. Optimizing the lightweight control model in an actual scenario to obtain an optimized lightweight control model includes: According to the perturbation strategy of the lightweight control model in the actual scenario, optimize the coefficients to be optimized through a reinforcement learning algorithm to obtain an optimized lightweight control model.

4. The method for high-speed flow interpretable artificial intelligence control based on knowledge distillation according to claim 1, wherein Using the high-speed flow control model to distill a lightweight control model in the form of a symbolic expression includes: Generate a plurality of symbolic expressions; Evaluate a plurality of preferred expressions from the current plurality of symbolic expressions using the labels generated by the high-speed flow control model; If there is no lightweight control model that meets the performance requirements among the plurality of preferred expressions, generate a plurality of new symbolic expressions according to the plurality of preferred expressions, and return to evaluating a plurality of preferred expressions from the current plurality of symbolic expressions using the labels generated by the high-speed flow control model until a lightweight control model that meets the performance requirements is obtained.

5. The method for high-speed flow interpretable artificial intelligence control based on knowledge distillation according to claim 4, characterized in that Generating a plurality of new symbolic expressions according to the plurality of preferred expressions includes: Perform crossover and mutation processing on the plurality of preferred expressions to obtain the plurality of new symbolic expressions.

6. The method for high-speed flow interpretable artificial intelligence control based on knowledge distillation according to claim 1, wherein Obtain a pre-trained high-speed flow control model, including: Run the current policy model to be trained through a Field-Programmable Gate Array (FPGA) to obtain a sample perturbation strategy; Obtain an experience sample including the sample perturbation strategy; According to the experience sample, use a Graphics Processing Unit (GPU) to update the current policy model to be trained to obtain a new policy model to be trained; If the new policy model to be trained does not meet the conditions for being the high-speed flow control model, deploy the new policy model to the FPGA, and return to running the current policy model to be trained through the Field-Programmable Gate Array (FPGA) to obtain a sample perturbation strategy until the high-speed flow control model is obtained.

7. The method for high-speed flow interpretable artificial intelligence control based on knowledge distillation according to claim 6, characterized in that Running the current policy model to be trained through a Field-Programmable Gate Array (FPGA) to obtain a sample perturbation strategy includes: Run the current policy model to be trained through the FPGA to process a plurality of consecutive historical flow field state information to obtain the sample perturbation strategy.

Citation Information

Patent Citations

  • Data intelligent model training and hardware acceleration method and system based on artificial intelligence

    CN118657186A

  • Model lightweight application deployment method based on knowledge distillation and related device

    CN119067161A

  • Genetic programming and reinforcement learning fused interpretable intelligent flow control method

    CN119439745A

  • Progressive knowledge distillation lightweight method, device and equipment based on pruning network

    CN119538977A

  • KR20250070784A