A method and device for optimizing the cooling design of turbine blades based on reinforcement learning

By employing a reinforcement learning-based method for optimizing turbine blade cooling efficiency, and utilizing the reinforcement learning environment and policy network training of the agent, the problem of long design cycles for turbine blade cooling structures is solved, achieving rapid and efficient optimization of cooling structures.

CN117725811BActive Publication Date: 2026-07-21BEIHANG UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIHANG UNIV
Filing Date
2023-11-02
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing turbine blade cooling structure designs suffer from long design cycles and low efficiency, especially when optimizing complex structures, making it difficult to quickly find the optimal solution.

Method used

A reinforcement learning-based approach is adopted. By creating a reinforcement learning environment for the agent, defining action, state, and reward data, the policy network is trained to minimize the average temperature of the turbine blade surface, and automatic optimization is performed in conjunction with turbine blade cooling effect design tools.

Benefits of technology

This greatly improves the efficiency of turbine blade cooling design, enabling the finding of optimal cooling structure parameters in a short time, reducing the design cycle, and improving design quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117725811B_ABST
    Figure CN117725811B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data processing, in particular to a turbo blade cold efficiency design optimization method and device based on reinforcement learning. An intelligent agent reinforcement learning environment is created based on a turbo blade cold efficiency design tool. The reinforcement learning environment defines the action of the intelligent agent, and state and reward data obtained from the turbo blade cold efficiency design tool. The state includes the average temperature of the surface of a turbo blade. The intelligent agent inputs the action to the turbo blade cold efficiency design tool, and obtains the state and the reward data generated by the turbo blade cold efficiency design tool in response to the action. The policy network of the intelligent agent is trained based on the state and the reward data, so as to minimize the average temperature of the surface of the turbo blade. Through the reinforcement learning training of the intelligent agent, the turbo blade cold efficiency design parameters are quickly optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method and apparatus for optimizing the cooling effect design of turbine blades based on reinforcement learning. Background Technology

[0002] Turbine blades operate at speeds of thousands or even tens of thousands of revolutions per minute, driven by the high-temperature, high-pressure gas flow from the combustion chamber. Therefore, fluid simulation is necessary to obtain the blade surface temperature. Fluid simulation typically follows a fixed process: first, a 3D geometric model is established; then, the solid model is meshed; next, CFD simulation software is used for solving the problem; and finally, the simulation results are post-processed. This process is also largely used for turbine blade cooling structure design. After each software calculation, intermediate data is manually processed before being transferred to the next software step. This design handover and communication process significantly increases the design cycle and hinders design preservation. Although multi-parameter, multi-disciplinary optimization tools exist that can build simulation optimization workflows, engineers from different departments often rely on experience and manual "trial calculation-evaluation-correction" methods to optimize engine design. Furthermore, engine performance evaluation involves hundreds of complex simulation programs with extremely stringent optimization constraints and objectives; even after months of optimization, an ideal solution may not be obtained.

[0003] Current research on turbine blade cooling structure design mainly focuses on methods for solving combinatorial optimization problems. Optimization problems refer to finding the optimal solution among numerous parameter values ​​under certain conditions to achieve the best performance for one or more functional indicators. Currently, heuristic algorithms such as genetic algorithms and particle swarm optimization, or neural networks with strong representational capabilities, are commonly used to seek optimal and near-feasible solutions in the design domain. However, these algorithms mostly only apply to relatively simple geometric cooling structures or local parts of the cooling structure, lacking research on complex cooling structures and exhibiting low efficiency. Summary of the Invention

[0004] To overcome the shortcomings of the prior art, this application provides a method and apparatus for optimizing the cooling effect design of turbine blades based on reinforcement learning, which can greatly improve the efficiency of optimizing the cooling effect design of turbine blades.

[0005] In a first aspect, embodiments of this application provide a turbine blade cooling effect design optimization method based on reinforcement learning, the method comprising the following steps:

[0006] A reinforcement learning environment for an agent is created based on a turbine blade cooling effect design tool; wherein, the reinforcement learning environment defines the agent's actions, and state and reward data obtained from the turbine blade cooling effect design tool, the state including the average temperature of the turbine blade surface;

[0007] The agent inputs the action into the turbine blade cooling effect design tool and obtains the state and reward data generated by the turbine blade cooling effect design tool in response to the action.

[0008] The agent's policy network is trained using reinforcement learning based on the state and reward data to minimize the average temperature on the turbine blade surface.

[0009] In some embodiments, the change in the turbine blade cooling efficiency design parameters is defined as the action of the agent, and action parameter values ​​at different scales are set to represent the magnitude of the change.

[0010] In some embodiments, the turbine blade cooling efficiency design parameters are defined by the following steps:

[0011] The turbine blade cooling effect design tool is used to construct all design parameters for the turbine blade structure. First design parameters that have no effect on the average surface temperature of the turbine blade are then removed from all design parameters, resulting in second design parameters. The first design parameters include the wall structure and the inner baffle structure. The second design parameters include the turbulence column, the layered baffle structure, and multiple partition structures.

[0012] The layer partition structure is fixedly installed; wherein, the number of layer partition structures directly affects the number of partition structures;

[0013] The turbulence column and the multiple partition structures are defined as the cooling efficiency design parameters of the turbine blade; wherein, the partition structure includes film vent diameter, film vent inclination angle, impact vent diameter, vent row distance and vent column distance.

[0014] In some embodiments, the diameter of the turbulence column ranges from 0.5 to 1.5 mm, the diameter of the air film pores ranges from 0.4 to 0.8 mm, the inclination angle of the air film pores ranges from 30° to 90°, the diameter of the impact pores ranges from 0.9 to 1.5 mm, the spacing between the pore rows ranges from 3 to 5 mm, and the spacing between the pore columns ranges from 3 to 7 mm.

[0015] In some embodiments, reward data is defined by setting at least one reward function, and the reward function is set in the following manner, including the following steps:

[0016] A first reward function is set based on the average and variance values ​​of the temperature on the turbine blade surface; wherein, the lower the average and variance values ​​of the temperature on the turbine blade surface, the higher the reward data obtained.

[0017] A second reward function is set based on the average temperature of the turbine blade surface; wherein, the lower the average temperature of the turbine blade surface, the higher the reward data is obtained.

[0018] In some embodiments, training the agent's policy network using reinforcement learning based on the state and the reward data to minimize the average temperature of the turbine blade surface includes the following steps:

[0019] Based on the state and reward data, and using different reinforcement learning algorithms to train the policy network of the agent with different network structures, the average temperature and tuning time of the turbine blade surface are minimized accordingly.

[0020] Based on the obtained minimum average temperature of the turbine blade surface and the tuning time, the target network structure and target reinforcement learning algorithm of the policy network are determined.

[0021] In some embodiments, the SAC / PPO reinforcement learning algorithm is used to train the agent's policy network for reinforcement learning; the network structure of the policy network consists of fully connected layers or a combination of fully connected layers and convolutional layers.

[0022] Secondly, embodiments of this application provide a turbine blade cooling efficiency design optimization device based on reinforcement learning, the device comprising:

[0023] A module is created to create a reinforcement learning environment for an agent based on a turbine blade cooling effect design tool; wherein the reinforcement learning environment defines the agent's actions and state and reward data obtained from the turbine blade cooling effect design tool, the state including the average temperature of the turbine blade surface;

[0024] The acquisition module is used to input the action into the turbine blade cooling effect design tool using the intelligent agent, and to acquire the state and reward data generated by the turbine blade cooling effect design tool in response to the action;

[0025] A reinforcement learning module is used to train the agent's policy network based on the state and the reward data to minimize the average temperature on the turbine blade surface.

[0026] Thirdly, an electronic device provided in this application includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, the steps of the reinforcement learning-based turbine blade cooling efficiency design optimization method described in any of the first aspects are executed.

[0027] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the reinforcement learning-based turbine blade cooling efficiency design optimization method described in any of the first aspects.

[0028] This application discloses a reinforcement learning-based method and apparatus for optimizing turbine blade cooling design. It creates a reinforcement learning environment for an agent based on a turbine blade cooling design tool. The reinforcement learning environment defines the agent's actions and state and reward data obtained from the turbine blade cooling design tool, where the state includes the average temperature of the turbine blade surface. The agent inputs the actions into the turbine blade cooling design tool and obtains the state and reward data generated by the tool in response to the actions. The agent's policy network is trained using reinforcement learning based on the state and reward data to minimize the average temperature of the turbine blade surface. Thus, through reinforcement learning training of the agent, rapid optimization of turbine blade cooling design parameters is achieved. Attached Figure Description

[0029] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 This document shows a flowchart of an embodiment of the reinforcement learning-based turbine blade cooling efficiency design optimization method described in this application.

[0031] Figure 2 A schematic diagram of the turbine blade described in an embodiment of this application is shown;

[0032] Figure 3 This diagram illustrates the interaction between the turbine blade cooling efficiency design tool described in an embodiment of this application and the intelligent agent.

[0033] Figure 4This illustration shows a schematic diagram of the change curve of reward data during the training of the agent using the SAC / PPO reinforcement learning algorithm in an embodiment of this application.

[0034] Figure 5 This illustration shows a schematic diagram of the average temperature change curve of the turbine blade surface during the training of the agent using the SAC / PPO reinforcement learning algorithm in an embodiment of this application.

[0035] Figure 6 This paper illustrates a schematic diagram of the change curves of reward data during the training of agents using different network structures according to embodiments of this application.

[0036] Figure 7 The diagram illustrates the change curve of reward data during the training of an agent using genetic algorithm, simulated annealing algorithm, particle swarm optimization algorithm, differential evolution algorithm and SAC in an embodiment of this application.

[0037] Figure 8 This diagram illustrates the structure of the reinforcement learning-based turbine blade cooling efficiency design optimization device described in an embodiment of this application.

[0038] Figure 9 A structural block diagram of the electronic device described in an embodiment of this application is shown. Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.

[0040] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0041] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.

[0042] In view of the technical problems raised in the background art, this application provides a method, device, electronic device and storage medium for turbine blade cooling effect design optimization based on reinforcement learning, which can greatly improve the efficiency of turbine blade cooling effect design optimization.

[0043] See the instruction manual appendix Figure 1 The application provides a reinforcement learning-based method for optimizing the cooling effect design of turbine blades, comprising the following steps:

[0044] S1. Create a reinforcement learning environment for the agent based on the turbine blade cooling effect design tool; wherein, the reinforcement learning environment defines the agent's actions and state and reward data obtained from the turbine blade cooling effect design tool, the state including the average temperature of the turbine blade surface;

[0045] S2. Using the intelligent agent, input the action into the turbine blade cooling effect design tool, and obtain the state and reward data generated by the turbine blade cooling effect design tool in response to the action;

[0046] S3. Based on the state and the reward data, perform reinforcement learning training on the agent's policy network to minimize the average temperature on the turbine blade surface.

[0047] To clearly understand the technical solutions of the embodiments of the present invention, please refer to the appendix to the specification. Figure 2First, let's illustrate the structure of a turbine blade with an example. The turbine blade profile refers to the cross-section of a blade with a specific aerodynamic shape. The entire blade profile is divided into four parts: the blade base, the blade back, the trailing edge, and the leading edge. The blade back is the convex surface of the blade body, where the pressure is relatively low. The blade base is the concave surface of the blade body, where the pressure is relatively high. The leading edge guides the airflow to the blade surface and is the part connecting the blade base and blade back at the air intake edge. The trailing edge guides the airflow away from the blade surface and is the part connecting the blade base and blade back at the air exhaust edge. Furthermore, for ease of research, automated modeling, and production, the blade structure needs to be parametrically modeled. The trailing edge is used as the starting position (0), and the blade returns to the trailing edge after passing the blade back, leading edge, and blade base, serving as the ending position (1). This definition allows for a convenient description of the positions of the inner and outer partitions, which are all between 0 and 1. The two ends of the inner partition are located at different positions, and the inner partition has three attributes: starting position, ending position, and partition thickness. The spacers are located at the same position on both the inner and outer walls; therefore, they only have two attributes: position and thickness. The blade wall structure includes three attributes: inner wall thickness, outer wall thickness, and distance between the inner and outer walls. Two spacers and the area between the trailing edge and the spacers constitute a partition, and the partition's extent can be described by its position. For each partition structure, the inclination angle of the film cooling holes, the diameter of the film cooling holes, the diameter of the impingement holes, the hole row spacing, and the hole column spacing within its region need to be defined. The turbulence columns within the blade are generally cylindrical structures; therefore, their diameter attribute needs to be defined.

[0048] In step S1, the turbine blade cooling effect design tool can be a turbine blade cooling effect design software, which includes functions such as automatic parametric modeling, automatic simulation calculation, mesh generation, intelligent recommendation, and intelligent evaluation. In this application, the turbine blade cooling effect design tool is combined with a reinforcement learning algorithm to achieve automatic optimization of the turbine blade cooling effect design.

[0049] The reinforcement learning mainly consists of an agent, an environment, a state, an action, and a reward. After the agent performs an action, the environment will transition to a new state. The environment will provide a reward signal (positive or negative reward) for the new state. Subsequently, the agent will execute a new action according to a certain strategy based on the new state and the reward feedback from the environment.

[0050] In this application, the interaction method between the turbine blade cooling effect design tool and the intelligent agent in steps S2-S3 can be found in the appendix of the specification. Figure 3That is, the turbine blade cold effect design tool is used as the environment for reinforcement learning of the agent, so that the agent can obtain all or part of the data in the turbine blade cold effect design tool as its state, including operating conditions, structural design parameters, model files, evaluation results, and intermediate results and files. Since the purpose of this application is to find a set of design parameters that can minimize the surface temperature of the turbine blade, the evaluation results of the turbine blade cold effect design tool can be used as the standard for judging the design effect, and the agent can be given certain rewards or penalties. The turbine blade structural design needs to meet the operating conditions under the maximum operating conditions, i.e., the climbing conditions, which are generally fixed values. The intermediate result files and 3D models are not directly related to the evaluation results, so the agent's actions only need to consider the setting of the structural design parameters. When the agent performs an action... i When applied to the turbine blade cooling effect design tool, the state of the turbine blade cooling effect design tool also changes from state to state. i Transform into state i+1 At the same time, a corresponding reward is given. i The agent learns and updates its policy network through continuous interaction with the turbine blade cooling effect design tool.

[0051] Specifically, the change (increase or decrease) in the turbine blade cooling efficiency design parameters is defined as the action of the intelligent agent, and action parameter values ​​of different scales are set to represent the magnitude of the change. As can be seen from the turbine blade structure and its parameterized form, a complete turbine blade cooling structure can be designed by providing the parameters of the wall structure, inner baffle structure, layer baffle structure, and each partition structure. The wall structure and inner baffle structure mainly serve a mechanical support function and have little impact on the blade surface temperature; therefore, the action space is only designed for the layer baffle structure and each partition structure parameter. In actual design, the number of layer baffles is variable, and the number of layer baffles directly affects the number of partitions. However, considering that the input and output dimensions of the reinforcement learning algorithm network cannot be changed during operation once designed, it is necessary to set the number of layer baffles to a fixed value. Therefore, in this application, the turbulence column and multiple partition structures are defined as the turbine blade cooling efficiency design parameters.

[0052] In one embodiment, the spacers are fixed at [0.27, 0.4, 0.64], resulting in a 21-dimensional action space. This includes the diameter of the baffle columns, the positions of the three spacers, and the diameter, inclination angle, impact orifice diameter, row spacing, and column spacing of the four zones' film membrane pores. Generally, the diameter of the baffle columns ranges from 0.5 to 1.5 mm, the diameter of the film membrane pores ranges from 0.4 to 0.8 mm, the inclination angle of the film membrane pores ranges from 30° to 90°, the impact orifice diameter ranges from 0.9 to 1.5 mm, the row spacing of the pores ranges from 3 to 5 mm, and the column spacing of the pores ranges from 3 to 7 mm. Furthermore, it is reasonable that the accuracy of all parameters, except for the inclination angle which is 1, is 0.01.

[0053] It should be noted that the ultimate goal of optimization is to obtain a set of turbine blade cooling efficiency design parameters that minimize the average surface temperature of the blades. Considering that users will observe the magnitude of the turbine blade cooling efficiency design parameters to determine the adjustment range when manually optimizing, in this application, the state includes not only the average surface temperature of the turbine blades but also the turbine blade cooling efficiency design parameters.

[0054] The reward data is defined by setting at least one reward function. In one embodiment, three forms of reward functions are set, as shown in the following formulas:

[0055]

[0056]

[0057]

[0058] T is the surface temperature of the blade, a = 0.1. Formula (1) can be used as the main reward function. The lower the average temperature and variance of the turbine blade surface, the more reward should be given to the agent. Formula (2) can be used as the additional gradient reward function. The lower the average temperature of the turbine blade surface, the greater the magnitude of the additional gradient reward. Formula (3) is obtained by combining formula (1) and formula (2).

[0059] Furthermore, in this application, based on the state and the reward data, different reinforcement learning algorithms are used to train the policy networks of the agents with different network structures to obtain the corresponding minimized average temperature and tuning time of the turbine blade surface; then, based on the obtained minimized average temperature and tuning time of the turbine blade surface, the target network structure and target reinforcement learning algorithm of the policy network are determined.

[0060] The following section primarily compares the tuning performance of agents trained using two classic reinforcement learning algorithms, PPO and SAC. (The instruction manual includes...) Figure 4The diagram shows the change curve of reward data during the training of the agent using the SAC / PPO reinforcement learning algorithm. (See attached instruction manual.) Figure 5 The curves showing the change in the average temperature of the turbine blade surface during the training of the agent using the SAC / PPO reinforcement learning algorithm are presented. It can be seen that in the initial training phase, the reward data value of both agents continuously increases while the average temperature of the blade surface continuously decreases, then stabilizes. The PPO agent's reward increase speed is faster than that of the SAC agent, and the temperature reached during training is lower, around 840℃, compared to 900℃ for the SAC agent. In an experiment, five sets of random initial cooling structure design parameters were set to explore the optimization effects of the two agents, as shown in Table 1. The SAC algorithm can reduce the average temperature by 258.82℃ with an average time of 135.5s; the PPO algorithm reduces the average temperature by 135.94℃ with an average time of 89.2s. The SAC agent has a better optimization effect, and in multiple experiments, it can optimize the average temperature of the blade surface to below 930℃, but it takes longer. During training, the PPO agent learns and converges faster, but its optimization effect is not as good as that of the SAC agent.

[0061]

[0062] Table 1

[0063] Next, the main focus is on testing the performance tuning of different network structures under the SAC algorithm. Commonly used policy network structures include fully connected layers, convolutional neural networks (CNNs), and recurrent neural networks (RNNs). Fully connected layers can handle linear relationships between states and actions, CNNs can extract feature information from image states, and RNNs can handle state information with temporal relationships. Using overly complex policy networks in reinforcement learning may lead to problems such as the optimization algorithm getting stuck in local optima and instability during training. This experiment tests the tuning performance of SAC agents with a three-layer fully connected policy network structure and a network structure combining fully connected layers and convolutional layers. (See attached instruction manual.) Figure 6 The table shows the reward data variation curves during the training process of agents using the two different network structures described above. It can be seen that the agent with a fully connected policy network experiences faster reward learning and eventually stabilizes around the same reward value. As shown in Table 2 (Comparison of the tuning performance of the two policy networks), the average temperature optimization of the policy network combining fully connected and convolutional layers decreased to 260.88℃, while the average temperature optimization of the policy network consisting of three fully connected layers decreased to 258.82℃. This indicates that the policy network combining fully connected and convolutional layers has a 0.7% improvement in tuning performance compared to the network with three fully connected layers, but the training time increases by 31.2%.

[0064]

[0065] Table 2

[0066] Furthermore, the optimization performance of the SAC agent with a fully connected policy network layer was compared with that of four heuristic algorithms, verifying that the auto-tuning agent can achieve rapid optimization within 200 seconds and has good optimization results. Five sets of random initial cooling structure design parameters were set, and the performance of the four heuristic algorithms and reinforcement learning algorithms for parameter tuning was compared. The parameter settings of the heuristic algorithms are shown in Table 3.

[0067]

[0068] Table 3

[0069] Among them, the five optimization paths of genetic algorithm, simulated annealing algorithm, particle swarm optimization algorithm, differential evolution algorithm and SAC agent are as follows: Figure 7 As shown in the table, all five algorithms demonstrate optimization effects. The PSO and SA algorithms show good results in all five optimizations, achieving temperatures below 900℃. The DE and SA algorithms can optimize the average blade surface temperature to between 920-950℃. Combined with the SAC auto-tuning agent, the final temperature optimization can reach 900-910℃. Table 4 records the average time and optimization effect of the five algorithms after five optimizations. The SA algorithm shows the best optimization effect, averaging 883.16℃; the SAC agent is second best. The optimization performance of the other five algorithms is in the order of SA > SAC agent > PSO > DE > GA. The optimization time of the five algorithms is in the order of PSO > DE > GA > SA > SAC agent. Table 5 shows that the SAC agent's optimization time is within 200s, with an average time of 145.16s, while the other four algorithms' optimization times all exceed 20 minutes. This indicates that the SAC agent proposed in this application has a good optimization effect and can complete optimization in a very short time.

[0070]

[0071]

[0072] Table 4

[0073] Tuning time (s) 111.02 186.24 187.39 119.97 121.17 145.16

[0074] Table 5

[0075] As can be seen, the turbine blade cooling effect design optimization method provided in this application combines turbine blade cooling effect design tools with reinforcement learning algorithms. By defining appropriate environment, state, action, and agent policy network structure and reinforcement learning algorithm, the efficiency of turbine blade cooling effect design optimization is greatly improved.

[0076] Based on the same inventive concept, this application also provides a turbine blade cooling effect design optimization device based on reinforcement learning. Since the principle of the device in this application is similar to the turbine blade cooling effect design optimization method based on reinforcement learning described above, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0077] As per the instruction manual Figure 8 As shown, this application also provides a reinforcement learning-based turbine blade cooling efficiency design optimization device, the device comprising:

[0078] A creation module 801 is used to create a reinforcement learning environment for an agent based on a turbine blade cooling effect design tool; wherein, the reinforcement learning environment defines the agent's actions and state and reward data obtained from the turbine blade cooling effect design tool, the state including the average temperature of the turbine blade surface;

[0079] The acquisition module 802 is used to input the action into the turbine blade cooling effect design tool using the intelligent agent, and to acquire the state and reward data generated by the turbine blade cooling effect design tool in response to the action;

[0080] The reinforcement learning module 803 is used to perform reinforcement learning training on the agent's policy network based on the state and the reward data, so as to minimize the average temperature on the surface of the turbine blade.

[0081] In some embodiments, the change in the turbine blade cooling efficiency design parameters is defined as the action of the agent, and action parameter values ​​at different scales are set to represent the magnitude of the change.

[0082] In some embodiments, the creation module 801 defines the turbine blade cooling efficiency design parameters, including:

[0083] The turbine blade cooling effect design tool is used to construct all design parameters for the turbine blade structure. First design parameters that have no effect on the average surface temperature of the turbine blade are then removed from all design parameters, resulting in second design parameters. The first design parameters include the wall structure and the inner baffle structure. The second design parameters include the turbulence column, the layered baffle structure, and multiple partition structures.

[0084] The layer partition structure is fixedly installed; wherein, the number of layer partition structures directly affects the number of partition structures;

[0085] The turbulence column and the multiple partition structures are defined as the cooling efficiency design parameters of the turbine blade; wherein, the partition structure includes film vent diameter, film vent inclination angle, impact vent diameter, vent row distance and vent column distance.

[0086] In some embodiments, the creation module 801 defines reward data by setting at least one reward function, including:

[0087] A first reward function is set based on the average and variance values ​​of the temperature on the turbine blade surface; wherein, the lower the average and variance values ​​of the temperature on the turbine blade surface, the higher the reward data obtained.

[0088] A second reward function is set based on the average temperature of the turbine blade surface; wherein, the lower the average temperature of the turbine blade surface, the higher the reward data is obtained.

[0089] In some embodiments, the reinforcement learning module 803 trains the agent's policy network based on the state and reward data to minimize the average temperature of the turbine blade surface, including:

[0090] Based on the state and reward data, and using different reinforcement learning algorithms to train the policy network of the agent with different network structures, the average temperature and tuning time of the turbine blade surface are minimized accordingly.

[0091] Based on the obtained minimum average temperature of the turbine blade surface and the tuning time, the target network structure and target reinforcement learning algorithm of the policy network are determined.

[0092] In some embodiments, the reinforcement learning module 803 uses the SAC / PPO reinforcement learning algorithm to train the agent's policy network; the network structure of the policy network consists of fully connected layers or a combination of fully connected layers and convolutional layers.

[0093] This application provides a reinforcement learning-based turbine blade cooling effect design optimization device. A creation module creates a reinforcement learning environment for an agent based on a turbine blade cooling effect design tool. This environment defines the agent's actions and state and reward data obtained from the tool, including the average temperature of the turbine blade surface. An acquisition module uses the agent to input the actions into the tool and acquires the state and reward data generated by the tool in response. The reinforcement learning module trains the agent's policy network based on the state and reward data to minimize the average temperature of the turbine blade surface. Thus, rapid optimization of turbine blade cooling effect design parameters is achieved through reinforcement learning training of the agent.

[0094] Based on the same concept of the present invention, the specification is attached. Figure 9As shown in the figure, an embodiment of this application provides the structure of an electronic device 900, which includes: at least one processor 901, at least one network interface 904 or other user interface 903, a memory 905, and at least one communication bus 902. The communication bus 902 is used to realize the connection and communication between these components. The electronic device 900 may optionally include a user interface 903, including a display (e.g., touch screen, LCD, CRT, holographic imaging, or projector, etc.), a keyboard, or a clicking device (e.g., mouse, trackball, touchpad, or touch screen, etc.).

[0095] Memory 905 may include read-only memory and random access memory, and provides instructions and data to processor 901. A portion of memory 905 may also include non-volatile random access memory (NVRAM).

[0096] In some implementations, memory 905 stores elements that can protect modules or data structures, or subsets thereof, or extended sets thereof:

[0097] The 9051 operating system contains various system programs used to implement various basic business functions and handle hardware-based tasks.

[0098] Application module 9052 contains various applications, such as desktop launcher, media player, and browser, to implement various application functions.

[0099] In this embodiment of the application, by calling the program or instructions stored in the memory 905, the processor 901 is used to execute steps such as in a reinforcement learning-based turbine blade cooling effect design optimization method, which can greatly improve the efficiency of turbine blade cooling effect design optimization.

[0100] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs steps such as those in a reinforcement learning-based turbine blade cooling efficiency design optimization method.

[0101] Specifically, the storage medium can be a general-purpose storage medium, such as a portable disk or hard disk. When the computer program on the storage medium is run, it can execute the above-mentioned reinforcement learning-based turbine blade cooling efficiency design optimization method.

[0102] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there may be other division methods in actual implementation. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interface; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0103] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0104] In addition, the functional units in the embodiments provided in this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0105] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0106] Finally, it should be noted that the above embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be determined by the protection scope of the claims.

Claims

1. A method for designing and optimizing the cooling effect of turbine blades based on reinforcement learning, characterized in that, The method includes the following steps: A reinforcement learning environment is created based on a turbine blade cooling effect design tool. This environment defines the agent's actions and state and reward data obtained from the tool, including the average temperature of the turbine blade surface. Changes in turbine blade cooling effect design parameters are defined as actions of the agent, and different scales of action parameter values ​​represent the magnitude of these changes. The turbine blade cooling effect design parameters are defined by the following steps: determining all design parameters for constructing the turbine blade structure using the tool, and removing a first design parameter that has no effect on the average temperature of the turbine blade surface from all design parameters, resulting in a second design parameter. The first design parameter includes a wall structure and an inner baffle structure. The second design parameter includes a baffle column, a layer baffle structure, and multiple partition structures. The layer baffle structure is fixed; the number of layer baffle structures directly affects the number of partition structures. The baffle column and multiple partition structures are defined as turbine blade cooling effect design parameters. The partition structure includes film vent diameter, film vent angle, impact vent diameter, vent row distance, and vent column distance. The agent inputs the action into the turbine blade cooling effect design tool and obtains the state and reward data generated by the turbine blade cooling effect design tool in response to the action. Training the agent's policy network using reinforcement learning based on the state and reward data to minimize the average temperature of the turbine blade surface includes the following steps: training the agent's policy network with different network structures using reinforcement learning algorithms based on the state and reward data to obtain the corresponding minimized average temperature of the turbine blade surface and tuning time; determining the target network structure and target reinforcement learning algorithm of the policy network based on the obtained minimized average temperature of the turbine blade surface and tuning time.

2. The method for designing and optimizing the cooling effect of turbine blades based on reinforcement learning according to claim 1, characterized in that, in, The diameter of the turbulence column ranges from 0.5 to 1.5 mm, the diameter of the air film pores ranges from 0.4 to 0.8 mm, the inclination angle of the air film pores ranges from 30° to 90°, the diameter of the impact pores ranges from 0.9 to 1.5 mm, the spacing between the pore rows ranges from 3 to 5 mm, and the spacing between the pore columns ranges from 3 to 7 mm.

3. The method for designing and optimizing the cooling effect of turbine blades based on reinforcement learning according to claim 2, characterized in that, in, Define reward data by setting at least one reward function, and set the reward function in the following manner, including the following steps: A first reward function is set based on the average and variance values ​​of the temperature on the turbine blade surface; wherein, the lower the average and variance values ​​of the temperature on the turbine blade surface, the higher the reward data obtained. A second reward function is set based on the average temperature of the turbine blade surface; wherein, the lower the average temperature of the turbine blade surface, the higher the reward data obtained.

4. The method for designing and optimizing the cooling effect of turbine blades based on reinforcement learning according to claim 3, characterized in that, in, The SAC / PPO reinforcement learning algorithm is used to train the policy network of the agent; the network structure of the policy network consists of fully connected layers or a combination of fully connected layers and convolutional layers.

5. A turbine blade cooling efficiency design optimization device based on reinforcement learning, characterized in that, The device includes: A module is created to create a reinforcement learning environment for an agent based on a turbine blade cooling effect design tool. The reinforcement learning environment defines the agent's actions and state and reward data obtained from the turbine blade cooling effect design tool, where the state includes the average temperature of the turbine blade surface. Changes in turbine blade cooling effect design parameters are defined as actions of the agent, and different scales of action parameter values ​​are set to represent the magnitude of these changes. The turbine blade cooling effect design parameters are defined as follows: all design parameters for constructing the turbine blade structure using the turbine blade cooling effect design tool are determined, and a first design parameter that has no effect on the average temperature of the turbine blade surface is removed from all design parameters, resulting in a second design parameter. The first design parameter includes the wall structure and the inner baffle structure. The second design parameter includes a baffle column, a layer baffle structure, and multiple partition structures. The layer baffle structure is fixedly set; the number of layer baffle structures directly affects the number of partition structures. The baffle column and multiple partition structures are defined as turbine blade cooling effect design parameters; the partition structure includes film vent diameter, film vent angle, impact vent diameter, vent row distance, and vent column distance. The acquisition module is used to input the action into the turbine blade cooling effect design tool using the intelligent agent, and to acquire the state and reward data generated by the turbine blade cooling effect design tool in response to the action; A reinforcement learning module is used to train the agent's policy network based on the state and reward data to minimize the average temperature of the turbine blade surface. This includes: training the agent's policy network with different network structures using different reinforcement learning algorithms based on the state and reward data to obtain the corresponding minimized average temperature of the turbine blade surface and tuning time; and determining the target network structure and target reinforcement learning algorithm for the policy network based on the obtained minimized average temperature of the turbine blade surface and tuning time.

6. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, they perform the steps of the reinforcement learning-based turbine blade cooling efficiency design optimization method as described in any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the reinforcement learning-based turbine blade cooling efficiency design optimization method as described in any one of claims 1 to 4.