Ventilator operating parameter optimization method based on reinforcement learning

By using a reinforcement learning-based method, the state parameters of the ventilation fan are obtained and forward reasoning is performed to generate optimized adjustment parameters. This solves the problem of insufficient adaptability of the ventilation fan under complex operating conditions in the existing technology, and achieves energy saving, consumption reduction and stability improvement by maximizing the global reward expectation.

CN122260876APending Publication Date: 2026-06-23TIANJIN GENTECH POWER EQUIP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-26
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing ventilation fan operating parameter adjustment schemes are based on static models and fixed rule control logic, which makes it difficult to achieve adaptive optimization under complex operating conditions. This causes the fans to deviate from the optimal economic operating range for a long time, reducing energy saving and consumption reduction as well as operational safety.

Method used

A reinforcement learning-based approach is adopted. The state parameters of the ventilator are obtained and input into the reinforcement learning network model for forward inference to generate exploration action parameters. Markov experience tuples are constructed to update the network connection weights until the expected global reward is maximized, and the optimized adjustment parameters are output.

Benefits of technology

It achieves the global reward expectation maximization adjustment of the ventilation fan under complex dynamic operating conditions, improves energy saving and consumption reduction level and operational stability, and solves the problem of insufficient adaptability of traditional control logic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122260876A_ABST
    Figure CN122260876A_ABST
Patent Text Reader

Abstract

The application provides a ventilation fan operation parameter optimization method based on reinforcement learning. State variables of the ventilation fan are acquired, including aerodynamic variables, thermal variables and equipment operation feedback variables; the state variables are input into a constructed reinforcement learning network model for forward reasoning, and an exploration action variable containing adjustment control parameters of an actuator of the ventilation fan is output; a plurality of Markov experience tuples are generated based on the state variables and the exploration action variable to construct a strategy training batch; network connection weights of the reinforcement learning network model are updated based on the strategy training batch, until a weight update difference value between adjacent batches is less than or equal to a preset model convergence threshold, the training is ended, and the reinforcement learning network model is used as a ventilation fan operation parameter optimization decision model; the ventilation fan operation parameter optimization decision model is used for forward reasoning based on real-time collected state variables, and ventilation fan operation adjustment parameters reaching global reward expectation maximization are output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent control technology for industrial equipment, and more specifically, to a method for optimizing the operating parameters of a ventilation fan based on reinforcement learning. Background Technology

[0002] With the deepening of energy-saving transformation across the entire process in heavy industries such as steel and coking, high-energy-consuming ventilation equipment such as large-scale integrated desulfurization and denitrification fans and sintering main exhaust fans play a crucial role in the process flow. These fans typically operate under complex conditions such as high temperature and variable load, and their aerodynamic and thermal parameters exhibit highly dynamic changes during actual operation. To achieve better energy conservation and cost reduction, adaptive optimization control of the operating parameters of the fans under dynamic conditions has become a core application requirement in this field.

[0003] In existing ventilation fan operating parameter adjustment schemes, a system control mechanism based on static empirical curves combined with multi-loop PID feedback regulation is typically adopted. This scheme first uses field sensors to collect real-time operating status data such as inlet and outlet pressure and temperature of the ventilation fan; then, these static operating condition data are substituted into a pre-set mapping model, and the target response action is calculated by looking up tables or comparing with fixed logic; finally, control commands are sent to the actuator through the underlying controller to mechanically follow and adjust the operating frequency of the variable frequency motor and the adjustment angle of the moving blades.

[0004] However, this control scheme based on static models and fixed rules has obvious technical defects. Due to the strong nonlinearity and time-varying nature of the multi-dimensional parameter coupling of the ventilation fan under complex actual operating conditions, the existing static control logic lacks the ability to self-evolve and proactively learn based on real-time acquired state parameters. It is difficult to build an experience sequence to optimize the global control strategy in a continuously changing environment, which leads to the system's inability to adaptively output the best adjustment command to maximize the expected global reward. As a result, the fan deviates from the optimal economic operating range for a long time, which not only limits the overall energy saving and consumption reduction potential, but also reduces the safety and stability of the equipment under various changing operating conditions. Summary of the Invention

[0005] This application provides a method for optimizing the operating parameters of a ventilation fan based on reinforcement learning, so as to at least alleviate the above-mentioned technical problems.

[0006] A method for optimizing the operating parameters of a ventilation fan based on reinforcement learning, comprising: The state parameters of the ventilator are obtained, including the aerodynamic parameters, thermal parameters, and equipment operation feedback parameters of the ventilator. The state parameters are input into the constructed reinforcement learning network model for forward inference to output exploration action parameters, which include the adjustment and control parameters of the ventilator actuator. Based on the state parameters and exploration action parameters, several Markov experience tuples are generated to construct policy training batches; Based on the policy training batch, the network connection weights of the reinforcement learning network model are updated until the difference in network connection weight updates between adjacent policy training batches is less than or equal to a preset model convergence threshold. Then, the training of the reinforcement learning network model is terminated and it is used as a decision model for optimizing the operating parameters of the ventilation fan. The decision-making model for optimizing the operating parameters of the ventilator is used to perform forward reasoning based on the real-time collected state parameters of the ventilator, and output the ventilator operating adjustment parameters that maximize the expected global reward.

[0007] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the ventilation fan operating parameter optimization method according to any one of the claims in this application.

[0008] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the ventilation fan operating parameter optimization method according to any one of the claims in this application.

[0009] A device for optimizing the operating parameters of a ventilation fan based on reinforcement learning, comprising: The state parameter sensing module is used to acquire the state parameters of the ventilator, wherein the state parameters include the aerodynamic parameters, thermal parameters and equipment operation feedback parameters of the ventilator; The action exploration reasoning module is used to input the state parameters into the constructed reinforcement learning network model for forward reasoning, so as to output exploration action parameters, which include the adjustment and control parameters of the ventilator actuator; The batch construction module is used to generate several Markov experience tuples based on the state parameters and exploration action parameters to construct policy training batches. The model optimization training module is used to update the network connection weights of the reinforcement learning network model based on the training batch of the policy until the difference in network connection weight updates between adjacent training batches of the policy meets the condition of being less than or equal to a preset model convergence threshold. Then, the training of the reinforcement learning network model is terminated and used as the optimization decision model for the ventilation fan operating parameters. The adaptive optimization execution module is used to optimize the decision model using the fan operating parameters. Based on the real-time collected state parameters of the fan, it performs forward reasoning and outputs the fan operation adjustment parameters that maximize the expected global reward.

[0010] The technical advantages of the technical solution provided in this application are: This application presents a reinforcement learning-based method for optimizing the operating parameters of a ventilation fan. This addresses the technical deficiency of existing static control logic, which lacks the ability to self-evolve and proactively learn based on real-time acquired state parameters, resulting in the system's inability to adaptively output globally optimal adjustment commands. Compared to traditional control methods based on static preset parameter curves combined with multi-loop PID local feedback, this application acquires the state parameters of the ventilation fan, including aerodynamic parameters, thermal parameters, and equipment operation feedback parameters. These state parameters are input into a constructed reinforcement learning network model for forward inference to output exploratory action parameters, which include the adjustment and control parameters of the ventilation fan's actuator. Based on the state parameters and exploratory action parameters, several Markov experience tuples are generated to construct a strategy training batch. This mechanism eliminates the mechanical dependence on fixed rules and manual experience lookup tables, endowing the system with the ability to autonomously accumulate experience and interact with the environment under complex nonlinear conditions, thus rationally and effectively improving the dynamic representation level of the coupling characteristics of the multi-dimensional state parameters of the ventilation fan.

[0011] Simultaneously, based on the policy training batches, this application updates the network connection weights of the reinforcement learning network model until the difference in network connection weight updates between adjacent policy training batches is less than or equal to a preset model convergence threshold. At this point, the training of the reinforcement learning network model ends and it is used as the optimization decision model for the ventilation fan's operating parameters. Using this optimization decision model, forward inference is performed based on the real-time collected state parameters of the ventilation fan to output ventilation fan operation adjustment parameters that maximize the global reward expectation. Compared to traditional mechanical feedback regulation lacking global optimization capabilities, this application establishes a forward-looking policy optimization closed loop under continuously changing time-varying conditions through the extraction of Markov empirical tuples and dynamic model training convergence. This allows control commands to continuously learn and evolve adaptively in complex and dynamic operating environments, thus clearly solving the problem of the fan deviating from the optimal economic operating zone for a long time. This not only achieves a high level of global energy efficiency optimization and energy saving, but also better ensures the stability and safety of large industrial equipment operating under diverse abnormal operating conditions. Attached Figure Description

[0012] Figure 1 This application provides an embodiment of a method for optimizing the operating parameters of a ventilation fan based on reinforcement learning. Figure 2 This application provides an embodiment of a computer device. Figure 3 This application provides an embodiment of a computer-readable storage medium. Figure 4This application provides an embodiment of a ventilation fan operation parameter optimization device based on reinforcement learning. Detailed Implementation

[0013] like Figure 1 As shown, this embodiment of the present application provides a method for optimizing the operating parameters of a ventilation fan based on reinforcement learning, which includes: The state parameters of the ventilator are obtained, including the aerodynamic parameters, thermal parameters, and equipment operation feedback parameters of the ventilator. The state parameters are input into the constructed reinforcement learning network model for forward inference to output exploration action parameters, which include the adjustment and control parameters of the ventilator actuator. Based on the state parameters and exploration action parameters, several Markov experience tuples are generated to construct policy training batches; Based on the policy training batch, the network connection weights of the reinforcement learning network model are updated until the difference in network connection weight updates between adjacent policy training batches is less than or equal to a preset model convergence threshold. Then, the training of the reinforcement learning network model is terminated and it is used as a decision model for optimizing the operating parameters of the ventilation fan. The decision-making model for optimizing the operating parameters of the ventilator is used to perform forward reasoning based on the real-time collected state parameters of the ventilator, and output the ventilator operating adjustment parameters that maximize the expected global reward.

[0014] Optionally, the reinforcement learning network model includes: a feature perception fusion layer, a policy optimization decision layer, and an action instruction mapping layer. Correspondingly, the state parameters are input into the constructed reinforcement learning network model for forward inference to output exploration action parameters. The exploration action parameters include the adjustment and control parameters of the ventilator actuator, including: The state parameters are input into the feature perception fusion layer to extract the operating condition features that characterize the real-time operating status of the ventilator from the state parameters, and the operating condition features are vectorized to generate a global operating condition status feature vector of the ventilator. The global operating condition feature vector of the ventilator is passed to the strategy optimization decision layer for state feature mapping processing to output the hidden variables of action decision; The action instruction mapping layer maps the latent variables of the action decision to exploration action parameters.

[0015] Preferably, in the specific technical implementation of this application, the reinforcement learning network model performs forward inference according to the sequential processing relationship of the feature perception fusion layer, the strategy optimization decision layer, and the action instruction mapping layer; the feature perception fusion layer receives the state parameters and merges the aerodynamic parameters from the airflow process of the ventilation fan, the thermal parameters from the high-temperature flue gas process, and the equipment operation feedback parameters from the equipment operation process at the same operating moment to form a merged state parameter record; the merged state parameter record continues to enter the feature perception fusion layer for dimensional coordination processing to obtain a dimensionally coordinated merged state parameter record. The dimensionally coordinated merged state parameter record enables the inlet flow rate, total pressure rise, flue gas inlet temperature, bearing vibration value, bearing temperature rise, and single-unit operating power to participate in subsequent operating condition feature extraction under the same operating condition description framework. The same operating condition description framework is used to receive the dimensionally coordinated merged state parameter record and constrain the parameter reading order of the operating condition feature extraction.

[0016] Preferably, in the specific technical implementation of this application, the feature perception fusion layer performs component identification processing on the merged record of the dimensionally coordinated state parameters to obtain the aerodynamic load description component, the thermal drift description component, and the equipment feedback description component respectively; the aerodynamic load description component is characterized by the inlet flow rate and the total pressure rise, and is used to reflect the aerodynamic load state of the fan under the current duct resistance and airflow delivery requirements; the thermal drift description component is characterized by the flue gas inlet temperature, and is used to reflect the thermal drift influence of the high-temperature medium entering the fan on the aerodynamic density, bearing thermal state, and actuator response state; The equipment feedback description component is jointly characterized by the bearing vibration value, bearing temperature rise, and single-unit operating power, reflecting the mechanical response and energy consumption response of the fan under the current adjustment and control parameters. The aerodynamic load description component, the thermal drift description component, and the equipment feedback description component are further correlated and combined in the feature perception fusion layer to extract the operating condition features characterizing the real-time operating status of the fan. The operating condition features also carry the coupling information between the aerodynamic load status, the thermal drift effect, the actuator response status, the mechanical response status, and the energy consumption response status.

[0017] Preferably, the specific implementation process of this application is as follows: When the feature perception fusion layer extracts the operating condition features, it does not only perform threshold judgment on a single component of the state parameters, but first identifies whether the ventilator is currently in a high flow rate and low pressure rise aerodynamic load state, a high pressure rise and low flow rate aerodynamic load state, or an airflow delivery near the rated area aerodynamic load state based on the aerodynamic load description component, and takes the high flow rate and low pressure rise aerodynamic load state, the high pressure rise and low flow rate aerodynamic load state, or the airflow delivery near the rated area aerodynamic load state as the classification result of the aerodynamic load state; then, it identifies the smoke based on the thermal drift description component. The air inlet temperature affects the classification result of the aerodynamic load state by causing thermal drift. Then, the equipment feedback description component is superimposed onto the classification result of the aerodynamic load state affected by the thermal drift to obtain the operating condition feature reflecting the coupling relationship between the aerodynamic load state, the thermal drift effect, the mechanical response state, and the energy consumption response state. This operating condition feature is then used as input for vectorization processing, enabling it to be transformed from a combined expression of the aerodynamic load description component, the thermal drift description component, and the equipment feedback description component into a global operating condition feature vector of the ventilation fan that can participate in state feature mapping processing.

[0018] Preferably, in the specific technical implementation of this application, when the operating condition features are vectorized, the feature perception fusion layer performs scale normalization processing on the aerodynamic load description component, the thermal drift description component, and the equipment feedback description component in the operating condition features according to a pre-configured operating condition parameter range, so as to form scale-normalized operating condition features. The scale-normalized operating condition features are used to reduce the interference of different dimensions on the state feature mapping processing. Each dimension element in the global operating condition state feature vector of the ventilation fan corresponds to the inlet flow load occupancy degree, the total pressure rise response degree, the flue gas inlet temperature thermal drift degree, and the bearing vibration, respectively. The dynamic mechanical disturbance degree, bearing temperature rise thermal safety deviation degree, and single-unit operating power energy consumption response degree are all considered. The global operating condition feature vector of the ventilation fan is formed by arranging the inlet flow load occupancy degree, the total pressure rise response degree, the flue gas inlet temperature thermal drift degree, the bearing vibration mechanical disturbance degree, the bearing temperature rise thermal safety deviation degree, and the single-unit operating power energy consumption response degree in a pre-configured dimensional order, and is then passed to the strategy optimization decision layer, so that the strategy optimization decision layer can read the real-time operating status of the ventilation fan based on the global operating condition feature vector of the ventilation fan.

[0019] Preferably, in this application, after the global operating condition feature vector of the ventilation fan is passed to the strategy optimization decision layer, the strategy optimization decision layer first reads the pre-configured central kernel feature and calculates the operating condition difference between the global operating condition feature vector of the ventilation fan and the central kernel feature. The central kernel feature is used to characterize the typical operating condition position formed by the reinforcement learning network model during training, and the operating condition difference is used to characterize the closeness between the real-time operating state of the ventilation fan and the corresponding typical operating condition position. The strategy optimization decision layer then combines the pre-configured activation width parameter to perform nonlinear activation mapping on the operating condition difference to convert the operating condition difference into operating condition response intensity. The operating condition response intensity is used to express the degree of response of the real-time operating state of the ventilation fan to the typical operating condition position. The operating condition response intensity is further combined in the strategy optimization decision layer to form the action decision latent variable, so that the action decision latent variable can carry the state feature mapping processing result between the global operating condition feature vector of the ventilation fan and the central kernel feature.

[0020] Preferably, in the specific technical implementation of this application, the action decision latent variable is not a regulation and control parameter directly issued to the fan actuator, but rather an intermediate decision expression formed by the strategy optimization decision layer after performing state feature mapping on the global operating condition feature vector of the fan; the intermediate decision expression is used to transmit the action decision latent variable between the strategy optimization decision layer and the action command mapping layer; the action decision latent variable is used to express the action tendency degree corresponding to the real-time operating state of the fan under different typical operating condition positions, wherein the central core feature closer to the real-time operating state of the fan will correspond to a higher operating condition response intensity, and the central core feature farther away from the real-time operating state of the fan will correspond to a lower operating condition response intensity; thus, the action decision latent variable can transform the coupled changes between the inlet flow rate, the total pressure rise, the flue gas inlet temperature, the bearing vibration value, the bearing temperature rise, and the single-unit operating power into an action tendency degree that can be read by the subsequent action command mapping layer.

[0021] Preferably, in this application, when the action decision latent variables are mapped to exploratory action parameters through the action command mapping layer, the action command mapping layer reads the action decision latent variables and the network connection weight parameters between the strategy optimization decision layer and the action command mapping layer, and performs weighted mapping processing according to the action contribution relationship of each action decision latent variable to different ventilation fan actuators; the action contribution relationship is used to characterize the influence direction and influence intensity of each action decision latent variable on the variable frequency motor operating frequency and the moving blade adjustment angle; the action command mapping layer then reads the bias scalar corresponding to each ventilation fan actuator, and uses the bias scalar as the adjustment benchmark offset in the current operating scenario to participate in the weighted mapping processing to generate the exploratory action parameters; the action output values ​​of each dimension in the exploratory action parameters correspond to the variable frequency motor operating frequency and the moving blade adjustment angle, respectively, so that the exploratory action parameters can simultaneously reflect the adjustment direction of the variable frequency motor operating frequency, the adjustment range of the variable frequency motor operating frequency, the adjustment direction of the moving blade adjustment angle, and the adjustment range of the moving blade adjustment angle.

[0022] Preferably, in the operation scenarios of sintering main exhaust fan, desulfurization and denitrification booster fan, or coke oven gas dry quenching circulating fan, when the state parameters show increased inlet flow demand, insufficient total pressure rise response, increased flue gas inlet temperature, and increased single-unit operating power, the feature perception fusion layer groups the state parameters corresponding to the increased inlet flow demand, insufficient total pressure rise response, increased flue gas inlet temperature, and increased single-unit operating power into operating condition features that can reflect high-load thermal disturbances, and vectorizes the operating condition features that can reflect high-load thermal disturbances into the global operating condition state feature vector of the ventilation fan; the strategy optimization decision layer is based on the ventilation fan The difference in operating conditions between the global operating condition feature vector and the central kernel feature forms the action decision latent variable; the action command mapping layer then generates the exploratory action parameter based on the action decision latent variable, so that the exploratory action parameter can provide a joint adjustment expression for the operating frequency of the variable frequency motor and the adjustment angle of the moving blade; the joint adjustment expression belongs to the action output content in the exploratory action parameter, and the joint adjustment expression enables the operating frequency of the variable frequency motor and the adjustment angle of the moving blade to jointly participate in the generation of the adjustment control parameters of the fan actuator, rather than mechanically following the adjustment based solely on a single pressure deviation or a single flow deviation.

[0023] Preferably, the exploration action parameter obtained in this application also maintains a correspondence with the state parameter. The state parameter represents the real-time operating state of the ventilator before the generation of the exploration action parameter, and the exploration action parameter represents the candidate output of the adjustment control parameter formed by the reinforcement learning network model based on the real-time operating state of the ventilator. When generating several Markov experience tuples in the future, the state parameter and the exploration action parameter jointly participate in the construction of the Markov experience tuple, so that the exploration action parameter obtained by forward inference not only serves as a candidate output of adjustment control parameter but also as the source of action parameters for subsequent policy training batches. Through the above-mentioned connection, the feature perception fusion layer, the policy optimization decision layer, and the action instruction mapping layer form a continuous technical processing relationship from state parameter parsing, operating condition feature generation, ventilator global operating condition state feature vector expression, action decision latent variable generation to exploration action parameter output. The continuous technical processing relationship is used to ensure that the exploration action parameter maintains a clear source relationship in the subsequent construction of Markov experience tuples and policy training batches.

[0024] Preferably, in the specific technical implementation of this application, the feature-aware fusion layer is not a convolutional layer oriented towards image pixel neighborhood scanning, but rather adopts a grouped gated normalization fusion structure oriented towards multi-source state parameters of industrial large wind turbines; the feature-aware fusion layer takes the state parameters as input, and performs parameter grouping and reading of the state parameters according to the physical source of the aerodynamic parameters, the thermal parameters, and the equipment operation feedback parameters, to form aerodynamic parameter reading results, thermal parameter reading results, and equipment operation feedback parameter reading results; the aerodynamic parameter reading results, the thermal parameter reading results, and the equipment operation feedback parameter reading results continue to enter the feature-aware fusion layer for merging processing at the same operating moment, so as to... The state parameter merging record is formed; the state parameter merging record is then processed by the feature perception fusion layer for dimensional coordination to obtain the dimensionally coordinated state parameter merging record. This allows the feature perception fusion layer to unify parameters with different physical dimensions, different acquisition sources, and different response speeds in the state parameters into a state expression that can participate in state feature mapping processing before entering the strategy optimization decision layer. The dimensionally coordinated state parameter merging record continues to participate in component identification processing and association combination in the feature perception fusion layer to form the operating condition feature. The operating condition feature is then vectorized by the feature perception fusion layer to form the global operating condition state feature vector of the ventilation fan.

[0025] Preferably, the feature perception fusion layer includes a parameter source identification channel, an operating time alignment channel, a dimension coordination channel, and an operating condition gating fusion channel; the parameter source identification channel reads the inlet flow rate, total pressure rise, flue gas inlet temperature, bearing vibration value, bearing temperature rise, and single-unit operating power from the state parameters, and forms the aerodynamic parameter reading results, the thermal parameter reading results, and the equipment operating feedback parameter reading results according to the source relationship of the aerodynamic parameters, the thermal parameters, and the equipment operating feedback parameter reading results; the operating time alignment channel performs merging processing on the aerodynamic parameter reading results, the thermal parameter reading results, and the equipment operating feedback parameter reading results under the same operating time to form the state parameter merging record; the dimension coordination channel... The state parameter merging record is subjected to dimensional coordination processing to obtain the dimensionally coordinated state parameter merging record. The operating condition gating fusion channel then performs component identification processing based on the dimensionally coordinated state parameter merging record to obtain the aerodynamic load description component, the thermal drift description component, and the equipment feedback description component, respectively. The aerodynamic load description component, the thermal drift description component, and the equipment feedback description component are then correlated and combined to form the operating condition feature, which can simultaneously express the correlation between the fan aerodynamic load change, high-temperature flue gas disturbance, bearing vibration response, bearing temperature rise response, and single-unit operating power response. The operating condition feature is further vectorized in the feature perception fusion layer to form the global operating condition state feature vector of the fan.

[0026] Preferably, the operating condition gating fusion channel does not simply concatenate the aerodynamic load description component, the thermal drift description component, and the equipment feedback description component. Instead, it first generates an aerodynamic gating coefficient based on the aerodynamic load description component, then generates a thermal drift gating coefficient based on the thermal drift description component, and finally generates an equipment feedback gating coefficient based on the equipment feedback description component. The aerodynamic gating coefficient is used to adjust the proportion of the aerodynamic load description component in the operating condition characteristics, the thermal drift gating coefficient is used to adjust the correction ratio of the thermal drift description component to the aerodynamic load state, the aerodynamic load state being characterized by the aerodynamic load description component, and the equipment feedback gating coefficient is used to adjust the correction ratio of the equipment feedback description component to the mechanical response state. The correction ratio of the mechanical response state and the energy consumption response state, wherein the mechanical response state and the energy consumption response state are characterized by the equipment feedback description component; the operating condition gating fusion channel applies the aerodynamic gating coefficient, the thermal drift gating coefficient, and the equipment feedback gating coefficient to the dimensionally coordinated state parameters and records them to form the operating condition feature, so that the operating condition feature is not a static parameter set, but a fusion expression that can change the component contribution relationship as the real-time operating state of the fan changes; the operating condition feature is further vectorized through the feature perception fusion layer to form the global operating condition state feature vector of the fan, so that the operating condition feature can serve as the pre-input source for the state feature mapping processing of the strategy optimization decision layer.

[0027] Preferably, the strategy optimization decision layer is not a convolutional layer, nor is it a simple mapping layer formed by stacking fully connected neurons. Instead, it adopts a radial basis function strategy optimization structure based on the operating condition kernel response. The strategy optimization decision layer takes the global operating condition state feature vector of the ventilator as input and reads the pre-configured central kernel feature and the pre-configured activation width parameter. The central kernel feature corresponds to the typical operating condition position formed by the ventilator during training, and the activation width parameter corresponds to the operating condition neighborhood range that different typical operating condition positions can cover. The strategy optimization decision layer calculates the operating condition difference between the global operating condition state feature vector of the ventilator and multiple central kernel features, and converts the operating condition difference into the operating condition response intensity based on the activation width parameter. The operating condition response intensity is further combined in the strategy optimization decision layer to form the action decision latent variable, so that the action decision latent variable can express the degree of action tendency after the current real-time operating state of the ventilator falls into the neighborhood of different typical operating condition positions. The action decision latent variable is further passed to the action command mapping layer as the input source for the action command mapping layer to generate the exploration action parameters.

[0028] Preferably, the central kernel feature in the strategy optimization decision layer is not an arbitrarily set fixed node, but is formed according to the distribution of the global operating condition feature vector of the ventilation fan in the strategy training batch; in the strategy training batch, the global operating condition feature vector of the ventilation fan corresponds to different inlet flow load occupancy, total pressure rise response, flue gas inlet temperature thermal drift, bearing vibration mechanical disturbance, bearing temperature rise thermal safety offset, and single-unit operating power consumption response; the strategy optimization decision layer is based on the inlet flow load occupancy, total pressure rise response, flue gas inlet temperature thermal drift, and bearing vibration mechanical disturbance. The clustering locations of the bearing temperature rise thermal safety offset and the single-unit operating power energy consumption response under different operating conditions form the central core feature, and the activation width parameter is formed based on the range of operating condition changes around the different clustering locations. The central core feature and the activation width parameter jointly participate in the calculation of the operating condition difference and the generation of the operating condition response intensity, enabling the strategy optimization decision layer to form a state feature mapping process with operating condition neighborhood discrimination capability for the variable load scenarios of sintering main exhaust fan, desulfurization and denitrification booster fan or coke oven gas dry quenching circulating fan. The processing result of the state feature mapping process is further transmitted to the action command mapping layer in the form of the action decision latent variable.

[0029] Preferably, when forming the latent variables for action decisions, the strategy optimization decision layer further groups the operating condition response intensity according to the source relationships of the aerodynamic load state, the thermal drift effect, the mechanical response state, and the energy consumption response state to obtain the aerodynamic response intensity, thermal drift response intensity, mechanical response intensity, and energy consumption response intensity. The aerodynamic response intensity is used to express the degree of proximity between the current real-time operating state of the ventilator and the typical operating condition position related to the aerodynamic load; the thermal drift response intensity is used to express the degree of proximity between the current real-time operating state of the ventilator and the typical operating condition position related to high-temperature flue gas disturbance; and the mechanical response intensity is used to express the degree of proximity between the current real-time operating state of the ventilator and the typical operating condition position related to high-temperature flue gas disturbance. The degree of proximity between the real-time operating status and the bearing vibration value and bearing temperature rise related to typical operating conditions is measured. The energy consumption response intensity is used to express the degree of proximity between the current real-time operating status of the fan and the typical operating conditions related to the single-unit operating power. The strategy optimization decision layer combines the aerodynamic response intensity, the thermal drift response intensity, the mechanical response intensity, and the energy consumption response intensity to form the action decision latent variable, which can distinguish the source of influence of different types of operating condition changes on the adjustment and control parameters of the fan actuator. The action decision latent variable continues to serve as the input of the action command mapping layer to participate in the generation of the exploration action parameters.

[0030] Preferably, the action command mapping layer is neither a regular classification output layer nor a regression layer that only outputs a single continuous value, but rather adopts a dual-channel constrained projection structure oriented towards the ventilation fan actuator. The action command mapping layer takes the action decision latent variables as input and reads the pre-configured network connection weight parameters and the pre-configured bias scalar updated during training. According to the pre-configured action contribution relationship updated during training, the action command mapping layer projects the action decision latent variables onto the variable frequency motor operating frequency channel and the moving blade adjustment angle channel, respectively, to form candidate outputs for the variable frequency motor operating frequency in the variable frequency motor operating frequency channel and candidate outputs for the moving blade adjustment angle in the moving blade adjustment angle channel. The candidate outputs for the variable frequency motor operating frequency and the candidate outputs for the moving blade adjustment angle continue to enter the action command mapping layer for channel collaborative processing. This channel collaborative processing coordinates the output ratio between the candidate outputs for the variable frequency motor operating frequency and the candidate outputs for the moving blade adjustment angle to form the exploratory action parameters, ensuring that the action output values ​​of each dimension in the exploratory action parameters simultaneously correspond to the variable frequency motor operating frequency and the moving blade adjustment angle.

[0031] Preferably, when generating the exploratory action parameters, the action command mapping layer sets the variable frequency motor operating frequency channel and the moving blade adjustment angle channel as two mutually constraining output channels. When the action decision latent variable indicates that the single-machine operating power energy consumption response is too high and the aerodynamic load state still needs to meet the airflow delivery requirements, the action command mapping layer enables the candidate output of the variable frequency motor operating frequency and the candidate output of the moving blade adjustment angle to jointly participate in the generation of the exploratory action parameters, rather than simply increasing or decreasing the action output value of a single channel. When the action decision latent variable indicates that the bearing vibration value mechanical disturbance or the bearing temperature rise thermal safety offset is too high, the action command mapping layer adjusts the proportional relationship between the candidate output of the variable frequency motor operating frequency and the candidate output of the moving blade adjustment angle according to the action contribution relationship, and forms the exploratory action parameters that can simultaneously reflect energy consumption response, mechanical response, and thermal safety offset through the channel collaborative processing. The exploratory action parameters continue to serve as candidate outputs of the adjustment control parameters of the fan actuator to participate in the subsequent construction of the Markov empirical tuple.

[0032] Preferably, the action command mapping layer further includes action boundary projection processing; the action boundary projection processing reads the candidate output of the variable frequency motor operating frequency, the candidate output of the moving blade adjustment angle, and the pre-configured actuator action boundary, and projects the candidate output of the variable frequency motor operating frequency and the candidate output of the moving blade adjustment angle onto the action range defined by the actuator action boundary, so as to form the variable frequency motor operating frequency candidate output and the moving blade adjustment angle candidate output after boundary constraint; the variable frequency motor operating frequency candidate output and the moving blade adjustment angle candidate output after boundary constraint are further combined in the action command mapping layer to form the exploratory action parameters, so that the exploratory action parameters can be used as the adjustment control parameters of the ventilator actuator to participate in the subsequent construction of the Markov experience tuple; the Markov experience tuple is then connected with the strategy training batch, so that the exploratory action parameters can continue to be used as the source of action parameters in the construction of the subsequent strategy training batch.

[0033] Optionally, the global operating condition state feature vector of the ventilator is passed to the strategy optimization decision layer for state feature mapping processing to output latent variables for action decisions, including: Configure the Gaussian radial basis function of the strategy optimization decision layer, and determine the central kernel feature corresponding to the Gaussian radial basis function; Calculate the Euclidean distance between the global operating condition feature vector of the ventilation fan and the central kernel feature; The action decision latent variables are calculated by performing a nonlinear activation mapping based on the Euclidean distance and activation width parameter.

[0034] Preferably, in the specific technical implementation of this application, the strategy optimization decision layer adopts a Gaussian radial basis mapping structure for neighborhood identification of industrial large fan operating conditions. This Gaussian radial basis mapping structure does not directly perform a normal fully connected mapping on the global operating condition feature vector of the fan. Instead, it first establishes the current real-time operating state of the fan and typical operating conditions based on the inlet flow load occupancy, total pressure rise response, flue gas inlet temperature thermal drift, bearing vibration mechanical disturbance, bearing temperature rise thermal safety offset, and single-unit operating power consumption response expressed by the global operating condition feature vector of the fan. The distance response relationship between operating condition positions; the distance response relationship is represented by the Euclidean distance between the global operating condition state feature vector of the ventilator and the central kernel feature. The distance response relationship continues to serve as the input basis for the nonlinear activation mapping of the Gaussian radial basis function, enabling the strategy optimization decision layer to first determine which type of typical operating condition position the current real-time operating state of the ventilator is close to based on the distance response relationship, and then output the action decision latent variable based on the degree of proximity of the operating condition corresponding to the distance response relationship, instead of directly generating the action decision latent variable based solely on the threshold change of a single state parameter.

[0035] Preferably, when configuring the Gaussian radial basis function of the strategy optimization decision layer in this application, the system first reads the sample record of the global operating condition feature vector of the ventilation fan formed by the strategy training batch, and extracts from the sample record the inlet flow load occupancy, total pressure rise response, flue gas inlet temperature thermal drift, bearing vibration mechanical disturbance, bearing temperature rise thermal safety offset, and single-unit operating power consumption response under different operating conditions. The system divides the operating condition neighborhood into candidate records to form a central kernel. These candidate records then participate in the central kernel screening process, which selects records from the candidate records that can characterize typical operating conditions to determine the central kernel features corresponding to the Gaussian radial basis function. These central kernel features are derived from the actual operating condition distribution of the fan under different loads, different flue gas inlet temperatures, different bearing vibration values, different bearing temperature rises, and different single-unit operating power. The actual operating condition distribution continues to participate in the subsequent Euclidean distance calculation through the central kernel features.

[0036] Preferably, in the specific implementation of this application, the central kernel feature is used to characterize the typical operating condition position that can be activated by the current real-time operating state of the ventilator in the strategy optimization decision layer; the central kernel feature is not an isolated fixed parameter, but maintains the same dimensional arrangement relationship with the global operating condition feature vector of the ventilator. The position of each dimension in the central kernel feature corresponds to the inlet flow load occupancy, the total pressure rise response, the flue gas inlet temperature thermal drift, the bearing vibration value mechanical disturbance, the bearing temperature rise thermal safety offset, and the single-unit operating power energy consumption response; when the global operating condition feature vector of the ventilator is transmitted to the strategy optimization decision layer, the strategy optimization decision layer reads the global operating condition feature vector of the ventilator and the central kernel feature according to the same dimensional arrangement relationship, so that the Euclidean distance calculation between the global operating condition feature vector of the ventilator and the central kernel feature is established between dimensional elements with the same physical meaning, avoiding direct misalignment comparison of dimensional elements with different physical meanings such as the inlet flow load occupancy and the bearing vibration value mechanical disturbance.

[0037] Preferably, in this application, when calculating the Euclidean distance between the global operating condition feature vector of the ventilation fan and the central kernel feature, the strategy optimization decision layer first performs a dimension-by-dimensional difference calculation on each dimension element of the global operating condition feature vector of the ventilation fan and the corresponding dimension element of the central kernel feature to form a dimension-by-dimensional operating condition deviation. The dimension-by-dimensional operating condition deviation respectively characterizes the deviation of the current inlet flow load occupancy degree from the corresponding typical operating condition position, the deviation of the current total pressure rise response degree from the corresponding typical operating condition position, and the deviation of the current flue gas inlet temperature thermal drift degree from the corresponding typical operating condition position. The deviations of the current bearing vibration value (mechanical disturbance) from the corresponding typical operating condition position, the deviations of the current bearing temperature rise thermal safety deviation from the corresponding typical operating condition position, and the deviations of the current single-unit operating power consumption response from the corresponding typical operating condition position are calculated. These deviations are then processed by distance convergence, which synthesizes the deviations according to the correspondence of elements in each dimension to form the Euclidean distance. This Euclidean distance serves as the overall approximation of the current real-time operating state of the ventilator with its corresponding central core feature and continues to participate in the nonlinear activation mapping.

[0038] Preferably, before calculating the Euclidean distance, the strategy optimization decision layer further performs a scale-consistent reading on the global operating condition feature vector of the ventilation fan. This scale-consistent reading receives the global operating condition feature vector of the ventilation fan output by the feature perception fusion layer and verifies that each dimension element in the global operating condition feature vector of the ventilation fan is in a normalized expression state suitable for distance calculation. This normalized expression state originates from the expression result formed by the feature perception fusion layer after performing dimensional coordination and vectorization processing on the state parameters. The normalized expression state continues to be processed... This forms the numerical basis for the Euclidean distance calculation. After the inlet flow load occupancy, total pressure rise response, flue gas inlet temperature thermal drift, bearing vibration mechanical disturbance, bearing temperature rise thermal safety offset, and single-unit operating power energy consumption response in the global operating condition feature vector of the ventilation fan have been read in a consistent manner, the strategy optimization decision layer then calculates the Euclidean distance between the global operating condition feature vector of the ventilation fan and the central kernel feature. This ensures that the Euclidean distance mainly reflects the differences in operating conditions, rather than numerical differences affected by different dimensions or different acquisition scales.

[0039] Preferably, in this application, when performing nonlinear activation mapping based on the Euclidean distance and the activation width parameter, the activation width parameter is used to express the coverage range of the central kernel feature over the surrounding operating conditions, and the activation width parameter is set correspondingly to the central kernel feature. For typical operating conditions where load changes are slow, flue gas inlet temperature fluctuations are small, and equipment operation feedback parameters change relatively smoothly, the activation width parameter causes the Gaussian radial basis function to respond to Euclidean distance changes within a small range. For typical operating conditions where load changes are rapid, flue gas inlet temperature disturbances are strong, and equipment operation feedback parameters change significantly, the activation width parameter causes the Gaussian radial basis function to cover a wider operating condition neighborhood. The Euclidean distance and the activation width parameter are jointly incorporated into the nonlinear activation mapping to form the operating condition response intensity. The operating condition response intensity characterizes the degree of response of the typical operating condition location corresponding to the central kernel feature to the current real-time operating state of the ventilator, and the operating condition response intensity continues to participate in the calculation of the latent variables of the action decision.

[0040] Preferably, the nonlinear activation mapping does not directly use the Euclidean distance as the latent variable for action decision. Instead, it performs response attenuation processing on the typical operating condition position corresponding to the central kernel feature based on the relative relationship between the Euclidean distance and the activation width parameter. The response attenuation processing converts the Euclidean distance into the operating condition response intensity, causing the operating condition response intensity to decrease as the Euclidean distance increases and to increase as the Euclidean distance decreases. When the Euclidean distance is small, the nonlinear activation mapping forms a high operating condition response intensity through the response attenuation processing, indicating that the current real-time operating state of the ventilator is close to the typical operating condition position represented by the corresponding central kernel feature. When the Euclidean distance is large, the nonlinear activation mapping forms a low operating condition response intensity through the response attenuation processing, indicating that the current real-time operating state of the ventilator deviates from the typical operating condition position represented by the corresponding central kernel feature. The operating condition response intensity then participates in the generation of the latent variable for action decision, enabling the latent variable for action decision to reflect the response differences of different typical operating condition positions to the current real-time operating state of the ventilator.

[0041] Preferably, in this application, when calculating the action decision latent variables, the strategy optimization decision layer performs Euclidean distance calculation and nonlinear activation mapping on multiple central kernel features to form multiple operating condition response intensities; the multiple operating condition response intensities are combined according to the order of their corresponding central kernel features in the strategy optimization decision layer to form a sequence of action decision latent variable components; each action decision latent variable component in the sequence of action decision latent variable components receives an operating condition response intensity corresponding to a central kernel feature; the sequence of action decision latent variable components continues to serve as the action decision latent variables, enabling the action decision latent variables to express the distribution relationship of the current real-time operating status of the ventilation fan among multiple typical operating condition positions; the action decision latent variables are further passed to the action command mapping layer, enabling the action command mapping layer to generate the exploratory action parameters based on the action decision latent variables.

[0042] Preferably, in the operating scenarios of the sintering main exhaust fan, desulfurization and denitrification booster fan, or coke oven gas dry quenching circulating fan, when the strategy optimization decision layer determines the central core characteristics, it considers high flow rate and low pressure rise operating conditions, high pressure rise and low flow rate operating conditions, high flue gas inlet temperature operating conditions, high bearing vibration value operating conditions, high bearing temperature rise operating conditions, and high single-unit operating power operating conditions as operating condition types that can form central core candidate records; the operating condition types are used to limit the source range of the central core candidate records, and the central core candidate records are formed after being screened by the central core. The central kernel feature is then compared with the real-time input global operating condition feature vector of the ventilation fan to calculate the Euclidean distance, thereby determining which typical operating condition the current real-time operating state of the ventilation fan is close to. The operating condition response intensity formed by the Euclidean distance and the activation width parameter further participates in the generation of the action decision latent variables, enabling the action decision latent variables to reflect the state feature mapping results of the large ventilation fan under conditions of high temperature, variable load, and fluctuation of equipment operation feedback parameters. The state feature mapping results are then passed to the action command mapping layer in the form of the action decision latent variables.

[0043] Optionally, the mathematical expression of the Gaussian radial basis function of the strategy optimization decision layer is: in, To optimize the strategy, the latent variables of the action decision output by the j-th node of the decision layer are used. The feature vector representing the global operating status of the ventilation fan is passed to the feature perception fusion layer. The central kernel feature of the j-th node, The activation width parameter for the j-th node. Let i be an exponential function with the natural constant e as the base, where i and j are both integers.

[0044] Preferably, in the specific technical implementation of the strategy optimization decision layer, the Gaussian radial basis function is not simply a function for general numerical fitting, but is configured as a nonlinear activation function for identifying the operating condition neighborhood of industrial large fans. The Gaussian radial basis function takes the global operating condition state feature vector of the fan as input and uses the Euclidean distance between the global operating condition state feature vector of the fan and the central kernel feature as the distance input of the nonlinear activation mapping, so as to generate action decision latent variable components according to the degree of deviation of the current real-time operating state of the fan from the typical operating condition position corresponding to the central kernel feature. The action decision latent variable components continue to participate in the formation of the action decision latent variables, so that the strategy optimization decision layer does not directly linearly convert the global operating condition state feature vector of the fan into the exploration action parameter, but first determines the operating condition neighborhood of the current real-time operating state of the fan according to the Euclidean distance, and then passes the action decision latent variables to the action command mapping layer, so that the action command mapping layer generates the exploration action parameter.

[0045] Preferably, in the technical meaning of the Gaussian radial basis function, the global operating condition feature vector of the ventilation fan originates from the extraction and vectorization processing of the state parameters by the feature perception fusion layer. Each dimension element in the global operating condition feature vector of the ventilation fan corresponds to the inlet flow load occupancy, the total pressure rise response, the flue gas inlet temperature thermal drift, the bearing vibration mechanical disturbance, the bearing temperature rise thermal safety offset, and the single-unit operating power consumption response. The central kernel feature maintains the same dimensional arrangement relationship as the global operating condition feature vector of the ventilation fan. The position of each dimension in the central kernel feature also corresponds to the inlet flow load occupancy, the total pressure rise response, the flue gas inlet temperature thermal drift, the bearing vibration mechanical disturbance, the bearing temperature rise thermal safety offset, and the single-unit operating power consumption response. Thus, the Euclidean distance is the operating condition deviation formed between dimension elements with the same physical meaning, rather than a misaligned comparison of parameters with different dimensions or different sources. The operating condition deviation continues to participate in the formation of the Euclidean distance.

[0046] Preferably, the central kernel feature in the Gaussian radial basis function is used to characterize the typical operating condition positions formed in the strategy training batch. The typical operating condition positions are derived from the actual operating condition distribution formed by the sintering main exhaust fan, desulfurization and denitrification booster fan, or coke oven gas dry quenching circulating fan under variable load, high temperature flue gas, and equipment operation feedback parameter fluctuation conditions. The strategy optimization decision layer reads the global operating condition state feature vector sample records of the ventilation fans in the strategy training batch, extracts the actual operating condition distribution with a high degree of aggregation, and uses the vector expression corresponding to the actual operating condition distribution as the central kernel feature. The central kernel feature continues to participate in the calculation of the Euclidean distance, so that the Gaussian radial basis function can perform nonlinear activation mapping around the more representative aerodynamic load state, thermal drift effect, mechanical response state, and energy consumption response state in the actual operating condition distribution, rather than performing nonlinear activation mapping around arbitrarily set abstract nodes.

[0047] Preferably, the Euclidean distance in the Gaussian radial basis function is used to express the overall degree of similarity between the current real-time operating state of the ventilator and the central kernel feature. When calculating the Euclidean distance, the strategy optimization decision layer first calculates the difference between each dimension element in the global operating state feature vector of the ventilator and the dimension element at the same dimension position in the central kernel feature to form a dimension-by-dimensional operating condition deviation. The dimension-by-dimensional operating condition deviation is then squared and converged to form the Euclidean distance. Through the above processing, the influence of the inlet flow load occupancy, the total pressure rise response, the flue gas inlet temperature thermal drift, the bearing vibration mechanical disturbance, the bearing temperature rise thermal safety offset, and the single-unit operating power energy consumption response on the current real-time operating state of the ventilator are all included in the formation of the Euclidean distance with the corresponding dimension-by-dimensional operating condition deviation, so that the Euclidean distance can reflect the operating condition differences under the combined action of multiple source state parameters.

[0048] Preferably, the squaring of the dimension-by-dimensional operating condition deviation in the Gaussian radial basis function is not simply to amplify the numerical difference, but to ensure that both positive and negative deviations of each dimension element can be converted into non-negative operating condition deviation contributions. For example, if the current inlet flow load occupancy is higher than or lower than that in the central kernel feature, corresponding dimension-by-dimensional operating condition deviations will be generated, and these deviations will be squared to form corresponding operating condition deviation contributions. Similarly, when the mechanical disturbance degree of the bearing vibration value, the thermal safety deviation degree of the bearing temperature rise, and the energy consumption response degree of the single-unit operation power consumption shift upward or downward relative to the central kernel feature, they will also be converted into comparable operating condition deviation contributions. These operating condition deviation contributions will then undergo convergence processing to form the Euclidean distance, thereby enabling the Euclidean distance to express the comprehensive deviation degree of the current real-time operating state of the ventilator relative to the typical operating condition position.

[0049] Preferably, the activation width parameter is used to control the coverage range of the typical operating condition location corresponding to the central kernel feature to the surrounding operating conditions. The activation width parameter is set corresponding to the central kernel feature and participates in the response attenuation processing of the Gaussian radial basis function to the Euclidean distance. When the typical operating condition location corresponding to a certain central kernel feature originates from the actual operating condition distribution with slow load changes, small flue gas inlet temperature fluctuations, and relatively stable equipment operation feedback parameters, the activation width parameter causes the Gaussian radial basis function to respond to the Euclidean distance changes within a narrower operating condition neighborhood. When the typical operating condition location corresponding to a certain central kernel feature originates from the actual operating condition distribution with rapid load changes, strong flue gas inlet temperature disturbances, and large fluctuations in equipment operation feedback parameters, the activation width parameter causes the Gaussian radial basis function to respond to the Euclidean distance changes within a wider operating condition neighborhood. The activation width parameter thus enables different typical operating condition locations to have different operating condition neighborhood coverage capabilities, and the operating condition neighborhood coverage capability continues to affect the formation of the operating condition response intensity.

[0050] Preferably, the Gaussian radial basis function performs a nonlinear activation mapping on the Euclidean distance and the activation width parameter using an exponential function with a base of the natural constant. This nonlinear activation mapping converts the Euclidean distance into a working condition response intensity. When the Euclidean distance is small, the overall working condition of the global operating condition feature vector of the ventilator is highly similar to that of the central kernel feature, resulting in a high working condition response intensity from the nonlinear activation mapping. Conversely, when the Euclidean distance is large, the overall working condition of the global operating condition feature vector of the ventilator is less similar to that of the central kernel feature, resulting in a low working condition response intensity from the nonlinear activation mapping. The working condition response intensity continues to participate in the formation of the action decision latent variable as a component of the action decision latent variable, enabling the action decision latent variable to express the response differences of the current real-time operating state of the ventilator to different typical operating condition positions.

[0051] Preferably, the exponential decay relationship in the Gaussian radial basis function enables the strategy optimization decision layer to form a continuous operating condition neighborhood response among multiple central kernel features, rather than hard switching between multiple typical operating condition positions. When the current real-time operating state of the ventilator is between two typical operating condition positions, the strategy optimization decision layer calculates the Euclidean distance between the global operating condition state feature vector of the ventilator and the two central kernel features, and forms two operating condition response intensities respectively through the corresponding activation width parameters. The two operating condition response intensities continue to enter the action decision latent variable, so that the action decision latent variable can simultaneously retain the response information of adjacent typical operating condition positions. This processing method is suitable for state feature mapping processing of industrial large fans under conditions of continuous load change, continuous fluctuation of flue gas inlet temperature, and continuous change of equipment operation feedback parameters.

[0052] Preferably, the action decision latent variable components output by the Gaussian radial basis function are not directly equivalent to the exploratory action parameters, but rather serve as intermediate decision expressions passed from the policy optimization decision layer to the action command mapping layer. Multiple central kernel features correspond to multiple action decision latent variable components, which are combined according to the node arrangement in the policy optimization decision layer to form the action decision latent variables. These action decision latent variables then enter the action command mapping layer and are mapped to the exploratory action parameters. Thus, the Gaussian radial basis function in the reinforcement learning network model plays a role in generating the working condition neighborhood response, and the action command mapping layer transforms the action decision latent variables into adjustment and control parameters for the ventilation fan actuator. These adjustment and control parameters, as components of the exploratory action parameters, continue to participate in the subsequent construction of Markov empirical tuples.

[0053] Preferably, in the operation scenarios of sintering main exhaust fan, desulfurization and denitrification booster fan, or coke oven gas dry quenching circulating fan, when the global operating condition feature vector of the ventilation fan shows an increase in the inlet flow load occupancy, insufficient total pressure rise response, increased thermal drift of flue gas inlet temperature, and increased single-unit operating power energy consumption response, the strategy optimization decision layer will calculate the Euclidean distance between the global operating condition feature vector of the ventilation fan and the central kernel features corresponding to the high flow low pressure rise operating condition, high flue gas inlet temperature operating condition, and high single-unit operating power operating condition, respectively. Each Euclidean distance is entered into the Gaussian radial basis function for nonlinear activation mapping with the corresponding activation width parameter to form the corresponding operating condition response intensity. After multiple operating condition response intensities are combined into the action decision latent variable, they provide the operating condition response basis for the action command mapping layer to generate exploration action parameters including the variable frequency motor operating frequency and the moving blade adjustment angle. The operating condition response basis is passed to the action command mapping layer by the action decision latent variable.

[0054] Preferably, the technical role of the Gaussian radial basis function is also reflected in the identification of the operating condition neighborhood of the equipment operating feedback parameters related to operational safety; the equipment operating feedback parameters related to operational safety include the bearing vibration value and the bearing temperature rise; when the mechanical disturbance degree of the bearing vibration value or the thermal safety offset degree of the bearing temperature rise in the global operating condition feature vector of the ventilator is close to the typical operating condition position represented by the corresponding central kernel feature, the Euclidean distance will decrease, the nonlinear activation mapping will form a higher operating condition response intensity, and the operating condition response intensity will further participate in the action decision latent variable; after the action decision latent variable enters the action command mapping layer, it can affect the joint adjustment direction of the variable frequency motor operating frequency and the moving blade adjustment angle in the exploration action parameters, so that the generation process of the exploration action parameters simultaneously considers the single machine operating power energy consumption response degree, the mechanical disturbance degree of the bearing vibration value, and the thermal safety offset degree of the bearing temperature rise, rather than adjusting only around the single deviation of flow rate or pressure.

[0055] Preferably, there is a continuous technical relationship between the Gaussian radial basis function and the subsequent policy training batch construction; the policy optimization decision layer forms the action decision latent variables based on the Gaussian radial basis function, the action instruction mapping layer generates the exploration action parameters based on the action decision latent variables, the exploration action parameters and the state parameters jointly participate in the construction of Markov experience tuples, and the Markov experience tuples further participate in the construction of policy training batches; in subsequent training, the policy training batches are used to update the network connection weights of the reinforcement learning network model and affect the configuration of subsequent central kernel features and activation width parameters, so that the working condition neighborhood response of the Gaussian radial basis function can be gradually adjusted according to the actual operating condition distribution of the ventilation fan, thereby providing a state feature mapping basis that fits the operating scenario of industrial large fans for the ventilation fan operation parameter optimization decision model.

[0056] Optionally, the action instruction mapping layer maps the action decision latent variables to the output mathematical expression of the explored action parameters as follows: .in, This refers to the action output value of the k-th dimension in the exploration action parameters. To determine the network connection weight parameters between the j-th node in the strategy optimization decision layer and the k-th node in the action instruction mapping layer. is the bias scalar corresponding to the k-th node of the action instruction mapping layer, m is the total number of nodes in the policy optimization decision layer, and k is an integer.

[0057] Preferably, in the specific technical implementation of the action command mapping layer, the action command mapping layer maps the action decision latent variables to the output mathematical expression of the exploratory action parameters. Essentially, it weights and aggregates multiple action decision latent variable components output by the strategy optimization decision layer according to the action output dimension corresponding to the adjustment and control parameters of the ventilation fan actuator. The action decision latent variable components originate from the working condition response intensity corresponding to different central kernel features in the strategy optimization decision layer. The working condition response intensity already characterizes the degree of proximity between the current real-time operating state of the ventilation fan and different typical operating condition positions. After reading the action decision latent variable components, the action command mapping layer does not directly use a single action decision latent variable component as the exploratory action parameter. Instead, it incorporates multiple action decision latent variable components into the calculation of the same action output dimension to form the action output value of the corresponding action output dimension in the exploratory action parameter. This allows the exploratory action parameter to bear the common influence of multiple typical operating condition positions on the current real-time operating state of the ventilation fan. The action output value continues to participate in the subsequent construction of Markov empirical tuples as a component of the exploratory action parameter.

[0058] Preferably, when generating the exploratory action parameters, the action command mapping layer divides the exploratory action parameters into action output dimensions corresponding to the adjustment and control parameters of the fan actuator. These action output dimensions include at least a variable frequency motor operating frequency action output dimension and a moving blade adjustment angle action output dimension. For the variable frequency motor operating frequency action output dimension, the action command mapping layer reads each action decision latent variable component and performs weighted aggregation based on the action contribution relationship of each action decision latent variable component to the variable frequency motor operating frequency action output dimension, thereby forming a variable frequency motor operating frequency action output value. Similarly, for the moving blade adjustment angle action output dimension, the action command mapping layer reads each action decision latent variable component and performs weighted aggregation based on the action contribution relationship of each action decision latent variable component to the moving blade adjustment angle action output dimension, thereby forming a moving blade adjustment angle action output value. The variable frequency motor operating frequency action output value and the moving blade adjustment angle action output value together constitute the exploratory action parameters, enabling the exploratory action parameters to simultaneously correspond to both the variable frequency motor operating frequency and the moving blade adjustment angle.

[0059] Preferably, the network connection weight parameter is used to express the action contribution relationship between the action decision latent variable component in the strategy optimization decision layer and the action output dimension in the action instruction mapping layer. This action contribution relationship is not a fixed empirical ratio, but is adjusted as the network connection weights of the reinforcement learning network model are updated during the training of the strategy training batch. When a certain action decision latent variable component corresponds to the operating condition response intensity of a high-flow-low-pressure rise operating condition, the network connection weight parameter is used to express the dynamic response of that action decision latent variable component to the variable frequency motor operating frequency action output dimension and the moving blade adjustment angle action output dimension, respectively. The network connection weight parameter expresses the contribution relationship between the latent variable component of the action decision and the operating conditions when the bearing vibration value is too high or the bearing temperature rise is too high. This contribution is used to express the action contribution of the latent variable component to the variable frequency motor operating frequency action output dimension and the moving blade adjustment angle action output dimension, respectively. Thus, the network connection weight parameter transforms the operating condition response intensity into the node action contribution amount corresponding to the action output dimension, enabling the action command mapping layer to generate the variable frequency motor operating frequency action output value and the moving blade adjustment angle action output value for different operating condition response sources.

[0060] Preferably, the bias scalar is used to express the adjustment reference offset corresponding to each action output dimension in the action command mapping layer. The adjustment reference offset does not replace the action decision latent variable components, but participates in the adjustment reference correction of the same action output dimension after the weighted aggregation of each action decision latent variable component. For the variable frequency motor operating frequency action output dimension, the bias scalar enables the variable frequency motor operating frequency action output value to change around the variable frequency motor operating frequency adjustment reference. For the moving blade adjustment angle action output dimension, the bias scalar enables the moving blade adjustment angle action output value to change around the moving blade adjustment angle adjustment reference. The bias scalar and the network connection weight parameters jointly participate in the generation of the exploration action parameters, so that the action command mapping layer can express the action contribution relationship between different action decision latent variable components and different action output dimensions, and can also retain the adjustment reference offset related to the variable frequency motor operating frequency adjustment reference and the moving blade adjustment angle adjustment reference.

[0061] Preferably, the action command mapping layer adopts a mapping method combining weighted convergence and adjustment benchmark correction. This is to adapt to the technical scenario where the operating frequency of the variable frequency motor and the adjustment angle of the moving blades are simultaneously affected by multiple operating conditions in the adjustment of operating parameters of industrial large fans. During the operation of sintering main exhaust fans, desulfurization and denitrification booster fans, or coke oven gas dry quenching circulating fans, the degree of inlet flow load occupancy, the degree of total pressure rise response, the degree of thermal drift of flue gas inlet temperature, the degree of mechanical disturbance of bearing vibration value, the degree of thermal safety offset of bearing temperature rise, and the degree of energy consumption response of single-unit operation power often jointly affect the operation of large industrial fans. The variable frequency motor operating frequency action output value and the moving blade adjustment angle action output value are affected; if the exploration action parameters are generated only based on the response intensity of a single working condition, the exploration action parameters are easily biased towards a single target; the action command mapping layer maps multiple action decision hidden variable components to the same action output dimension through the network connection weight parameters, and completes the adjustment benchmark correction through the bias scalar, so that the exploration action parameters can simultaneously reflect the influence of aerodynamic load state, thermal drift effect, mechanical response state and energy consumption response state on the adjustment and control parameters of the ventilation fan actuator.

[0062] Preferably, in the processing corresponding to the output mathematical expression of the action instruction mapping layer, the total number of nodes in the strategy optimization decision layer is used to limit the number of action decision latent variable components participating in the calculation of the action output value. The action instruction mapping layer reads each action decision latent variable component sequentially according to the node arrangement relationship in the strategy optimization decision layer. Each action decision latent variable component is first converted with the corresponding network connection weight parameter to form the node action contribution. Multiple node action contributions are then converged according to the same action output dimension to form the node action contribution convergence. The node action contribution convergence continues to be adjusted and corrected with the bias scalar to form the action output value of the corresponding action output dimension in the exploration action parameter. The action output value includes the variable frequency motor operating frequency action output value and the moving blade adjustment angle action output value. The variable frequency motor operating frequency action output value and the moving blade adjustment angle action output value together constitute the exploration action parameter, so that the exploration action parameter has a clear source relationship and dimension correspondence.

[0063] Preferably, when the current real-time operating state of the ventilator is characterized by increased inlet flow load, insufficient total pressure rise response, and increased single-unit operating power consumption response, the strategy optimization decision layer forms action decision latent variable components related to the high flow low pressure rise operating condition and the high single-unit operating power condition. After reading the above action decision latent variable components, the action command mapping layer calculates the node action contribution of the above action decision latent variable components to the variable frequency motor operating frequency action output dimension and the moving blade adjustment angle action output dimension through the corresponding network connection weight parameters. After the node action contribution is converged and corrected by the adjustment benchmark, it forms the variable frequency motor operating frequency action output value and the moving blade adjustment angle action output value in the exploratory action parameters. Thus, the exploratory action parameters are not directly triggered by a single flow deviation or a single pressure deviation, but are jointly generated by multiple action decision latent variable components according to the action contribution relationship, and continue to participate in the subsequent Markov empirical tuple construction as candidate outputs of the ventilator actuator's adjustment control parameters.

[0064] Preferably, when the current real-time operating state of the ventilator is characterized by an increase in the mechanical disturbance of bearing vibration or an increase in the thermal safety deviation of bearing temperature rise, the strategy optimization decision layer forms latent variable components of action decisions related to the high bearing vibration or high bearing temperature operation conditions. The action command mapping layer maps these latent variable components of action decisions to the variable frequency motor operating frequency action output dimension and the moving blade adjustment angle action output dimension through the network connection weight parameters, so that the variable frequency motor operating frequency action output value and the moving blade adjustment angle action output value in the exploratory action parameters can reflect the changes in equipment operation feedback parameters related to operational safety. The bias scalar continues to adjust the variable frequency motor operating frequency action output value and the moving blade adjustment angle action output value to correct the adjustment benchmark, so that the exploratory action parameters can maintain an action output range that matches the variable frequency motor operating frequency adjustment benchmark and the moving blade adjustment angle adjustment benchmark while reflecting the mechanical response state and thermal safety deviation. The action output range further constrains the exploratory action parameters as candidate outputs of the ventilator actuator's adjustment control parameters to participate in the subsequent construction of Markov empirical tuples.

[0065] Preferably, the mathematical expression of the action command mapping layer is calculated separately for each action output dimension, enabling different fan actuators to share the same set of action decision latent variable components, while forming different action output values ​​through different network connection weight parameters. For the variable frequency motor operating frequency action output dimension, the network connection weight parameter corresponding to the same action decision latent variable component is used to express the action contribution relationship of that action decision latent variable component to the variable frequency motor operating frequency action output dimension. For the moving blade adjustment angle action output dimension, the network connection weight parameter corresponding to the same action decision latent variable component is used to express the action contribution relationship of that action decision latent variable component to the moving blade adjustment angle action output dimension. This mapping method of setting network connection weight parameters separately for each action output dimension enables the action command mapping layer to form the variable frequency motor operating frequency action output value for the variable frequency motor operating frequency action output dimension and the moving blade adjustment angle action output value for the moving blade adjustment angle action output dimension under the same current real-time operating state of the fan, avoiding the mixing of the adjustment control parameters of different fan actuators into a single action quantity.

[0066] Preferably, after the action instruction mapping layer generates the exploratory action parameters, the exploratory action parameters continue to participate in the construction of Markov experience tuples together with the state parameters. The Markov experience tuples then participate in the construction of policy training batches. In subsequent training, the policy training batches are used to update the network connection weight parameters and the bias scalar, so that the network connection weight parameters can adjust the contribution relationship of each action decision latent variable component to different action output dimensions according to the actual operating conditions of the ventilator, and the bias scalar can adjust the adjustment reference offset of the corresponding action output dimension according to the changes in the variable frequency motor operating frequency adjustment reference and the moving blade adjustment angle adjustment reference. Thus, the output mathematical expression is not an isolated output calculation formula, but a technical processing link that takes over the action decision latent variables, generates the exploratory action parameters, and participates in the construction of subsequent policy training batches.

[0067] Optionally, based on the training batch of the policy, the network connection weights of the reinforcement learning network model are updated, including: The global reward error for optimizing the parameters of the reinforcement learning network model is calculated using the training batches of the strategy, and the error gradient for updating the network parameters is calculated by backchain differentiation of the global reward error. Based on the error gradient of the network parameter update, the network connection weights of the action instruction mapping layer, policy optimization decision layer, and feature perception fusion layer are updated in batches through backpropagation.

[0068] Preferably, in a specific implementation, when calculating the global reward error using the policy training batch, the state parameters, exploration action parameters, equipment operation feedback parameters after executing the exploration action parameters, and state parameters at adjacent operating times are first read from each Markov experience tuple in the policy training batch. The state parameters, exploration action parameters, equipment operation feedback parameters after executing the exploration action parameters, and state parameters at adjacent operating times are then organized into batch training records that can participate in the parameter optimization iteration of the reinforcement learning network model. The state parameters in the batch training records are used to characterize the ventilation before the ventilator executes the exploration action parameters. The real-time operating status of the machine is recorded. The exploration action parameters in the batch training records are used to characterize the adjustment and control parameters of the fan actuator under the real-time operating status of the fan. The equipment operation feedback parameters after executing the exploration action parameters in the batch training records are used to characterize the changes in bearing vibration value, bearing temperature rise, and single-unit operating power after executing the exploration action parameters. The batch training records continue to participate in the calculation of the comprehensive expected return to form the predicted action value and target action value corresponding to each Markov empirical tuple. The difference between the predicted action value and the target action value is used to calculate the global reward error.

[0069] Preferably, in the calculation of the global reward error, the strategy training batch does not only evaluate the single-objective power of a single machine, but also incorporates bearing vibration value, bearing temperature rise, and single-machine power into the comprehensive expected return calculation. Specifically, the bearing vibration value reflects the mechanical response state of the fan rotor, bearing housing, and transmission components under the influence of the exploration action parameters; the bearing temperature rise reflects the thermal safety deviation of the fan under high-temperature flue gas and variable load conditions; and the single-machine power reflects the energy consumption response state of the fan actuator under the influence of the exploration action parameters. The mechanical response state, the thermal safety deviation, and the energy consumption response state all participate in the comprehensive expected return calculation, enabling the global reward error to simultaneously reflect deviations in the energy consumption direction, mechanical response direction, and thermal safety direction, rather than just the deviation in single-machine power.

[0070] Preferably, the predicted action value is calculated by the reinforcement learning network model based on the state parameters and the exploration action parameters. The predicted action value is used to characterize the expected comprehensive benefit of the exploration action parameters to the real-time operating state of the ventilator under the current network connection weights. The target action value is calculated based on the equipment operation feedback parameters after executing the exploration action parameters and the state parameters at adjacent operating times. The target action value is used to characterize the expected comprehensive benefit reference value corresponding to the exploration action parameters under actual operation feedback. The difference between the predicted action value and the target action value is calculated to form a temporal difference error. The temporal difference error is further aggregated within the policy training batch to form the global reward error, so that the global reward error can bear the actual operation feedback of multiple Markov empirical tuples in the policy training batch.

[0071] Preferably, when performing a reverse chain derivative on the global reward error, the global reward error is first differentiated with the exploration action parameters output by the action command mapping layer to form a local error gradient corresponding to the action command mapping layer. The local error gradient corresponding to the action command mapping layer is used to characterize the direction and magnitude of the influence of each action output value in the exploration action parameters on the global reward error. The action output values ​​in the exploration action parameters include the variable frequency motor operating frequency action output value and the moving blade adjustment angle action output value. The local error gradient corresponding to the action command mapping layer continues to act on the network connection weight parameters and bias scalar in the action command mapping layer, so that the action command mapping layer can form the network connection weight update amount corresponding to the action command mapping layer according to the local error gradient corresponding to the action command mapping layer.

[0072] Preferably, after the action command mapping layer completes the calculation of the local error gradient corresponding to the action command mapping layer, the local error gradient corresponding to the action command mapping layer is further propagated back to the strategy optimization decision layer through the network connection weight parameters between the action command mapping layer and the strategy optimization decision layer to form the error gradient corresponding to the strategy optimization decision layer. The error gradient corresponding to the strategy optimization decision layer is used to characterize the direction and magnitude of the influence of the action decision latent variable on the global reward error. Since the action decision latent variable is formed by the combination of the working condition response intensity corresponding to multiple central kernel features, the error gradient corresponding to the strategy optimization decision layer continues to act on the working condition response intensity, the central kernel feature, and the network connection weights corresponding to the activation width parameter, so that the strategy optimization decision layer can form the network connection weight update amount corresponding to the strategy optimization decision layer according to the error gradient corresponding to the strategy optimization decision layer.

[0073] Preferably, the error gradient corresponding to the strategy optimization decision layer continues to be transmitted to the feature perception fusion layer along the source direction of the global operating condition state feature vector of the ventilation fan, so as to form the error gradient corresponding to the feature perception fusion layer; the error gradient corresponding to the feature perception fusion layer is used to characterize the direction and magnitude of the influence of the operating condition features and the global operating condition state feature vector of the ventilation fan on the global reward error; since the global operating condition state feature vector of the ventilation fan comes from the aerodynamic parameters, thermal parameters and equipment operation feedback parameters in the state parameters, the error gradient corresponding to the feature perception fusion layer continues to act on the network connection weights in the feature perception fusion layer used for parameter grouping reading, dimensional coordination processing, operating condition feature extraction and vectorization processing, so that the feature perception fusion layer can form the network connection weight update amount corresponding to the feature perception fusion layer according to the error gradient corresponding to the feature perception fusion layer.

[0074] Preferably, when the network connection weights of the action command mapping layer, the strategy optimization decision layer, and the feature perception fusion layer are updated in batches through backpropagation based on the error gradient updated by the network parameters, the network connection weights of the action command mapping layer are first updated according to the update amount of the network connection weights corresponding to the action command mapping layer, so that the output values ​​of the variable frequency motor operating frequency and the adjustment angle of the moving blades in the exploration action parameters can be adjusted in the direction of reducing the global reward error; then, the network connection weights of the strategy optimization decision layer are updated according to the update amount of the network connection weights corresponding to the strategy optimization decision layer, so that the action decision latent variables can adjust the working condition response intensity of different typical operating condition positions in the direction of reducing the global reward error; finally, the network connection weights of the feature perception fusion layer are updated according to the update amount of the network connection weights corresponding to the feature perception fusion layer, so that the expression contribution of each dimension element of the global operating condition feature vector of the ventilation fan can be adjusted in the direction of reducing the global reward error.

[0075] Preferably, the batch collaborative update does not separate the action command mapping layer, the strategy optimization decision layer, and the feature perception fusion layer into unrelated single-layer training. Instead, within the same strategy training batch, the local error gradient corresponding to the action command mapping layer, the error gradient corresponding to the strategy optimization decision layer, and the error gradient corresponding to the feature perception fusion layer are calculated sequentially according to the backpropagation direction from the output end to the input end. The network connection weight update of the action command mapping layer, the strategy optimization decision layer, and the feature perception fusion layer is completed within the network connection weight update cycle corresponding to the same strategy training batch. The network connection weight update result of the action command mapping layer affects the output direction of the exploration action parameter, the network connection weight update result of the strategy optimization decision layer affects the working condition response distribution of the action decision latent variable, and the network connection weight update result of the feature perception fusion layer affects the state expression content of the global operating condition state feature vector of the ventilator. Thus, the action command mapping layer, the strategy optimization decision layer, and the feature perception fusion layer form a continuous parameter optimization relationship around the global reward error within the same strategy training batch.

[0076] Preferably, in the operation scenarios of sintering main exhaust fan, desulfurization and denitrification booster fan, or coke oven gas dry quenching circulating fan, when the Markov empirical tuple in the strategy training batch indicates that a certain exploration action parameter, although reducing the single-unit operating power, causes an increase in bearing vibration value or an upward shift in bearing temperature rise, the comprehensive benefit expectation calculation will simultaneously read the single-unit operating power change result, the bearing vibration value change result, and the bearing temperature rise change result, thereby ensuring that the global reward error retains the constraint information of mechanical response state and thermal safety offset on parameter optimization; after the global reward error forms the error gradient for updating the network parameters through reverse chain differentiation, the action command mapping layer adjusts the variable frequency motor operating frequency action output value and the moving blade adjustment angle action output value, the strategy optimization decision layer adjusts the operating condition response intensity corresponding to the typical operating condition position, and the feature perception fusion layer adjusts the contribution of the equipment operating feedback parameter in the global operating condition state feature vector of the ventilation fan.

[0077] Preferably, when the Markov empirical tuples in the strategy training batch indicate that a certain exploratory action parameter reduces the single-unit operating power while meeting the airflow delivery requirements, and keeps the bearing vibration value and bearing temperature rise within the operating range allowed by the equipment safety operation procedures, the comprehensive benefit expectation calculation will use the actual operating feedback corresponding to the exploratory action parameter as the feedback direction with a higher comprehensive benefit expectation in the global reward error calculation; the global reward error, after being differentiated by the back chain, forms the error gradient of the network parameter update, which enables the action instruction mapping layer to enhance the action contribution relationship of the latent variable component of the action decision corresponding to the exploratory action parameter to the action output dimension, enables the strategy optimization decision layer to enhance the working condition response intensity of the typical operating condition position corresponding to the exploratory action parameter, and enables the feature perception fusion layer to maintain the working condition feature expression that matches the actual operating feedback corresponding to the exploratory action parameter, thereby enabling subsequent forward inference to form ventilation fan operation adjustment parameters similar to the exploratory action parameter under similar operating conditions.

[0078] Preferably, the network connection weight update amount continues to participate in the calculation of the network connection weight update difference value between adjacent policy training batches; after each policy training batch completes the network connection weight update of the action instruction mapping layer, the policy optimization decision layer, and the feature perception fusion layer, the network connection weight update amount corresponding to that policy training batch is recorded, and the difference between the network connection weight update amount corresponding to that policy training batch and the network connection weight update amount corresponding to adjacent policy training batches is calculated to form the network connection weight update difference value; the network connection weight update difference value is further used to determine whether the reinforcement learning network model has reached the model convergence threshold requirement, so that a continuous training judgment relationship is formed between the global reward error, the error gradient of the network parameter update, the network connection weight update amount, and the network connection weight update difference value.

[0079] Preferably, during the training process of the reinforcement learning network model, the policy training batch is used to calculate the global reward error on the one hand, and to generate the error gradient for updating the network parameters on the other hand through the global reward error, and further generate the network connection weight update amount for the action instruction mapping layer, the policy optimization decision layer, and the feature perception fusion layer; after the network connection weight update amount is applied to the reinforcement learning network model, the reinforcement learning network model continues to perform forward reasoning on the state parameters of the ventilator collected in real time to output the subsequent exploration action parameters; the subsequent exploration action parameters continue to generate subsequent Markov experience tuples with the corresponding state parameters, and the subsequent Markov experience tuples continue to participate in the construction of subsequent policy training batches, so that the network connection weights of the reinforcement learning network model can be continuously adjusted with the actual operating conditions of the ventilator, and finally form the ventilator operating parameter optimization decision model.

[0080] Optionally, the training of the reinforcement learning network model is terminated and used as a decision model for optimizing the operating parameters of the ventilation fan when the difference in network connection weight updates between adjacent policy training batches is less than or equal to a preset model convergence threshold, including: Based on the network connection weights of the action instruction mapping layer, the policy optimization decision layer, and the feature perception fusion layer, the difference between the network connection weight update amounts corresponding to adjacent policy training batches is calculated. The training of the reinforcement learning network model ends when the difference is less than or equal to a preset model convergence threshold, and it is used as the ventilation fan operation parameter optimization decision model.

[0081] Preferably, in this specific implementation, when determining whether the reinforcement learning network model has ended training, not only is the trend of the global reward error read, but also the difference in network connection weight updates between adjacent policy training batches is read. The difference in network connection weight updates originates from the amount of network connection weight updates of the action instruction mapping layer, the policy optimization decision layer, and the feature perception fusion layer in adjacent policy training batches. The amount of network connection weight updates is used to characterize the parameter adjustment magnitude of the action instruction mapping layer, the policy optimization decision layer, and the feature perception fusion layer within a policy training batch due to the backpropagation of the global reward error. The amount of network connection weight updates continues to participate in the difference calculation between adjacent policy training batches to form the difference in network connection weight updates, so that the determination of training termination can reflect the stability of the network connection weight changes of the reinforcement learning network model in consecutive policy training batches, rather than relying solely on the number of training iterations or the magnitude of a single global reward error.

[0082] Preferably, after each strategy training batch completes the network connection weight update, the network connection weight update amounts of the action command mapping layer, the strategy optimization decision layer, and the feature perception fusion layer are recorded respectively. The network connection weight update amount of the action command mapping layer is used to characterize the change in the action mapping relationship between the variable frequency motor operating frequency action output value and the moving blade adjustment angle action output value. The network connection weight update amount of the strategy optimization decision layer is used to characterize the change in the state feature mapping relationship between the central kernel feature, the activation width parameter, and the operating condition response intensity. The network connection weight update amount of the feature perception fusion layer is used to characterize the change in the state expression relationship between the state parameter and the operating condition feature and the global operating condition state feature vector of the ventilation fan. The network connection weight update amounts of the action command mapping layer, the strategy optimization decision layer, and the feature perception fusion layer are further merged according to their hierarchical sources to form the batch weight update record corresponding to the strategy training batch.

[0083] Preferably, the batch weight update record does not simply record the update amount of a single network connection weight in a certain layer, but simultaneously includes the network connection weight update amounts of the action command mapping layer, the strategy optimization decision layer, and the feature perception fusion layer. The network connection weight update amounts of the action command mapping layer, the strategy optimization decision layer, and the feature perception fusion layer in the batch weight update record maintain their respective hierarchical source relationships and continue to be compared with the batch weight update records corresponding to adjacent strategy training batches. Through the above processing, the network connection weight update difference value can simultaneously reflect the changes in the action mapping relationship corresponding to the exploration action parameters in the action command mapping layer, the changes in the state feature mapping relationship corresponding to the action decision latent variables in the strategy optimization decision layer, and the changes in the state expression relationship corresponding to the global operating condition state feature vector of the ventilator in the feature perception fusion layer, avoiding the neglect of the continuous fluctuations in the state expression relationship of the feature perception fusion layer by using only the network connection weight update amount of the action command mapping layer as the basis for training termination.

[0084] Preferably, when calculating the difference between the network connection weight update amounts corresponding to adjacent policy training batches, the batch weight update record corresponding to the current policy training batch is first read, and the batch weight update record corresponding to the previous policy training batch is also read. Then, the network connection weight update amounts of the action command mapping layer, the policy optimization decision layer, and the feature perception fusion layer are calculated to form the update differences for the action command mapping layer, the policy optimization decision layer, and the feature perception fusion layer. These differences are then merged to form the network connection weight update difference value, ensuring that the network connection weight update difference value has a clear hierarchical source and calculation destination.

[0085] Preferably, in the calculation of the network connection weight update difference value, the network connection weight update amounts of the action command mapping layer, the strategy optimization decision layer, and the feature perception fusion layer are calculated using the same-layer difference calculation. This is because the network connection weights corresponding to the action command mapping layer, the strategy optimization decision layer, and the feature perception fusion layer differ in quantity, physical meaning, and position of action. The network connection weights of the action command mapping layer mainly affect the variable frequency motor operating frequency action output value and the moving blade adjustment angle action output value in the exploration action parameters. The network connection weights of the strategy optimization decision layer mainly affect the working condition response intensity in the action decision latent variables. The network connection weights of the feature perception fusion layer mainly affect the working condition characteristics and the global operating condition state feature vector of the ventilation fan. Therefore, the update differences of the action command mapping layer, the strategy optimization decision layer, and the feature perception fusion layer are first formed within their respective layers, and then jointly participate in the formation of the network connection weight update difference value, so that the network connection weight update difference value can correspond to the respective technical functions of the action command mapping layer, the strategy optimization decision layer, and the feature perception fusion layer.

[0086] Preferably, the model convergence threshold is not an immediate judgment threshold for whether a certain exploration action parameter reduces the single-machine operating power, but a training stability judgment threshold set for the difference value of network connection weight updates. When the difference value of network connection weight updates is greater than the model convergence threshold, it indicates that the action command mapping layer, the policy optimization decision layer, and the feature perception fusion layer still have a large need for network connection weight adjustment between adjacent policy training batches, and the network connection weights of the reinforcement learning network model continue to be updated based on subsequent policy training batches. When the difference value of network connection weight updates is less than or equal to the model convergence threshold, it indicates that the amount of network connection weight updates of the reinforcement learning network model between adjacent policy training batches has entered a low change range, and the training of the reinforcement learning network model ends and it is used as the optimization decision model for the ventilator operating parameters.

[0087] Preferably, in the operation scenarios of sintering main exhaust fan, desulfurization and denitrification booster fan, or coke oven gas dry quenching circulating fan, the network connection weight update difference value can reflect the impact of high-temperature flue gas, variable load, and equipment operation feedback parameter fluctuations on the training of the reinforcement learning network model; for example, when the Markov experience tuples in a certain policy training batch mainly reflect the operating state of flue gas inlet temperature rising and bearing temperature rising, the network connection weight update amount of the feature perception fusion layer will reflect the thermal parameters and the equipment operation feedback parameters in the global operating condition characteristics of the ventilation fan. The changes in the contribution of the expression in the vector, the update of the network connection weights of the strategy optimization decision layer will reflect the changes in the operating condition response intensity related to the high flue gas inlet temperature operating condition and the high bearing temperature rise operating condition, and the update of the network connection weights of the action command mapping layer will reflect the changes in the adjustment direction of the variable frequency motor operating frequency action output value and the moving blade adjustment angle action output value. After the changes in the expression contribution, the changes in the operating condition response intensity, and the changes in the adjustment direction are incorporated into the network connection weight update difference value, the training end judgment can be connected with the actual operating condition changes of the industrial large fan.

[0088] Preferably, the introduction of the network connection weight update difference value differs from the processing method that uses only a single point decrease in the global reward error as the basis for training termination. In the scenario of optimizing the operating parameters of industrial large fans, the global reward error may temporarily decrease due to a decrease in the single-unit operating power of a certain strategy training batch, but at the same time, the corresponding bearing vibration value or bearing temperature rise is still changing. The network connection weight update amounts of the action command mapping layer, the strategy optimization decision layer, and the feature perception fusion layer will still change significantly in subsequent strategy training batches. By continuing to calculate the network connection weight update difference value, it is possible to identify whether the reinforcement learning network model has formed a relatively stable network connection weight adjustment relationship between the energy consumption response state, the mechanical response state, and the thermal safety offset, thereby making the training termination judgment more in line with the scenario requirements of fan operating parameter optimization.

[0089] Preferably, the network connection weight update difference value is also used to distinguish between short-term operating condition disturbances and continuous network connection weight adjustment needs. When short-term inlet flow fluctuations or short-term flue gas inlet temperature fluctuations occur in the policy training batch, the global reward error will be affected by the corresponding Markov empirical tuple. However, if the batch weight update records corresponding to adjacent policy training batches change little, the network connection weight update difference value will be in a low range of change, indicating that the reinforcement learning network model does not need to make large network connection weight adjustments due to short-term operating condition disturbances. When consecutive policy training batches all reflect high-pressure low-flow operation conditions, high bearing vibration values, or high bearing temperature rise operation conditions, the batch weight update records corresponding to adjacent policy training batches will change continuously, and the network connection weight update difference value will continue to indicate that the reinforcement learning network model still needs to update the network connection weights.

[0090] Preferably, after each generation of the network connection weight update difference value, the network connection weight update difference value is compared with the model convergence threshold; if the network connection weight update difference value is greater than the model convergence threshold, the batch weight update record corresponding to the current policy training batch is retained, and subsequent policy training batches are received to continue updating the network connection weights of the action command mapping layer, the policy optimization decision layer, and the feature perception fusion layer; if the network connection weight update difference value is less than or equal to the model convergence threshold, the current network connection weights of the action command mapping layer, the policy optimization decision layer, and the feature perception fusion layer are fixed, and the reinforcement learning network model with fixed network connection weights is used as the ventilation fan operating parameter optimization decision model; the ventilation fan operating parameter optimization decision model is then used for forward inference based on the real-time collected ventilation fan state parameters.

[0091] Preferably, after the ventilation fan operating parameter optimization decision model is formed, the feature perception fusion layer continues to receive the real-time collected state parameters of the ventilation fan and transforms the real-time collected state parameters of the ventilation fan into a global operating condition state feature vector of the ventilation fan; the strategy optimization decision layer continues to generate action decision latent variables based on the global operating condition state feature vector of the ventilation fan; the action command mapping layer continues to map the action decision latent variables into ventilation fan operation adjustment parameters; since the ventilation fan operating parameter optimization decision model is formed after the network connection weight update difference value is less than or equal to the model convergence threshold, the ventilation fan operation adjustment parameters can inherit the state expression relationship, state feature mapping relationship and action mapping relationship that have entered a low range of change during the training phase, thereby outputting ventilation fan operation adjustment parameters that maximize the global reward expectation during the real-time operation phase.

[0092] Preferably, there is a continuous technical processing relationship between the network connection weight update difference value, the model convergence threshold, and the ventilation fan operating parameter optimization decision model: the network connection weight update difference value is calculated from the difference between the network connection weight update amounts corresponding to adjacent policy training batches; the model convergence threshold is used to determine whether the network connection weight update difference value has entered a low change range; and the ventilation fan operating parameter optimization decision model is formed by a reinforcement learning network model that satisfies the model convergence threshold. Through the above relationship, the training termination condition of the reinforcement learning network model is no longer limited to abstract training rounds or a single error value, but is directly related to the network connection weight update status of the action instruction mapping layer, the policy optimization decision layer, and the feature perception fusion layer under the operating conditions of industrial large fans.

[0093] Optionally, performing a backchain derivative on the global reward error to calculate the error gradient for network parameter updates includes: The value function is configured with positive excitation factors such as reducing bearing vibration, controlling bearing temperature rise to a safe range, and reducing single-unit operating power. With the goal of maximizing the expected overall return of the value function output, a gradient descent optimizer is used to calculate the local error gradient of the objective function relative to the connection weights of the action instruction mapping layer network. The local error gradient is sequentially passed in reverse to the policy optimization decision layer and the feature perception fusion layer to calculate the error gradient of the network parameter update between the layers respectively. Wherein, the objective function is the mean square loss function characterizing the temporal difference error; the temporal difference error is obtained by subtracting the predicted action value from the target action value; the predicted action value is the expected comprehensive return calculated based on the state parameters and the exploration action parameters; the target action value is calculated based on the decay term of the predicted action value.

[0094] Preferably, in this specific implementation, when configuring the value function, the state parameters, exploration action parameters, equipment operation feedback parameters after executing the exploration action parameters, and state parameters at adjacent operating times are first read from each Markov experience tuple in the strategy training batch. The bearing vibration value, bearing temperature rise, and single-unit operating power in the equipment operation feedback parameters after executing the exploration action parameters are used as the feedback input of the value function. The bearing vibration value in the feedback input is used to characterize the mechanical response state of the fan after the exploration action parameters are applied. The bearing temperature rise in the feedback input is used to characterize the thermal safety offset of the fan under high-temperature flue gas and variable load conditions. The single-unit operating power in the feedback input is used to characterize the energy consumption response state after the fan actuator is adjusted. The value function incorporates the mechanical response state, the thermal safety offset, and the energy consumption response state into the comprehensive expected return calculation, so that the comprehensive expected return is simultaneously affected by the positive excitation factors of reducing the bearing vibration value, controlling the bearing temperature rise to within the safe range, and reducing the single-unit operating power.

[0095] Preferably, the positive incentive factor is not an independent evaluation label, but an operational feedback constraint that changes together with the state parameter and the exploration action parameter. When the exploration action parameter reduces the single-machine operating power but causes the bearing vibration value to increase or the bearing temperature rise to increase, the value function will simultaneously read the decreasing trend of the single-machine operating power, the increasing trend of the bearing vibration value, and the increasing trend of the bearing temperature rise when calculating the expected comprehensive benefit. The expected comprehensive benefit reflects the combined effect of the exploration action parameter on the energy consumption response state, mechanical response state, and thermal safety offset. When the exploration action parameter reduces the single-machine operating power and the bearing vibration value and bearing temperature rise are within the safe range, the value function takes the equipment operation feedback parameter after executing the exploration action parameter as the feedback direction with a higher expected comprehensive benefit, so that the global reward error can be calculated around the common needs of reducing the single-machine operating power, reducing the bearing vibration value, and controlling the bearing temperature rise to within the safe range.

[0096] Preferably, the predicted action value is calculated by the reinforcement learning network model based on the state parameters and the exploration action parameters. The predicted action value is used to characterize the expected comprehensive return that the reinforcement learning network model can generate for the exploration action parameters under the current network connection weights. The target action value is calculated based on the device operation feedback parameters after executing the exploration action parameters, the state parameters at adjacent operation times, and the decay term of the predicted action value. The target action value is used to characterize the expected comprehensive return reference value of the exploration action parameters after being corrected by the device operation feedback parameters after executing the exploration action parameters. The predicted action value and the target action value continue to enter the objective function, which adopts a mean square loss function that characterizes the temporal difference error, so that the difference between the predicted action value and the target action value can continue to participate in the calculation of the global reward error.

[0097] Preferably, the temporal difference error is obtained by subtracting the predicted action value from the target action value. The temporal difference error is used to characterize the deviation between the predicted comprehensive expected return under the current network connection weight and the reference comprehensive expected return after correction by the device operation feedback parameters after executing the exploration action parameters. The mean square loss function squares the temporal difference error so that the deviation formed by the predicted action value being higher than the target action value and the deviation formed by the predicted action value being lower than the target action value can both be converted into error contributions in the same direction. The mean square loss function continues to summarize the error contributions corresponding to multiple Markov empirical tuples within the policy training batch to form the global reward error.

[0098] Preferably, when the goal is to maximize the expected comprehensive return of the value function output, the global reward error does not simply reflect the mathematical difference between the predicted action value and the target action value, but rather reflects the deviation of the exploration action parameters from the target action value in multiple directions, including reducing bearing vibration, controlling bearing temperature rise to a safe range, and reducing single-machine operating power. The gradient descent optimizer reads the global reward error and establishes a derivative relationship between the global reward error and the network connection weights of the action command mapping layer to calculate the local error gradient of the objective function relative to the network connection weights of the action command mapping layer. The local error gradient is used to characterize the direction and magnitude of the influence of the network connection weights of the action command mapping layer on the global reward error.

[0099] Preferably, the network connection weights of the action command mapping layer have a direct mapping relationship with the action output values ​​in the exploration action parameters. The action output values ​​include the action output values ​​corresponding to the operating frequency of the variable frequency motor and the action output values ​​corresponding to the adjustment angle of the moving blade. When the gradient descent optimizer calculates the local error gradient, it first determines the contribution relationship between the action output values ​​corresponding to the operating frequency of the variable frequency motor and the action output values ​​corresponding to the adjustment angle of the moving blade to the global reward error based on the influence direction of the exploration action parameters on the predicted action value. Then, it passes the contribution relationship to the network connection weights of the action command mapping layer to obtain the local error gradient. The local error gradient is further used to update the network connection weights of the action command mapping layer, so that the subsequent exploration action parameters output by the action command mapping layer are adjusted in the direction of reducing the global reward error.

[0100] Preferably, when the local error gradient is backpropagated to the strategy optimization decision layer, the local error gradient is transmitted to the working condition response intensity corresponding to the action decision latent variable through the network connection weight parameters between the action instruction mapping layer and the strategy optimization decision layer; the action decision latent variable is formed by the combination of working condition response intensities corresponding to multiple central kernel features. Therefore, the strategy optimization decision layer calculates the error gradient of the network parameter update corresponding to the strategy optimization decision layer based on the local error gradient; the error gradient of the network parameter update corresponding to the strategy optimization decision layer is used to characterize the direction and magnitude of the influence of the central kernel feature, the Euclidean distance, the activation width parameter, and the working condition response intensity on the global reward error.

[0101] Preferably, the error gradient of the network parameter update corresponding to the strategy optimization decision layer continues to be propagated backward along the generation source of the global operating condition state feature vector of the ventilator to the feature perception fusion layer; after receiving the error gradient of the network parameter update corresponding to the strategy optimization decision layer, the feature perception fusion layer calculates the error gradient of the network parameter update corresponding to the feature perception fusion layer; the error gradient of the network parameter update corresponding to the feature perception fusion layer is used to characterize the direction and magnitude of the influence of the operating condition features and the global operating condition state feature vector of the ventilator on the global reward error, so that the aerodynamic parameters, thermal parameters and equipment operation feedback parameters in the state parameters can participate in the parameter optimization iteration of the reinforcement learning network model through the error gradient of the network parameter update corresponding to the feature perception fusion layer.

[0102] Preferably, in the error gradient calculation process of the network parameter update corresponding to the feature perception fusion layer, the inlet flow rate and total pressure rise in the state parameters participate in the expression of aerodynamic load state through the aerodynamic parameters, the flue gas inlet temperature participates in the expression of thermal safety offset through the thermal parameters, and the bearing vibration value, the bearing temperature rise, and the single-unit operating power participate in the expression of mechanical response state and energy consumption response state through the equipment operation feedback parameters. The error gradient of the network parameter update corresponding to the feature perception fusion layer adjusts the network connection weight of the feature perception fusion layer according to the influence of the aerodynamic load state expression, the thermal safety offset expression, the mechanical response state, and the energy consumption response state on the global reward error, so that the subsequently formed global operating condition feature vector of the ventilation fan can better fit the operating state of the ventilation fan under high temperature, variable load, and equipment operation feedback parameter fluctuation conditions.

[0103] Preferably, in the operation scenarios of sintering main exhaust fan, desulfurization and denitrification booster fan, or coke oven gas dry quenching circulating fan, when the state parameters show an increase in inlet flow load occupancy, insufficient total pressure rise response, increased thermal drift of flue gas inlet temperature, and increased single-unit operating power consumption response, the value function will simultaneously read the single-unit operating power, the bearing vibration value, and the feedback changes corresponding to the bearing temperature rise to form the comprehensive expected return. The comprehensive expected return continues to participate in the calculation of the objective function and the global reward error. The global reward error, after being differentiated by a reverse chain, forms the error gradient for updating the network parameters, enabling the action command mapping layer, the strategy optimization decision layer, and the feature perception fusion layer to update the network connection weights around the action mapping relationship, the state feature mapping relationship, and the state expression relationship, respectively. The action mapping relationship is represented by the network connection weights of the action command mapping layer, the state feature mapping relationship is represented by the network connection weights of the strategy optimization decision layer, and the state expression relationship is represented by the network connection weights of the feature perception fusion layer.

[0104] Preferably, when the Markov empirical tuples in the strategy training batch indicate that the exploration action parameter causes the bearing vibration value to shift upward or the bearing temperature rise to approach the safe range boundary, the target action value will be corrected according to the attenuation term of the predicted action value based on the equipment operation feedback parameter after executing the exploration action parameter, so that the target action value is lower than the target action value formed when only considering the decrease in single-machine operating power; the temporal difference error increases accordingly, and the mean square loss function will incorporate the increased temporal difference error into the global reward error; the global reward error continues to form the error gradient of the network parameter update through back-chain differentiation, so that the action instruction mapping layer reduces the action output tendency corresponding to the exploration action parameter in subsequent training, so that the strategy optimization decision layer reduces the working condition response intensity corresponding to the equipment operation feedback parameter after executing the exploration action parameter, and so that the feature perception fusion layer enhances the expression contribution of the equipment operation feedback parameter in the working condition features.

[0105] Preferably, when the Markov empirical tuples in the strategy training batch indicate that the exploration action parameter reduces the single-machine operating power while keeping the bearing vibration value and bearing temperature rise within a safe range, the target action value will be corrected based on the equipment operation feedback parameter after executing the exploration action parameter, thereby correcting the attenuation term of the predicted action value. This ensures that the target action value reflects the higher overall expected return corresponding to the exploration action parameter. Consequently, the temporal difference error decreases, and the mean square loss function incorporates the reduced temporal difference error into the global reward error. The global reward error continues to form the error gradient for updating the network parameters through back-chain differentiation, enabling the action instruction mapping layer to retain the action output tendency corresponding to the exploration action parameter in subsequent training, enabling the strategy optimization decision layer to retain the operating condition response intensity corresponding to the equipment operation feedback parameter after executing the exploration action parameter, and enabling the feature perception fusion layer to retain the expression contribution of the equipment operation feedback parameter in the operating condition features.

[0106] Preferably, after the error gradient of the network parameter update is formed, the action command mapping layer updates the network connection weights of the action command mapping layer according to the local error gradient, the policy optimization decision layer updates the network connection weights of the policy optimization decision layer according to the error gradient of the network parameter update corresponding to the policy optimization decision layer, and the feature perception fusion layer updates the network connection weights of the feature perception fusion layer according to the error gradient of the network parameter update corresponding to the feature perception fusion layer. The updated network connection weights of the action command mapping layer, the updated network connection weights of the policy optimization decision layer, and the updated network connection weights of the feature perception fusion layer are continued to be used for forward inference and backward chain differentiation of subsequent policy training batches, so that the reinforcement learning network model can continuously adjust the network connection weights according to the actual operation feedback of the ventilator.

[0107] Preferably, the value function, the expected comprehensive return, the predicted action value, the target action value, the objective function, the temporal difference error, the global reward error, and the error gradient of the network parameter update have a continuous technical processing relationship: the value function forms the expected comprehensive return based on the bearing vibration value, the bearing temperature rise, and the single-unit operating power; the expected comprehensive return participates in the formation of the predicted action value and the target action value; the difference between the predicted action value and the target action value forms the temporal difference error; the objective function forms the global reward error based on the temporal difference error; and the global reward error, through backchain differentiation, forms the error gradient of the network parameter update. Through the above continuous technical processing relationship, the training process of the reinforcement learning network model can revolve around the operational needs of industrial large fans under high temperature, variable load, and equipment operation feedback parameter fluctuation conditions, rather than simply performing error backpropagation training of a general neural network.

[0108] Preferably, to facilitate the description of the calculation process of the value function, the i-th Markov empirical tuple in the policy training batch is represented as... Where i represents the Markov empirical tuple index in the training batch of the policy. This represents the state parameters before executing the aforementioned exploration action parameters. This represents the parameters of the exploration action. This represents the immediate comprehensive benefit expectation formed based on the equipment operation feedback parameters after executing the aforementioned exploration action parameters. These represent the state parameters at adjacent runtime moments. Further expressed as ,in, Indicates import flow. Indicates total pressure rise, Indicates the flue gas inlet temperature. Indicates the bearing vibration value. Indicates the temperature rise of the bearing bush. Indicates the single-machine operating power; the exploration action parameters Further expressed as ,in, Indicates the operating frequency of the variable frequency motor. This represents the adjustment angle of the moving blade. Through this representation, the source relationships between the state parameters, the exploration action parameters, and the equipment operation feedback parameters can remain consistent in the value function, the objective function, and the error gradient of the network parameter updates.

[0109] Preferably, before constructing the value function, the bearing vibration value, bearing temperature rise, and single-machine operating power after executing the exploration action parameters are first processed for dimensional consistency to avoid inconsistencies in dimensions caused by directly weighting different physical quantities. The bearing vibration value after executing the exploration action parameters can be expressed as... The bearing temperature rise after performing the exploration action parameters is expressed as: The single-machine operating power after executing the exploration action parameters is expressed as: The apostrophe in the upper right corner of the symbol indicates the feedback result after executing the parameters of the exploratory action at adjacent running times. Correspondingly, the positive excitation component of the bearing vibration value can be expressed as: The positive excitation component of the bearing temperature rise can be expressed as: The positive excitation component of the single-machine operating electrical power can be expressed as: ;in, and These represent the upper safety limit and lower reference limit corresponding to the bearing vibration value, respectively. and These represent the upper and lower safety limits for bearing temperature rise, respectively. and These represent the upper and lower reference limits for the operating power of a single unit, respectively. This indicates a small positive number used to avoid a denominator of zero. This represents a truncation function that restricts the calculation results to a preset range. Through the above dimensional coordination processing, the bearing vibration value, the bearing shell temperature rise, and the single-machine operating electrical power are all converted into dimensionless positive excitation components.

[0110] Preferably, the value function can be expressed as: ;in, This represents the expected instantaneous aggregate return corresponding to the i-th Markov experience tuple. This indicates the weight of the positive excitation component of the bearing vibration value in the expected instantaneous comprehensive return. This indicates the weight of the positive excitation component of the bearing temperature rise in the expected instantaneous overall return. This represents the weight of the positive excitation component of the single-machine operating electrical power in the expected instantaneous comprehensive return, and , and All weights are non-negative. The value function adopts the above form because optimizing the operating parameters of the ventilation fan is not solely aimed at reducing the single-unit operating power, but simultaneously reads the bearing vibration value, bearing temperature rise, and equipment operation feedback parameters corresponding to the single-unit operating power; when the exploratory action parameters result in a lower bearing vibration value, a bearing temperature rise within a safe range, and a lower single-unit operating power,... , and Together The trend is towards higher overall returns; when the single-unit operating power decreases but the bearing vibration value or bearing temperature rise increases, or This will reduce the expected immediate comprehensive benefit, thereby enabling the value function to express the constraints between the energy response state, the mechanical response state, and the thermal safety offset.

[0111] Preferably, the predicted action value under the current network connection weight is represented as ;in, This represents the predicted action value corresponding to the i-th Markov experience tuple. This indicates the reinforcement learning network model's position in the current network connection weights. The resulting action value mapping function This refers to the collective network connection weights of the action instruction mapping layer, the policy optimization decision layer, and the feature perception fusion layer. This represents the state parameters before executing the aforementioned exploration action parameters. This represents the parameters of the exploration action. The predicted action value... Used to express the current network connection weight Before the update based on the i-th Markov experience tuple is completed, the reinforcement learning network model's parameters for the exploration action... The expected value of the overall return that can be generated.

[0112] Preferably, the value of the target action is represented as ;in, This represents the target action value corresponding to the i-th Markov experience tuple. This represents the immediate overall expected return calculated based on the equipment operation feedback parameters after executing the aforementioned exploration action parameters. Indicates the attenuation coefficient. State parameters at adjacent running times and their corresponding actions The decay term of the predicted action value formed below, This indicates that the reinforcement learning network model is based on the state parameters at adjacent runtimes. Action parameters obtained from adjacent runtime moments through forward reasoning. The target action value. pass The device operation feedback parameters are received after the current exploration action parameters are executed, and then... By incorporating the influence of state parameters at adjacent operating times on subsequent adjustment benefits, the value of the target action includes both the current operating feedback and the subsequent benefit attenuation effect corresponding to the state parameters at adjacent operating times.

[0113] Preferably, the timing difference error can be expressed as: ;in, This represents the time-series difference error corresponding to the i-th Markov empirical tuple. Indicates the value of the target action. Indicates the predicted value of an action. When... When the absolute value is large, it indicates a significant deviation between the expected comprehensive revenue reference value after correction by the device operation feedback parameters and the predicted comprehensive revenue under the current network connection weight; when When the absolute value is small, it indicates that the predicted comprehensive revenue under the current network connection weight is relatively close to the reference value of the comprehensive revenue after the device operation feedback parameter correction. The time-series difference error Therefore, it can reflect the correction requirements of each Markov empirical tuple for the current reinforcement learning network model and continue to participate in the calculation of the objective function.

[0114] Preferably, the objective function is the mean square loss function characterizing the time-series difference error, which can be expressed as: ;in, Indicates the current network connection weight The objective function is given below, where N represents the number of Markov experience tuples in the training batch of the policy. This represents the squared term of the temporal difference error corresponding to the i-th Markov empirical tuple. The objective function... Simultaneously, the temporal difference errors corresponding to multiple Markov empirical tuples within the policy training batch are summarized, so that the global reward error can be expressed as: Where G represents the global reward error. Since By converting positive and negative deviations into error contributions in the same direction, the global reward error G can uniformly express the degree of training deviation in two cases: when the predicted action value is higher than the target action value and when the predicted action value is lower than the target action value.

[0115] Preferably, the action output relationship of the action command mapping layer can be represented as follows: ;in, This represents the k-th action output value of the exploration action parameter in the i-th Markov empirical tuple, where k represents the action output dimension. When k corresponds to the operating frequency of the frequency-controlled motor, This represents the output value corresponding to the operating frequency of the variable frequency motor. When k corresponds to the adjustment angle of the moving blade... This represents the action output value corresponding to the adjustment angle of the moving blade; j represents the node number in the strategy optimization decision layer, and m represents the total number of nodes in the strategy optimization decision layer. This represents the latent variable component of the action decision output by the i-th Markov empirical tuple at the j-th node of the policy optimization decision layer. This represents the network connection weight between the j-th node of the strategy optimization decision layer and the k-th action output dimension of the action instruction mapping layer. This represents the bias scalar corresponding to the k-th action output dimension of the action command mapping layer. Through this action output relationship, the action decision latent variable can be mapped to exploratory action parameters including the operating frequency of the variable frequency motor and the adjustment angle of the moving blades.

[0116] Preferably, when calculating the local error gradient of the objective function relative to the network connection weights of the action instruction mapping layer, the target action value can be... The reference value, corrected by the equipment operation feedback parameters, is read fixedly, and the value of the predicted action is determined accordingly. The derivative is calculated along the output path of the action command mapping layer. Correspondingly, the local error gradient can be expressed as... ;in, This represents the network connection weights in the action instruction mapping layer. The corresponding local error gradient, This indicates the sensitivity of the predicted action value to the output value of the k-th action. This indicates the source of contribution of the latent variable component of the action decision to the action output value. The formula shows that when a certain latent variable component of the action decision... When the action output value has a significant impact on the predicted action value and contributes significantly to the operating frequency of the variable frequency motor or the adjustment angle of the moving blade, the corresponding network connection weight in the action command mapping layer will obtain a large local error gradient, thereby participating in the subsequent network connection weight update.

[0117] Preferably, the network connection weights of the action command mapping layer can be determined according to... Update; among them, the left side This indicates the updated network connection weights, on the right. This indicates the network connection weights before the update. Indicates the learning rate. This represents the local error gradient. The update relationship does not directly adjust the network connection weights of the action command mapping layer based on whether the single-machine operating power decreases. Instead, it adjusts the network connection weights of the action command mapping layer based on the local error gradient obtained by the backchain derivative of the global reward error G. This ensures that the subsequent outputs of the variable frequency motor operating frequency and the moving blade adjustment angle are simultaneously influenced by the bearing vibration value, bearing temperature rise, and the equipment operation feedback parameters corresponding to the single-machine operating power.

[0118] Preferably, the latent variable components of action decision in the strategy optimization decision layer can be represented as follows: ;in, This represents the latent variable component of the action decision output by the i-th Markov empirical tuple at the j-th node of the policy optimization decision layer. This represents an exponential function with the natural constant as its base. This represents the global operating condition feature vector of the ventilation fan generated by the feature perception fusion layer. This represents the central kernel feature corresponding to the j-th node in the strategy optimization decision layer. This represents the activation width parameter corresponding to the j-th node in the strategy optimization decision layer. This represents the squared Euclidean distance between the global operating condition feature vector of the ventilation fan and the central kernel feature. This formula is used to express the degree of proximity between the current real-time operating state of the ventilation fan and different typical operating condition positions, and to enable the local error gradient to continue to be transmitted to the strategy optimization decision layer.

[0119] Preferably, when the local error gradient is backpropagated to the strategy optimization decision layer, the latent variable error propagation amount corresponding to the latent variable component of the action decision can be calculated first, expressed as follows: ;in, K represents the latent variable error propagation amount corresponding to the i-th Markov empirical tuple at the j-th node of the policy optimization decision layer, and K represents the total number of action output dimensions of the action instruction mapping layer. This indicates that the objective function outputs the value for the k-th action. Error propagation amount, This represents the network connection weights of the action command mapping layer. The error propagation amount... can be Indicates; among which, Indicates timing difference error. This indicates the sensitivity of the predicted action value to the action output value. Through the above calculation, the local error gradient can be transmitted to the latent variable components of the action decision along the network connection weights between the action instruction mapping layer and the policy optimization decision layer.

[0120] Preferably, the error gradient of the network parameter update corresponding to the strategy optimization decision layer can be further expressed as: ,as well as ;in, Indicates the central kernel feature The corresponding error gradient, Indicates the activation width parameter The corresponding error gradient, This represents the amount of error propagation by latent variables. Represent the latent variable components of action decision. This represents the feature vector of the overall operating condition of the ventilation fan. Indicates central kernel characteristics, This represents the activation width parameter. Using the above formula, the strategy optimization decision layer can adjust the central kernel feature and activation width parameter based on the global reward error, enabling the latent variable components of subsequent action decisions to form an appropriate operating condition response intensity under conditions of high temperature, variable load, and fluctuations in equipment operation feedback parameters.

[0121] Preferably, the error gradient of the network parameter update corresponding to the feature-aware fusion layer can be expressed as: ,in, This represents the error gradient of the network parameter update corresponding to the feature-aware fusion layer. This represents the network connection weights of the feature-aware fusion layer. This represents the amount of error propagation by latent variables. This indicates the sensitivity of the latent variable components of the action decision to the feature vector of the overall operating condition of the ventilation fan. This indicates the sensitivity of the global operating condition state feature vector of the ventilation fan to the network connection weights of the feature perception fusion layer. This formula shows that the update of the network connection weights of the feature perception fusion layer is not performed in isolation, but is the result of the global reward error G being transmitted through the action command mapping layer and the strategy optimization decision layer before being applied to the global operating condition state feature vector of the ventilation fan.

[0122] Preferably, the network connection weights of the feature-aware fusion layer can be determined according to... The central kernel features can be updated according to... The activation width parameter can be updated according to... Update; among them, Indicates the learning rate. This represents the error gradient of the network parameter update corresponding to the feature-aware fusion layer. This represents the error gradient corresponding to the central kernel feature. This represents the error gradient corresponding to the activation width parameter. Through the above update relationship, the network connection weights of the feature perception fusion layer, the strategy optimization decision layer, and the action instruction mapping layer can be updated hierarchically around the same global reward error G. This allows the influence of inlet flow rate, total pressure rise, flue gas inlet temperature, bearing vibration value, bearing temperature rise, and single-unit operating power on the ventilation fan's operating adjustment parameters to be expressed along the forward inference direction and participate in the error gradient calculation of network parameter updates along the backward chain derivative direction.

[0123] Preferably, the value function, the objective function, and the error gradient of the network parameter update can be expressed by the following continuous mathematical relationship: firstly, by... , and Forming immediate comprehensive expected return Then by Forming target action value , and by Forming predictive action value ; then by Formation of timing difference error ,Depend on Forming the objective function And the global reward error G, then by , , as well as Error gradients for updating network parameters are formed for the action command mapping layer, the strategy optimization decision layer, and the feature perception fusion layer, respectively. Through the above formula relationship, the training process of the reinforcement learning network model can transform the three positive excitation factors—reducing bearing vibration value, controlling bearing temperature rise to a safe range, and reducing single-machine operating power—into a mathematical processing process that is differentiable, can be backpropagated, and can update network connection weights.

[0124] Additionally, in the above formula, the subscript... The subscript is used to represent the parameter components corresponding to the bearing vibration value. Used to represent parameter components corresponding to bearing temperature rise, subscript Used to represent parameter components corresponding to the operating power of a single unit; superscript Used to indicate that the corresponding parameter belongs to the action instruction mapping layer, superscript Used to indicate that the corresponding parameter belongs to the central kernel feature, superscript This is used to indicate that the corresponding parameter belongs to the feature-aware fusion layer. Therefore, , All correspond to the bearing vibration values. , All correspond to the temperature rise of the bearing bush. , All correspond to the single-unit operating power; , and All correspond to the action instruction mapping layer. Corresponding to the central kernel features, and Each corresponds to a feature-aware fusion layer, enabling a one-to-one correspondence between the feedback parameters of different network layers and different devices at the symbol level.

[0125] Preferably, in the above formula, the subscript The subscript is used to represent the parameter components corresponding to the bearing vibration value. Used to represent parameter components corresponding to bearing temperature rise, subscript Used to represent parameter components corresponding to the operating power of a single unit; superscript Used to indicate that the corresponding parameter belongs to the action instruction mapping layer, superscript Used to indicate that the corresponding parameter belongs to the central kernel feature, superscript This is used to indicate that the corresponding parameter belongs to the feature-aware fusion layer. Therefore, , All correspond to the bearing vibration values. , All correspond to the temperature rise of the bearing bush. , All correspond to the single-unit operating power; , and All correspond to the action instruction mapping layer. Corresponding to the central kernel features, and Each corresponds to a feature-aware fusion layer, enabling a one-to-one correspondence between the feedback parameters of different network layers and different devices at the symbol level.

[0126] Preferably, , , , , as well as All of these are derived from equipment safety operation procedures, historical safety operation records of ventilation fans, or trial operation calibration records of ventilation fans; among them, This indicates the safe upper limit corresponding to the bearing vibration value. This indicates the lower reference limit corresponding to the bearing vibration value. This indicates the safe upper limit for the temperature rise of the bearing bush. This indicates the lower reference limit corresponding to the bearing temperature rise. This indicates the upper limit of the operating power corresponding to the single unit's operating power. This represents the lower reference limit corresponding to the single-unit operating electrical power. The aforementioned upper and lower limits are used to convert the bearing vibration value, bearing temperature rise, and single-unit operating electrical power into dimensionless positive excitation components, avoiding the direct inclusion of equipment operation feedback parameters with different dimensions into the same value function calculation.

[0127] Preferably, in the formula This indicates the summation of multiple Markov experience tuples or multiple action output dimensions within a policy training batch. This represents the differentiation operation. This represents an exponential function with the natural constant as its base. This represents the truncation function. This represents the squared Euclidean distance.

[0128] Optionally, a policy training batch is constructed by: evaluating the priority weight of each Markov empirical tuple based on the temporal difference algorithm, and extracting Markov empirical tuples with priority weights greater than the weight threshold to construct the policy training batch.

[0129] Optionally, the priority weights of each Markov empirical tuple are evaluated based on a temporal difference algorithm to extract Markov empirical tuples with priority weights greater than a weight threshold to construct policy training batches, including: Determine the target action value and predicted action value for each matching Markov empirical tuple and calculate the temporal difference error between the target action value and the predicted action value; Based on the absolute value of the time-series difference error, the importance sampling probability of each Markov empirical tuple is calculated, and the importance sampling probability is used as the priority weight of the corresponding Markov empirical tuple for sorting. Markov empirical tuples with priority weights greater than the weight threshold are selected to construct the policy training batch.

[0130] Preferably, in the specific implementation of constructing the strategy training batch, several Markov experience tuples are first read, and from each Markov experience tuple, state parameters, exploration action parameters, equipment operation feedback parameters after executing the exploration action parameters, and state parameters at adjacent operating times are read respectively; the state parameters are used to characterize the operating conditions of the fan before executing the exploration action parameters, the exploration action parameters are used to characterize the adjustment and control parameters of the fan actuator under the operating conditions, and the equipment operation feedback parameters after executing the exploration action parameters are used to characterize the bearing vibration value and bearing bearing value formed after the fan executes the exploration action parameters. Temperature rise and single-unit operating power changes, the state parameters at adjacent operating moments are used to characterize the subsequent operating conditions entered by the ventilator after executing the exploration action parameters; the state parameters, the exploration action parameters, the equipment operation feedback parameters after executing the exploration action parameters, and the state parameters at adjacent operating moments together form an empirical evaluation record, which continues to participate in the determination of predicted action value and target action value, so that each Markov empirical tuple is transformed into a training candidate that can be evaluated by the temporal difference algorithm, and the training candidate continues to participate in subsequent importance sampling probability calculation and strategy training batch construction.

[0131] Preferably, when determining the predicted action value for each matching Markov empirical tuple, the state parameters and exploration action parameters in the empirical evaluation record are input into the reinforcement learning network model under the current network connection weights to calculate the expected comprehensive return prediction value corresponding to the state parameters and the exploration action parameters, and the expected comprehensive return prediction value is used as the predicted action value. The predicted action value is used to express the expected comprehensive return prediction relationship of the energy consumption response state, mechanical response state, and thermal safety offset that the reinforcement learning network model can bring to the exploration action parameters under the current network connection weights. The expected comprehensive return prediction relationship is passed to the subsequent difference calculation process through the predicted action value, so that the predicted action value can serve as a calculation source for the time-series difference error value, and participate in the formation of the time-series difference error value together with the target action value.

[0132] Preferably, when determining the target action value for each Markov empirical tuple, the equipment operation feedback parameters after executing the exploratory action parameters and the state parameters at adjacent operating times are read from the empirical evaluation record. The instantaneous comprehensive benefit expectation corresponding to the exploratory action parameter is calculated based on the equipment operation feedback parameters after executing the exploratory action parameters. The instantaneous comprehensive benefit expectation simultaneously reads the bearing vibration value, bearing temperature rise, and single-unit operating power, so that the instantaneous comprehensive benefit expectation can reflect the feedback changes of the fan in mechanical response state, thermal safety offset, and energy consumption response state after executing the exploratory action parameters. The state parameters at adjacent operating times are further used to determine the attenuation term of the predicted action value under subsequent operating conditions. The instantaneous comprehensive benefit expectation and the attenuation term of the predicted action value together form the target action value, thus ensuring that the target action value includes both the equipment operation feedback parameters after executing the exploratory action parameters and the influence of the state parameters at adjacent operating times on subsequent adjustment benefits.

[0133] Preferably, when calculating the temporal difference error between the target action value and the predicted action value, the target action value is used as a reference value for the expected comprehensive return after correction of the device operation feedback parameters after executing the exploration action parameters, and the predicted action value is used as the predicted comprehensive return under the current network connection weights. The difference between the reference value and the predicted comprehensive return is calculated to obtain the temporal difference error value. The temporal difference error value characterizes the degree of deviation between the device operation feedback parameters carried by the Markov empirical tuple after executing the exploration action parameters and the existing prediction results of the current reinforcement learning network model. When the temporal difference error value is large, it indicates that the state parameters, exploration action parameters, device operation feedback parameters after executing the exploration action parameters, and state parameters at adjacent running times corresponding to the Markov empirical tuple have high correction value for the current reinforcement learning network model. When the temporal difference error value is small, it indicates that the device operation feedback parameters corresponding to the Markov empirical tuple after executing the exploration action parameters are already close to the prediction results of the current reinforcement learning network model.

[0134] Preferably, when calculating the importance sampling probability of each Markov empirical tuple based on the absolute value of the temporal difference error, the tuples are not directly extracted according to their acquisition order. Instead, the temporal difference error value corresponding to each Markov empirical tuple is converted into a non-negative error amplitude, and the importance sampling probability of the corresponding Markov empirical tuple is determined based on the non-negative error amplitude. The non-negative error amplitude can simultaneously cover the deviation caused by the predicted action value being higher than the target action value and the deviation caused by the predicted action value being lower than the target action value. This allows the importance sampling probability to be allocated according to the device operation feedback parameters after executing the exploration action parameters to correct the current reinforcement learning network model. The importance sampling probability continues to serve as the priority weight of the corresponding Markov empirical tuple, so that the priority weight can directly bear the evaluation result of the temporal difference error value.

[0135] Preferably, when using the importance sampling probability as the priority weight for sorting the corresponding Markov empirical tuples, a correspondence between the Markov empirical tuple and the priority weight is first established for each Markov empirical tuple. Then, the Markov empirical tuples are sorted in descending order of the priority weight to form a priority empirical ranking result. The ranking position in the priority empirical ranking result is used to express the parameter optimization value of the corresponding Markov empirical tuple for the current reinforcement learning network model. The Markov empirical tuples ranked higher usually correspond to larger temporal difference error values, indicating that the state parameters, exploration action parameters, device operation feedback parameters after executing the exploration action parameters, and state parameters at adjacent operation times recorded in the Markov empirical tuple are more likely to cause network connection weight correction in the reinforcement learning network model. The priority empirical ranking result continues to participate in weight threshold screening to form a policy training batch.

[0136] Preferably, when screening Markov empirical tuples with priority weights greater than the weight threshold, the weight threshold is a screening criterion pre-configured based on the strategy training batch capacity and the industrial large wind turbine operating condition coverage requirements. The industrial large wind turbine operating condition coverage requirements are used to limit the high-temperature flue gas disturbance, load change, bearing vibration value change, bearing temperature rise change, and single-unit operating power change that need to be covered in the strategy training batch. The weight threshold is used to remove Markov empirical tuples with low correction effect on the current reinforcement learning network model from the priority empirical ranking results, and retain Markov empirical tuples with high correction effect on the current reinforcement learning network model. The Markov empirical tuples with priority weights greater than the weight threshold continue to enter the strategy training batch according to their ranking position in the priority empirical ranking results, so that the strategy training batch can preferentially include Markov empirical tuples whose high-temperature flue gas disturbance, load change, bearing vibration value change, bearing temperature rise change, and single-unit operating power change can better reflect the current training requirements.

[0137] Preferably, in the operation scenarios of sintering main exhaust fan, desulfurization and denitrification booster fan, or coke oven gas dry quenching circulating fan, when a Markov empirical tuple indicates that the exploration action parameter reduces the single-machine operating power but causes the bearing vibration value to increase or the bearing temperature rise to increase, the target action value will deviate from the predicted action value due to the equipment operation feedback parameter after executing the exploration action parameter, thus forming a large temporal difference error value. The large temporal difference error value will cause the corresponding Markov empirical tuple to obtain a higher importance sampling probability and place the corresponding Markov empirical tuple in a higher position in the priority empirical ranking result. When the priority weight of the corresponding Markov empirical tuple is greater than the weight threshold, the corresponding Markov empirical tuple enters the policy training batch, so that the subsequent reinforcement learning network model update can read the technical correlation between the exploration action parameter and the increase in equipment operation feedback parameter.

[0138] Preferably, when a Markov empirical tuple indicates that the exploration action parameter reduces the single-machine operating power while meeting the airflow transport requirements, and keeps the bearing vibration value and bearing temperature rise within the safe range allowed by the equipment safety operation procedures, the target action value and the predicted action value will form a time-series difference error value that reflects the joint satisfaction of energy-saving operation and equipment operation safety; the importance sampling probability corresponding to the time-series difference error value participates in the formation of the priority experience ranking result, so that the Markov empirical tuple can enter the priority experience ranking result according to its contribution relationship to the calculation of the comprehensive benefit expectation; when the priority weight of the Markov empirical tuple is greater than the weight threshold, the Markov empirical tuple enters the policy training batch, so that the subsequent reinforcement learning network model update can retain the correspondence between the exploration action parameter and the lower single-machine operating power, the lower bearing vibration value, and the controlled bearing temperature rise.

[0139] Preferably, the policy training batch does not simply collect all Markov empirical tuples, but is formed by orderly filtering Markov empirical tuples based on the temporal difference error value, the importance sampling probability, the priority weight, and the weight threshold. The temporal difference error value is used to evaluate the degree of correction of the current reinforcement learning network model by the Markov empirical tuples. The importance sampling probability is used to convert the temporal difference error value into a sortable extraction criterion. The priority weight is used to take the importance sampling probability and participate in the sorting. The weight threshold is used to determine the range of Markov empirical tuples entering the policy training batch. Through the above processing, the policy training batch can more centrally include Markov empirical tuples that have training value for optimizing the operating parameters of the ventilation fan, rather than being occupied by Markov empirical tuples with low change and low correction requirements.

[0140] Preferably, after the policy training batch is formed, it continues to be used to calculate the global reward error, and the error gradient for updating network parameters is formed in reverse through the global reward error. Since the Markov experience tuples in the policy training batch have been filtered by the priority weights, the global reward error can better reflect the changes in equipment operation feedback parameters after executing the exploration action parameters corresponding to larger temporal difference error values. This makes the network connection weight updates of the action instruction mapping layer, the policy optimization decision layer, and the feature perception fusion layer more in line with the training needs of industrial large wind turbines under high temperature, variable load, and equipment operation feedback parameter fluctuation conditions. The policy training batch thus forms a clear technical connection between the construction of Markov experience tuples and the updating of reinforcement learning network models.

[0141] Preferably, the Markov experience tuple, the experience evaluation record, the predicted action value, the target action value, the temporal difference error value, the importance sampling probability, the priority weight, the priority experience ranking result, the weight threshold, and the policy training batch have a continuous technical processing relationship: the Markov experience tuple provides state parameters, exploration action parameters, equipment operation feedback parameters after executing the exploration action parameters, and state parameters at adjacent operation times. The state parameters, the exploration action parameters, the equipment operation feedback parameters after executing the exploration action parameters, and the state parameters at adjacent operation times together form... The experience evaluation record contains state parameters and exploration action parameters used to form the predicted action value. The equipment operation feedback parameters after executing the exploration action parameters and the state parameters at adjacent operation times in the experience evaluation record are used to form the target action value. The predicted action value and the target action value form the temporal difference error value. The temporal difference error value forms the importance sampling probability. The importance sampling probability is used as the priority weight to participate in the sorting to form the priority experience sorting result. In the priority experience sorting result, Markov experience tuples with priority weights greater than the weight threshold constitute the policy training batch.

[0142] Optionally, the state parameters of the ventilator are obtained, including: extracting the inlet flow rate and total pressure rise of the ventilator to construct the aerodynamic parameters, extracting the flue gas inlet temperature to construct the thermal parameters, and extracting the bearing vibration value, bearing temperature rise, and single-unit operating power to construct the equipment operation feedback parameters; the adjustment and control parameters of the ventilator actuator in the exploration action parameters include the variable frequency motor operating frequency and the moving blade adjustment angle.

[0143] Preferably, in the specific implementation of acquiring the state parameters of the ventilation fan, the inlet flow rate, total pressure rise, flue gas inlet temperature, bearing vibration value, bearing temperature rise, and single-unit operating power are first read from the measurement signals at the ventilation fan operation site. The inlet flow rate, total pressure rise, flue gas inlet temperature, bearing vibration value, bearing temperature rise, and single-unit operating power are then aligned under the same acquisition sequence to form a ventilation fan state acquisition record. This ventilation fan state acquisition record, as the source record before state parameter grouping, is further grouped according to aerodynamic parameters, thermal parameters, and equipment operation feedback parameters. This ensures that the state parameters are not composed of a single pressure or a single flow rate, but simultaneously contain multi-source operating information that can express aerodynamic load status, flue gas inlet temperature thermal drift status, and equipment operation feedback status. The state parameters are then input into the reinforcement learning network model, enabling the reinforcement learning network model to perform forward inference based on the multi-source operating information at the same operating time.

[0144] Preferably, when constructing the aerodynamic parameters, the inlet flow rate and the total pressure rise are used as the source content for constructing the aerodynamic parameters; the inlet flow rate is used to characterize the airflow transport load undertaken by the fan at the corresponding operating time, and the total pressure rise is used to characterize the pressure rise capability of the fan to the airflow at the corresponding operating time. The airflow transport load and the pressure rise capability together reflect the aerodynamic load state of the fan; after the inlet flow rate and the total pressure rise together constitute the aerodynamic parameters, the aerodynamic parameters can reflect the high flow rate and low pressure rise operating conditions, the high pressure rise and low flow rate operating conditions, and the operating conditions where the flow rate and pressure rise fluctuate simultaneously; the aerodynamic parameters continue to enter the feature perception fusion layer, so that the feature perception fusion layer can extract the operating condition features characterizing the aerodynamic load state, and make the operating condition features further participate in the generation of the global operating condition state feature vector of the fan.

[0145] Preferably, when constructing the thermal parameters, the flue gas inlet temperature is used as the source content for constructing the thermal parameters; the flue gas inlet temperature is used to characterize the inlet flue gas heat load borne by the fan in the operating scenarios of sintering main exhaust fan, desulfurization and denitrification booster fan or coke oven gas dry quenching circulating fan, and the inlet flue gas heat load is used to reflect the thermal drift state of the flue gas inlet temperature formed on the inlet side of the fan; after the flue gas inlet temperature constitutes the thermal parameters, the thermal parameters can be used together with the aerodynamic parameters to express the influence of flow rate and pressure rise changes on the operating state of the fan under high temperature flue gas conditions; the thermal parameters continue to enter the feature perception fusion layer, so that the feature perception fusion layer can extract the operating condition features characterizing the thermal drift state of the flue gas inlet temperature, and make the operating condition features further participate in the generation of the global operating condition feature vector of the fan.

[0146] Preferably, when constructing the equipment operation feedback parameters, the bearing vibration value, the bearing shell temperature rise, and the single-unit operating power are used as the source content for constructing the equipment operation feedback parameters; the bearing vibration value is used to characterize the mechanical response state of the fan rotor, bearing housing, and transmission components during operation; the bearing shell temperature rise is used to characterize the thermal safety offset formed by the fan support under high-temperature flue gas and variable load conditions; and the single-unit operating power is used to characterize the energy consumption response state formed after the fan actuator is adjusted; after the bearing vibration value, bearing shell temperature rise, and single-unit operating power jointly constitute the equipment operation feedback parameters, the equipment operation feedback parameters can reflect the combined influence of actuator adjustment on the mechanical response state, thermal safety offset, and energy consumption response state; the equipment operation feedback parameters continue to enter the feature perception fusion layer, so that the feature perception fusion layer can extract the operating condition features characterizing the equipment operation feedback state, which includes the mechanical response state, the thermal safety offset, and the energy consumption response state.

[0147] Preferably, after the aerodynamic parameters, thermal parameters, and equipment operation feedback parameters form the state parameters, the state parameters are not directly used as single values ​​in control judgment. Instead, the feature perception fusion layer first performs parameter grouping and scale coordination processing. The parameter grouping is used to maintain the source relationship of the inlet flow rate, total pressure rise, flue gas inlet temperature, bearing vibration value, bearing temperature rise, and single-unit operating power in the state parameters. The scale coordination processing is used to ensure that the aerodynamic parameters, thermal parameters, and equipment operation feedback parameters are within the expression range that can be processed by state feature mapping. The state parameters after parameter grouping and scale coordination processing continue to participate in the operating condition feature extraction to form a global operating condition state feature vector of the ventilation fan that can be transmitted to the strategy optimization decision layer.

[0148] Preferably, in the process of forming the global operating condition feature vector of the ventilation fan, the inlet flow corresponds to the inlet flow load occupancy level, the total pressure rise corresponds to the total pressure rise response level, the flue gas inlet temperature corresponds to the flue gas inlet temperature thermal drift level, the bearing vibration value corresponds to the bearing vibration value mechanical disturbance level, the bearing bush temperature rise corresponds to the bearing bush temperature rise thermal safety offset level, and the single-unit operating power corresponds to the single-unit operating power energy consumption response level. The inlet flow load occupancy level, the total pressure rise response level, the flue gas inlet temperature thermal drift level, the bearing vibration value mechanical disturbance level, the bearing bush temperature rise thermal safety offset level, and the single-unit operating power energy consumption response level are arranged in the same dimension to form the global operating condition feature vector of the ventilation fan. The global operating condition feature vector of the ventilation fan is further passed to the strategy optimization decision layer to participate in the neighborhood identification processing between the central kernel feature and the global operating condition feature vector of the ventilation fan, and to participate in the calculation of the latent variables of the action decision.

[0149] Preferably, in the operation scenario of industrial large fans, the construction method of the state parameters differs from the single monitoring method that only collects inlet and outlet pressures or only collects motor power; when the inlet flow rate increases but the total pressure rise response is insufficient, the aerodynamic parameters can express the high flow rate and low pressure rise operating condition; when the flue gas inlet temperature increases and the bearing temperature rises, the thermal parameters and the equipment operation feedback parameters can jointly express the influence of high temperature flue gas on the thermal safety deviation; when the single-unit operating power decreases but the bearing vibration value increases, the equipment operation feedback parameters can express the constraint relationship between the energy consumption response state and the mechanical response state; after the state parameters enter the reinforcement learning network model, the generation of the exploration action parameters can simultaneously read the aerodynamic load state, the flue gas inlet temperature thermal drift state, the mechanical response state, and the energy consumption response state.

[0150] Preferably, the adjustment and control parameters of the fan actuator in the exploration action parameters include the variable frequency motor operating frequency and the moving blade adjustment angle; the variable frequency motor operating frequency is used to characterize the adjustment amount of the fan drive side to the speed output, and the moving blade adjustment angle is used to characterize the adjustment amount of the fan aerodynamic flow channel side to the blade opening; after the variable frequency motor operating frequency and the moving blade adjustment angle together constitute the exploration action parameters, the exploration action parameters can act on both the fan drive side and the aerodynamic flow channel side; the exploration action parameters and the state parameters together generate the Markov empirical tuple, so that the Markov empirical tuple can record the equipment operation feedback parameters formed after executing the variable frequency motor operating frequency and the moving blade adjustment angle under a certain state parameter.

[0151] Preferably, when the action command mapping layer outputs the exploration action parameters, the action command mapping layer does not combine the operating frequency of the variable frequency motor and the adjustment angle of the moving blade into a single action quantity. Instead, it forms corresponding action output values ​​according to the adjustment and control parameters of the fan actuator. The action output value corresponding to the operating frequency of the variable frequency motor is used to participate in the adjustment and control of the fan drive side, and the action output value corresponding to the adjustment angle of the moving blade is used to participate in the adjustment and control of the fan aerodynamic flow channel side. The action output value corresponding to the operating frequency of the variable frequency motor and the action output value corresponding to the adjustment angle of the moving blade together constitute the exploration action parameters. The exploration action parameters continue to participate in the construction of the Markov empirical tuple together with the state parameters, so that subsequent strategy training batches can evaluate the impact of the adjustment and control parameters of different fan actuators on the bearing vibration value, the bearing temperature rise, and the single-unit operating power.

[0152] Preferably, in the operation scenarios of sintering main exhaust fan, desulfurization and denitrification booster fan, or coke oven gas dry quenching circulating fan, when the state parameters indicate that the inlet flow rate is at a high load, the total pressure rise response is insufficient, the flue gas inlet temperature is rising, and the single-unit operating power is rising, the reinforcement learning network model can generate a corresponding global operating condition state feature vector of the ventilation fan based on the aerodynamic parameters, the thermal parameters, and the equipment operation feedback parameters; the strategy optimization decision layer forms action decision latent variables based on the global operating condition state feature vector of the ventilation fan, and the action command mapping layer outputs exploratory action parameters containing the operating frequency of the variable frequency motor and the adjustment angle of the moving blades based on the action decision latent variables; the bearing vibration value, the bearing temperature rise, and the single-unit operating power generated after the exploratory action parameters are executed continue to constitute the equipment operation feedback parameters, so that there is a continuous training data source relationship between the state parameters, the exploratory action parameters, and the equipment operation feedback parameters.

[0153] Preferably, when the state parameters indicate that the bearing vibration value has increased or the bearing temperature rise is close to the boundary of the safe range, the equipment operation feedback parameters not only enter the reinforcement learning network model as part of the state parameters, but also participate in the calculation of the comprehensive expected return in the subsequent value function calculation; after the bearing vibration value and the bearing temperature rise enter the global operating condition feature vector of the ventilator through the equipment operation feedback parameters, they can affect the action decision latent variables formed by the strategy optimization decision layer; the action decision latent variables continue to affect the variable frequency motor operating frequency and the moving blade adjustment angle output by the action command mapping layer, so that the process of generating the exploration action parameters can read the equipment operation feedback parameters related to the equipment operation safety, instead of adjusting only according to the deviation of the inlet flow rate or the total pressure rise.

[0154] Preferably, the state parameters, aerodynamic parameters, thermal parameters, equipment operation feedback parameters, global operating condition state feature vector of the ventilation fan, exploratory action parameters, and Markov experience tuples have a continuous technical processing relationship: the inlet flow rate and the total pressure rise constitute the aerodynamic parameters, the flue gas inlet temperature constitutes the thermal parameters, the bearing vibration value, the bearing temperature rise, and the single-unit operating power constitute the equipment operation feedback parameters, and the aerodynamic parameters, thermal parameters, and equipment operation feedback parameters together constitute the state parameters; after the state parameters enter the feature perception fusion layer, they form the global operating condition state feature vector of the ventilation fan, and the global operating condition state feature vector of the ventilation fan continues to generate the exploratory action parameters through the strategy optimization decision layer and the action command mapping layer, and the exploratory action parameters and the state parameters jointly participate in the construction of the Markov experience tuples.

[0155] Optionally, after outputting the fan operation adjustment parameters that maximize the expected global reward, the method further includes: pre-setting hard constraints on safety boundaries based on the equipment safety operation procedures; performing safety boundary verification on the fan operation adjustment parameters obtained by forward reasoning of the fan operation parameter optimization decision model; when it is predicted that executing the fan operation adjustment parameters will trigger the hard constraints on safety boundaries, truncating the output of the fan operation adjustment parameters and enabling the fan safety back-off operation mechanism; when it is predicted that executing the fan operation adjustment parameters will not trigger the hard constraints on safety boundaries, using the fan operation adjustment parameters as the output.

[0156] Preferably, after outputting the fan operation adjustment parameters that maximize the global reward expectation, the equipment safety operation procedure is first read, and the safety operation boundary content related to the fan operation adjustment parameters is extracted from the equipment safety operation procedure. The safety operation boundary content includes the drive-side operation boundary corresponding to the variable frequency motor operating frequency, the aerodynamic flow channel-side operation boundary corresponding to the moving blade adjustment angle, the mechanical response boundary corresponding to the bearing vibration value, the thermal safety boundary corresponding to the bearing temperature rise, and the energy consumption operation boundary corresponding to the single unit operating power. The safety operation boundary content is constrained and organized to form the safety boundary hard constraint conditions. The safety boundary hard constraint conditions are used to constrain the fan operation adjustment parameters themselves, as well as to constrain the equipment operation feedback parameters formed after executing the fan operation adjustment parameters.

[0157] Preferably, the hard constraint condition of the safety boundary is not a fixed boundary set for a single fan actuator, but corresponds to the technical relationship between the state parameters, the fan operation adjustment parameters, and the equipment operation feedback parameters. The aerodynamic parameters in the state parameters are used to characterize the current aerodynamic load state of the fan, the thermal parameters in the state parameters are used to characterize the current thermal drift state of the flue gas inlet temperature of the fan, and the equipment operation feedback parameters in the state parameters are used to characterize the current mechanical response state, the current thermal safety offset, and the current energy consumption response state of the fan. After reading the state parameters and the fan operation adjustment parameters, the safety boundary verification process performs a safety boundary verification on the execution risk of the fan operation adjustment parameters corresponding to the current operation state represented by the state parameters, based on the current aerodynamic load state, the current thermal drift state of the flue gas inlet temperature, the current mechanical response state, the current thermal safety offset, and the current energy consumption response state of the fan. This ensures that the execution risk of the fan operation adjustment parameters corresponds to the current operating state represented by the state parameters.

[0158] Preferably, when performing safety boundary verification on the ventilation operating adjustment parameters obtained by forward reasoning of the ventilation operating parameter optimization decision model, the safety boundary verification process first decomposes the ventilation operating adjustment parameters into adjustment parameters corresponding to the operating frequency of the variable frequency motor and adjustment parameters corresponding to the adjustment angle of the moving blades; then, the adjustment parameters corresponding to the operating frequency of the variable frequency motor are compared with the driving-side operating boundary, and the adjustment parameters corresponding to the adjustment angle of the moving blades are compared with the aerodynamic flow channel-side operating boundary to form adjustment parameter boundary verification results; the adjustment parameter boundary verification results continue to participate in the verification of predicted operating feedback parameters, so that the safety boundary verification process verifies both the adjustment control parameters of the ventilation actuator itself and the impact of executing the ventilation operating adjustment parameters on the ventilation operating state.

[0159] Preferably, after forming the boundary verification result of the adjustment parameters, the safety boundary verification process uses the state parameters and the fan operation adjustment parameters together as prediction input to predict the predicted operation feedback parameters after executing the fan operation adjustment parameters. The predicted operation feedback parameters include predicted bearing vibration value, predicted bearing temperature rise, and predicted single-unit operating power. The predicted bearing vibration value is used to characterize the change in mechanical response state after executing the fan operation adjustment parameters, the predicted bearing temperature rise is used to characterize the change in thermal safety offset after executing the fan operation adjustment parameters, and the predicted single-unit operating power is used to characterize the change in energy consumption response state after executing the fan operation adjustment parameters. The predicted operation feedback parameters are then compared with the safety boundary hard constraint conditions to form the boundary verification result of the predicted operation feedback parameters.

[0160] Preferably, the predicted operation feedback parameter boundary verification result is used to determine whether executing the fan operation adjustment parameters triggers the safety boundary hard constraint condition; when the predicted bearing vibration value exceeds the mechanical response boundary, or the predicted bearing temperature rise exceeds the thermal safety boundary, or the predicted single-unit operating power exceeds the energy consumption operation boundary, the predicted operation feedback parameter boundary verification result indicates that executing the fan operation adjustment parameters will trigger the safety boundary hard constraint condition; when the predicted bearing vibration value does not exceed the mechanical response boundary, the predicted bearing temperature rise does not exceed the thermal safety boundary, and the predicted single-unit operating power does not exceed the energy consumption operation boundary, the predicted operation feedback parameter boundary verification result indicates that executing the fan operation adjustment parameters does not trigger the safety boundary hard constraint condition.

[0161] Preferably, when it is predicted that executing the fan operation adjustment parameters will trigger the safety boundary hard constraint condition, the output of the fan operation adjustment parameters is first truncated to prevent the fan operation adjustment parameters from entering the fan actuator. Truncating the output of the fan operation adjustment parameters does not discard all control information, but rather transfers the fan operation adjustment parameters to the safety backoff determination process. The safety backoff determination process reads the state parameters, the adjustment parameter boundary verification result, and the predicted operation feedback parameter boundary verification result, and determines the safety backoff operation parameters according to the safety boundary hard constraint condition. The safety backoff operation parameters continue to participate in the fan safety backoff operation mechanism as safety boundary verification output parameters, enabling the fan actuator to enter an operating state matching the equipment safety operation procedure when not executing the fan operation adjustment parameters that trigger the safety boundary hard constraint condition.

[0162] Preferably, the fan safety reversal operation mechanism includes variable frequency motor operating frequency safety reversal adjustment and moving blade adjustment angle safety reversal adjustment; when the predicted bearing vibration value exceeds the mechanical response boundary, the safety reversal determination process determines safety reversal operation parameters to reduce mechanical response state disturbance based on the bearing vibration value in the state parameters, the predicted bearing vibration value, and the fan operating adjustment parameters; when the predicted bearing temperature rise exceeds the thermal safety boundary, the safety reversal determination process determines safety reversal operation parameters to reduce thermal safety offset based on the flue gas inlet temperature in the state parameters, the bearing temperature rise in the state parameters, the predicted bearing temperature rise, and the fan operating adjustment parameters; the safety reversal operation parameters then replace the truncated fan operating adjustment parameters in the fan actuator adjustment, and act on the fan actuator through the variable frequency motor operating frequency safety reversal adjustment and the moving blade adjustment angle safety reversal adjustment.

[0163] Preferably, in the operation scenarios of sintering main exhaust fan, desulfurization and denitrification booster fan, or coke oven gas dry quenching circulating fan, when the fan operation adjustment parameters indicate that the variable frequency motor operating frequency is increased and the moving blade adjustment angle is increased, while the state parameters simultaneously indicate that the flue gas inlet temperature is rising, the bearing temperature rise is approaching the thermal safety boundary, and the bearing vibration value is shifting upward, the safety boundary verification process will form the predicted operation feedback parameters based on the state parameters and the fan operation adjustment parameters. If the predicted bearing temperature rise in the predicted operation feedback parameters exceeds the thermal safety boundary, or the predicted bearing vibration value in the predicted operation feedback parameters exceeds the mechanical response boundary, the output of the fan operation adjustment parameters is cut off, and the fan safety retreat operation mechanism is activated. Thus, the fan operation adjustment parameters output by the fan operation parameter optimization decision model can be constrained again by the safety boundary hard constraint conditions before entering the fan actuator.

[0164] Preferably, when it is predicted that the execution of the fan operation adjustment parameters will not trigger the safety boundary hard constraint condition, the safety boundary verification process uses the fan operation adjustment parameters as the safety boundary verification output parameters. After the fan operation adjustment parameters are used as the safety boundary verification output parameters, the adjustment parameters corresponding to the variable frequency motor operating frequency are entered into the fan actuator corresponding to the variable frequency motor operating frequency, and the adjustment parameters corresponding to the moving blade adjustment angle are entered into the fan actuator corresponding to the moving blade adjustment angle. After the fan actuator executes the fan operation adjustment parameters, it generates updated bearing vibration values, updated bearing temperature rise, and updated single-unit operating power. The updated bearing vibration values, updated bearing temperature rise, and updated single-unit operating power continue to constitute updated equipment operation feedback parameters. The updated equipment operation feedback parameters and the fan operation adjustment parameters used as safety boundary verification output parameters jointly participate in the subsequent construction of Markov empirical tuples.

[0165] Preferably, the safety boundary hard constraint and the global reward expectation maximization are not mutually exclusive, but rather act sequentially on different processing stages of the fan operation adjustment parameters; the fan operation parameter optimization decision model performs forward reasoning based on the real-time collected fan state parameters to output the fan operation adjustment parameters that maximize the global reward expectation; the safety boundary verification process continues to read the fan operation adjustment parameters and determines whether executing the fan operation adjustment parameters will cause the equipment operation feedback parameters to exceed the corresponding boundary based on the safety boundary hard constraint; therefore, the fan operation adjustment parameters are subject to both the global reward expectation maximization result of the reinforcement learning network model and the safety boundary hard constraint corresponding to the equipment safety operation procedure.

[0166] Preferably, the safety boundary verification process also writes the verification result of triggering the safety boundary hard constraint condition into the safety verification feedback record. The verification result of triggering the safety boundary hard constraint condition is used to characterize the correspondence between the boundary verification result of the predicted operation feedback parameter and the safety boundary hard constraint condition. The safety verification feedback record includes the truncated fan operation adjustment parameters, the predicted operation feedback parameters that trigger the safety boundary hard constraint condition, the corresponding state parameters, and the safety rollback operation parameters. The safety verification feedback record continues to participate in the subsequent Markov experience tuple construction together with the corresponding state parameters, the safety boundary verification output parameters, and the equipment operation feedback parameters formed after executing the safety boundary verification output parameters. This enables subsequent policy training batches to read the safety verification feedback record and reduces the output of fan operation adjustment parameters that are close to the safety boundary hard constraint condition in subsequent training.

[0167] Preferably, the safety boundary hard constraint conditions, the safety boundary verification process, the predicted operation feedback parameters, the predicted operation feedback parameter boundary verification results, the safety backoff operation parameters, the safety boundary verification output parameters, and the ventilation fan safety backoff operation mechanism have a continuous technical processing relationship: the equipment safety operation procedure is used to form the safety boundary hard constraint conditions, the safety boundary hard constraint conditions are used to perform safety boundary verification on the ventilation fan operation adjustment parameters obtained by forward reasoning of the ventilation fan operation parameter optimization decision model, the safety boundary verification process forms the predicted operation feedback parameters based on the state parameters and the ventilation fan operation adjustment parameters, and the predicted operation feedback parameters are... The operational feedback parameters are used to form the predicted operational feedback parameter boundary verification result. The predicted operational feedback parameter boundary verification result is used to determine whether the execution of the fan operation adjustment parameters triggers the safety boundary hard constraint condition. When the execution of the fan operation adjustment parameters will trigger the safety boundary hard constraint condition, the output of the fan operation adjustment parameters is truncated and the safety fallback operational parameters are formed. The safety fallback operational parameters are used as the safety boundary verification output parameters to continue enabling the fan safety fallback operational mechanism. When the execution of the fan operation adjustment parameters does not trigger the safety boundary hard constraint condition, the fan operation adjustment parameters are used as the safety boundary verification output parameters.

[0168] like Figure 2 As shown, a computer device according to an embodiment of this application includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the ventilation fan operating parameter optimization method according to any one of the claims in this application.

[0169] like Figure 3As shown, a computer-readable storage medium is provided in an embodiment of this application, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the ventilation fan operating parameter optimization method described in any one of this application.

[0170] like Figure 4 As shown, this is an embodiment of a ventilation fan operating parameter optimization device based on reinforcement learning, which includes: The state parameter sensing module is used to acquire the state parameters of the ventilator, wherein the state parameters include the aerodynamic parameters, thermal parameters and equipment operation feedback parameters of the ventilator; The action exploration reasoning module is used to input the state parameters into the constructed reinforcement learning network model for forward reasoning, so as to output exploration action parameters, which include the adjustment and control parameters of the ventilator actuator; The batch construction module is used to generate several Markov experience tuples based on the state parameters and exploration action parameters to construct policy training batches. The model optimization training module is used to update the network connection weights of the reinforcement learning network model based on the training batch of the policy until the difference in network connection weight updates between adjacent training batches of the policy meets the condition of being less than or equal to a preset model convergence threshold. Then, the training of the reinforcement learning network model is terminated and used as the optimization decision model for the ventilation fan operating parameters. The adaptive optimization execution module is used to optimize the decision model using the fan operating parameters. Based on the real-time collected state parameters of the fan, it performs forward reasoning and outputs the fan operation adjustment parameters that maximize the expected global reward.

[0171] Figures 2-4 For an exemplary description, please refer to the above. Figure 1 This will not be elaborated upon here.

Claims

1. A method for optimizing the operating parameters of a ventilation fan based on reinforcement learning, characterized in that, include: The state parameters of the ventilator are obtained, including the aerodynamic parameters, thermal parameters, and equipment operation feedback parameters of the ventilator. The state parameters are input into the constructed reinforcement learning network model for forward inference to output exploration action parameters, which include the adjustment and control parameters of the ventilator actuator. Based on the state parameters and exploration action parameters, several Markov experience tuples are generated to construct policy training batches; Based on the policy training batch, the network connection weights of the reinforcement learning network model are updated until the difference in network connection weight updates between adjacent policy training batches is less than or equal to a preset model convergence threshold. Then, the training of the reinforcement learning network model is terminated and it is used as a decision model for optimizing the operating parameters of the ventilation fan. The decision-making model for optimizing the operating parameters of the ventilator is used to perform forward reasoning based on the real-time collected state parameters of the ventilator, and output the ventilator operating adjustment parameters that maximize the expected global reward.

2. The method according to claim 1, characterized in that, The reinforcement learning network model includes: a feature perception fusion layer, a policy optimization decision layer, and an action instruction mapping layer. Correspondingly, the state parameters are input into the constructed reinforcement learning network model for forward inference to output exploration action parameters. These exploration action parameters include the adjustment and control parameters of the ventilation fan actuator, including: The state parameters are input into the feature perception fusion layer to extract the operating condition features that characterize the real-time operating status of the ventilator from the state parameters, and the operating condition features are vectorized to generate a global operating condition status feature vector of the ventilator. The global operating condition feature vector of the ventilator is passed to the strategy optimization decision layer for state feature mapping processing to output the hidden variables of action decision; The action instruction mapping layer maps the latent variables of the action decision to exploration action parameters.

3. The method according to claim 2, characterized in that, The global operating condition state feature vector of the ventilation fan is passed to the strategy optimization decision layer for state feature mapping processing to output latent variables for action decisions, including: Configure the Gaussian radial basis function of the strategy optimization decision layer, and determine the central kernel feature corresponding to the Gaussian radial basis function; Calculate the Euclidean distance between the global operating condition feature vector of the ventilation fan and the central kernel feature; The action decision latent variables are calculated by performing a nonlinear activation mapping based on the Euclidean distance and activation width parameter.

4. The method for optimizing ventilation fan operating parameters based on reinforcement learning according to claim 3, characterized in that, The mathematical expression of the Gaussian radial basis function of the strategy optimization decision layer is: in, To optimize the strategy, the latent variables of the action decision output by the j-th node of the decision layer are used. The feature vector representing the global operating status of the ventilation fan is passed to the feature perception fusion layer. The central kernel feature of the j-th node, The activation width parameter for the j-th node. It is an exponential function with the natural constant e as its base.

5. The method for optimizing ventilation fan operating parameters based on reinforcement learning according to claim 4, characterized in that, The action instruction mapping layer maps the action decision latent variables to the output mathematical expression of the explored action parameters as follows: in, This refers to the action output value of the k-th dimension in the exploration action parameters. To determine the network connection weight parameters between the j-th node in the strategy optimization decision layer and the k-th node in the action instruction mapping layer. is the bias scalar corresponding to the k-th node of the action instruction mapping layer, and m is the total number of nodes in the policy optimization decision layer.

6. The method according to claim 2, characterized in that, Based on the training batch of the aforementioned policy, the network connection weights of the reinforcement learning network model are updated, including: The global reward error for optimizing the parameters of the reinforcement learning network model is calculated using the training batches of the strategy, and the error gradient for updating the network parameters is calculated by backchain differentiation of the global reward error. Based on the error gradient of the network parameter update, the network connection weights of the action instruction mapping layer, policy optimization decision layer, and feature perception fusion layer are updated in batches through backpropagation.

7. The method according to claim 6, characterized in that, The training of the reinforcement learning network model ends when the difference in network connection weight updates between adjacent policy training batches is less than or equal to a preset model convergence threshold, and it is then used as a decision model for optimizing the operating parameters of the ventilation fan. This includes: Based on the network connection weights of the action instruction mapping layer, the policy optimization decision layer, and the feature perception fusion layer, the difference between the network connection weight update amounts corresponding to adjacent policy training batches is calculated. The training of the reinforcement learning network model ends when the difference is less than or equal to a preset model convergence threshold, and it is used as the ventilation fan operation parameter optimization decision model.

8. The method for optimizing ventilation fan operating parameters based on reinforcement learning according to claim 7, characterized in that, The backchain derivative of the global reward error is used to calculate the error gradient for network parameter updates, including: The value function is configured with positive excitation factors such as reducing bearing vibration, controlling bearing temperature rise to a safe range, and reducing single-unit operating power. With the goal of maximizing the expected overall return of the value function output, a gradient descent optimizer is used to calculate the local error gradient of the objective function relative to the connection weights of the action instruction mapping layer network. The local error gradient is sequentially passed in reverse to the policy optimization decision layer and the feature perception fusion layer to calculate the error gradient of the network parameter update between the layers respectively. Wherein, the objective function is the mean square loss function characterizing the temporal difference error; the temporal difference error is obtained by subtracting the predicted action value from the target action value; the predicted action value is the expected comprehensive return calculated based on the state parameters and the exploration action parameters; the target action value is calculated based on the decay term of the predicted action value.

9. The method according to claim 8, characterized in that, The policy training batch is constructed by evaluating the priority weight of each Markov empirical tuple based on the temporal difference algorithm, and extracting the Markov empirical tuples with priority weights greater than the weight threshold to construct the policy training batch.

10. The method for optimizing ventilation fan operating parameters based on reinforcement learning according to claim 9, characterized in that, The priority weights of each Markov empirical tuple are evaluated based on the temporal difference algorithm. Markov empirical tuples with priority weights greater than a weight threshold are extracted to construct policy training batches, including: Determine the target action value and predicted action value for each matching Markov empirical tuple and calculate the temporal difference error between the target action value and the predicted action value; Based on the absolute value of the time-series difference error, the importance sampling probability of each Markov empirical tuple is calculated, and the importance sampling probability is used as the priority weight of the corresponding Markov empirical tuple for sorting. Markov empirical tuples with priority weights greater than the weight threshold are selected to construct the policy training batch.