Intelligent fan control method and system for ventilation system of nuclear power station

Through semi-supervised comparative learning and multi-condition data enhancement technology, combined with multi-agent system and safety constraint model, the problem of insufficient control accuracy and response speed of the ventilation system of the nuclear power plant is solved, and efficient identification and adaptation of multiple ventilation conditions is achieved, ensuring the long-term reliable operation of the system and nuclear safety.

CN119989103AActive Publication Date: 2025-05-13DONGGUAN FOERSHENG M&E TECH CO LTD

Patent Information

Application Number
CN202510437775.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-13
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

When facing a complex and changing operating environment of nuclear power plant ventilation systems, traditional fan control methods cannot meet the requirements of control accuracy and response speed, and it is difficult to obtain labeled samples in special environments of nuclear power plants, which limits the application of advanced control methods.

Method used

Semi-supervised contrast learning and multi-case data enhancement technology are used to train with a small amount of labeled data and a large amount of unlabeled data, and combined with multi-agent system and attention communication mechanism to achieve coordination among different control units. Through time causal graph modeling and aggregation function design, the time sequence causal dependence between key state variables in fan control is captured, and a four-layer cascaded security constraint model is constructed to ensure the security of the control strategy.

Benefits of technology

It improves the ability to identify and adapt to a variety of ventilation conditions, ensures that the ventilation system is efficient and reliable in the long term, can adapt to long-term changes, and strictly follows nuclear safety requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989103A_ABST
    Figure CN119989103A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent fan control, and discloses an intelligent fan control method and system for a ventilation system of a nuclear power station. The method comprises the following steps: acquiring and preprocessing differential pressure data, radioactivity monitoring data, fan operation parameters and environmental data of a nuclear power plant ventilation system to obtain a characteristic representation space; according to the feature representation space and the target monitoring parameter combination, establishing an action space including main exhaust fan control, auxiliary exhaust fan control and filter system control, and executing multi-agent analysis to obtain an initial fan control strategy; a four-layer cascade security constraint model is constructed based on the initial fan control strategy, the constraint optimization problem is solved through a projection gradient method, and the optimal fan control strategy is obtained.According to the method, the recognition and adaptability to multiple ventilation working conditions are improved, the method can adapt to long-term change factors, and it is ensured that a ventilation system operates efficiently and reliably for a long time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent fan control, and in particular to an intelligent fan control method and system for a ventilation system of a nuclear power plant. Background Art

[0002] As the scale of nuclear power plants expands and safety standards improve, ventilation systems, as a key component of nuclear power plant safety management, bear the important responsibilities of maintaining the negative pressure state of the plant, controlling the spread of radioactive materials, and ensuring the normal operation environment of the equipment. Traditional fan control methods mainly rely on empirical models and simple PID control. Faced with the complex and changeable operating environment of nuclear power plants, the control accuracy and response speed often cannot meet the requirements. Although the intelligent control method of fans based on deep learning has obvious advantages over traditional methods in terms of control accuracy, such methods usually rely on a large number of labeled samples for training, which are difficult to obtain in the special environment of nuclear power plants, limiting the application of advanced control methods.

[0003] The coordination of ventilation and safety systems in nuclear power plants involves a variety of uncertainties, including changes in environmental parameters, fluctuations in radioactivity levels, and adjustments to plant pressure requirements. These daily changes require the fan control system to respond quickly and re-optimize the control strategy. Traditional methods require frequent re-implementation of complex control algorithms, resulting in a huge computational burden. In addition, the ventilation requirements of different areas (such as reactor buildings, auxiliary buildings, and fuel buildings) vary greatly, and there are complex airflow interactions between areas, making the global optimization of fan control strategies extremely challenging. Simple independent control schemes are difficult to achieve the optimization of the overall system performance. Summary of the invention

[0004] The present invention provides a method and system for intelligently controlling fans of a ventilation system of a nuclear power plant. The present invention improves the ability to identify and adapt to various ventilation conditions, can adapt to long-term changing factors, and ensures long-term efficient and reliable operation of the ventilation system.

[0005] In a first aspect, the present invention provides a method for intelligently controlling a fan of a ventilation system of a nuclear power plant, the method comprising: The pressure difference data, radioactivity monitoring data, fan operating parameters and environmental data of the nuclear power plant ventilation system are collected and preprocessed to obtain labeled data sets and unlabeled data sets; Performing semi-supervised contrastive learning and multi-condition data enhancement processing on the labeled data set and the unlabeled data set to obtain a feature representation space; An action space including the control of the main exhaust fan, the auxiliary exhaust fan and the filtration system is established according to the combination of the feature representation space and the target monitoring parameters, and a multi-agent analysis is performed to obtain an initial fan control strategy; A four-layer cascade safety constraint model is constructed based on the initial wind turbine control strategy, and the constraint optimization problem is solved by the projected gradient method to obtain the optimal wind turbine control strategy.

[0006] In a second aspect, the present invention provides an intelligent control system for fans of a nuclear power plant ventilation system, the intelligent control system for fans of the nuclear power plant ventilation system comprising: The acquisition module is used to collect and preprocess the pressure difference data, radioactivity monitoring data, fan operation parameters and environmental data of the nuclear power plant ventilation system to obtain labeled data sets and unlabeled data sets; A data enhancement module, used for performing semi-supervised contrastive learning and multi-condition data enhancement processing on the labeled data set and the unlabeled data set to obtain a feature representation space; An execution module, used to establish an action space including the main exhaust fan, the auxiliary exhaust fan and the filter system control according to the feature representation space and the target monitoring parameter combination, and perform multi-agent analysis to obtain an initial fan control strategy; A construction module is used to construct a four-layer cascade safety constraint model based on the initial wind turbine control strategy, and solve the constraint optimization problem through the projected gradient method to obtain the optimal wind turbine control strategy.

[0007] In the technical solution provided by the present invention, through semi-supervised contrastive learning and multi-condition data enhancement technology, a small amount of labeled data and a large amount of unlabeled data can be effectively used for training, solving the problem of difficulty in obtaining labeled samples in the nuclear power plant environment. The contrastive learning mechanism constructed based on the NT-Xent loss function is combined with four data enhancement methods to improve the model's recognition and adaptability to various ventilation conditions. A multi-agent system composed of five specialized agents is adopted, and the attention communication mechanism is used to achieve effective coordination between different control units. Through time causal graph modeling and aggregation function design, the temporal causal dependency between key state variables in fan control is accurately captured, and the original state space is compressed into an aggregated state space. The four-layer cascade safety constraint model is combined with a dual-value network architecture to ensure that all control strategies strictly follow nuclear safety requirements. Approximate dynamic programming and experience replay technology are used to enable the control system to generate the optimal control strategy directly from environmental monitoring data. The hierarchical control coordination mechanism is combined with progressive control adjustment to automatically decompose the large changes in fan speed into small steps to prevent system instability. The present invention can adapt to long-term changing factors and ensure long-term efficient and reliable operation of the ventilation system. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0009] Figure 1 A schematic diagram of a flow chart of a method for intelligently controlling a fan of a nuclear power plant ventilation system provided in an embodiment of the present application; Figure 2 A schematic block diagram of the structure of an intelligent control system for fans of a nuclear power plant ventilation system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0010] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0011] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may also be decomposed, combined or partially merged, so the actual execution order may change based on actual conditions.

[0012] It should also be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit the application. As used in this application specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.

[0013] It should be further understood that the term “and / or” used in the specification and appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0014] In conjunction with the accompanying drawings, some embodiments of the present application are described in detail below. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.

[0015] See also Figure 1 , Figure 1 A flow chart of a method for intelligently controlling a fan in a ventilation system of a nuclear power plant provided in an embodiment of the present application is shown in FIG. Figure 1As shown, the intelligent control method for fans of a nuclear power plant ventilation system provided in an embodiment of the present application includes steps S100 to S600.

[0016] Step S100, collecting and preprocessing the pressure difference data, radioactivity monitoring data, fan operation parameters and environmental data of the nuclear power plant ventilation system to obtain a labeled data set and an unlabeled data set; It is understandable that the execution subject of the present invention may be a fan intelligent control system of a nuclear power plant ventilation system, or a terminal or a server, which is not limited here. The embodiment of the present invention is described by taking a server as the execution subject as an example.

[0017] Specifically, a multi-source data acquisition network system is built in the reactor building, auxiliary building and fuel building of the nuclear power plant. The system monitors the internal environmental parameters, operating status and safety monitoring indicators of the building in real time and continuously. The data acquisition network includes a variety of monitoring sensors and instrumentation equipment arranged at different locations in each building, which are used to obtain the pressure difference data inside and outside the building, radioactivity monitoring data (covering α, β, γ rays and aerosol concentrations, etc.), fan operating parameters (speed, power, flow and vibration, etc.) and ventilation system filter pressure difference, ambient temperature and humidity data, forming an original monitoring data set containing all key operating parameters. The collected original monitoring data are preliminarily sorted and preprocessed. Special outlier detection and missing value filling processing are performed on the multiple types of parameters in the original data set, and the time interpolation method based on local linear regression is used to process missing values ​​to ensure the integrity of the data; outliers are identified by the improved Z-score method, and the specific threshold is set to ±3.5. Data points beyond the threshold range are replaced by local medians, thereby effectively ensuring the reliability and accuracy of the data. After preliminary processing, the data is standardized to eliminate the differences in dimensions and magnitudes of different monitoring data. The Min-Max normalization method is used for standardization. By uniformly mapping all data values ​​to the standard interval of [0,1], standard monitoring data is obtained. At the same time, considering that the data comes from multiple monitoring devices and there is a time series deviation between the acquisition devices, in order to ensure the accuracy of data analysis, the timestamps of the above standard monitoring data are unified to the unified standard time coordinate of the nuclear power plant to achieve the time series consistency of the monitoring data and form a target data set. The unified target data set is divided into data. In order to train the intelligent model more effectively, experts confirm and mark various operating conditions in the data, and clearly define the optimal fan control parameters corresponding to each operating condition. Based on the marking situation, the above target data set is divided into a labeled data set at a ratio of about 10%, which is used for model supervised learning tasks; the remaining approximately 90% of the data only contains various monitoring parameters, but does not contain clear control labels, and is defined as an unlabeled data set for semi-supervised learning tasks. After division, the labeled data set and the unlabeled data set need to be stored in independent databases to support the efficient implementation of the next step of intelligent learning and model training. To ensure the reliability and validity of the divided data set, after completing the above data division, the data quality is evaluated and checked for the labeled and unlabeled data respectively, and a data quality report is generated to clearly define the completeness rate, abnormal data rate and timestamp alignment of the data at each monitoring point in the report.

[0018] Step S200, performing semi-supervised contrastive learning and multi-condition data enhancement processing on the labeled data set and the unlabeled data set to obtain a feature representation space; Specifically, corresponding data enhancement operations are performed on labeled data sets and unlabeled data sets to improve the diversity of training data and enhance the robustness and generalization performance of the model to changes in monitoring data under different working conditions. For the operation data of the ventilation system of nuclear power plants, several different data enhancement strategies are adopted, including adding additive Gaussian noise to the original data, and the noise intensity is strictly controlled to about 5% of the standard deviation of the original signal, which can effectively simulate the random fluctuation of the data and avoid excessive noise from destroying the original distribution characteristics of the data; random time window translation is performed according to the characteristics of time series data, that is, the data window is randomly selected to move the position within a certain time window (for example, ±60 seconds), so as to increase the adaptability of the model to slight drift of data on the time axis; at the same time, some key monitoring parameters (such as plant pressure difference, radioactivity level and other important indicators) are selected, and the amplitude is randomly amplified. The amplitude amplification factor is generally selected between 1.1 and 1.3 to simulate the local changes that may occur in the actual monitoring data under different working conditions; in order to further improve the generalization ability of the model, some non-critical monitoring parameters (about 15%) are randomly masked and their values ​​are reset to zero to simulate the situation of partial data missing or sensor failure in the data collection process. In the above way, each original data sample obtains several enhanced samples to form an enhanced data set. Based on the above enhanced dataset, a 5-layer deep neural network encoder for feature learning is constructed. The network structure consists of multiple fully connected layers. The specific structure is designed as follows: the input data dimension is mapped to a 512-dimensional hidden layer feature space through a fully connected layer, and then batch normalization and ReLU activation functions are introduced, and the Dropout layer is used to reduce the risk of overfitting. After that, it is mapped to a 256-dimensional feature space through the second fully connected layer, and batch normalization, ReLU activation and Dropout processing are also performed. Then, the feature dimension is further compressed to 128 dimensions through the third fully connected network, batch normalization is performed again, and the fourth fully connected layer is used to reduce the features to the final required 64-dimensional feature representation space. Through deep and gradual dimensionality reduction and nonlinear transformation, the model can capture the inherent differences and potential connections between data of different operating conditions, forming a good feature representation effect. On this basis, an unsupervised contrast loss function is constructed, and a normalized temperature-adjusted cross entropy loss function is used. This loss function adjusts the sensitivity of the feature distance between different data points through the temperature coefficient, and the temperature coefficient is set to about 0.07 to ensure that the model can more effectively distinguish the differences between different samples, so that the enhanced sample pairs generated from the same original sample are closer in the feature space, while the different sample pairs are farther away. At the same time, considering that some samples in the data used have clear working condition labels, a supervised contrast loss function is additionally introduced for these labeled data samples, that is, by explicitly constraining the positions of data samples of the same working condition category in the feature space as close as possible, a clear working condition clustering effect is formed.Then, the above-mentioned supervised contrast loss and unsupervised contrast loss functions are combined according to a certain weight coefficient to form a joint semi-supervised contrast learning loss function, in which the weight coefficient gradually increases linearly with the training process to gradually strengthen the influence of the supervisory signal. The established semi-supervised contrast learning loss function is used to train the above-mentioned 5-layer deep feature encoder network, and the network parameters are adjusted by the Adam optimizer. The learning rate adjustment strategy of cosine annealing is combined to achieve effective convergence of the encoder parameters. After the training is fully completed, the final feature encoder model is obtained. All the labeled data sets and unlabeled data sets are input into the trained feature encoder model for forward propagation mapping, and finally all data points are projected into a 64-dimensional feature space. In this space, data points of similar working conditions will naturally gather into tight clusters, while data points of different working conditions are separated from each other, forming a high-quality working condition feature representation space that can be used for subsequent multi-agent reinforcement learning decision-making.

[0019] Step S300: establishing an action space including the main exhaust fan, the auxiliary exhaust fan and the filter system control according to the feature representation space and the target monitoring parameter combination, and performing multi-agent analysis to obtain an initial fan control strategy; Specifically, based on the feature representation space, combined with several target monitoring parameters monitored in real time during the operation of the nuclear power plant, including the reactor building pressure difference, auxiliary building pressure difference, fuel building pressure difference, and key indicators such as the maximum radioactivity level in each area of ​​the building, ambient temperature and humidity. The real-time target monitoring parameters reflect the key state information of the nuclear power plant ventilation system during operation, which has an important impact on the decision-making of the control strategy. Therefore, they are effectively combined and spliced ​​with the feature representation space obtained in the early stage to construct an overall state vector that can fully reflect the current operating state of the system. Define the action space of the system, which covers the entire adjustable range of the main exhaust fan, auxiliary exhaust fan and filtration system of the nuclear power plant ventilation system, so the action space needs to be reasonably quantified and clearly expressed. For the main exhaust fan control command, the fan speed range is carefully discretized from 0% to 100% and quantized into M discrete value points; the auxiliary exhaust fan not only involves the adjustment of the speed, but also involves the switch state of the fan start and stop, so a clear distinction is made between the two states of 0 (stop) and 1 (start), and the speed is also subdivided into N value points, so that the auxiliary exhaust fan action space reaches a higher dimension; at the same time, for the filtration system, which is crucial in the ventilation system, its start and stop combination is logically quantified according to the design of the actual system, and a total of F effective start and stop state combinations are obtained. Through the above action space definition and combination, a clear and executable discrete action space is obtained, which covers all the control situations that may need to be considered in the actual operation of the ventilation system. According to the structure of the ventilation system of the nuclear power plant and its control logic characteristics, five dedicated agents are designed to be responsible for ventilation control tasks in different areas, including the main exhaust fan control agent of the reactor building, the main exhaust fan control agent of the auxiliary building, the main exhaust fan control agent of the fuel building, the auxiliary exhaust fan control agent of the building, and the filter system control agent. Each agent is designed based on the Actor-Critic structure. The Actor network maps the state vector constructed above into specific action outputs such as the fan speed and the start and stop status of the filter system in a specific area; the Critic network is used to accurately evaluate the value of the current state, thereby providing effective value feedback for the agent's actions. The network architecture adopts a multi-layer fully connected network structure and uses nonlinear activation functions (such as LeakyReLU) to ensure the stability of training and the accuracy of the strategy. In order to achieve collaborative control and information sharing between agents, an agent communication method based on the attention mechanism is designed and implemented. This communication mechanism allows each agent to share the current state information in real time, and at the same time uses the attention weight mechanism to effectively extract and aggregate key state information from other agents, thereby achieving efficient fusion of collaborative control strategies between agents. The Proximal Policy Optimization (PPO) algorithm is used to train the entire multi-agent system.In the process of policy gradient optimization, the PPO algorithm clips the objective function to avoid instability caused by excessive policy updates and improve the robustness of the training process. During the training process, the discount factor γ is explicitly set to 0.95, the generalized advantage estimation (GAE) parameter λ is set to 0.97, and the clipping parameter ε is set to 0.2. These parameters ensure the stability and efficiency of the training process. With the continuous iteration and optimization of training, each agent gradually learns how to make more reasonable wind turbine control actions based on state information, so that the multi-agent system gradually converges and finally generates an initial wind turbine control strategy with good performance and high coordination.

[0020] Step S400: construct a four-layer cascade safety constraint model based on the initial wind turbine control strategy, and solve the constraint optimization problem by projected gradient method to obtain the optimal wind turbine control strategy.

[0021] Specifically, the state variables in the operation process of the initial fan control strategy are analyzed for their temporal causal dependence. The state variables mainly include important parameters of the nuclear power plant ventilation system operation, such as fan speed, plant pressure difference, radioactivity level, and aerosol concentration. In the specific analysis, based on the actual operation experience and historical data of the nuclear power plant, correlation analysis and statistical methods are used to clarify the causal relationship between different variables and the specific delay characteristics between variables in time. For example, there is usually a response delay of about 15 to 30 seconds between the adjustment of fan speed and the change of plant pressure difference, a time delay relationship of 30 to 60 seconds between the change of radioactivity level and the switching of ventilation path, and a delay of about 45 to 120 seconds between the start of the filter and the change of aerosol concentration. After clarifying the temporal dependency between these variables, a temporal causal graph with a clear structure is constructed. By setting the causal strength weight and the specific delay time in the graph, the temporal causal relationship between the state variables is expressed in the form of a graph. In order to reasonably reduce the dimension and aggregate the high-dimensional state information with time series characteristics, the corresponding time causal aggregation model is developed based on the time causal graph constructed above, and the specific mathematical expression of the aggregation function is clarified. The aggregation state space needs to take into account the current state, historical state and the prediction of future state change trend. In the specific calculation, the aggregation function considers the current value of each state variable, the exponential weighted average of the historical state value and the time series change trend of the variable. After aggregation, a state vector with smaller dimension but more compact information is obtained, which more effectively characterizes the dynamic characteristics and time series characteristics of the ventilation system. The specific aggregation operation is realized by the linear combination method of weighted summation, and the weight matrix is ​​determined by model optimization to ensure that the characteristic state after aggregation fully reflects the causal dependency between variables, while reducing the complexity of the model state space. Based on the existing initial fan control strategy, a four-layer cascade safety constraint model for the strict safety requirements of the nuclear power plant ventilation system is established. The first layer is the basic operation constraint, which specifies that the equipment operation parameters such as fan speed and filtration system status must always be kept within the safe and stable physical limit range; the second layer is the working condition constraint, which sets special ventilation requirements for different operation modes of nuclear power plants, such as normal operation mode, shutdown maintenance mode, and accident response mode, including the minimum negative pressure maintenance level, specific air volume, and filtration efficiency requirements; the third layer is the radioactive protection constraint, which clearly requires that any control strategy must ensure that the plant maintains a continuous negative pressure state inside and outside, and minimize the risk of radioactive material leakage; the fourth layer is the fault response constraint. When the monitoring system identifies that the equipment is abnormal or fails, it automatically activates the predefined safety response control strategy to minimize the risk of accidents and ensure the safe and stable operation of the nuclear power plant. These four layers of safety constraints work together in a combination of hard constraints and soft constraints. Hard constraints limit the scope of the action space, while soft constraints are reflected in the optimization objective function of the control strategy through additional penalty terms.The constructed time causal aggregation state space is combined with the four-layer cascade safety constraint model, and the optimization of the control strategy of the nuclear power plant fan is taken as the goal. The solution is solved by the projected gradient method, an efficient numerical optimization algorithm. The projected gradient method is to perform constrained projection processing on the gradient update during the optimization process to ensure that the solution obtained at each step of optimization always meets the aforementioned safety constraints, thereby achieving the optimization of performance goals (such as ventilation efficiency, energy consumption reduction, and accurate maintenance of plant pressure difference) while strictly meeting the constraints of nuclear safety requirements. In the process of solving the projected gradient method, the optimal solution is gradually approached through iterative calculations, and it is guaranteed that the iteration converges to the optimal solution. Through this series of steps, the optimal fan control strategy is finally obtained.

[0022] Based on the aggregate state space, an initial value function approximator for control strategy optimization is constructed. The value function approximator aims to estimate the long-term rewards that the system can obtain under a given aggregate state to guide the selection of the wind turbine control strategy. When constructing it specifically, a multi-layer neural network structure is selected, and a three-layer network structure is used as the basis, that is, the input layer receives the feature vector of the aggregate state space, passes through multiple hidden layers (for example, the first hidden layer uses 64 nodes, the second hidden layer uses 32 nodes, and the hidden layer usually uses a nonlinear activation function such as Tanh), and then outputs a scalar state value estimate. In order to efficiently and stably update the parameters of the value function, a value function parameter update mechanism based on a combination of time difference learning and random approximation is designed. By calculating the difference between the real observed immediate reward of the current state and the value function prediction of the next moment state in each step of training, a time difference error is formed to guide the gradient update of the network parameters; at the same time, a random approximation method is introduced to determine that the learning rate gradually decreases with the number of training steps, so that the network parameters can converge to the real long-term reward value efficiently and stably. In order to improve the efficiency and stability of the training process, a corresponding experience replay training system is constructed according to the above-designed value function parameter update mechanism. The main function of this system is to store the transfer samples composed of historical states, actions, rewards, and next states in a special experience buffer, which is set as a buffer with a large capacity (for example, hundreds of thousands of records). Each time the network parameters are updated, a batch of data is randomly extracted from the experience buffer for training to reduce the correlation between the training data, thereby preventing the network from falling into a local optimum. In order to reduce the overestimation or instability problems caused by the training process of a single value function approximator, a dual network architecture consisting of a target network and an evaluation network is constructed at the same time, in which the evaluation network is used to update parameters in real time, while the target network slowly copies parameters from the evaluation network at a lower frequency to stabilize the training process. Based on the above-mentioned stable value function training framework, according to the strict safety constraint requirements in the actual operation of the nuclear power plant ventilation system, the four-layer cascade safety constraint model established above is converted into a safety probability function that can be used for policy optimization. This probability function can quantitatively express the possibility that a given action satisfies all safety constraints in the current state, and combines the above-mentioned value function to form a safety-conscious policy optimizer. In the specific implementation process, the objective function of the safety constraint policy optimizer combines the action value function and the safety probability function, and dynamically adjusts the importance of the safety constraint through the safety weight factor to form a constrained optimization problem that simultaneously optimizes long-term returns and strictly satisfies nuclear safety conditions. The specific method for solving this optimization problem is the projected gradient method, that is, in the gradient update process of the strategy parameters, the projection method is used to always project the results of the parameter update into the feasible area defined by the safety constraint, thereby effectively ensuring that each optimization step meets the safe operation conditions.In order to obtain the optimal fan control strategy that can be applied in the end, based on the strategy optimizer under the above safety constraints, large-scale offline training is carried out using the historical operating data accumulated over a long period of time under various operating conditions of the nuclear power plant ventilation system. During the training process, not only the conventional operating condition data under normal operating conditions are covered, but also the data of special operating conditions such as shutdown maintenance, accident response, and power adjustment, so that the training process can cover all possible operating scenarios during the actual system operation. After a large number of iterations and optimization training in this way, the fan control strategy finally obtained strictly complies with and meets the various safety requirements of nuclear power plant operation while maintaining efficient control performance.

[0023] According to the optimal fan control strategy, a fan control execution system including a data interface layer, a state representation layer, a decision layer and an execution layer is constructed. Among them, the data interface layer obtains the latest environmental data, pressure difference data, radioactivity monitoring data and fan operating parameters from the on-site monitoring network in real time and continuously to ensure the real-time and accuracy of the input data; the state representation layer quickly converts the real-time monitoring data into a state vector through the previously trained feature encoder and time causal aggregation model, reflecting the actual operating conditions of the current nuclear power plant ventilation system; next, the decision layer generates the corresponding fan control instructions in real time through the well-trained multi-agent reinforcement learning model and the optimized optimal fan control strategy, combined with the current state vector; the execution layer effectively converts the control instructions into specific equipment operation commands, and accurately transmits them to the actual control equipment of the fan, thus completing the closed-loop control process. After establishing the above-mentioned fan control execution system, a hierarchical control coordination mechanism is implemented to divide the intelligent control strategy of the fan of the nuclear power plant ventilation system into three different control levels according to the control accuracy and scope of action, including the strategic layer for operating condition identification and overall ventilation strategy formulation. This layer is mainly responsible for identifying the overall operating status and determining the ventilation strategy for the entire plant area, guiding and coordinating the realization of global ventilation goals; the next is the tactical layer for coordinated control of the plant ventilation system. This layer is responsible for coordinated control of ventilation equipment in different plant areas, and is decomposed into specific execution strategies for each plant area according to the goals of the strategic layer; and finally, the operational layer for precise control of individual fans. This layer is directly responsible for detailed and accurate real-time control of the speed, start and stop status of individual fans and the start and stop actions of the filtration system, so as to accurately implement the operating goals and safety requirements formulated by the strategic and tactical layers. At the same time, in order to ensure the stability and safety of the system during the fan control adjustment process, a progressive control adjustment mechanism is designed. When the fan speed adjustment command generated by the control system exceeds a specific threshold (for example, 20%), the system automatically decomposes the large-scale adjustment command into a series of progressive adjustment steps with smaller amplitudes (such as no more than 5%). Each step is separated by a fixed short time period, so that the fan state transitions smoothly and effectively avoids large fluctuations in system pressure or operating status. In addition, a safety protection fallback mechanism is established based on the progressive control adjustment mechanism to continuously monitor key safety parameters such as plant pressure difference and radioactivity level in real time. Once the system detects that the parameters are abnormal or exceed the set safety threshold, the pre-designed emergency response strategy is immediately activated to quickly restore the system to a safe operating state and ensure that the ventilation safety of the nuclear power plant is not affected. In order to achieve continuous optimization and performance improvement of the fan intelligent control system in actual applications, an online learning and model update mechanism is constructed. This mechanism relies on the real operating data such as state-action-reward that the wind turbine control system continuously collects during operation, performs online processing of incremental data at a fixed period (for example, every 7 days), and fine-tunes the value function network and strategy network parameters through a small learning rate to continuously improve the control model's ability to adapt to real environmental changes.At the same time, a complete model performance monitoring and diagnosis module has been implemented, which regularly evaluates the control effect of the system from multiple dimensions, including pressure difference control accuracy, radioactivity level management effect, fan energy consumption, safety and other indicators. Once any indicator is found to drop below the preset warning threshold during the monitoring process, the control system will automatically trigger the abnormal diagnosis process, quickly locate the cause of the performance degradation and the specific model components, and adjust and update the corresponding model in real time to quickly restore the optimal performance state of the fan control system. Through the above process, a set of intelligent control execution solutions for fans in the ventilation system of nuclear power plants with real-time perception capabilities, active adjustment and adaptive continuous optimization characteristics is finally obtained, ensuring safe, reliable, stable and efficient operation performance in the actual complex operating environment.

[0024] In the embodiment of the present invention, through semi-supervised contrastive learning and multi-condition data enhancement technology, a small amount of labeled data and a large amount of unlabeled data can be effectively used for training, solving the problem of difficulty in obtaining labeled samples in the nuclear power plant environment. The contrastive learning mechanism constructed based on the NT-Xent loss function is combined with four data enhancement methods to improve the model's recognition and adaptability to various ventilation conditions. A multi-agent system composed of five specialized agents is adopted, and the attention communication mechanism is used to achieve effective coordination between different control units. Through time causal graph modeling and aggregation function design, the temporal causal dependency between key state variables in fan control is accurately captured, and the original state space is compressed into an aggregated state space. The four-layer cascade safety constraint model is combined with a dual-value network architecture to ensure that all control strategies strictly follow nuclear safety requirements. Approximate dynamic programming and experience replay technology are used to enable the control system to generate the optimal control strategy directly from environmental monitoring data. The hierarchical control coordination mechanism is combined with progressive control adjustment to automatically decompose the large changes in fan speed into small steps to prevent system instability. The present invention can adapt to long-term changing factors and ensure long-term efficient and reliable operation of the ventilation system.

[0025] In a specific embodiment, the process of executing step S100 may specifically include the following steps: Set up a multi-source data acquisition network in the reactor building, auxiliary building and fuel building of the nuclear power plant, and collect the original monitoring data set; The pressure difference between inside and outside the plant, α, β, γ ray and aerosol concentration, fan speed, power, flow, vibration, filter pressure difference and ambient temperature and humidity data in the original monitoring data set are classified and sorted to obtain classified monitoring data; Normalize the classified monitoring data to obtain standard monitoring data, and unify the timestamps of the standard monitoring data to the standard time of the nuclear power plant to obtain the target data set; The data are divided into 10% of the target data set containing the optimal fan control parameters under various operating conditions as the labeled data set, and the remaining 90% of the target data set containing only monitoring parameters as the unlabeled data set.

[0026] Specifically, a multi-source data acquisition network is established in the reactor building, auxiliary building and fuel building of the nuclear power plant. The network consists of sensors and monitoring equipment of multiple types and locations, such as differential pressure sensors, radioactive detectors (capable of detecting α, β, and γ rays), aerosol concentration monitors, fan speed sensors, power meters, flow meters, vibration sensors, filter differential pressure sensors, and ambient temperature and humidity sensors. Various types of monitoring equipment continuously obtain real-time information on the operating environment inside and outside the nuclear power plant and the operating status of the fan according to the predetermined acquisition cycle, and generate an original monitoring data set covering all key operating parameters. For example, multiple groups of differential pressure monitoring points are set up in the reactor building to monitor the pressure difference inside and outside the building in real time to ensure that the building is always in a safe negative pressure state, while the radiation monitoring equipment monitors the real-time dose rate of α, β, and γ rays at different locations inside the building, as well as the concentration of aerosol particles in the air, so as to ensure that the radioactivity level is always within the nuclear power safety standard. The characteristics of the original monitoring data set are professionally classified and sorted to clearly distinguish the category differences between different monitoring parameters and form a unified and standardized classified monitoring data set. During the classification process, the pressure difference data inside and outside the plant, radioactive monitoring data (α, β, γ rays and aerosol concentrations), fan operation data (speed, power, flow, vibration), filter pressure difference and environmental data (temperature, humidity) are classified and stored separately for subsequent unified data processing. For example, the data are clearly numbered and named according to the data source, parameter nature and monitoring location, such as defined as pressure difference data sets. , Radioactivity Dataset , Wind turbine operation data set In order to eliminate the differences in dimensions and scales between data, the classified monitoring data are uniformly normalized using the normalization method after Min-Max transformation, that is, any data value is converted to Mapped to the interval [0,1]. The specific formula in the normalization process is: in, represents the normalized data, is the original data, and Respectively represent the maximum and minimum values ​​of this type of parameter data. Align the timestamps of all normalized classified monitoring data to the unified standard time of the nuclear power plant, eliminate the micro-time series differences generated by different monitoring devices during the sampling process, and ensure the uniformity and time series accuracy of the data in the subsequent model training process. For example, calibrate the timestamps of data collected by different devices to the same nuclear power plant clock server to ensure that each data sample is accurately aligned on the time axis and form a precisely synchronized target data set. The target data set is divided into supervised and semi-supervised learning. Based on expert evaluation and historical experience, about 10% of the data in the target data set that clearly records the optimal wind turbine control parameters under different operating conditions is separated to form a labeled data set. The remaining 90% of the data that do not have clear control labels and only contain environmental monitoring parameters are classified as unlabeled data sets. , which is used for subsequent semi-supervised contrastive learning and unsupervised feature extraction. Through a systematic data quality assessment and diagnosis mechanism, the labeled and unlabeled data sets are evaluated for data integrity, abnormal data rate, and time consistency to ensure that the data quality used for the training of the intelligent control model meets the strict nuclear safety operation standards.

[0027] In a specific embodiment, the process of executing step S200 may specifically include the following steps: Perform data enhancement on labeled data sets and unlabeled data sets to generate enhanced data sets; Construct a 5-layer neural network encoder based on the enhanced dataset and use the 5-layer neural network encoder as the feature encoder network; An unsupervised contrastive loss function is constructed for the feature encoder network based on the normalized temperature-adjusted cross entropy loss function; Introducing supervised contrast loss to samples from labeled datasets to cluster features of samples under the same working conditions, and combining supervised contrast loss function with unsupervised contrast loss function into a semi-supervised contrast learning loss function; The feature encoder network is trained based on the semi-supervised contrastive learning loss function to obtain a trained feature encoder model; The labeled data set and the unlabeled data set are input into the trained feature encoder model for mapping, the data points of similar working conditions are clustered, and the data points of different working conditions are separated to obtain the feature representation space.

[0028] Specifically, effective data enhancement processing is performed on the labeled and unlabeled data sets after classification and sorting to expand the capacity of the data sets and improve the generalization and stability of the model for various operating conditions. For example, controlled additive Gaussian noise is added to the original data to simulate the random errors of measuring instruments and sensors, and the variance of the noise is set to about 5% of the standard deviation of the original signal; at the same time, for monitoring data with time series characteristics, a small translation operation of the time window is randomly performed, such as randomly translating the original data fragments within the range of ±60 seconds, to enhance the robustness of the model to data acquisition equipment delays or clock asynchronization problems. For key radioactive monitoring parameters (such as gamma ray intensity, aerosol concentration) or plant pressure difference, a local amplitude random amplification strategy is implemented, such as randomly selecting an amplification factor of 1.1 to 1.3 times for local amplification to reflect the diversity of parameter fluctuations under different operating conditions; for non-critical parameters, about 15% of the data is randomly masked to zero values ​​to simulate local sensor failure phenomena that occur during actual operation. Through the combined application of these strategies, each original data sample can derive multiple enhanced data samples to form a richer enhanced data set. Based on the enhanced data set obtained above, a deep neural network encoder with a 5-hidden layer structure is designed and built. The encoder is used to learn the potential feature information in the data and realize the efficient mapping of high-dimensional data to low-dimensional feature representation. For example, if the input data dimension is , the output dimension of the first hidden layer of the network is 512, which is mapped to 256 and 128 layer by layer after activation function, batch normalization and Dropout, and the final output dimension is 64. The network is expressed as the following inter-layer nonlinear transformation formula: Among them, the variable Indicates The feature output of the layer network, For the The weight matrix of the layer, For the Layer bias vector, ReLU is a nonlinear activation function, BN stands for batch normalization, and Mask Represents the Dropout function, parameter Indicates that 30% of neurons are randomly turned off to prevent overfitting, and finally an encoded 64-dimensional feature vector is obtained. In the network structure, an unsupervised loss function is constructed through an unsupervised contrastive learning method, and the normalized temperature is used to adjust the cross entropy loss. Its expression is: Among them, the variable , , Represents the feature vector of the data augmented sample after mapping through the feature encoder. The function sim Represents the feature vector and The cosine similarity between variables is used to measure the similarity between features of data; is the temperature coefficient, which is used to adjust the sensitivity of the distance between feature vectors and is set to 0.07 to ensure training stability; Indicates the number of data pairs used in each training. At the same time, in order to make full use of the working condition information carried in the labeled data set, a supervised contrast loss function is additionally introduced for samples with clear labels, and its formula is defined as: The variables Representation and Sample The index set of all positive samples of the same working condition, is a temperature parameter dedicated to supervised contrastive learning, and The values ​​are similar, but can be slightly adjusted according to the data clustering situation to better cluster data of similar working conditions. Combining the above supervised contrast loss function and unsupervised contrast loss function, with adjustable coefficients, a complete semi-supervised contrast learning loss function is formed: L Among them, the variable The initial value is 0.5, and gradually increases to 2.0 during the training process to gradually enhance the constraint effect of supervisory information on the construction of feature space. Through the semi-supervised contrastive learning loss function, the Adam optimization algorithm is used to iteratively train the neural network encoder for multiple times until it converges to a stable state, and a trained feature encoder model is obtained. The labeled data set and the unlabeled data set are respectively input into the trained encoder model, and mapped to a 64-dimensional feature space through network forward propagation. In this space, the feature encoder will automatically aggregate data points of similar working conditions and clearly separate data points of different working conditions, thereby constructing an effective and highly generalizable feature representation space for the ventilation system of a nuclear power plant.

[0029] In a specific embodiment, the process of executing step S300 may specifically include the following steps: The characteristic representation space is combined and spliced ​​with the target monitoring parameters to obtain the state vector. The target monitoring parameters include the pressure difference of the reactor building, the pressure difference of the auxiliary building, the pressure difference of the fuel building, the maximum radioactivity level of each building, the ambient temperature and the ambient humidity; The main exhaust fan speed percentage is quantified into M discrete values, the auxiliary exhaust fan switch state and speed percentage are quantified into N discrete values, and the start and stop state of the filtration system is quantified into F combinations, and the action space is obtained through the combination; Based on the state vector and action space, five agents are designed to be responsible for the control of the main exhaust fan of the reactor building, the main exhaust fan of the auxiliary building, the main exhaust fan of the fuel building, the auxiliary exhaust fan of the building and the filtration system, respectively, to obtain a multi-agent control system. At the same time, a strategy generation and value evaluation network is constructed for each agent in the multi-agent control system. According to the multi-agent control system, a communication mechanism between agents is realized through state sharing and attention weighting, and the PPO algorithm is used for training to obtain the initial fan control strategy.

[0030] Specifically, by splicing the 64-dimensional feature representation space vector obtained in the previous step with the real-time key parameters actually monitored by the nuclear power plant, the target monitoring parameters include the reactor building pressure difference, the auxiliary building pressure difference, the fuel building pressure difference, the maximum radioactivity level of each building, the ambient temperature and the ambient humidity, a complete state description is formed. In order to achieve the refined description and efficient execution of specific control instructions, the action space is constructed. The action space consists of three parts, namely the discretization space of the main exhaust fan speed percentage, the discretization space of the auxiliary exhaust fan start-stop switch state and the speed percentage, and the start-stop state combination space of the filtration system. Taking the main exhaust fan speed percentage as an example, the speed percentage range (0% to 100%) is subdivided into M equally spaced discrete values, each of which represents a specific fan speed setting value; and the action space of the auxiliary exhaust fan, in addition to the fan speed percentage being subdivided into N discrete values, also introduces additional switch start and stop states (start is 1, stop is 0). For example, when the auxiliary exhaust fan is on state 1, its speed can be selected from a discrete speed percentage in the range of 0% to 100%, thereby clarifying the actual control strategy of the auxiliary exhaust fan; for the filtration system in the nuclear power plant ventilation system, since it involves multiple independent devices, the start and stop states of the filtration system are discretized into F combinations. If there are 3 groups of filtration systems in the nuclear power plant, then through different combinations of switch states, 2³=8 different start and stop combinations are obtained. The above parts are combined to form a total action space, where the total action dimension The innovative expression of is: Among them, the variable M represents the number of discrete levels of the main exhaust fan speed, N represents the number of combinations of the auxiliary exhaust fan speed and switch state, and F represents the number of combinations of the start and stop states of the filtration system. This action space can accurately and flexibly cover all possible control actions of the ventilation system of the nuclear power plant. After obtaining the state vector and the clear action space, a multi-agent control system is established based on the actual operation requirements of the nuclear power plant. The multi-agent system specifically includes 5 specially designed agents, which are responsible for the operation tasks of the five subsystems: the main exhaust fan control of the reactor building, the main exhaust fan control of the auxiliary building, the main exhaust fan control of the fuel building, the auxiliary exhaust fan control of the building, and the filtration system control. A special strategy generation network and value evaluation network are designed and constructed inside each agent. The strategy generation network (Actor network) is used to generate the probability of specific action distribution, and the value evaluation network (Critic network) is used to evaluate the value of the current state. Taking the main exhaust fan control agent of the reactor building as an example, its policy network is designed to input a 70-dimensional state vector, which is mapped to a probability distribution on a specific action space dimension (such as 21 discrete speed values) through a multi-layer fully connected network; the value network outputs a scalar for the state vector, representing the expected value of the long-term return of the current state. In order to improve the coordination ability and decision consistency between agents, a communication mechanism based on the attention mechanism is introduced between agents. This communication mechanism expresses the importance of each agent in the form of an attention coefficient by weighted fusion of the state information of other agents. The attention mechanism formula is expressed as: in, represents the aggregated information received by agent i from other agents, Represents the state information held by agent j, and denote the attention query matrix and the key matrix respectively, is the state vector of agent i itself, Used to scale attention weights. Through this mechanism, the agent can selectively pay attention to the state information of other agents, thereby effectively realizing the interaction and fusion of information between agents. For example, when an abnormal pressure occurs in the reactor building, the agents in the auxiliary building and the fuel building quickly perceive the change in this state through the attention mechanism, and actively coordinate the corresponding fans for dynamic adjustment to jointly ensure the overall safety of the nuclear power plant. The strategies and communication mechanisms constructed by the above agents are effectively trained and optimized through the proximal policy optimization algorithm (PPO algorithm). The PPO algorithm optimizes the objective function through the clipped probability ratio constraint, so that each update of the strategy does not deviate too far, avoiding training instability, and thus obtaining a stable and reliable initial fan control strategy. After training iterations, the control strategies generated by each agent can work together efficiently to achieve goals such as precise fan control, stable pressure difference maintenance, and radioactive safety control under different working conditions.

[0031] In a specific embodiment, the process of executing step S400 may specifically include the following steps: Perform time causal dependency analysis on the state variables during the operation of the initial fan control strategy to obtain the temporal causal relationship between the state variables; Construct a temporal causal graph based on temporal causal relationships, and design an aggregated state space that considers the current state value, the exponential weighted average of historical state values, and predicts future state changes based on the temporal causal graph; A four-layer cascade safety constraint model is established according to the safety requirements of the ventilation system of nuclear power plants. The first layer of the four-layer cascade safety constraint model is the basic operation constraint to ensure that the fan parameters are within the physical limit range. The second layer is the working condition constraint to set the ventilation requirements for different nuclear power plant operation modes. The third layer is the radioactive protection constraint to ensure the maintenance of negative pressure in the plant and control the release of radioactive materials. The fourth layer is the fault response constraint of the preset response strategy activated when equipment failure or abnormality is detected. Based on the aggregated state space and four-layer cascade safety constraint model, the constrained optimization problem is solved by the projected gradient method to obtain the optimal wind turbine control strategy.

[0032] Specifically, a time causal dependency analysis is conducted on various key state variables involved in the actual operation of the initial control strategy to clarify the causal mechanism of changes between system variables over time. For example, when the speed of the main exhaust fan increases, the pressure difference between the inside and outside of the plant will not reach a new stable value instantaneously, but there is a typical response delay period. For example, there is usually a delay of about 15-30 seconds from the change in the speed of the main exhaust fan to the pressure difference change and tending to a stable state. This relationship is defined as a temporal causal relationship. Similarly, there is a delayed response relationship of about 30-60 seconds between the adjustment of the ventilation path and the change in radioactive concentration, and there is a correlation delay of 45-120 seconds from the activation of the filter to the significant decrease in aerosol concentration. Causal relationship analysis can clearly reveal the dynamic evolution logic between state variables. Based on the above causal relationship analysis results, a time causal graph is constructed. Each edge in the graph defines two key parameters: causal strength weight and time delay parameters , where the causal strength weight Representation variables For variables The degree of influence of the time delay parameter Seconds describe the specific delayed effect of this influence in the time dimension. Based on the temporal causal relationship and the constructed temporal causal graph, a state aggregation function is designed to achieve effective dimensionality reduction and aggregation expression of the original high-dimensional state space. The specific expression of this aggregation function is: in, is the state space after aggregation, It is an innovative time series feature extraction function, which is specifically defined as: In the above formula, is a state variable At the present moment The real-time value of Represents the impact of other variables’ historical states on the current variable determined by the time series causal diagram, where the coefficient , , Represent the current value, historical weighted average and weight of trend change respectively, so as to effectively integrate the state information of past, present and future trends and obtain the aggregated state space. On this basis, a four-layer cascade safety constraint model is established to ensure the safety and reliability of the ventilation system of nuclear power plants. The first layer is the basic operation constraint, which clearly defines the physical limit constraints of basic equipment parameters such as fan speed percentage, power, flow and filter status; the second layer is the operating condition constraint, which sets specific ventilation goals for different operation modes of nuclear power plants, such as normal operation, maintenance status and accident operation conditions; the third layer is the radioactive protection constraint, which strictly ensures the continuous stability of the negative pressure state inside the plant and effectively controls the release level of radioactive substances from exceeding the limit; the fourth layer is the fault response constraint. Once the system equipment is detected to be abnormal or faulty, the predefined emergency control response plan is immediately activated. These four layers of constraint systems work together to ensure the safety of system operation with a comprehensive constraint mechanism. After defining the aggregated state space and the multi-level safety constraint model, the projected gradient method is used to solve the optimization problem with safety constraints to obtain the optimal fan control strategy. The objective function of this constraint optimization is expressed as: Among them, the variable represents a specific control action in the action space, and the variable Represents the aggregate state vector of the current system, function represents the long-term value function of executing actions in the current state, It means that the current action is in state The security probability function that satisfies the four-layer security constraints is It is a safety weight factor used to balance the trade-off between system performance optimization and safety constraints. The initial value is high and gradually decreases to 1.0 during training to achieve the best balance between safety and performance. When solving this optimization problem, the projected gradient method is used. The characteristic of this method is that the parameters or strategy update values ​​obtained at each gradient update are projected back to the safe action space defined by the constraints, so that the optimization results continue to be in the feasible domain defined by the safety constraints, thereby ensuring that the strategy generated during the optimization process always meets all safety requirements. Through this method, continuous iterative updates eventually converge to the optimal solution. The optimal wind turbine control strategy is obtained.

[0033] In a specific embodiment, the execution step is based on the aggregated state space and the four-layer cascade safety constraint model, and the process of solving the constraint optimization problem by the projected gradient method to obtain the optimal wind turbine control strategy can specifically include the following steps: An initial value function approximator is constructed based on the aggregate state space, and a value function parameter update mechanism based on time difference learning and random approximation is designed for the initial value function approximator. Construct an experience replay training system based on the value function parameter update mechanism, and introduce a value function training framework consisting of a target network and an evaluation network into the experience replay training system; According to the value function training framework, the four-layer cascade safety constraint model is transformed into a safety probability function. The constraint optimization target is constructed by combining the safety weight factor with the value function to obtain the policy optimizer under safety constraints. The strategy optimizer under safety constraints uses historical data under different operating conditions of nuclear power plants for offline training to generate the optimal wind turbine control strategy.

[0034] Specifically, an initial value function approximator is designed based on the aggregate state space constructed in the previous step. The purpose of this approximator is to accurately estimate the value function corresponding to the long-term control performance of the nuclear power plant ventilation system under a given aggregate state, so as to effectively guide the subsequent intelligent decision-making process of the system. During the specific construction, a deep neural network composed of multiple hidden layers is used as the network structure of the initial value function approximator. For example, the dimension of the input aggregate state space is 25 dimensions. The first hidden layer is designed to have a Tanh nonlinear activation function with 64 nodes, and the second hidden layer uses 32 nodes. The final output is a scalar state value estimate. This network structure not only helps to capture the nonlinear relationship between states, but also enhances the generalization ability of the value function approximation. After the initial value function approximator is constructed, a parameter update mechanism based on time difference learning and random approximation is designed to ensure that the value function parameters converge to the true optimal value function efficiently and stably during the training process. Through the value function parameter update formula: In the formula, is the value function network in The parameter vector updated by step, is the value function estimate of the current state, is the discount factor, set to 0.95, represents the state space representation after aggregation, Indicates the gradient update direction of the current state value function. The random approximation method introduced in the parameter update process is reflected in the gradual decrease of the learning rate with the number of training steps, thereby avoiding drastic fluctuations in the value function parameters and achieving smooth update and stable convergence of the parameters. In order to improve the stability of the value function network training and the efficiency of data use, an experience replay training system is established based on the above-mentioned value function parameter update mechanism. The training system creates a large-capacity experience buffer (for example, a capacity of 100,000 transfer samples) to store state, action, reward, and next state transfer data samples in real time, and randomly extracts training samples from the experience buffer to update the value function network to avoid excessive data correlation. At the same time, in order to avoid the overestimation or instability of the value function estimation during the training of a single network, a value function training framework consisting of a target network and an evaluation network is introduced. The target network parameters are soft-updated, namely: in, represents the parameters of the evaluation network, Represents the parameters of the target network, parameters is a soft update coefficient, and a small value such as 0.005 is usually set to stabilize the training process. In order to enable the model to meet the strict safety constraint requirements of nuclear power plants, the previously constructed four-layer cascade safety constraint model is innovatively transformed into a safety probability function , specifically defined as an action In Status The probability estimation function that satisfies all safety constraints is expressed as follows: Among them, the variable Indicates Hierarchical security constraints in state Next action The safety margin value after the test is completed. The larger the margin is, the more safety conditions are met. It is used to adjust the sensitivity of different security levels. The sigmoid function converts the safety margin into a safety probability expression between [0,1], so that the safety probability function It intuitively reflects the possibility of the action satisfying all safety constraints. Based on the above safety probability function, the strategy optimization objective function of safety perception is constructed, and the specific formula is expressed as: In the above formula, Representative strategy The expected value of the comprehensive optimization objective, Represents the state of a given aggregate Next select action Long-term value estimation at is the safety weight factor, which is initially high, such as 5.0, and gradually decreases to 1.0 as the training progresses to adjust the balance between safety and performance, so that the strategy can maximize system performance while strictly satisfying safety constraints. Based on the constructed strategy optimization objective function under safety constraints and the experience replay training system, offline training is performed using the historical operating data accumulated over a long period of time in the nuclear power plant (covering different operating conditions such as normal operation, power changes, equipment maintenance, and accident response). During the training iteration process, the projected gradient method is continuously used to constrain the optimization results of each step within the safety domain until the control strategy converges stably. After sufficient training, the final intelligent control strategy for wind turbines can significantly improve the stability and accuracy of the control strategy while strictly following nuclear safety constraints.

[0035] In a specific embodiment, the method for intelligently controlling fans of a nuclear power plant ventilation system further includes the following steps: According to the optimal wind turbine control strategy, a wind turbine control execution system including a data interface layer, a state representation layer, a decision layer and an execution layer is constructed; According to the fan control execution system, a hierarchical control coordination mechanism is implemented, and fan control is divided into a strategic layer for working condition identification and overall ventilation strategy formulation, a tactical layer for coordinated control of ventilation systems in each plant, and an operational layer for precise control of individual fans, thus obtaining a hierarchical control system. Based on the hierarchical control system, a progressive control adjustment mechanism is designed, and based on the progressive control adjustment mechanism, a safety assurance fallback mechanism is implemented to obtain an emergency response strategy; Through emergency response strategies, an online learning and model updating mechanism is built, and model performance monitoring and diagnosis are performed to obtain multi-dimensional evaluation indicators; When it is detected that the multi-dimensional evaluation indicators are lower than the warning threshold, the abnormal diagnosis process is automatically triggered to identify the root cause of the problem and adjust the model components to obtain the intelligent control execution plan of the wind turbine.

[0036] Specifically, a fan control execution system is constructed based on the optimal control strategy. The system as a whole forms a top-down data flow and execution mechanism, which specifically includes four important functional modules: data interface layer, state representation layer, decision layer and execution layer. The data interface layer collects actual data such as fan operating parameters, plant pressure difference, radioactivity level, ambient temperature and humidity from each plant in the nuclear power plant in real time, and transmits them to the state representation layer for unified processing through a unified data format and communication protocol; the state representation layer uses the feature encoder model and time causal aggregation function established in the early stage to convert the real-time collected environmental monitoring parameters into an aggregated state vector that can directly reflect the current state of the system. This state representation combines the implicit characteristics driven by data and the explicit characteristic information of the real-time monitoring parameters; the decision layer uses the trained optimal fan control strategy model to receive the aggregated state vector input and gives specific fan action control instructions through the strategy generation network; and the execution layer directly executes the above action instructions, adjusts the fan speed and filter switch status in real time, and completes the complete closed-loop intelligent control process. According to the fan control execution system, a hierarchical control coordination mechanism is implemented, and the fan control tasks of the entire nuclear power plant ventilation system are divided into strategic layer, tactical layer and operational layer, so that the system can operate in an orderly and coordinated manner at different levels; the main task of the strategic layer is real-time working condition identification and overall ventilation strategy decision-making, such as judging whether the current operating state is normal, negative pressure abnormality or special working conditions, and determining the strategic goal of overall ventilation; the tactical layer coordinates the refined allocation of ventilation subsystems in each plant according to the overall ventilation goal of the strategic layer, for example, the reactor plant actively coordinates the adjustment of fans in other auxiliary plants and fuel plants when the pressure difference fluctuates; and the operational layer is directly responsible for the precise execution of a specific single fan, such as adjusting the speed percentage of a main exhaust fan to 70%, turning on a filter system and other action details, and the three layers work together to achieve the overall optimization control goal. In order to avoid drastic changes in fan control actions during actual operation, resulting in a decrease in system stability or triggering safety hazards, a progressive control adjustment mechanism is designed based on the hierarchical control system, and the progressive adjustment formula for fan action adjustment amplitude control is defined as follows: In this formula, the variable The variable is the change in the fan speed or state in the current control step. represents the optimal setting value of the current fan control command given by the strategy network, and the variable Indicates the actual operating value of the fan at the last moment, the coefficient Indicates the number of subdivision steps of gradual adjustment (for example, set to 4, indicating that each larger action will be broken down into 4 steps of gradual adjustment), function sign is used to maintain the consistency of the direction of the action change, and The maximum change allowed at a time is to prevent the control action from being too drastic, which may cause negative pressure in the nuclear power plant or unstable environmental conditions. In order to effectively deal with abnormal operating conditions and equipment failures during the operation of the ventilation system, a safety protection fallback mechanism is designed based on the above-mentioned progressive control adjustment mechanism. This mechanism defines a real-time safety fallback trigger function: The variables Represents in state Next action Security response score, parameters For the The weight coefficient of each safety monitoring variable is Indicates the actual safety status value monitored after the action is executed, and is the safety threshold standard for each parameter. For example, if the current negative pressure difference of the plant exceeds the safety threshold, the exponential term in the above formula will generate a large penalty value, automatically triggering the emergency response strategy of the safety assurance fallback mechanism, thereby quickly falling back to a safe state. In order to achieve continuous optimization and online stable operation of the intelligent control system of nuclear power plant fans, on the basis of the aforementioned emergency response strategy, an online learning and model incremental update mechanism is constructed to continuously store the new data generated by the system in real time in the experience playback database, and use the accumulated data to fine-tune and update the online model; the model performance monitoring and diagnosis module tracks the evaluation indicators of multiple dimensions in real time. When it is detected that the multi-dimensional evaluation indicators are lower than the warning threshold, the abnormal diagnosis process is automatically activated to quickly analyze the cause of the problem and adjust the relevant model components, so as to effectively correct the deviations in the model control strategy and quickly restore the system to the optimal performance state.

[0037] See also Figure 2 , Figure 2 A schematic block diagram of the structure of the fan intelligent control system 200 of the nuclear power plant ventilation system provided in the embodiment of the present application is shown in FIG. Figure 2 As shown, the fan intelligent control system 200 of the nuclear power plant ventilation system includes: The acquisition module 210 is used to collect and preprocess the pressure difference data, radioactivity monitoring data, fan operation parameters and environmental data of the nuclear power plant ventilation system to obtain a labeled data set and an unlabeled data set; A data enhancement module 220 is used to perform semi-supervised contrastive learning and multi-condition data enhancement processing on the labeled data set and the unlabeled data set to obtain a feature representation space; An execution module 230 is used to establish an action space including the main exhaust fan, the auxiliary exhaust fan and the filter system control according to the feature representation space and the target monitoring parameter combination, and perform multi-agent analysis to obtain an initial fan control strategy; The construction module 240 is used to construct a four-layer cascade safety constraint model based on the initial wind turbine control strategy, and solve the constraint optimization problem by the projected gradient method to obtain the optimal wind turbine control strategy.

[0038] Through the synergy of the above components, through semi-supervised contrastive learning and multi-condition data enhancement technology, a small amount of labeled data and a large amount of unlabeled data can be effectively used for training, solving the problem of difficulty in obtaining labeled samples in nuclear power plant environments. The contrastive learning mechanism constructed based on the NT-Xent loss function is combined with four data enhancement methods to improve the model's recognition and adaptability to various ventilation conditions. A multi-agent system composed of five specialized agents is adopted, and the attention communication mechanism is used to achieve effective coordination between different control units. Through temporal causal graph modeling and aggregation function design, the temporal causal dependency between key state variables in fan control is accurately captured, and the original state space is compressed into an aggregated state space. The four-layer cascade safety constraint model is combined with a dual-value network architecture to ensure that all control strategies strictly follow nuclear safety requirements. Approximate dynamic programming and experience replay technology are used to enable the control system to generate the optimal control strategy directly from environmental monitoring data. The hierarchical control coordination mechanism is combined with progressive control adjustment to automatically decompose large changes in fan speed into small steps to prevent system instability. The present invention can adapt to long-term changing factors and ensure long-term efficient and reliable operation of the ventilation system.

[0039] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0040] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.

[0041] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for intelligently controlling fans of a nuclear power plant ventilation system, characterized in that: include: The pressure difference data, radioactivity monitoring data, fan operating parameters and environmental data of the nuclear power plant ventilation system are collected and preprocessed to obtain labeled data sets and unlabeled data sets; Performing semi-supervised contrastive learning and multi-condition data enhancement processing on the labeled data set and the unlabeled data set to obtain a feature representation space; An action space including the control of the main exhaust fan, the auxiliary exhaust fan and the filtration system is established according to the combination of the feature representation space and the target monitoring parameters, and a multi-agent analysis is performed to obtain an initial fan control strategy; A four-layer cascade safety constraint model is constructed based on the initial wind turbine control strategy, and the constraint optimization problem is solved by the projected gradient method to obtain the optimal wind turbine control strategy.

2. The intelligent control method for fans of a nuclear power plant ventilation system according to claim 1, characterized in that: The pressure difference data, radioactivity monitoring data, fan operation parameters and environmental data of the nuclear power plant ventilation system are collected and preprocessed to obtain a labeled data set and an unlabeled data set, including: Set up a multi-source data acquisition network in the reactor building, auxiliary building and fuel building of the nuclear power plant, and collect the original monitoring data set; Classify and sort the data of the pressure difference between inside and outside the plant, α, β, γ rays and aerosol concentrations, fan speed, power, flow, vibration, filter pressure difference, and ambient temperature and humidity in the original monitoring data set to obtain classified monitoring data; Normalizing the classified monitoring data to obtain standard monitoring data, and unifying the timestamps of the standard monitoring data to the standard time of the nuclear power plant to obtain a target data set; The data are divided into 10% of the target data set containing the optimal fan control parameters under various operating conditions as the labeled data set, and the remaining 90% of the target data set containing only the monitoring parameters is divided into the unlabeled data set.

3. The intelligent control method for fans of a nuclear power plant ventilation system according to claim 1, characterized in that: The semi-supervised contrastive learning and multi-condition data enhancement processing are performed on the labeled data set and the unlabeled data set to obtain a feature representation space, including: Performing data enhancement on the labeled data set and the unlabeled data set to generate an enhanced data set; Construct a 5-layer neural network encoder according to the enhanced data set, and use the 5-layer neural network encoder as a feature encoder network; constructing an unsupervised contrastive loss function for the feature encoder network based on a normalized temperature adjusted cross entropy loss function; Introducing supervised contrast loss to samples from the labeled data set to cluster features of samples under the same working condition, and combining the supervised contrast loss function with the unsupervised contrast loss function into a semi-supervised contrast learning loss function; Training the feature encoder network based on the semi-supervised contrastive learning loss function to obtain a trained feature encoder model; The labeled data set and the unlabeled data set are input into the trained feature encoder model for mapping, data points of similar working conditions are clustered, and data points of different working conditions are separated to obtain a feature representation space.

4. The intelligent control method for fans of a nuclear power plant ventilation system according to claim 1, characterized in that: The step of establishing an action space including the main exhaust fan, the auxiliary exhaust fan and the filter system control according to the feature representation space and the target monitoring parameter combination, and performing multi-agent analysis to obtain an initial fan control strategy includes: The characteristic representation space is combined and spliced ​​with target monitoring parameters to obtain a state vector, wherein the target monitoring parameters include a reactor building pressure difference, an auxiliary building pressure difference, a fuel building pressure difference, a maximum radioactivity level of each building, an ambient temperature, and an ambient humidity; The main exhaust fan speed percentage is quantified into M discrete values, the auxiliary exhaust fan switch state and speed percentage are quantified into N discrete values, and the start and stop state of the filtration system is quantified into F combinations, and the action space is obtained through the combination; Based on the state vector and the action space, five agents are designed to be responsible for the control of the main exhaust fan of the reactor building, the main exhaust fan of the auxiliary building, the main exhaust fan of the fuel building, the auxiliary exhaust fan of the building and the filtration system, respectively, to obtain a multi-agent control system, and at the same time, a strategy generation and value evaluation network is constructed for each agent in the multi-agent control system; According to the multi-agent control system, a communication mechanism between agents is implemented through state sharing and attention weighting, and the PPO algorithm is used for training to obtain an initial fan control strategy.

5. The intelligent control method for fans of a nuclear power plant ventilation system according to claim 4, characterized in that: The four-layer cascade safety constraint model is constructed based on the initial wind turbine control strategy, and the constraint optimization problem is solved by the projected gradient method to obtain the optimal wind turbine control strategy, including: Performing a time causal dependency analysis on the state variables during the operation of the initial fan control strategy to obtain a temporal causal relationship between the state variables; Constructing a temporal causal graph based on the temporal causal relationship, and designing an aggregated state space that considers the exponential weighted average of current state values ​​and historical state values ​​and predicts future state changes according to the temporal causal graph; A four-layer cascade safety constraint model is established according to the safety requirements of the ventilation system of the nuclear power plant. The first layer of the four-layer cascade safety constraint model is a basic operation constraint to ensure that the fan parameters are within the physical limit range, the second layer is an operating condition constraint to set ventilation requirements for different nuclear power plant operation modes, the third layer is a radioactive protection constraint to ensure the maintenance of negative pressure in the plant and control the release of radioactive materials, and the fourth layer is a fault response constraint of a preset response strategy activated when a device fault or abnormality is detected; Based on the aggregated state space and the four-layer cascade safety constraint model, the constraint optimization problem is solved by the projected gradient method to obtain the optimal wind turbine control strategy.

6. The intelligent control method for fans of a nuclear power plant ventilation system according to claim 5, characterized in that: The method of solving the constraint optimization problem based on the aggregated state space and the four-layer cascade safety constraint model by the projected gradient method to obtain the optimal wind turbine control strategy includes: An initial value function approximator is constructed based on the aggregate state space, and a value function parameter updating mechanism based on time difference learning and random approximation is designed for the initial value function approximator; Constructing an experience replay training system according to the value function parameter updating mechanism, and introducing a value function training framework consisting of a target network and an evaluation network into the experience replay training system; According to the value function training framework, the four-layer cascade safety constraint model is converted into a safety probability function, and the constraint optimization target is constructed by combining the safety weight factor and the value function to obtain a policy optimizer under safety constraints; The strategy optimizer under the safety constraints uses historical data under different operating conditions of the nuclear power plant to perform offline training to generate an optimal fan control strategy.

7. The intelligent control method for fans of a nuclear power plant ventilation system according to claim 1, characterized in that: The fan intelligent control method of the nuclear power plant ventilation system also includes: Constructing a fan control execution system including a data interface layer, a state representation layer, a decision layer and an execution layer according to the optimal fan control strategy; According to the fan control execution system, a hierarchical control coordination mechanism is implemented, and fan control is divided into a strategic layer for working condition identification and overall ventilation strategy formulation, a tactical layer for coordinated control of ventilation systems in each plant, and an operational layer for precise control of a single fan, thereby obtaining a hierarchical control system; Designing a progressive control adjustment mechanism based on the hierarchical control system, and implementing a safety assurance fallback mechanism based on the progressive control adjustment mechanism to obtain an emergency response strategy; Through the emergency response strategy, an online learning and model updating mechanism is constructed, and model performance monitoring and diagnosis are performed to obtain multi-dimensional evaluation indicators; When it is detected that the multi-dimensional evaluation index is lower than the warning threshold, the abnormal diagnosis process is automatically triggered to identify the root cause of the problem and adjust the model components to obtain the wind turbine intelligent control execution plan.

8. An intelligent control system for fans of a nuclear power plant ventilation system, characterized in that: A method for intelligently controlling a fan of a nuclear power plant ventilation system according to any one of claims 1 to 7, wherein the intelligent control system for a fan of the nuclear power plant ventilation system comprises: The acquisition module is used to collect and preprocess the pressure difference data, radioactivity monitoring data, fan operation parameters and environmental data of the nuclear power plant ventilation system to obtain labeled data sets and unlabeled data sets; A data enhancement module, used for performing semi-supervised contrastive learning and multi-condition data enhancement processing on the labeled data set and the unlabeled data set to obtain a feature representation space; An execution module, used to establish an action space including the main exhaust fan, the auxiliary exhaust fan and the filter system control according to the feature representation space and the target monitoring parameter combination, and perform multi-agent analysis to obtain an initial fan control strategy; A construction module is used to construct a four-layer cascade safety constraint model based on the initial wind turbine control strategy, and solve the constraint optimization problem through the projected gradient method to obtain the optimal wind turbine control strategy.

Citation Information

Patent Citations

  • Cable channel edge Internet of Things terminal and method based on ubiquitous power Internet of Things

    CN111107675A

  • Ventilation system autonomous optimization operation regulation and control platform and method based on digital twinning

    CN114322199A

  • Nuclear power station intelligent supervision information system and method

    CN116705363A

  • Energy consumption monitoring and optimizing method and system based on large model and multiple agents

    CN118916778A

  • Nuclear power station failure diagnosis and state monitoring system based on wireless sensor network

    CN206413023U

Cited By

  • Deep well multistage intelligent ventilation control method and system

    CN120315292A

  • Grinding parameter optimization method of full-automatic coarse and fine grinding all-in-one machine based on particle swarm optimization

    CN120610467A

  • Nuclear power plant distributed control system and control method

    CN121165677A