Intelligent control method and system for fans of ventilation system in nuclear power plant

Through semi-supervised comparative learning and multi-condition data enhancement technology, combined with multi-agent system and four-layer cascade safety constraint model, fan control strategy is optimized, and the accuracy and response speed of fan control in complex environments of nuclear power plants are solved, and global optimization and nuclear safety are achieved, ensuring the efficient and reliable operation of the ventilation system of the nuclear power plant.

CN119989103BActive Publication Date: 2025-08-12DONGGUAN FOERSHENG M&E TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510437775.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-08-12
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

Traditional fan control methods are difficult to meet the requirements of control accuracy and response speed in the complex and changing operating environment of nuclear power plants, and are difficult to achieve global optimization. Especially in different areas, ventilation requirements vary greatly and airflows affect each other in complex ways, making it difficult to optimize the overall performance of the system.

Method used

Semi-supervised comparative learning and multi-condition data enhancement technology, combined with multi-agent system and four-layer cascade safety constraint model, fan control strategies are optimized through projection gradient method to realize the identification and adaptability of multiple ventilation conditions, ensuring efficient and reliable operation of the system.

Benefits of technology

Effectively using a small amount of labeled data and a large amount of unlabeled data improves the model's ability to identify and adapt to a variety of ventilation conditions, realizes the global optimization of fan control strategies and strictly follow the nuclear safety requirements, and ensures the long-term efficient and reliable operation of the ventilation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989103B_ABST
    Figure CN119989103B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of intelligent fan control technology, and discloses a method and system for intelligent fan control of a nuclear power plant ventilation system. The method comprises: collecting and preprocessing pressure differential data, radioactivity monitoring data, fan operating parameters, and environmental data of the nuclear power plant ventilation system to obtain a feature representation space; establishing an action space encompassing the control of the main exhaust fan, auxiliary exhaust fan, and filtration system based on the feature representation space and target monitoring parameters, and performing multi-agent analysis to obtain an initial fan control strategy; constructing a four-layer cascade safety constraint model based on the initial fan control strategy, and solving the constraint optimization problem using the projected gradient method to obtain an optimal fan control strategy. The present invention improves the ability to identify and adapt to various ventilation conditions, can adapt to long-term changing factors, and ensures the long-term efficient and reliable operation of the ventilation system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent fan control, and in particular to an intelligent fan control method and system for a ventilation system of a nuclear power plant. Background Art

[0002] As nuclear power plants expand in size and safety standards improve, ventilation systems, as a key component of nuclear power plant safety management, bear the important responsibility of maintaining negative pressure in the plant, controlling the spread of radioactive materials, and ensuring the normal operation of equipment. Traditional fan control methods rely primarily on empirical models and simple PID control. Given the complex and changing operating environment of nuclear power plants, control accuracy and response speed often fall short of requirements. While deep learning-based intelligent fan control methods offer significant advantages over traditional methods in terms of control accuracy, such methods typically rely on large numbers of labeled samples for training. These samples are difficult to obtain in the unique environment of nuclear power plants, limiting the application of advanced control methods.

[0003] Coordination of ventilation and safety systems in nuclear power plants involves numerous uncertainties, including changes in environmental parameters, fluctuations in radioactivity levels, and adjustments in building pressure requirements. These daily fluctuations require the fan control system to rapidly respond and reoptimize the control strategy. Traditional approaches require frequent reimplementation of complex control algorithms, resulting in a significant computational burden. Furthermore, ventilation requirements vary significantly across different areas (such as the reactor building, auxiliary buildings, and fuel building), and complex airflow interactions between these areas make global optimization of fan control strategies extremely challenging. Simple, standalone control schemes struggle to optimize overall system performance. Summary of the Invention

[0004] The present invention provides a method and system for intelligently controlling fans in a nuclear power plant ventilation system. The present invention improves the ability to identify and adapt to various ventilation conditions, can adapt to long-term changing factors, and ensures the long-term efficient and reliable operation of the ventilation system.

[0005] In a first aspect, the present invention provides a method for intelligently controlling fans of a nuclear power plant ventilation system, the method comprising:

[0006] The pressure difference data, radioactivity monitoring data, fan operating parameters and environmental data of the nuclear power plant ventilation system are collected and preprocessed to obtain labeled and unlabeled data sets;

[0007] Performing semi-supervised contrastive learning and multi-condition data enhancement processing on the labeled dataset and the unlabeled dataset to obtain a feature representation space;

[0008] An action space including the control of the main exhaust fan, the auxiliary exhaust fan, and the filtration system is established based on the combination of the feature representation space and the target monitoring parameters, and a multi-agent analysis is performed to obtain an initial fan control strategy;

[0009] A four-layer cascade safety constraint model is constructed based on the initial wind turbine control strategy, and the constraint optimization problem is solved by the projected gradient method to obtain the optimal wind turbine control strategy.

[0010] In a second aspect, the present invention provides an intelligent control system for fans of a nuclear power plant ventilation system, the intelligent control system for fans of the nuclear power plant ventilation system comprising:

[0011] The acquisition module is used to collect and preprocess the pressure difference data, radioactivity monitoring data, fan operating parameters and environmental data of the nuclear power plant ventilation system to obtain labeled data sets and unlabeled data sets;

[0012] a data enhancement module, configured to perform semi-supervised contrastive learning and multi-condition data enhancement processing on the labeled dataset and the unlabeled dataset to obtain a feature representation space;

[0013] an execution module, configured to establish an action space including the control of the main exhaust fan, the auxiliary exhaust fan, and the filtration system based on the combination of the feature representation space and the target monitoring parameters, and perform multi-agent analysis to obtain an initial fan control strategy;

[0014] A construction module is used to construct a four-layer cascade safety constraint model based on the initial wind turbine control strategy, and solve the constraint optimization problem through the projected gradient method to obtain the optimal wind turbine control strategy.

[0015] The technical solution provided by the present invention utilizes semi-supervised contrastive learning and multi-condition data augmentation techniques to effectively utilize a small amount of labeled data and a large amount of unlabeled data for training, addressing the difficulty of obtaining labeled samples in nuclear power plant environments. A contrastive learning mechanism based on the NT-Xent loss function, combined with four data augmentation methods, improves the model's ability to recognize and adapt to various ventilation conditions. A multi-agent system consisting of five specialized agents, combined with an attention communication mechanism, achieves effective coordination between different control units. Through temporal causal graph modeling and aggregation function design, the temporal causal dependencies between key state variables in fan control are accurately captured, compressing the original state space into an aggregated state space. A four-layer cascaded safety constraint model, combined with a dual-value network architecture, ensures that all control strategies strictly adhere to nuclear safety requirements. Approximate dynamic programming and experience replay techniques enable the control system to generate optimal control strategies directly from environmental monitoring data. A hierarchical control coordination mechanism, combined with progressive control adjustment, automatically decomposes large changes in fan speed into small steps, preventing system instability. The present invention can adapt to long-term changing factors, ensuring the long-term efficient and reliable operation of the ventilation system. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0017] Figure 1 A flow chart of a method for intelligently controlling fans in a ventilation system of a nuclear power plant provided in an embodiment of the present application;

[0018] Figure 2 A schematic block diagram of the structure of the fan intelligent control system of the nuclear power plant ventilation system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0020] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may change based on actual circumstances.

[0021] It should also be understood that the terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0022] It should be further understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0023] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features in the embodiments may be combined with each other.

[0024] See also Figure 1 , Figure 1A flow chart of a method for intelligently controlling fans in a ventilation system of a nuclear power plant provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the intelligent control method for fans of a nuclear power plant ventilation system provided in an embodiment of the present application includes steps S100 to S600.

[0025] Step S100: collecting and preprocessing the pressure difference data, radioactivity monitoring data, fan operating parameters, and environmental data of the nuclear power plant ventilation system to obtain a labeled data set and an unlabeled data set;

[0026] It is understandable that the execution subject of the present invention may be a fan intelligent control system of a nuclear power plant ventilation system, or a terminal or a server, which is not limited here. The embodiment of the present invention is described by taking a server as the execution subject as an example.

[0027] Specifically, a multi-source data acquisition network system was established within the nuclear power plant's reactor building, auxiliary building, and fuel building. This system provides real-time and continuous monitoring of internal environmental parameters, operating status, and safety indicators. The data acquisition network includes various sensors and instrumentation located at various locations within the building. These sensors collect data on internal and external pressure differentials, radioactivity monitoring data (including alpha, beta, and gamma radiation and aerosol concentrations), fan operating parameters (speed, power, flow, and vibration), ventilation system filter pressure differentials, and ambient temperature and humidity data, forming a raw monitoring dataset encompassing all key operating parameters. The collected raw monitoring data undergoes preliminary organization and preprocessing. Specialized outlier detection and missing value imputation are performed on multiple parameter categories within the raw dataset. Missing values are handled using a time interpolation method based on local linear regression to ensure data integrity. Outliers are identified using a modified Z-score method with a threshold of ±3.5. Data points outside the threshold are replaced with the local median, effectively ensuring data reliability and accuracy. After initial data processing, the data is standardized to eliminate differences in dimension and magnitude between different monitoring data. The Min-Max normalization method is used for standardization, mapping all data values to the standard interval [0, 1] to obtain standard monitoring data. Furthermore, considering that the data originates from multiple monitoring devices and may have timing deviations between acquisition devices, to ensure accurate data analysis, the timestamps of the standard monitoring data are unified to a unified standard time coordinate across the nuclear power plant to achieve temporal consistency and form a target dataset. This unified target dataset is then partitioned. To more effectively train the intelligent model, experts identify and label various operating conditions in the data, identifying the optimal wind turbine control parameters for each condition. Based on the labeling, approximately 10% of the target dataset is partitioned into a labeled dataset for supervised model learning. The remaining approximately 90% of the data, consisting solely of monitoring parameters without explicit control labels, is defined as an unlabeled dataset for semi-supervised learning. After partitioning, the labeled and unlabeled datasets are stored in separate databases to support efficient intelligent learning and model training. To ensure the reliability and validity of the divided data set, after completing the above data division, the data quality of the labeled and unlabeled data is evaluated and checked separately, and a data quality report is generated to clearly state the completeness rate, abnormal data rate and timestamp alignment of the data at each monitoring point in the report.

[0028] Step S200: performing semi-supervised contrastive learning and multi-condition data enhancement processing on the labeled dataset and the unlabeled dataset to obtain a feature representation space;

[0029] Specifically, corresponding data augmentation operations are performed on labeled and unlabeled datasets to improve the diversity of training data and enhance the robustness and generalization performance of the model to changes in monitoring data under different working conditions. For the operational data of the nuclear power plant's ventilation system, several different data augmentation strategies were employed. These included adding additive Gaussian noise to the raw data, with the noise intensity strictly controlled to approximately 5% of the original signal standard deviation. This effectively simulates random fluctuations in the data while preventing excessive noise from disrupting the data's distributional characteristics. Furthermore, random time window shifting was performed to address the characteristics of time series data. This involves randomly shifting the data window within a certain time window (e.g., ±60 seconds), thereby increasing the model's adaptability to slight drift in the data's time axis. Furthermore, a selection of key monitoring parameters (such as plant pressure differentials and radioactivity levels) was randomly amplified by a factor generally between 1.1 and 1.3 to simulate the local variations that can occur in actual monitoring data under different operating conditions. To further improve the model's generalization, approximately 15% of non-critical monitoring parameters were randomly masked and reset to zero to simulate data loss or sensor failure during data acquisition. Through these methods, each raw data sample was augmented with several augmented samples, forming an augmented dataset. Based on the above-mentioned enhanced dataset, a 5-layer deep neural network encoder for feature learning is constructed. The network structure consists of multiple fully connected layers. The specific structure is designed as follows: the input data dimension is mapped to a 512-dimensional hidden layer feature space through a first fully connected layer, and then batch normalization and ReLU activation functions are introduced, combined with a Dropout layer to reduce the risk of overfitting. After that, it is mapped to a 256-dimensional feature space through a second fully connected layer, and also processed by batch normalization, ReLU activation and Dropout. Then, it is further compressed to 128 dimensions through a third fully connected network, and batch normalization is performed again. The fourth fully connected layer is used to reduce the features to the final required 64-dimensional feature representation space. Through deep and gradual dimensionality reduction and nonlinear transformation, the model can capture the inherent differences and potential connections between data of different operating conditions, forming a good feature representation effect. On this basis, an unsupervised contrastive loss function is constructed, employing a normalized temperature-adjusted cross-entropy loss function. This loss function adjusts the sensitivity of feature distances between different data points via a temperature coefficient, set to approximately 0.07. This ensures that the model more effectively distinguishes between different samples, bringing enhanced sample pairs generated from the same original sample closer together in feature space, while pairs of different samples are further apart. Furthermore, considering that some samples in the data used carry clear operating condition labels, an additional supervised contrastive loss function is introduced for these labeled data samples. This explicitly constrains data samples of the same operating condition category to be as close as possible in feature space, thereby forming a clear operating condition clustering effect.The supervised contrastive loss and unsupervised contrastive loss functions are then combined according to certain weight coefficients to form a joint semi-supervised contrastive learning loss function. The weight coefficients increase linearly with the training process to gradually strengthen the influence of the supervisory signal. The established semi-supervised contrastive learning loss function is used to train the aforementioned five-layer deep feature encoder network. The network parameters are adjusted using the Adam optimizer, combined with a cosine annealing learning rate adjustment strategy to achieve effective convergence of the encoder parameters. After sufficient training, the final feature encoder model is obtained. All labeled and unlabeled datasets are input into the trained feature encoder model for forward propagation mapping. Ultimately, all data points are projected into a 64-dimensional feature space. Within this space, data points with similar working conditions will naturally cluster into tight clusters, while data points with different working conditions will be separated from each other, forming a high-quality working condition feature representation space that can be used for subsequent multi-agent reinforcement learning decision-making.

[0030] Step S300: Establish an action space including the main exhaust fan, auxiliary exhaust fan, and filtration system control based on the feature representation space and the target monitoring parameters, and perform multi-agent analysis to obtain an initial fan control strategy;

[0031] Specifically, based on the feature representation space, the system combines several target monitoring parameters monitored in real time during nuclear power plant operation, including the reactor building pressure differential, auxiliary building pressure differential, and fuel building pressure differential, as well as key indicators such as the maximum radioactivity level in each area of the building, ambient temperature, and humidity. These real-time target monitoring parameters reflect the critical state information of the nuclear power plant ventilation system during operation and have a significant impact on control strategy decisions. Therefore, they are effectively combined and spliced with the feature representation space obtained earlier to construct an overall state vector that fully reflects the current operating state of the system. The system's action space is defined, covering the entire controllable range of the nuclear power plant ventilation system's main exhaust fans, auxiliary exhaust fans, and filtration system. Therefore, the action space needs to be reasonably quantified and clearly expressed. For the main exhaust fan control instructions, the fan speed range is carefully discretized from 0% to 100%, quantized into M discrete value points. For the auxiliary exhaust fan, not only speed adjustment is involved, but also the fan start and stop states. Therefore, a clear distinction is made between the two states 0 (off) and 1 (on), and the speed is similarly subdivided into N value points, bringing the auxiliary exhaust fan action space to a higher dimension. At the same time, for the filtration system, a crucial component of the ventilation system, its start and stop combinations are logically quantized according to the actual system design, resulting in a total of F valid start and stop state combinations. Through the above action space definition and combination, a clear and executable discrete action space is obtained, which encompasses all control situations that may need to be considered in the actual operation of the ventilation system. Based on the structure and control logic of the nuclear power plant's ventilation system, five specialized agents were designed, each responsible for ventilation control tasks in a different area. These agents include a reactor building main exhaust fan control agent, an auxiliary building main exhaust fan control agent, a fuel building main exhaust fan control agent, a building auxiliary exhaust fan control agent, and a filtration system control agent. Each agent is designed based on an actor-critic architecture. The actor network maps the constructed state vector into specific action outputs, such as the fan speed and filtration system start / stop status for a specific area. The critic network accurately assesses the value of the current state, providing effective value feedback for the agent's actions. This network architecture utilizes a multi-layer, fully connected network structure and nonlinear activation functions (such as LeakyReLU) to ensure training stability and policy accuracy. To achieve collaborative control and information sharing among agents, an agent communication scheme based on an attention mechanism was designed and implemented. This communication mechanism allows each agent to share its current state information in real time. Attention weighting is also used to effectively extract and aggregate key state information from other agents, thereby enabling efficient integration of collaborative control strategies among agents. The proximal policy optimization (PPO) algorithm is used to train the entire multi-agent system.During policy gradient optimization, the PPO algorithm clips the objective function to avoid instability caused by excessive policy updates and improve the robustness of the training process. During training, the discount factor γ is explicitly set to 0.95, the generalized advantage estimate (GAE) parameter λ is set to 0.97, and the clipping parameter ε is set to 0.2. These parameters ensure a stable and efficient training process. With continuous iterations and optimizations in training, each agent gradually learns how to make more appropriate wind turbine control actions based on state information, leading to gradual convergence of the multi-agent system and ultimately generating an initial wind turbine control policy with good performance and high coordination.

[0032] Step S400: construct a four-layer cascade safety constraint model based on the initial wind turbine control strategy, and solve the constraint optimization problem by the projected gradient method to obtain the optimal wind turbine control strategy.

[0033] Specifically, a temporal causal dependency analysis was conducted on the state variables during the initial fan control strategy operation. These state variables primarily included key parameters of the nuclear power plant's ventilation system, such as fan speed, building pressure differential, radioactivity level, and aerosol concentration. Based on actual nuclear power plant operating experience and historical data, correlation analysis and statistical methods were used to clarify the causal relationships between different variables and the specific temporal delay characteristics between them. For example, there is typically a response delay of approximately 15 to 30 seconds between fan speed adjustment and changes in building pressure differential, a 30 to 60 second delay between changes in radioactivity level and ventilation path switching, and a 45 to 120 second delay between filter activation and changes in aerosol concentration. After clarifying the temporal dependencies between these variables, a well-structured temporal causal graph was constructed. By assigning causal strength weights and specific delay times to the graph, the temporal causal relationships between the state variables were graphically expressed. To rationally reduce and aggregate high-dimensional, time-series state information, a corresponding time-series aggregation model is developed based on the temporal causal graph constructed above. The specific mathematical expression of the aggregation function is clarified. The aggregated state space must take into account the current state, historical state, and predictions of future state trends. The aggregation function considers the exponentially weighted average of each state variable's current value and historical state values, as well as the variable's temporal trend. The resulting state vector is smaller in dimension but more compact in information, effectively representing the dynamic and temporal characteristics of the ventilation system. The aggregation operation is implemented via a linear combination of weighted summations. The weight matrix is determined through model optimization to ensure that the aggregated feature state fully reflects the causal dependencies between variables while reducing the complexity of the model state space. Based on the existing initial fan control strategy, a four-layer cascade safety constraint model is established to meet the stringent safety requirements of nuclear power plant ventilation systems. The first layer consists of basic operating constraints, which specify that equipment operating parameters such as fan speed and filtration system status must always remain within safe and stable physical limits. The second layer consists of operating condition constraints, which set specific ventilation requirements for different nuclear power plant operating modes, such as normal operation, outage and maintenance, and accident response. These requirements include minimum negative pressure maintenance, specific air volume, and filtration efficiency. The third layer consists of radiological protection constraints, which explicitly require that any control strategy maintain a continuous negative pressure inside and outside the plant and minimize the risk of radioactive material leakage. The fourth layer consists of fault response constraints. When the monitoring system identifies an equipment anomaly or fault, it automatically activates a predefined safety response control strategy to minimize the risk of an accident and ensure the safe and stable operation of the nuclear power plant. These four layers of safety constraints work together through a combination of hard and soft constraints. Hard constraints define the scope of the action space, while soft constraints are reflected in the control strategy's optimization objective function through additional penalty terms.The constructed temporal causal aggregation state space is combined with a four-layer cascade safety constraint model to optimize the control strategy for nuclear power plant fans. The projected gradient method, a highly efficient numerical optimization algorithm, is used to solve the problem. The projected gradient method applies constrained projections to the gradient updates during the optimization process, ensuring that the solution obtained at each optimization step consistently satisfies the aforementioned safety constraints. This achieves both performance goals (such as ventilation efficiency, energy consumption reduction, and precise maintenance of plant pressure differentials) and strict nuclear safety constraints. The projected gradient method solves the problem by iteratively approaching the optimal solution and ensuring that the iterations converge to the optimal solution. This series of steps ultimately results in the optimal fan control strategy.

[0034] Based on the aggregated state space, an initial value function approximator is constructed for control strategy optimization. This value function approximator aims to estimate the long-term reward that the system can obtain under a given aggregated state to guide the selection of wind turbine control strategies. Specifically, a multi-layer neural network architecture is selected, with a three-layer network structure as the foundation. The input layer receives the feature vector of the aggregated state space, passes through multiple hidden layers (e.g., the first hidden layer uses 64 nodes, the second hidden layer uses 32 nodes, and the hidden layers typically use nonlinear activation functions such as Tanh), and then outputs a scalar state value estimate. To efficiently and stably update the parameters of this value function, a value function parameter update mechanism based on a combination of temporal difference learning and stochastic approximation is designed. At each training step, the difference between the observed immediate reward of the current state and the value function prediction of the next state is calculated to form a temporal difference error, which is used to guide the gradient update of the network parameters. Furthermore, a stochastic approximation method is introduced to ensure that the learning rate decreases gradually with the number of training steps, enabling the network parameters to converge efficiently and stably toward the true long-term reward value. To improve the efficiency and stability of the training process, an experience replay training system is constructed based on the aforementioned value function parameter update mechanism. This system primarily stores historical transition samples consisting of states, actions, rewards, and next states in a dedicated experience buffer, designed to hold a large capacity (e.g., hundreds of thousands of records). Each time the network parameters are updated, a batch of data is randomly sampled from the experience buffer for training, reducing the correlation between the training data and preventing the network from falling into local optima. To mitigate overestimation and instability during training of a single value function approximator, a dual-network architecture consisting of a target network and an evaluation network is constructed. The evaluation network updates parameters in real time, while the target network slowly copies parameters from the evaluation network at a lower frequency to stabilize the training process. Building on this stable value function training framework, and in accordance with the stringent safety constraints required in the actual operation of nuclear power plant ventilation systems, the previously established four-layer cascaded safety constraint model is converted into a safety probability function suitable for policy optimization. This probability function quantitatively expresses the probability that a given action satisfies all safety constraints in the current state. Combined with the aforementioned value function, a safety-aware policy optimizer is formed. During the specific implementation process, the objective function of the safety constraint policy optimizer combines the action value function and the safety probability function, and dynamically adjusts the importance of the safety constraint through the safety weight factor to form a constrained optimization problem that simultaneously optimizes long-term returns and strictly meets nuclear safety conditions. The specific method for solving this optimization problem is the projected gradient method, that is, during the gradient update process of the policy parameters, the projection method is used to always project the results of the parameter update into the feasible area defined by the safety constraints, thereby effectively ensuring that each optimization step meets the safe operation conditions.To ultimately achieve the optimal, applicable fan control strategy, a large-scale offline training program was conducted using the strategy optimizer under the aforementioned safety constraints, utilizing historical operating data accumulated over a long period of time from the nuclear power plant's ventilation system under various operating conditions. This training process encompassed not only routine operating data under normal operation, but also data from special operating conditions such as outages for maintenance, accident response, and power adjustments, ensuring that the training process encompassed all possible operational scenarios during actual system operation. Through extensive iteration and optimization training, the resulting fan control strategy maintained efficient control performance while strictly adhering to and meeting all safety requirements for nuclear power plant operation.

[0035] Based on the optimal fan control strategy, a fan control execution system consisting of a data interface layer, a state representation layer, a decision layer, and an execution layer is constructed. The data interface layer continuously and in real time obtains the latest environmental data, pressure differential data, radioactivity monitoring data, and fan operating parameters from the on-site monitoring network to ensure the real-time and accuracy of the input data. The state representation layer, using a previously trained feature encoder and temporal causal aggregation model, rapidly converts real-time monitoring data into a state vector that reflects the actual operating conditions of the current nuclear power plant ventilation system. The decision layer then generates corresponding fan control instructions in real time using a well-trained multi-agent reinforcement learning model and the optimized optimal fan control strategy, combined with the current state vector. The execution layer effectively converts the control instructions into specific equipment operation commands and accurately transmits them to the fan's actual control equipment, thus completing the closed-loop control process. After establishing the above-mentioned fan control execution system, a hierarchical control coordination mechanism is implemented, and the intelligent control strategy of the nuclear power plant ventilation system fan is divided into three different control levels according to the control accuracy and scope of action, including the strategic layer for operating condition identification and overall ventilation strategy formulation. This layer is mainly responsible for identifying the overall operating status and determining the ventilation strategy for the entire plant area, guiding and coordinating the realization of global ventilation goals; next is the tactical layer for coordinated control of the plant ventilation system. This layer is responsible for coordinated control of ventilation equipment in different plant areas, and is decomposed into specific execution strategies for each plant area according to the goals of the strategic layer; and finally, the operational layer for precise control of individual fans. This layer is directly responsible for detailed and accurate real-time control of the speed, start and stop status of individual fans and the start and stop actions of the filtration system, so as to accurately implement the operating goals and safety requirements formulated by the strategic and tactical layers. To ensure system stability and safety during fan control adjustments, a progressive control adjustment mechanism was designed. When the fan speed adjustment command generated by the control system exceeds a specific threshold (e.g., 20%), the system automatically breaks down the large adjustment command into a series of smaller incremental adjustments (e.g., no more than 5%). Each step is separated by a fixed, short time interval, ensuring a smooth transition of the fan state and effectively avoiding large fluctuations in system pressure or operating status. Furthermore, a safety fallback mechanism was established based on the progressive control adjustment mechanism. This mechanism continuously monitors key safety parameters such as plant pressure differential and radioactivity levels in real time. If the system detects a parameter anomaly or exceeds a set safety threshold, a pre-designed emergency response strategy is immediately activated to quickly restore the system to a safe operating state, ensuring that ventilation safety in the nuclear power plant is not affected. To achieve continuous optimization and performance improvement of the fan intelligent control system in practical applications, an online learning and model update mechanism was constructed. This mechanism relies on the real operating data such as state-action-reward continuously collected by the wind turbine control system during operation, and performs online processing of incremental data at a fixed period (for example, every 7 days). It fine-tunes the value function network and policy network parameters through a small learning rate to continuously improve the control model's adaptability to real environmental changes.At the same time, a comprehensive model performance monitoring and diagnosis module has also been implemented. This module regularly evaluates the system's control effectiveness from multiple dimensions, including pressure difference control accuracy, radioactivity level management effectiveness, fan energy consumption, safety and other indicators. Once any indicator is found to drop below the preset warning threshold during the monitoring process, the control system will automatically trigger the abnormal diagnosis process, quickly locate the cause of the performance degradation and the specific model components, and adjust and update the corresponding model in real time to quickly restore the fan control system to the optimal performance state. Through the above process, a set of intelligent control execution solutions for nuclear power plant ventilation system fans with real-time perception capabilities, active adjustment and adaptive continuous optimization features are finally obtained, ensuring safe, reliable, stable and efficient operation performance in the actual complex operating environment.

[0036] In an embodiment of the present invention, semi-supervised contrastive learning and multi-condition data augmentation techniques enable effective training using a small amount of labeled data and a large amount of unlabeled data, addressing the difficulty of obtaining labeled samples in nuclear power plant environments. A contrastive learning mechanism based on the NT-Xent loss function, combined with four data augmentation methods, improves the model's ability to recognize and adapt to various ventilation conditions. A multi-agent system consisting of five specialized agents, coupled with an attention communication mechanism, achieves effective coordination between different control units. Through temporal causal graph modeling and aggregation function design, the temporal causal dependencies between key state variables in fan control are accurately captured, compressing the original state space into an aggregated state space. A four-layer cascaded safety constraint model, combined with a dual-value network architecture, ensures that all control strategies strictly adhere to nuclear safety requirements. Approximate dynamic programming and experience replay techniques enable the control system to generate optimal control strategies directly from environmental monitoring data. A hierarchical control coordination mechanism, combined with progressive control adjustment, automatically decomposes large changes in fan speed into small steps, preventing system instability. This invention can adapt to long-term changing factors, ensuring the long-term efficient and reliable operation of the ventilation system.

[0037] In a specific embodiment, the process of executing step S100 may specifically include the following steps:

[0038] Set up a multi-source data acquisition network in the reactor building, auxiliary building, and fuel building of the nuclear power plant and collect original monitoring data sets;

[0039] The original monitoring data sets, including the pressure difference between the inside and outside of the plant, α, β, γ radiation and aerosol concentrations, fan speed, power, flow, vibration, filter pressure difference, and ambient temperature and humidity, were classified and sorted to obtain classified monitoring data.

[0040] Normalize the classified monitoring data to obtain standard monitoring data, and unify the timestamps of the standard monitoring data to the standard time of the nuclear power plant to obtain the target data set;

[0041] The data are divided into 10% of the target data set containing the optimal wind turbine control parameters under various operating conditions as the labeled data set, and the remaining 90% of the target data set containing only monitoring parameters as the unlabeled data set.

[0042] Specifically, a multi-source data acquisition network is established within the nuclear power plant's reactor building, auxiliary building, and fuel building. This network comprises sensors and monitoring equipment of various types and locations, such as differential pressure sensors, radioactivity detectors (capable of detecting alpha, beta, and gamma rays), aerosol concentration monitors, fan speed sensors, power meters, flow meters, vibration sensors, filter differential pressure sensors, and ambient temperature and humidity sensors. These monitoring devices continuously acquire real-time information on the operating environment inside and outside the nuclear power plant, as well as the operating status of the fans, based on a predetermined collection cycle. This generates a raw monitoring dataset covering all key operating parameters. For example, multiple differential pressure monitoring points are installed within the reactor building to monitor the pressure difference between inside and outside the building in real time, ensuring a safe negative pressure inside the building. Radiation monitoring equipment monitors the real-time dose rates of alpha, beta, and gamma rays at various locations within the building, as well as the concentration of aerosol particles in the air, ensuring that radioactivity levels remain within nuclear power safety standards. The characteristics of the raw monitoring dataset are professionally classified and organized to clearly distinguish the categories of different monitoring parameters and form a unified and standardized classified monitoring dataset. During the classification and sorting process, the pressure difference data inside and outside the plant, radioactivity monitoring data (α, β, γ rays and aerosol concentrations), fan operation data (speed, power, flow, vibration), filter pressure difference and environmental data (temperature, humidity) are classified and stored separately for subsequent unified data processing. For example, the data are clearly numbered and named according to the data source, parameter nature and monitoring location, such as defining them as pressure difference data sets. , radioactivity dataset , wind turbine operation data set To improve the convenience of data management. In order to eliminate the differences in dimensions and scales between data, the classified monitoring data are normalized uniformly, and the normalization method after Min-Max transformation is used. Mapped to the interval [0,1]. The specific formula in the normalization process is:

[0043]

[0044] in, represents the normalized data, is the original data, and Respectively represent the maximum and minimum values of this type of parameter data. Align the timestamps of all normalized classified monitoring data to the unified standard time of the nuclear power plant, eliminating the slight time series differences generated by different monitoring devices during the sampling process, and ensuring the uniformity and time series accuracy of the data during subsequent model training. For example, the timestamps of data collected by different devices are uniformly calibrated to the same nuclear power plant clock server to ensure that each data sample is accurately aligned on the time axis, forming a precisely synchronized target data set. The target dataset is divided into supervised and semi-supervised learning. Based on expert evaluation and historical experience, about 10% of the data in the target dataset that clearly records the optimal wind turbine control parameters under different operating conditions is separated to form a labeled dataset. The remaining 90% of the data that do not have clear control labels and only contain environmental monitoring parameters are classified as unlabeled datasets. , which is used for subsequent semi-supervised contrastive learning and unsupervised feature extraction. Through a systematic data quality assessment and diagnosis mechanism, both labeled and unlabeled datasets are evaluated for data integrity, abnormal data rate, and temporal consistency, ensuring that the data quality ultimately used for intelligent control model training meets strict nuclear safety operation standards.

[0045] In a specific embodiment, the process of executing step S200 may specifically include the following steps:

[0046] Perform data enhancement on labeled datasets and unlabeled datasets to generate enhanced datasets;

[0047] Construct a 5-layer neural network encoder based on the enhanced dataset and use it as the feature encoder network;

[0048] An unsupervised contrastive loss function is constructed for the feature encoder network based on the normalized temperature-adjusted cross-entropy loss function;

[0049] Introducing supervised contrast loss to samples from labeled datasets to cluster features of samples under the same working conditions, and combining supervised contrast loss function with unsupervised contrast loss function to form semi-supervised contrast learning loss function;

[0050] The feature encoder network is trained based on the semi-supervised contrastive learning loss function to obtain a trained feature encoder model;

[0051] The labeled dataset and the unlabeled dataset are input into the trained feature encoder model for mapping, data points with similar working conditions are clustered, and data points with different working conditions are separated to obtain the feature representation space.

[0052] Specifically, effective data augmentation is performed on the classified and organized labeled and unlabeled datasets to expand the dataset capacity and improve the model's generalization and stability across a variety of operating conditions. For example, controlled additive Gaussian noise is added to the raw data to simulate random errors in measuring instruments and sensors, with the noise variance set to approximately 5% of the standard deviation of the original signal. Furthermore, for monitoring data with time series characteristics, small random shifts of the time window are performed, for example, randomly shifting raw data segments within a ±60-second range, to enhance the model's robustness to data acquisition device delays or clock asynchrony. For key radioactivity monitoring parameters (such as gamma-ray intensity and aerosol concentration) or plant pressure differentials, a localized random amplification strategy is implemented, for example, by randomly selecting amplification factors of 1.1 to 1.3 to reflect the diversity of parameter fluctuations under different operating conditions. For non-critical parameters, approximately 15% of the data is randomly masked to zero to simulate localized sensor failures that can occur during actual operation. Through the combined application of these strategies, each original data sample can be derived into multiple enhanced data samples, forming a richer enhanced data set. Based on the enhanced data set obtained above, a deep neural network encoder with a 5-hidden layer structure is designed and built. The encoder is used to learn the potential feature information in the data and realize the efficient mapping of high-dimensional data to low-dimensional feature representation. For example, if the input data dimension is , the output dimension of the first hidden layer of the network is 512, which is mapped to 256 and 128 layer by layer after activation function, batch normalization and Dropout, and the final output dimension is 64. The network is expressed as the following inter-layer nonlinear transformation formula:

[0053]

[0054] Among them, the variable Indicates the The feature output of the layer network, For the The weight matrix of the layer, For the Layer bias vector, ReLU is a nonlinear activation function, BN stands for batch normalization operation, and Mask Represents the Dropout function, parameter Indicates randomly shutting down 30% of neurons to prevent overfitting, and finally obtaining an encoded 64-dimensional feature vector. In the network structure, an unsupervised loss function is constructed through unsupervised contrastive learning method, and the normalized temperature is used to adjust the cross entropy loss. Its expression is:

[0055]

[0056] Among them, the variable 、 、 Represents the feature vector of the data-enhanced sample after mapping through the feature encoder, function sim Represents the feature vector and The cosine similarity between variables is used to measure the similarity between features of data; is the temperature coefficient, which is used to adjust the sensitivity of the distance between feature vectors and is set to 0.07 to ensure training stability; Indicates the number of data pairs used in each training. At the same time, in order to make full use of the working condition information carried in the labeled dataset, a supervised contrast loss function is additionally introduced for samples with clear labels, and its formula is defined as:

[0057]

[0058] The variables Representation and Sample The index set of all positive samples of the same working condition, A temperature parameter dedicated to supervised contrastive learning, The values are similar, but can be slightly adjusted according to the data clustering situation to better cluster data of similar working conditions. Combining the above supervised contrast loss function and unsupervised contrast loss function, with adjustable coefficients, a complete semi-supervised contrast learning loss function is formed:

[0059] L

[0060] Among them, the variable The initial value is 0.5 and gradually increases to 2.0 during training to gradually enhance the constraint effect of supervisory information on feature space construction. Using this semi-supervised contrastive learning loss function, the Adam optimization algorithm is used to iteratively train the neural network encoder until it converges to a stable state, obtaining a trained feature encoder model. Labeled and unlabeled datasets are input into the trained encoder model and mapped to a 64-dimensional feature space through network forward propagation. Within this space, the feature encoder automatically clusters data points with similar operating conditions and clearly separates data points with different operating conditions, thereby constructing an effective and highly generalizable feature representation space for nuclear power plant ventilation systems.

[0061] In a specific embodiment, the process of executing step S300 may specifically include the following steps:

[0062] The state vector is obtained by combining the feature representation space with the target monitoring parameters, including the pressure difference of the reactor building, the pressure difference of the auxiliary building, the pressure difference of the fuel building, the maximum radioactivity level of each building, the ambient temperature and the ambient humidity.

[0063] The main exhaust fan speed percentage is quantified into M discrete values, the auxiliary exhaust fan switch state and speed percentage are quantified into N discrete values, and the start and stop state of the filtration system are quantified into F combinations, and the action space is obtained through the combination;

[0064] Based on the state vector and action space, five agents were designed to control the reactor building main exhaust fan, auxiliary building main exhaust fan, fuel building main exhaust fan, auxiliary building exhaust fan, and filtration system, respectively. This resulted in a multi-agent control system. A strategy generation and value evaluation network was then constructed for each agent in the multi-agent control system.

[0065] According to the multi-agent control system, a communication mechanism between agents is implemented through state sharing and attention weighting, and the PPO algorithm is used for training to obtain the initial wind turbine control strategy.

[0066] Specifically, a complete state description is formed by combining the 64-dimensional feature representation space vector obtained in the previous step with the real-time key parameters actually monitored at the nuclear power plant. The target monitoring parameters include the reactor building pressure differential, auxiliary building pressure differential, fuel building pressure differential, the maximum radioactivity level in each building, ambient temperature, and ambient humidity. To achieve a refined description and efficient execution of specific control instructions, an action space is constructed. The action space consists of three parts: a discretized space for the main exhaust fan speed percentage, a discretized space for the auxiliary exhaust fan start and stop switch states and speed percentages, and a combined space for the start and stop states of the filtration system. Taking the main exhaust fan speed percentage as an example, the speed percentage range (0% to 100%) is subdivided into M equally spaced discrete values, each discrete value representing a specific fan speed setting value; and the action space of the auxiliary exhaust fan, in addition to the fan speed percentage being subdivided into N discrete values, also introduces additional switch start and stop states (start is 1, stop is 0). For example, when the auxiliary exhaust fan is on state 1, its speed can be selected from a discrete speed percentage in the range of 0% to 100%, thereby clarifying the actual control strategy of the auxiliary exhaust fan; for the filtration system in the nuclear power plant ventilation system, since it involves multiple independent devices, the start and stop states of the filtration system are discretized into F combinations. If there are 3 groups of filtration systems in the nuclear power plant, then through different combinations of switch states, 2³=8 different start and stop combinations are obtained. The above parts are combined to form a total action space, where the total action dimension The innovative expression is:

[0067]

[0068] The variable M represents the number of discrete levels of main exhaust fan speed, N represents the number of combinations of auxiliary exhaust fan speed and on / off states, and F represents the number of combinations of filtration system start / stop states. This action space accurately and flexibly covers all possible control actions for the nuclear power plant's ventilation system. With the state vector and a clear action space, a multi-agent control system is established based on the actual operational requirements of the nuclear power plant. The multi-agent system comprises five specially designed agents, each responsible for operating the five subsystems: reactor building main exhaust fan control, auxiliary building main exhaust fan control, fuel building main exhaust fan control, auxiliary building auxiliary exhaust fan control, and filtration system control. Within each agent, a dedicated policy generation network and value assessment network are designed and constructed. The policy generation network (actor network) generates the probability distribution of specific actions, while the value assessment network (critic network) evaluates the value of the current state. Taking the reactor building main exhaust fan control agent as an example, its policy network is designed to take as input a 70-dimensional state vector, which is mapped into a probability distribution over a specific action space dimension (for example, 21 discrete speed values) through a multi-layer fully connected network. The value network outputs a scalar for the state vector, representing the expected long-term reward of the current state. To improve the collaboration and decision consistency between agents, an attention-based communication mechanism is introduced between agents. This communication mechanism expresses the importance of each agent in the form of an attention coefficient by weighted fusion of the state information of other agents. The attention mechanism formula is expressed as:

[0069]

[0070] in, represents the aggregated information received by agent i from other agents, Represents the state information held by agent j, and denote the attention query matrix and key matrix respectively, is the state vector of agent i itself, Used to scale attention weights. Through this mechanism, agents can selectively attend to the state information of other agents, effectively enabling information interaction and fusion between agents. For example, when a pressure anomaly occurs in the reactor building, the agents in the auxiliary and fuel buildings quickly perceive this state change through the attention mechanism and proactively coordinate dynamic adjustments to the corresponding fans, jointly ensuring the overall safety of the nuclear power plant building. The strategies and communication mechanisms constructed by these agents are effectively trained and optimized using the Proximal Policy Optimization (PPO) algorithm. The PPO algorithm uses a clipped probability ratio constraint to optimize the objective function, ensuring that each policy update does not deviate too far, avoiding training instability, and thus obtaining a stable and reliable initial fan control strategy. After training iterations, the control strategies generated by each agent can effectively coordinate to achieve precise fan control, stable pressure differential maintenance, and radioactive safety control under different operating conditions.

[0071] In a specific embodiment, the process of executing step S400 may specifically include the following steps:

[0072] Perform temporal causal dependency analysis on the state variables during the operation of the initial wind turbine control strategy to obtain the temporal causal relationship between the state variables;

[0073] Construct a temporal causal graph based on temporal causal relationships, and design an aggregated state space based on the temporal causal graph that considers the exponential weighted average of current state values, historical state values, and predicts future state changes;

[0074] A four-layer cascade safety constraint model was established to meet the safety requirements of nuclear power plant ventilation systems. The first layer of the four-layer cascade safety constraint model is the basic operating constraint that ensures that fan parameters are within physical limits. The second layer is the operating condition constraint that sets ventilation requirements for different nuclear power plant operating modes. The third layer is the radioactive protection constraint that ensures the maintenance of negative pressure in the plant and controls the release of radioactive materials. The fourth layer is the fault response constraint that activates the preset response strategy when equipment failure or anomaly is detected.

[0075] Based on the aggregated state space and four-layer cascade safety constraint model, the constrained optimization problem is solved by the projected gradient method to obtain the optimal wind turbine control strategy.

[0076] Specifically, a time causal dependency analysis is conducted on various key state variables involved in the actual operation of the initial control strategy to clarify the causal mechanism of the time-varying changes between system variables. For example, when the speed of the main exhaust fan increases, the pressure difference between the inside and outside of the plant does not reach a new stable value instantaneously, but there is a typical response delay period. For example, there is usually a time delay of about 15-30 seconds from the change in the speed of the main exhaust fan to the occurrence of the pressure difference change and tending to a stable state. This relationship is defined as a temporal causal relationship. Similarly, there is a delayed response relationship of about 30-60 seconds between the adjustment of the ventilation path and the change in radioactive concentration, and there is a correlation delay of 45-120 seconds from the activation of the filter to the significant decrease in aerosol concentration. Causal analysis can clearly reveal the dynamic evolution logic between state variables. Based on the above causal analysis results, a time causal graph is constructed. Each edge in the graph defines two key parameters: the causal strength weight and time delay parameters , where the causal strength weight Representing variables For variables The degree of influence of the time delay parameter Seconds describe the specific delay effect of this influence in the time dimension. Based on the temporal causal relationship and the constructed temporal causal graph, a state aggregation function is designed to achieve effective dimensionality reduction and aggregation expression of the original high-dimensional state space. The aggregation function is specifically expressed as:

[0077]

[0078] in, is the state space after aggregation, It is an innovative time series feature extraction function, specifically defined as:

[0079]

[0080] In the above formula, is a state variable At the current moment The real-time value of Indicates the influence of the historical state of other variables on the current variable determined by the time series causal diagram, where the coefficient 、 、 Represent the current value, historical weighted average, and trend change weights respectively, thereby effectively integrating the state information of past, present, and future trends to obtain an aggregated state space. On this basis, a four-layer cascade safety constraint model is established to ensure the safety and reliability of the nuclear power plant ventilation system. The first layer is the basic operation constraint, which clearly defines the physical limit constraints of basic equipment parameters such as fan speed percentage, power, flow, and filter status; the second layer is the operating condition constraint, which sets specific ventilation goals for different operation modes of the nuclear power plant, such as normal operation, maintenance status, and accident operation conditions; the third layer is the radioactive protection constraint, which strictly ensures the continuous stability of the negative pressure state inside the plant and effectively controls the release level of radioactive substances from exceeding the limit; the fourth layer is the fault response constraint. Once an abnormality or fault is detected in the system equipment, the predefined emergency control response plan is immediately activated. These four layers of constraint systems work together to ensure the safety of system operation with a comprehensive constraint mechanism. After defining the aggregated state space and the multi-level safety constraint model, the projected gradient method is used to solve the optimization problem with safety constraints to obtain the optimal fan control strategy. The objective function of this constraint optimization is expressed as:

[0081]

[0082] Among them, the variable Represents a specific control action in the action space, the variable Represents the aggregate state vector of the current system, function represents the long-term value function of executing an action in the current state, It means the current action is in state The security probability function that satisfies the four-layer security constraints is is a safety weighting factor used to balance the trade-off between system performance optimization and safety constraints. It initially takes a high value and gradually decreases to 1.0 during training to achieve the optimal balance between safety and performance. The projected gradient method is used to solve this optimization problem. This method projects the resulting parameter or policy update value back into the safe action space defined by the constraints at each gradient update, ensuring that the optimization result remains within the feasible region defined by the safety constraints. This ensures that the strategy generated during the optimization process always meets all safety requirements. This method continuously iterates and updates, ultimately converging to the optimal solution. This yields the optimal wind turbine control strategy.

[0083] In a specific embodiment, the execution step is based on the aggregated state space and the four-layer cascade safety constraint model, and the process of solving the constrained optimization problem by the projected gradient method to obtain the optimal wind turbine control strategy can specifically include the following steps:

[0084] An initial value function approximator is constructed based on the aggregated state space, and a value function parameter update mechanism based on temporal difference learning and random approximation is designed for the initial value function approximator.

[0085] Construct an experience replay training system based on the value function parameter update mechanism, and introduce a value function training framework consisting of a target network and an evaluation network into the experience replay training system;

[0086] Based on the value function training framework, the four-layer cascade safety constraint model is transformed into a safety probability function. The constraint optimization objective is constructed by combining the safety weight factor with the value function, and a policy optimizer under safety constraints is obtained.

[0087] The strategy optimizer under safety constraints uses historical data under different operating conditions of nuclear power plants for offline training to generate the optimal wind turbine control strategy.

[0088] Specifically, an initial value function approximator is designed based on the aggregate state space constructed in the previous step. The purpose of this approximator is to accurately estimate the value function corresponding to the long-term control performance of the nuclear power plant ventilation system under a given aggregate state, so as to effectively guide the subsequent intelligent decision-making process of the system. During the specific construction, a deep neural network composed of multiple hidden layers is used as the network structure of the initial value function approximator. For example, if the dimension of the input aggregate state space is 25 dimensions, the first hidden layer is designed to have a Tanh nonlinear activation function with 64 nodes, and the second hidden layer uses 32 nodes. The final output is a scalar state value estimate. This network structure not only helps to capture the nonlinear relationship between states, but also enhances the generalization ability of the value function approximation. After the initial value function approximator is constructed, a parameter update mechanism based on time difference learning and random approximation is designed to ensure that the value function parameters converge to the true optimal value function efficiently and stably during the training process. The value function parameter update formula:

[0089]

[0090] In the formula, The value function network is The parameter vector of the step update, is the value function estimate of the current state, is the discount factor, set to 0.95, represents the state space representation after aggregation, Represents the gradient update direction of the current state value function. The random approximation method introduced in the parameter update process is reflected in the gradual decrease of the learning rate with the number of training steps, thereby avoiding drastic fluctuations in the value function parameters and achieving smooth parameter updates and stable convergence. In order to improve the stability of value function network training and data utilization efficiency, an experience replay training system is established based on the above-mentioned value function parameter update mechanism. This training system creates a large-capacity experience buffer (for example, a capacity of 100,000 transfer samples) to store state, action, reward, and next state transfer data samples in real time, and randomly extracts training samples from the experience buffer to update the value function network to avoid excessive data correlation. At the same time, in order to avoid overestimation or instability of value function estimation during single network training, a value function training framework consisting of a target network and an evaluation network is introduced. The target network parameters are soft-updated, namely:

[0091]

[0092] in, represents the parameters of the evaluation network, Represents the parameters of the target network, parameters is a soft update coefficient, which is usually set to a small value such as 0.005 to stabilize the training process. In order to enable the model to meet the strict safety constraints of nuclear power plants, the previously constructed four-layer cascade safety constraint model is innovatively transformed into a safety probability function , specifically defined as an action In state The probability estimation function that satisfies all safety constraints is expressed as follows:

[0093]

[0094] Among them, the variable Indicates the Hierarchical security constraints in state Next action The safety margin value after the test is larger, the greater the margin is, the more safety conditions are met. It is used to adjust the sensitivity of different security levels. The sigmoid function converts the safety margin into a safety probability expression with a value between [0,1], so that the safety probability function It intuitively reflects the possibility of the action satisfying all safety constraints. Based on the above safety probability function, the strategy optimization objective function of safety perception is constructed. The specific formula is expressed as:

[0095]

[0096] In the above formula, Representative Strategy The expected value of the comprehensive optimization objective, Represents the state of a given aggregate Select Action Long-term value estimation, The safety weighting factor, initially high, such as 5.0, is gradually reduced to 1.0 over the training process to adjust the balance between safety and performance, ensuring that the strategy maximizes system performance while strictly satisfying safety constraints. Based on the established safety-constrained strategy optimization objective function and experience replay training system, offline training is performed using the nuclear power plant's long-term accumulated historical operating data (covering various operating conditions such as normal operation, power fluctuations, equipment maintenance, and accident response). During the training iterations, the projected gradient method is continuously used to constrain the optimization results of each step within the safety domain until the control strategy converges stably. The resulting intelligent wind turbine control strategy, after sufficient training, can significantly improve the stability and accuracy of the control strategy while strictly adhering to nuclear safety constraints.

[0097] In a specific embodiment, the method for intelligently controlling fans in a nuclear power plant ventilation system further includes the following steps:

[0098] According to the optimal wind turbine control strategy, a wind turbine control execution system is constructed, which includes a data interface layer, a state representation layer, a decision layer, and an execution layer.

[0099] A hierarchical control coordination mechanism is implemented based on the fan control execution system. The fan control is divided into a strategic layer for identifying working conditions and formulating overall ventilation strategies, a tactical layer for coordinating the control of ventilation systems in each plant building, and an operational layer for precise control of individual fans, resulting in a hierarchical control system.

[0100] Based on the hierarchical control system, a progressive control adjustment mechanism is designed, and a safety fallback mechanism is implemented based on the progressive control adjustment mechanism to obtain an emergency response strategy.

[0101] Build an online learning and model update mechanism through emergency response strategies, and perform model performance monitoring and diagnosis to obtain multi-dimensional evaluation indicators;

[0102] When it is detected that the multi-dimensional evaluation indicators are lower than the warning threshold, the abnormal diagnosis process is automatically triggered to identify the root cause of the problem and adjust the model components to obtain the wind turbine intelligent control execution plan.

[0103] Specifically, a fan control execution system is constructed based on the optimal control strategy. The system as a whole forms a top-down data circulation and execution mechanism, which specifically includes four important functional modules: data interface layer, state representation layer, decision layer and execution layer. The data interface layer collects actual data such as fan operating parameters, plant pressure difference, radioactivity level, ambient temperature and humidity from each plant building of the nuclear power plant in real time, and transmits it to the state representation layer for unified processing through a unified data format and communication protocol; the state representation layer uses the feature encoder model and time causal aggregation function established in the early stage to convert the real-time collected environmental monitoring parameters into an aggregated state vector that can directly reflect the current state of the system. This state representation combines the implicit features driven by data with the explicit feature information of the real-time monitoring parameters; the decision layer uses the trained optimal fan control strategy model to receive the aggregated state vector input and gives specific fan action control instructions through the strategy generation network; and the execution layer directly executes the above action instructions, adjusts the fan speed and filter switch status in real time, and completes the complete closed-loop intelligent control process. A hierarchical control coordination mechanism is implemented based on the fan control execution system, dividing the fan control tasks of the entire nuclear power plant ventilation system into strategic, tactical, and operational layers, so that the system can operate in an orderly and coordinated manner at different levels. The main tasks of the strategic layer are real-time operating condition identification and overall ventilation strategy decision-making, such as determining whether the current operating state is normal, negative pressure abnormality, or special operating conditions, and determining the overall ventilation strategic goals. The tactical layer coordinates the refined allocation of ventilation subsystems among each plant building based on the overall ventilation goals of the strategic layer. For example, when the pressure differential fluctuates, the reactor building actively coordinates the fan adjustments of other auxiliary buildings and fuel buildings. The operational layer is directly responsible for the precise execution of specific individual fans, such as adjusting the speed percentage of a main exhaust fan to 70%, turning on a certain filtration system, and other action details. The three layers work together to achieve the overall optimized control goal. To avoid drastic changes in fan control actions during actual operation, which may lead to a decrease in system stability or trigger safety hazards, a progressive control adjustment mechanism is designed based on the hierarchical control system. The progressive adjustment formula for the fan action adjustment amplitude control is defined as follows:

[0104]

[0105] In this formula, the variable The variable is the change in fan speed or state in the current control step. Indicates the optimal setting value of the current wind turbine control instruction given by the strategy network, and the variable Indicates the actual operating value of the fan at the last moment, the coefficient Indicates the number of steps of gradual adjustment (for example, set to 4, indicating that each larger action will be broken down into 4 steps of gradual adjustment), the function sign is used to maintain the consistency of the direction of the action change, and The maximum change allowed at a time is to prevent excessive control action from causing negative pressure in the nuclear power plant or unstable environmental conditions. To effectively deal with abnormal operating conditions and equipment failures during ventilation system operation, a safety fallback mechanism is designed based on the above-mentioned progressive control adjustment mechanism. This mechanism defines a real-time safety fallback trigger function:

[0106]

[0107] The variables Represents the state Next action Security Response Score, Parameters For the The weight coefficient of each safety monitoring variable, Indicates the actual safety status value monitored after the action is executed, and is the safety threshold standard for each parameter. For example, if the current negative pressure difference in the plant exceeds the safety threshold, the exponential term in the above formula will generate a large penalty value, automatically triggering the emergency response strategy of the safety assurance fallback mechanism, thereby quickly falling back to a safe state. In order to achieve continuous optimization and stable online operation of the intelligent control system of nuclear power plant fans, based on the aforementioned emergency response strategy, an online learning and model incremental update mechanism is constructed to continuously store new data generated by the system in real time in the experience replay database, and use the accumulated data to fine-tune and update the online model; the model performance monitoring and diagnosis module tracks the evaluation indicators of multiple dimensions in real time. When it is monitored that the multi-dimensional evaluation indicators are lower than the warning threshold, the abnormal diagnosis process is automatically activated to quickly analyze the cause of the problem and adjust the relevant model components, so as to effectively correct the deviations in the model control strategy and quickly restore the system to the optimal performance state.

[0108] See also Figure 2 , Figure 2 A schematic block diagram of the structure of the fan intelligent control system 200 of the nuclear power plant ventilation system provided in the embodiment of the present application is shown in FIG. Figure 2 As shown, the fan intelligent control system 200 of the nuclear power plant ventilation system includes:

[0109] The acquisition module 210 is used to collect and pre-process the pressure difference data, radioactivity monitoring data, fan operating parameters and environmental data of the nuclear power plant ventilation system to obtain a labeled data set and an unlabeled data set;

[0110] A data enhancement module 220 is used to perform semi-supervised contrastive learning and multi-condition data enhancement processing on the labeled dataset and the unlabeled dataset to obtain a feature representation space;

[0111] An execution module 230 is configured to establish an action space including the control of the main exhaust fan, the auxiliary exhaust fan, and the filtration system based on the feature representation space and the target monitoring parameters, and perform multi-agent analysis to obtain an initial fan control strategy;

[0112] The construction module 240 is used to construct a four-layer cascade safety constraint model based on the initial wind turbine control strategy, and solve the constraint optimization problem by the projected gradient method to obtain the optimal wind turbine control strategy.

[0113] Through the collaborative efforts of the aforementioned components, semi-supervised contrastive learning and multi-condition data augmentation techniques enable effective training using a small amount of labeled data and a large amount of unlabeled data, addressing the difficulty of obtaining labeled samples in nuclear power plant environments. A contrastive learning mechanism based on the NT-Xent loss function, combined with four data augmentation methods, improves the model's ability to recognize and adapt to various ventilation conditions. A multi-agent system consisting of five specialized agents, coupled with an attention communication mechanism, achieves effective coordination between different control units. Through temporal causal graph modeling and aggregation function design, the temporal causal dependencies between key state variables in fan control are accurately captured, compressing the original state space into an aggregated state space. A four-layer cascaded safety constraint model, combined with a dual-value network architecture, ensures that all control strategies strictly adhere to nuclear safety requirements. Approximate dynamic programming and experience replay techniques enable the control system to generate optimal control strategies directly from environmental monitoring data. A hierarchical control coordination mechanism, combined with progressive control adjustments, automatically decomposes large changes in fan speed into small steps, preventing system instability. This invention can adapt to long-term changing factors, ensuring the long-term efficient and reliable operation of the ventilation system.

[0114] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0115] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.

[0116] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for intelligently controlling fans in a ventilation system of a nuclear power plant, characterized in that: include: The pressure difference data, radioactivity monitoring data, fan operating parameters and environmental data of the nuclear power plant ventilation system are collected and preprocessed to obtain labeled and unlabeled data sets; Performing semi-supervised contrastive learning and multi-condition data enhancement processing on the labeled dataset and the unlabeled dataset to obtain a feature representation space; An action space including the control of the main exhaust fan, the auxiliary exhaust fan, and the filtration system is established based on the combination of the feature representation space and the target monitoring parameters, and a multi-agent analysis is performed to obtain an initial fan control strategy; Based on the initial wind turbine control strategy, a four-layer cascade safety constraint model is constructed, and the constraint optimization problem is solved by the projected gradient method to obtain the optimal wind turbine control strategy. Specifically, the method includes: performing a time causal dependency analysis on the state variables during the operation of the initial wind turbine control strategy to obtain the temporal causal relationship between the state variables; A time causal graph is constructed based on the temporal causal relationship, and an aggregated state space is designed based on the time causal graph to consider the exponential weighted average of the current state value and the historical state value and to predict future state changes; a four-layer cascade safety constraint model is established for the safety requirements of the ventilation system of the nuclear power plant, the first layer of the four-layer cascade safety constraint model is to ensure that the fan parameters are within the physical limit range of the basic operation constraint, the second layer is the operating condition constraint for setting ventilation requirements for different nuclear power plant operation modes, the third layer is the radioactive protection constraint for ensuring the maintenance of negative pressure in the plant and the control of the release of radioactive substances, and the fourth layer is the fault response constraint of the preset response strategy activated when a device failure or abnormality is detected; based on the aggregated state An initial value function approximator is constructed in the state space, and a value function parameter updating mechanism based on time difference learning and random approximation is designed for the initial value function approximator; an experience replay training system is constructed according to the value function parameter updating mechanism, and a value function training framework consisting of a target network and an evaluation network is introduced into the experience replay training system; according to the value function training framework, the four-layer cascade safety constraint model is converted into a safety probability function, and a constraint optimization objective is constructed by combining the safety weight factor and the value function to obtain a policy optimizer under safety constraints; based on the policy optimizer under safety constraints, offline training is performed using historical data under different operating conditions of the nuclear power plant to generate an optimal wind turbine control strategy.

2. The intelligent control method for fans of a nuclear power plant ventilation system according to claim 1, characterized in that: The pressure difference data, radioactivity monitoring data, fan operating parameters and environmental data of the nuclear power plant ventilation system are collected and preprocessed to obtain a labeled data set and an unlabeled data set, including: Set up a multi-source data acquisition network in the reactor building, auxiliary building, and fuel building of the nuclear power plant and collect original monitoring data sets; Classify and organize the pressure difference between the inside and outside of the plant, α, β, γ radiation and aerosol concentrations, fan speed, power, flow, vibration, filter pressure difference, and ambient temperature and humidity data in the original monitoring data set to obtain classified monitoring data; Normalizing the classified monitoring data to obtain standard monitoring data, and unifying the timestamps of the standard monitoring data to the nuclear power plant standard time to obtain a target data set; The data are divided into 10% of the target data set containing the optimal wind turbine control parameters under various operating conditions as a labeled data set, and the remaining 90% of the target data set containing only monitoring parameters as an unlabeled data set.

3. The intelligent control method for fans of a nuclear power plant ventilation system according to claim 1, characterized in that: The semi-supervised contrastive learning and multi-condition data enhancement processing are performed on the labeled dataset and the unlabeled dataset to obtain a feature representation space, including: Performing data enhancement on the labeled dataset and the unlabeled dataset to generate an enhanced dataset; Constructing a 5-layer neural network encoder based on the enhanced data set, and using the 5-layer neural network encoder as a feature encoder network; constructing an unsupervised contrastive loss function for the feature encoder network based on a normalized temperature-adjusted cross-entropy loss function; Introducing a supervised contrast loss to samples from the labeled dataset to cluster features of samples under the same working conditions, and combining the supervised contrast loss function with the unsupervised contrast loss function to form a semi-supervised contrast learning loss function; Training the feature encoder network based on the semi-supervised contrastive learning loss function to obtain a trained feature encoder model; The labeled data set and the unlabeled data set are input into the trained feature encoder model for mapping, data points of similar working conditions are clustered, and data points of different working conditions are separated to obtain a feature representation space.

4. The intelligent control method for fans of a nuclear power plant ventilation system according to claim 1, characterized in that: The step of establishing an action space including the control of the main exhaust fan, the auxiliary exhaust fan, and the filtration system based on the combination of the feature representation space and the target monitoring parameters, and performing multi-agent analysis to obtain an initial fan control strategy includes: Combining and concatenating the characteristic representation space with target monitoring parameters to obtain a state vector, wherein the target monitoring parameters include the reactor building pressure difference, the auxiliary building pressure difference, the fuel building pressure difference, the maximum radioactivity level of each building, the ambient temperature, and the ambient humidity; The main exhaust fan speed percentage is quantified into M discrete values, the auxiliary exhaust fan switch state and speed percentage are quantified into N discrete values, and the start and stop state of the filtration system are quantified into F combinations, and the action space is obtained through the combination; Based on the state vector and the action space, five agents are designed to be responsible for the control of the reactor building main exhaust fan, the auxiliary building main exhaust fan, the fuel building main exhaust fan, the auxiliary building exhaust fan, and the filtration system, respectively, to obtain a multi-agent control system, and a strategy generation and value evaluation network is constructed for each agent in the multi-agent control system; According to the multi-agent control system, a communication mechanism through state sharing and attention weighting is implemented between agents, and the PPO algorithm is used for training to obtain an initial wind turbine control strategy.

5. The intelligent control method for fans of a nuclear power plant ventilation system according to claim 1, characterized in that: The intelligent control method for fans of the nuclear power plant ventilation system further includes: Constructing a wind turbine control execution system comprising a data interface layer, a state representation layer, a decision layer and an execution layer according to the optimal wind turbine control strategy; A hierarchical control coordination mechanism is implemented based on the fan control execution system, which divides fan control into a strategic layer for identifying working conditions and formulating an overall ventilation strategy, a tactical layer for coordinating the control of ventilation systems in each plant building, and an operational layer for precise control of individual fans, thereby obtaining a hierarchical control system. Designing a progressive control adjustment mechanism based on the hierarchical control system, and implementing a safety assurance fallback mechanism based on the progressive control adjustment mechanism to obtain an emergency response strategy; Through the emergency response strategy, an online learning and model updating mechanism is constructed, and model performance monitoring and diagnosis are performed to obtain multi-dimensional evaluation indicators; When it is detected that the multi-dimensional evaluation index is lower than the warning threshold, the abnormal diagnosis process is automatically triggered to identify the root cause of the problem and adjust the model components to obtain the wind turbine intelligent control execution plan.

6. An intelligent fan control system for a nuclear power plant ventilation system, characterized in that: A method for intelligently controlling a fan of a nuclear power plant ventilation system according to any one of claims 1 to 5, wherein the intelligent fan control system of the nuclear power plant ventilation system comprises: The acquisition module is used to collect and preprocess the pressure difference data, radioactivity monitoring data, fan operating parameters and environmental data of the nuclear power plant ventilation system to obtain labeled data sets and unlabeled data sets; a data enhancement module, configured to perform semi-supervised contrastive learning and multi-condition data enhancement processing on the labeled dataset and the unlabeled dataset to obtain a feature representation space; an execution module, configured to establish an action space including the control of the main exhaust fan, the auxiliary exhaust fan, and the filtration system based on the combination of the feature representation space and the target monitoring parameters, and perform multi-agent analysis to obtain an initial fan control strategy; A construction module is used to construct a four-layer cascade safety constraint model based on the initial wind turbine control strategy, and solve the constraint optimization problem through the projected gradient method to obtain the optimal wind turbine control strategy.

Citation Information

Patent Citations

  • Nuclear power station failure diagnosis and state monitoring system based on wireless sensor network

    CN206413023U

  • Nuclear reactor coolant system main circuit arrangement structure

    WO2017028201A1