Biological outbreak risk assessment method for water intake of nuclear power plant

By constructing a biohazard risk assessment model based on a random forest classification algorithm and utilizing marine environmental monitoring data, the problem of assessing the increased biodiversity at nuclear power plant intakes was solved. This enabled real-time identification and classification of biohazard risks, thereby improving the operational safety and reliability of nuclear power plant heat sinks.

CN121010212APending Publication Date: 2025-11-25TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511101319.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing technologies lack systematic modeling of the increased marine biodiversity at nuclear power plant intakes and its synergistic effects, resulting in delayed responses and high uncertainty in the face of sudden bioburden events. This makes it impossible to effectively identify and assess the risk of bioburden blockage, affecting the stable operation of heat sinks.

Method used

A biohazard risk assessment model was constructed using a random forest classification algorithm. Through data preprocessing, sample balancing, and model training optimization, environmental indicators such as chlorophyll, salinity, temperature, and dissolved oxygen were used to construct feature vectors to monitor and identify biodiversity levels in real time. Combined with Z-score normalization and SMOTE oversampling techniques, the model's ability to identify high-risk levels was improved.

Benefits of technology

It enables efficient identification and classification of biological outbreak risks at nuclear power plant intakes, provides reliable decision support, improves the safety and stability of heat sink operation, and can promptly identify high-risk conditions and issue alarms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121010212A_ABST
    Figure CN121010212A_ABST
Patent Text Reader

Abstract

The invention discloses a biological outbreak risk assessment method for a water intake of a nuclear power plant, and belongs to the field of nuclear power plant operation safety and risk management, and the method comprises the steps: collecting temperature, dissolved oxygen, chlorophyll, salinity and biological abundance monitoring data through a sensor arranged at the water intake of the nuclear power plant, and carrying out the standardization processing; the method comprises the following steps: balancing the number of samples with different biological abundance levels by adopting minority class oversampling, constructing a random forest classification model, optimizing hyper-parameter data preprocessing through random search and cross validation, training a random forest model with the best generalization performance, performing nuclear power hot trap risk assessment, and outputting a biological disaster risk level in the current time period. When the grade is higher, an alarm is given; the biological outbreak risk assessment method for the water intake of the nuclear power plant is higher in adaptability and practicability, the risk level can be recognized in time, the operation and maintenance decision can be assisted, the safety and toughness of a heat trap of the nuclear power plant can be improved, and good expandability and engineering applicability are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of nuclear power plant operation safety and risk management, and particularly relates to a biological outbreak risk assessment method for a nuclear power plant water intake. BACKGROUND

[0002] Nuclear power plants generally use seawater as a cooling medium, and the water intake thereof is directly connected with a sea area, and ecological disturbance caused by complex marine environment is faced in the operation process. Among them, a large number of marine organisms such as jellyfish are easy to form blockage and adhesion in the water intake area, thereby reducing the cooling water flow and heat exchange efficiency, increasing the operating load of the heat sink, and threatening the safe and stable operation of the nuclear power unit. As a core component of the cooling source system of the nuclear power plant, the heat sink undertakes the key task of discharging the residual heat of the reactor and regulating the environmental heat load, and the stability of the operation directly affects the thermal safety boundary and overall operation reliability of the nuclear power plant. Therefore, it is urgent to carry out biological monitoring and evaluation of the heat sink of the nuclear power plant, and to build a perfect risk identification and response mechanism to effectively prevent the occurrence of biological blockage events.

[0003] At present, related researches mainly explore the biological invasion mechanism through ocean current characteristics analysis, historical observation data statistics and engineering experience deduction. At the same time, the influence mechanism research on biological disturbance of the nuclear power plant water intake is relatively limited, and the related achievements mostly stay at the level of empirical analysis and case review. The core of the marine biological disaster risk lies in the biological abundance level in the water body. When the abundance increases sharply in a short time and exceeds the threshold value, it is in the state of biological outbreak, which is easy to cause heat sink blockage, operation obstruction and other problems. However, the biological disturbance factors at the water intake of the nuclear power plant have dynamic, nonlinear and coupling characteristics. The existing methods lack the system modeling of the increase of biological abundance and its synergistic effect, resulting in a large uncertainty and a lag in response when facing sudden biological outbreak events.

[0004] Therefore, it is urgent to establish a disaster-causing biological risk assessment model for the heat sink, which can integrate marine environment and water intake biological monitoring data, realize real-time identification and abundance grading of marine organisms, and timely judge whether it is in a high-risk biological outbreak state, so as to provide scientific and reliable decision support for the operation of the heat sink. SUMMARY

[0005] The purpose of the present application is to provide a biological outbreak risk assessment method for a nuclear power plant water intake. In order to achieve the above purpose, the present application provides a biological outbreak risk assessment method for a nuclear power plant water intake, comprising the following steps:

[0006] S1, data preprocessing, based on long-term monitoring data at the water intake of the nuclear power plant, four environmental indicators closely related to biological abundance, such as chlorophyll, salinity, temperature and dissolved oxygen, are selected as input features to form a feature vector x=(x1, x2, x3, x4), and each sample is labeled as "first abundance", "second abundance", "third abundance" and "fourth abundance" according to the historical observation of the abundance level, wherein the first abundance and the second abundance correspond to a higher biological abundance level, gradually approaching or even reaching the biological explosion state of the water intake, indicating a high risk of heat trap failure; the third abundance and the fourth abundance correspond to a lower biological abundance level, indicating a low risk of heat trap failure, to reflect the difference in the risk of heat trap failure at different abundance levels, and Z-score standardization method is used to standardize all features;

[0007] S2, sample balancing, minority oversampling technology is used to synthesize minority class samples in the feature space, new minority class samples are generated in the feature space by interpolation, the number of samples at each risk level is balanced, so as to improve the recognition ability of the model to high risk level and avoid the interference of unbalanced data on the classifier;

[0008] S3, model construction, a random forest classification algorithm is used to construct a biological disaster risk assessment model for nuclear heat trap;

[0009] S4, model training and optimization, define the search space of hyperparameters, use five-fold stratified cross-validation, take classification accuracy as evaluation index, select the optimal parameter combination after random search, and train the random forest model with the best generalization performance;

[0010] S5, risk assessment, the real-time monitored environmental data is processed according to the same standardization parameters during training, and the normalized four-dimensional evaluation index feature vector x'=(x'1, x'2, x'3, x'4) is obtained, which is input into the trained random forest model, and the biological abundance level of the current time period is output. When the level is first or second, it is determined as a higher risk, and an alarm is issued.

[0011] Preferably, the calculation formula of S1 using Z-score standardization method is:

[0012]

[0013] Wherein, x j is the value of the jth feature in the original data, x′ j is the standardized jth feature value, μ j and σ j are the mean and standard deviation of the jth feature, respectively.

[0014] Preferably, the minority class oversampling technique in S2 is the SMOTE (Synthetic Minority Over-sampling Technique) technique.

[0015] Preferably, the specific steps of S3 are as follows:

[0016] S31, Bootstrap resampling method is used to extract k sample subsets D from the standardized and balanced sample set D i (i = 1, 2,... k) from the standardized and balanced sample set D i (x, θ i );

[0017] S32, for a new sample x to be tested, the random forest aggregates the classification results of the k trees through the majority voting mechanism, and outputs the final biological abundance grade Y * .

[0018] Preferably, the calculation formula of Y * is as follows:

[0019]

[0020] Where Y is the candidate grade, I(·) is the indicator function, θ i is the random parameter vector of the i-th tree.

[0021] Preferably, the search space of the hyperparameters in S4 includes the number of decision trees n_estimators, the maximum depth of the tree max-depth, the minimum number of samples split min-samples-split, the minimum number of samples in the leaf node min_samples_leaf, and whether to use bootstrap sampling bootstrap.

[0022] Preferably, in the stratified cross-validation in S4, the proportion of samples of different abundance grades in the training set and the validation set of each fold is consistent with the SMOTE balanced data set.

[0023] Therefore, the biological outbreak risk assessment method for the water intake of the nuclear power plant has the following beneficial effects:

[0024] (1) The present application can identify the biological abundance state of different risk grades based on real-time monitoring data, has high overall classification accuracy and balanced performance, and performs outstandingly in identifying high biological outbreak risk grades;

[0025] (2) It shows stable evaluation effect, can identify and grade the biological outbreak risk state, and provides reliable auxiliary decision support for the safe operation and maintenance of the nuclear power plant heat trap.

[0026] The technical solutions of the present application are described in further detail below with reference to the accompanying drawings and examples. BRIEF DESCRIPTION OF DRAWINGS

[0027] Figure 1 A flowchart of a biological outbreak risk assessment method for a nuclear power plant intake. DETAILED DESCRIPTION

[0028] The following detailed description of embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of protection of the present application.

[0029] EMBODIMENT

[0030] As shown in the accompanying drawings, the present application provides a biological outbreak risk assessment method for a nuclear power plant intake, comprising the following steps: Figure 1

[0031] S1, data preprocessing, based on long-term monitoring data at the nuclear power plant intake, four environmental indicators closely related to biological abundance, namely chlorophyll, salinity, temperature and dissolved oxygen, are selected as input features to form a feature vector x = (x1, x2, x3, x4), and each sample is labeled as "first abundance", "second abundance", "third abundance" and "fourth abundance" according to the historical observed abundance level, wherein the first abundance and the second abundance correspond to a higher biological abundance level, gradually approaching or even reaching the biological outbreak state of the intake, indicating a high risk of heat trap failure; the third abundance and the fourth abundance correspond to a lower biological abundance level, indicating a low risk of heat trap failure, to reflect the difference in the risk of heat trap failure at different abundance levels.

[0032] Due to the significant difference in the dimensions of the indicators and the inconsistency of the feature scales, in order to improve the stability of the model and the comparability between different features, Z-score standardization method is used to standardize all features;

[0033]

[0034] wherein x j is the value of the jth feature in the original data, x j ' is the standardized jth feature value, μ j and σ j are the mean and standard deviation of the jth feature, respectively.

[0035] After standardization, the mean of each feature is 0 and the standard deviation is 1, effectively eliminating the dimensional difference.

[0036] ​S2, sample balancing, considering that in actual monitoring data, low-risk samples are much more than high-risk samples, which may lead the trained model to be biased towards the majority class and difficult to identify the minority high-risk samples; therefore, SMOTE is introduced to synthesize minority class samples in the training stage, new minority class samples are generated in the feature space by interpolation, the number of samples of each risk level is balanced, so as to improve the identification ability of the model to high-risk level and avoid the interference of unbalanced data on the classifier.

[0037] S3, model construction, after completing data standardization and sample balancing, a random forest classification algorithm is used to construct the biological disaster risk assessment model of nuclear heat sink.

[0038] S31, Bootstrap resampling method is used to extract k sample subsets D i (i = 1, 2, … k) from the standardized and balanced sample set D, and a decision tree classifier h i (x, θ i ) is trained on each subset.

[0039] S32, for a new sample x to be tested, the random forest aggregates the classification results of k trees by majority voting mechanism, and outputs the final biological abundance level Y * .

[0040] Y * is calculated as follows:

[0041]

[0042] Where Y is the candidate level; I(·) is the indicator function, which takes 1 when h i (x) = Y, otherwise 0; θ i is the random parameter vector of the i-th tree. The number of votes obtained by each candidate level is counted, and the level Y with the most votes is selected as the final output of the model.

[0043] S4, model training and optimization, after the random forest model is built, the key hyperparameters are trained and optimized; to prevent underfitting or overfitting caused by artificial parameter setting, define the search space of hyperparameters, including the number of decision trees n_estimators, the maximum depth of the tree max-depth, the minimum number of samples split min-samples-split, the minimum number of leaf nodes min_samples_leaf, and whether to bootstrap; adopt five-fold stratified cross-validation, and ensure that the proportion of samples of different abundance levels in the training set and the validation set of each fold is consistent with the SMOTE balanced dataset, take the classification accuracy as the evaluation index, select the optimal parameter combination after 100 random searches, train the random forest model with the best generalization performance, and improve the overall recognition effect.

[0044] S5, risk assessment, after training and optimization, the real-time monitored environmental data is processed according to the same standardized parameters as in training, to obtain the normalized four-dimensional evaluation index feature vector x'=(x'1,x'2,x'3,x'4), which is input into the trained random forest model, and the biological abundance level of the current time period is output. The output result can reflect the potential risk level of the hot trap in real time. When the model output level is level one or level two, it is determined that the level is high, and it is prompted that there is a large biological outbreak risk, and an automatic and real-time alarm is sent to the operation and maintenance personnel, thereby providing a reliable basis for timely response measures and operation and maintenance decision-making.

[0045] Embodiment 1

[0046] S1, the data acquisition work of the embodiment relies on multiple types of sensors and field sampling devices arranged at the water intake of Hainan Changjiang Nuclear Power Plant to obtain monitoring data. The system is equipped with temperature sensors, dissolved oxygen monitors, chlorophyll fluorescence instruments, salinometers and other ecological and hydrological environmental monitoring equipment, and combines with shallow net sampling technology to continuously and synchronously record key environmental parameters and biological disaster factors during the operation of the hot trap.

[0047] The original data covers all time periods from November 1, 2021 to November 30, 2021, with a monitoring frequency of once per hour, covering 24 hours a day. During this period, some data of November 23, 24 and 27 is missing due to equipment maintenance, and finally 648 valid observation records are obtained. Each record contains hourly synchronous monitoring records of temperature, dissolved oxygen, chlorophyll, salinity and biological abundance.

[0048] After the data collection is completed, time series synchronization processing is performed to align the data of each monitoring device according to a unified time axis. Then, obvious abnormalities and invalid observations are eliminated, and the nearest neighbor interpolation method is used to complete the missing data. Finally, to eliminate the influence of the dimensional difference and the too large numerical span between different characteristics on modeling, normalization preprocessing is uniformly performed to convert each characteristic into standardized dimensionless data. The preprocessed data set is complete and reliable in quality, and can be directly used for subsequent model training and verification. The specific data samples are shown in Table 1. After the system is arranged, invalid values are eliminated, and missing values are completed, the integrity and reliability of the data are ensured.

[0049] Table 1 Environmental data and part of biological abundance data

[0050]

[0051] In order to facilitate the construction of the random forest model and the classification of different abundance levels, the original biological abundance observation data is classified according to the interval characteristics, and is divided into four levels. The division standard is determined by considering historical monitoring experience and data distribution law. The high abundance value interval (first risk level) is significantly positively correlated with the biological outbreak risk. The levels and the corresponding abundance ranges are shown in Table 2.

[0052] Table 2 Biological abundance level division standard

[0053] Abundance rank Bioabundance range (ind / m 3 ) Primary abundance Abundance > 200 Secondary abundance 134 <Abundance < 200 Tertiary abundance 67 <Abundance < 134 Quaternary abundance Abundance < 67

[0054] In order to eliminate the dimensional difference between different characteristics, the input variables such as temperature, dissolved oxygen, chlorophyll and salinity are standardized. The Standard Scaler method is used to convert them into standard normal distribution with mean of zero and standard deviation of one, so as to improve the training stability and feature comparability of the model.

[0055] S2, on the standardized data, considering that the biological abundance level data distribution is significantly unbalanced, the proportion of high abundance level samples is low, which is easy to lead to model training bias to low risk level. The SMOTE algorithm is used to synthesize high risk level sample points in the feature space, and balance the sample proportion of each risk level. After standardization and balancing, the data set is more suitable for training of random forest model, which can effectively improve the identification ability and classification accuracy of high risk state.

[0056] S3、In the model construction stage, a classification model is trained based on the random forest algorithm, and hyperparameter optimization is performed through random search combined with stratified cross-validation. Random search randomly samples several combinations within the preset hyperparameter space, calculates the average classification performance of each combination through five-fold stratified cross-validation, and selects the combination with the best performance for final model training. For example, in this embodiment, the number of decision trees can be set between 100 and 400, the maximum depth can be selected between 10 and 40, the minimum number of samples for node partitioning can be set between 2 and 15, the minimum number of samples for leaf nodes can be selected between 1 and 6, and self-sampling can be turned on or off as needed. After 500 iterations of random search, the optimal parameter combination obtained is: 400 decision trees, maximum depth 25, minimum partition sample size 5, leaf node minimum sample size 2, and self-sampling enabled. This combination shows good classification performance and robustness on the validation set.

[0057] S4、The nuclear power plant heat trap monitoring data set after standardization and sample balancing is divided into training set and test set according to the ratio of about 8:2, and the test set accounts for about 20%. When the random forest model trained based on the training set is evaluated on the test set, the overall classification accuracy reaches about 82%, indicating that the model can effectively distinguish different biological abundance levels, and the classification results are highly consistent with the actual observation, and has good applicability in the risk assessment of biological disasters in the heat trap of the nuclear power plant.

[0058] The further calculated macro-average and weighted average indicators are both around 82%, indicating that the model has balanced classification performance between each abundance level without obvious bias, strong overall stability, and can well cope with the sample distribution difference of different risk states in the heat trap system. Even in the case of relatively scarce high-risk level samples, the model can still accurately identify high-abundance level states and effectively evaluate potential risks, which is of great significance to the safe operation of the heat trap of the nuclear power plant.

[0059] From the test results of each level, the identification of the second level abundance performs the best, with high accuracy and strong recall ability, which can reliably reflect the state of the water body at a medium-high risk level. The identification performance of the first level abundance is second, and it can also accurately judge the high-abundance and high-risk state, and timely prompt the biological outbreak risk. The identification performance of the third and fourth levels of abundance is relatively weak, and there is a certain misjudgment, but the overall performance is acceptable and does not affect the overall performance and response ability to high-risk states.

[0060] The overall evaluation results show that the model performs well in high-risk level identification, overall performance balance, result stability, etc., can effectively adapt to the complex biological abundance distribution at the water intake, and can timely prompt the potential biological outbreak risk, providing scientific decision-making basis for operation and maintenance. The classification report results are shown in Table 3, and the performance of each abundance level meets the expected requirements and can meet the actual application needs.

[0061] Classification report of Table 3

[0062]

[0063] In addition, to further verify the classification effect of the model at different abundance levels, the test results are shown in the form of a confusion matrix, as shown in Table 4. The confusion matrix clearly reflects the classification relationship between each risk level. The results show that the identification of secondary abundance is the most accurate, with almost no misclassification; primary abundance can also be well identified, with only a small amount of misclassification as the adjacent level. The identification of tertiary and quaternary abundance is relatively weak, with a certain degree of misclassification, but the overall prediction effect is still within an acceptable range.

[0064] Confusion matrix of Table 4

[0065] Actual / predicted Primary abundance Secondary abundance Tertiary abundance Quaternary abundance Primary abundance 66 3 5 2 Secondary abundance 4 95 0 0 Tertiary abundance 12 1 66 15 Quaternary abundance 1 2 18 66

[0066] Overall, the confusion matrix results further verify that the embodiment can accurately identify high-risk level states, and the classification results are stable, with errors mainly concentrated in the mutual confusion between low-risk levels, without affecting the ability to identify and classify high-risk states in a timely manner, meeting the application requirements of nuclear power plant heat trap biological disaster risk assessment.

[0067] Therefore, the biological outbreak risk assessment method for a nuclear power plant water intake port described above can effectively characterize the risk state of the system at different operating stages; by modeling and analyzing historical monitoring data, a model that can input key environmental indicators and output real-time risk levels is established; test results show that the model can accurately identify high-risk states and still maintain stable identification effects under complex and nonlinear interference conditions. Compared with traditional methods, this method is more adaptable and more practical, can identify risk levels in a timely manner, assist in operational decision-making, and improve the safety and resilience of nuclear power plant heat traps. In addition, this method has good scalability and engineering applicability, and can be applied to risk management of other water-related industrial systems or marine engineering facilities.

[0068] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application and not to limit them, although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that: it can still modify or equivalently replace the technical solutions of the present application, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present application.

Claims

1. A method for biological outbreak risk assessment for a nuclear power plant intake, characterized by, Comprise the following steps: S1, data preprocessing, based on long-term monitoring data at the water intake of nuclear power plant, select four environmental indicators closely related to biological abundance, namely chlorophyll, salinity, temperature and dissolved oxygen, as input features to form feature vector x=(x1, x2, x3, x4), label each sample according to the historical observed abundance level as "first level abundance", "second level abundance", "third level abundance" and "fourth level abundance", and use Z-score standardization method to standardize all features; S2, sample balancing, use minority oversampling technique to synthesize minority class samples in feature space, generate new minority class samples in feature space by interpolation method, and balance the sample number of each risk level; S3, model construction, use random forest classification algorithm to construct biological disaster risk assessment model of nuclear power heat trap; S4, model training and optimization, define the search space of hyperparameters, use five-fold stratified cross-validation, take classification accuracy as evaluation index, select the optimal parameter combination after random search, and train the random forest model with the best generalization performance; S5, risk assessment, the real-time monitored environmental data is processed according to the same standardized parameters in training to obtain a normalized four-dimensional evaluation index characteristic vector x ' =(x'1, x'2, x'3, x'4), input into the trained random forest model, and output the biological abundance level of the current time period. When the level is one or two, it is determined as a higher risk, and an alarm is issued.

2. A method for biological outbreak risk assessment for a nuclear power plant intake according to claim 1, characterized in that, The calculation formula of S1 using Z-score standardization method is: where x j is the value of the jth feature in the original data, x' j is the standardized jth feature value, μ j , and σ j are the mean and standard deviation of the jth feature, respectively.

3. The method for biological outbreak risk assessment for a nuclear power plant intake of claim 1, wherein: The minority class oversampling technique in S2 is SMOTE technique.

4. The method for biological outbreak risk assessment for a nuclear power plant intake of claim 1, wherein, The specific steps of S3 are as follows: S31, Bootstrap resampling method is used to extract k sample subsets D from the standardized and balanced sample set D i (i = 1, 2,... k), and train a decision tree classifier h on each subset i (x, θ i ); S32, for a new sample to be tested x, the random forest aggregates the classification results of the k trees by majority voting mechanism, and outputs the final biological abundance level Y * .

5. A method for biological outbreak risk assessment for a nuclear power plant intake according to claim 4, characterized in that, Y * The calculation formula is as follows: where Y is the candidate rank, I(·) is the indicator function, θ i is the random parameter vector of the i-th tree, and k is the number of decision trees.

6. A method for biological outbreak risk assessment for a nuclear power plant intake according to claim 1, characterized in that, The search space of hyperparameters in S4 includes: the number of decision trees n_estimators, the maximum depth of tree max-depth, the minimum number of samples split min-samples-split, the minimum number of leaf nodes min_samples_leaf, and whether to use bootstrap sampling bootstrap.

7. The method for biological outbreak risk assessment for a nuclear power plant intake of claim 1, wherein: In S4 stratified cross-validation, the proportion of samples of different abundance levels in each fold of training set and validation set is consistent with that of SMOTE balanced data set.