A method for predicting biological abundance at nuclear power water intakes based on deep learning

Through deep learning and swarm intelligence optimization technology, the biological population prediction model for nuclear power water intakes is dynamically selected and optimized, which solves the shortcomings of prediction accuracy and adaptability in traditional methods, realizes efficient and reliable biological population prediction, and ensures the safe and stable operation of nuclear power plants.

CN120409533BActive Publication Date: 2025-09-16自然资源部宁德海洋中心(自然资源部宁德海洋预报台)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510923792.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-09-16
Estimated Expiration
2045-07-04

AI Technical Summary

Technical Problem

Traditional methods for predicting the number of organisms at nuclear power water intakes are unable to accurately capture the complex nonlinear relationship between the number of organisms and environmental factors, and lack adaptability and optimization capabilities, resulting in low prediction accuracy and reliability, and an inability to adapt to environmental changes in a timely manner.

Method used

Using a deep learning-based method, by collecting environmental factor data, biological species distribution data and historical prediction deviation data, the final model selection correlation equation is constructed, the prediction model is dynamically selected, and the model structure is optimized using swarm intelligence optimization technology and feedforward neural networks to achieve self-adjustment and optimization.

Benefits of technology

It significantly improves the accuracy and reliability of biological population prediction at nuclear power water intakes, enhances the adaptability of the model, reduces manual intervention and maintenance costs, and ensures the safe and stable operation of nuclear power plants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120409533B_ABST
    Figure CN120409533B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of biological population prediction at nuclear power water intakes, and discloses a method for predicting biological populations at nuclear power water intakes based on deep learning. The method first collects data such as environmental factors, biological distribution, and historical prediction deviations, constructs a model selection correlation equation, selects the current prediction model based on this data, and then determines whether the model is optimized based on actual biological population data, ultimately obtaining an optimized prediction model structure. The method defines the monitoring area and target biological species, sets the statistical time period and key parameters, utilizes swarm intelligence optimization technology to construct correlation equations and weighted data sets, determines the model by comparing the success rate, and determines whether the model structure is optimized based on the critical value of the effect difference. This method can improve prediction accuracy and adaptability, providing an effective solution for predicting biological populations at nuclear power water intakes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of biological population prediction at nuclear power water intakes, and in particular to a method for predicting biological population at nuclear power water intakes based on deep learning. Background Art

[0002] Predicting the biomass at nuclear power plant water intakes is crucial during nuclear power plant operation. A nuclear power plant's cooling system draws large quantities of water from the outside to maintain normal operation. The number and distribution of biomass at the water intake significantly impacts the cooling system's efficiency. Excessive biomass at the water intake can clog pipes, affect water quality, and even cause cooling system failures, compromising the safe and stable operation of the nuclear power plant.

[0003] Traditional methods for predicting biomass at nuclear power plant water intakes have numerous shortcomings. These methods are often based on simple statistical models or empirical formulas, making it difficult to accurately capture the complex, nonlinear relationships between biomass and environmental factors. Changes in environmental factors such as water temperature, salinity, and dissolved oxygen affect the growth, reproduction, and distribution of organisms, and these factors interact with each other, making it difficult for traditional methods to fully and accurately describe these relationships. They also fail to fully utilize historical data. In practice, data on environmental factors, biomass distribution, and historical forecast deviations from different time periods contain a wealth of information, and traditional methods often fail to effectively integrate this data, resulting in low accuracy and reliability in forecast results.

[0004] Furthermore, traditional methods lack adaptability and optimization capabilities. When environmental conditions change, traditional methods struggle to automatically adjust model parameters or structures to accommodate the new situation. This requires extensive manual adjustment and optimization, which is not only time-consuming and labor-intensive, but also difficult to ensure accurate and timely adjustments.

[0005] With the continuous development of deep learning technology, its application in various fields is becoming increasingly widespread. Deep learning, with its powerful nonlinear fitting and self-learning capabilities, can effectively handle complex nonlinear problems, providing new ideas and methods for predicting biodiversity at nuclear power water intakes. However, research on the application of deep learning technology to biodiversity prediction at nuclear power water intakes is relatively limited, and a systematic and comprehensive deep learning-based prediction method is lacking.

[0006] Therefore, there is an urgent need for a deep learning-based method for predicting the number of organisms at nuclear power water intakes to address the shortcomings of traditional methods, improve the accuracy and reliability of predictions, and provide strong support for the safe and stable operation of nuclear power plants. Summary of the Invention

[0007] The purpose of the present invention is to provide a method for predicting the number of organisms at nuclear power water intakes based on deep learning to solve the problems raised in the above background technology.

[0008] To achieve the above objectives, the present invention provides a method for predicting the number of organisms at nuclear power water intakes based on deep learning, the method comprising:

[0009] S1. Collect environmental factor data, biological species distribution data, and historical prediction deviation data of the nuclear power water intake area at different time periods to obtain the current environmental factor dataset, current biological distribution dataset, and current prediction deviation dataset;

[0010] S2. Collect historical application success rate data of each candidate prediction model, environmental factor data of each candidate prediction model, biological distribution data at different time periods, and prediction deviation data, and construct the final model selection correlation equation;

[0011] S3, selecting a current prediction model based on the final model selection correlation equation, the current environmental factor data set, the current biological distribution data set, and the current prediction deviation data set;

[0012] S4. Determine whether parameter adjustment or structural optimization of the current prediction model is required by collecting actual biological population data after using the current prediction model, and obtain a determination result;

[0013] S5. Optimize the structure of the current prediction model according to the judgment result to obtain the final prediction model structure.

[0014] Preferably, said S1 comprises the following steps:

[0015] S11. Delineate the monitoring area of ​​the nuclear power water intake that needs to be predicted and the corresponding several types of biological species to obtain a target biological species set; then set several key parameter types that affect the selection of the prediction model to obtain a prediction model influencing parameter type set;

[0016] S12. Set the first current statistical period; in conjunction with the first current statistical period and the prediction model influencing parameter type set, obtain the environmental factor data of each biological species in the target biological species set, the biological distribution data of different time periods and the prediction deviation data, and obtain the current environmental factor data set, the current biological distribution data set and the current prediction deviation data set.

[0017] Preferably, said S2 comprises the following steps:

[0018] S21. Setting a historical statistical period; in conjunction with the target biological species set and the prediction model influencing parameter type set, collecting application success rate data of each candidate prediction model within the historical statistical period, environmental factor data of each candidate prediction model, biological distribution data at different time periods, and prediction deviation data, to obtain a historical candidate model success rate dataset, a historical environmental factor dataset, a historical biological distribution dataset, and a historical prediction deviation dataset;

[0019] S22. Use the historical candidate model success rate dataset, historical environmental factor dataset, historical biological distribution dataset, and historical prediction deviation dataset to construct the final model selection association equation and the final environmental factor weight dataset.

[0020] Preferably, the final model selection correlation equation and the final environmental factor weight data set are constructed in S22 using swarm intelligence optimization technology.

[0021] Preferably, said S3 comprises the following steps:

[0022] S31, setting the current initial prediction model; substituting each data in the current environmental factor dataset, the current biological distribution dataset, the current prediction deviation dataset, and the final environmental factor weight dataset into the final model selection correlation equation for correlation calculation to obtain the current candidate model success rate dataset;

[0023] S32. When the candidate prediction model corresponding to the highest success rate data in the current candidate model success rate data set is the same as the current initial prediction model, the current prediction model is not replaced; otherwise, the candidate prediction model corresponding to the highest success rate data in the current candidate model success rate data set is used as the current prediction model, and is recorded as the current prediction model.

[0024] Preferably, said S4 comprises the following steps:

[0025] S41, setting a second current statistical period; setting a number of time nodes within the second current statistical period to obtain a current time node set; setting a number of types of verification parameters that can reflect the application effect of the prediction model to obtain a model verification parameter type set;

[0026] S42. Collecting actual biological population data after the current prediction model is used, using the model validation parameter type set and the current time node set, to obtain a current validation parameter matrix; then collecting average biological population data at several time nodes before the current prediction model is used, to obtain a historical average population data set; and calculating the difference between the historical average population data set and each row of data in the current validation parameter matrix to obtain a current effect difference data set.

[0027] S43, setting a first effect difference critical value and a second effect difference critical value; when the current effect difference value in the current effect difference data set is less than the second effect difference critical value, switching the current prediction model back to the current initial prediction model; when the current effect difference value in the current effect difference data set is greater than or equal to the second effect difference critical value and less than the first effect difference critical value, entering S5; otherwise, entering S44;

[0028] S44. Predict the biological quantity data of future time nodes based on the current verification parameter matrix and adopt the feedforward neural network model to obtain the future verification parameter matrix; then calculate the difference value between the historical average quantity data set and each row of data in the future verification parameter matrix to obtain the future effect difference data set; when the future effect difference value in the future effect difference data set is less than the second effect difference critical value, enter S5; otherwise, no processing is performed.

[0029] Preferably, the S5 comprises the following steps:

[0030] S51, setting initial model structure optimization parameters; performing structural optimization processing on the current prediction model using the initial model structure optimization parameters and storing the result; after the structural optimization processing and storage of the current prediction model are completed, collecting current biological population data according to the model verification parameter type set to obtain a current optimized verification parameter set; calculating the difference value between the current optimized verification parameter set and the historical average population data set to obtain a current optimized difference value;

[0031] S52. When the current optimized difference value is greater than or equal to the first effect difference critical value, the initial model structure optimization parameters are used as the final prediction model structure; otherwise, the initial model structure optimization parameters are adjusted until the current optimized difference value is greater than or equal to the first effect difference critical value.

[0032] Preferably, adjusting the initial model structure optimization parameters in S52 includes the following steps:

[0033] S521, setting the value range of the initial model structure optimization parameter to obtain the current parameter value range; constructing a model structure optimization group; setting the maximum number of iterations and the current number of iterations of the model structure optimization group, which are recorded as the maximum number of iterations of structure optimization and the current number of iterations of structure optimization, respectively;

[0034] S522. Setting the initial position of each individual in the model structure optimization group according to the current parameter value range to obtain a second initial position set;

[0035] S523, constructing a fitness evaluation function for the model structure optimization group;

[0036] S524, start iteration, and set the current iteration number of the structure optimization to 1 before the iteration; in each round of iteration, use the fitness evaluation function of the model structure optimization group to calculate the fitness value of each individual position in the model structure optimization group updated in the previous round of iteration, and update the position of each individual in the model structure optimization group updated in the previous round of iteration;

[0037] S525. When the current number of iterations of the structural optimization is equal to the maximum number of iterations of the structural optimization, the iteration is stopped to obtain the second final global optimal fitness and the second final global optimal position; otherwise, the iteration is continued until the current number of iterations of the structural optimization is equal to the maximum number of iterations of the structural optimization; the second final global optimal fitness is used as the current post-optimization difference value after optimization; when the current post-optimization difference value after optimization is greater than or equal to the first effect difference critical value, the second final global optimal position is used to perform structural optimization processing on the current prediction model and store it to obtain the final prediction model structure; otherwise, return to S524 and continue to iterate until the current post-optimization difference value after optimization is greater than or equal to the first effect difference critical value.

[0038] Preferably, the model validation parameter type set includes water temperature parameters, salinity parameters, dissolved oxygen parameters and biological density parameters.

[0039] Preferably, the delineation of the target biological species set comprises the following steps:

[0040] S111. Conduct a biodiversity survey in the nuclear power plant water intake area, recording the species whose occurrence frequency exceeds the preset threshold within different time periods;

[0041] S112. Based on the operating characteristics of the nuclear power cooling system, screen out key biological species that may affect water extraction efficiency and summarize them to form a target biological species set.

[0042] Compared with the prior art, the present invention has the following beneficial effects:

[0043] By collecting environmental factor data, species distribution data, and historical forecast deviation data from nuclear power plant water intakes over different time periods, this method fully utilizes all types of data related to biomass prediction. This data collection lays a solid foundation for subsequent model construction and selection, enabling the model to better understand and capture the patterns of biomass variation.

[0044] In terms of model selection, this approach overcomes the limitations of traditional single models by constructing a final model selection equation and combining multiple datasets to select the current prediction model. It dynamically selects the most appropriate prediction model based on the actual environment and biodiversity distribution, significantly improving the model's adaptability to diverse scenarios. For example, in different seasons and hydrological conditions, this method automatically selects the model that most accurately reflects current biodiversity changes, thereby improving prediction accuracy.

[0045] By collecting actual biomass data after using the current prediction model, the system determines whether model parameters need adjustment or structural optimization. This process enables the model to self-optimize and improve. When the model's predictions deviate during actual application, these deviations are promptly detected and adjusted, allowing the model to continuously adapt to environmental changes and the dynamic development of biomass. This self-optimization mechanism avoids the tedious manual adjustments required by traditional methods and improves the model's prediction accuracy and reliability.

[0046] During the model structure optimization process, advanced technologies and methods such as swarm intelligence optimization and feedforward neural network models were used to further enhance model performance. Swarm intelligence optimization technology can find the optimal model structure parameters within the parameter space, making the model structure more reasonable. The feedforward neural network model can accurately predict future biomass data, providing strong support for model optimization.

[0047] Furthermore, the various parameters and critical values ​​set in this method, such as the first and second effect difference critical values, provide clear standards and basis for model adjustment and optimization. This makes the model optimization process more scientific and reasonable, avoids interference from subjective factors, and ensures the effectiveness and stability of model optimization.

[0048] The method of the present invention significantly improves the accuracy, reliability and adaptability of biological population prediction at nuclear power water intakes through comprehensive data collection, dynamic model selection, self-optimizing model adjustment and advanced technology application. It can provide more accurate biological population prediction information for the operation of the nuclear power plant's cooling system, ensuring the safe and stable operation of the nuclear power plant. At the same time, it also reduces manual intervention and maintenance costs, and has important practical application value and economic benefits. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 This is a working principle diagram of the method for predicting the number of organisms at nuclear power water intakes based on deep learning according to the present invention;

[0050] Figure 2 Flowchart for S1 data collection;

[0051] Figure 3Flowchart for selecting the current prediction model for S3;

[0052] Figure 4 Flowchart for adjusting the judgment of S4 model;

[0053] Figure 5 Flowchart for S52 parameter adjustment. DETAILED DESCRIPTION

[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0055] See also Figure 1-Figure 5 The present invention provides a method for predicting the number of organisms at a nuclear power water intake based on deep learning, comprising the following steps:

[0056] Environmental factor data, biological species distribution data and historical prediction deviation data of the nuclear power water intake area at different time periods are collected to obtain the current environmental factor dataset, current biological distribution dataset and current prediction deviation dataset.

[0057] The historical application success rate data of each candidate prediction model, the environmental factor data of each candidate prediction model, the biological distribution data in different time periods and the prediction deviation data are collected to construct the final model selection association equation.

[0058] The current prediction model is selected based on the final model selection correlation equation, the current environmental factor data set, the current biological distribution data set, and the current prediction deviation data set.

[0059] By collecting actual biological quantity data after using the current prediction model, it is determined whether the current prediction model needs to be parameter adjusted or the structure optimized to obtain a determination result.

[0060] According to the judgment results, the current prediction model is structurally optimized to obtain the final prediction model structure.

[0061] Example 1: This example refines step S1. First, the monitoring area for the nuclear power plant water intake to be predicted must be delineated. This area must be determined by comprehensively considering factors such as the actual geographic location of the nuclear power plant water intake, the impact range of water flow, and the primary areas of biological activity. For example, a specific geographic area can be reasonably defined based on the pipeline layout, water flow velocity distribution, and historical biological monitoring data at the nuclear power plant water intake to serve as the target area for subsequent monitoring and data collection.

[0062] Several corresponding biological species need to be identified to form a target species set. This process begins with a biodiversity survey of the nuclear power plant's water intake area. During the survey, detailed records of the various species present during different time periods are kept, and the frequency of each species' occurrence is calculated. These time periods can be divided based on factors such as the organisms' living habits and seasonal variations, such as by month, quarter, or hydrological cycle. A preset threshold is then set, and species whose occurrence frequency exceeds this threshold are recorded.

[0063] The recorded species were screened based on the operational characteristics of nuclear power cooling systems. During operation, parameters such as water temperature and flow rate fluctuate. These fluctuations can cause certain organisms to congregate near water intakes, impacting water extraction efficiency. Therefore, it is necessary to analyze each organism's living habits, adaptability to the environment, and potential impact on the cooling system. This allows us to identify key species that may impact water extraction efficiency. These key species are then aggregated to form a target species set.

[0064] After defining the target species set, you need to set the first current statistical period. This setting should take into account factors such as the cycle of biological population fluctuations, the feasibility of data collection, and the timeliness of the prediction. For example, if the population of the target species exhibits a clear seasonal cycle, the first current statistical period can be set to a full season or a combination of seasons to ensure that sufficient data reflecting the regularity of biological population fluctuations is collected.

[0065] After setting the first current statistical period and the predetermined set of influencing parameters for the prediction model, data collection begins for each species in the target set. The set of influencing parameters includes key parameters that influence biomass prediction, such as water temperature, salinity, dissolved oxygen, and other environmental parameters, as well as other factors that may affect biomass distribution.

[0066] For each target species, corresponding environmental factor data is collected at different time periods during the first current statistical period. This environmental factor data collection requires the use of specialized monitoring equipment. For example, water temperature can be measured using a temperature sensor, salinity can be tested using a salinometer, and dissolved oxygen content can be measured using a dissolved oxygen meter. Simultaneously, biological distribution data is collected at different time periods. This data can be collected using methods such as field sampling and video monitoring, recording information such as the distribution location and abundance of each target species over different time periods.

[0067] In addition, historical forecast deviation data must be collected. This data refers to the difference between the prediction model's prediction results and the actual biomass population during previous forecasts. This data can be obtained from historical forecast records. By comparing and analyzing historical prediction results with actual monitoring data, the forecast deviation value for each time period can be calculated.

[0068] Through the above series of operations, the collected environmental factor data, biological distribution data, and prediction deviation data are organized and stored, respectively, to obtain the current environmental factor dataset, the current biological distribution dataset, and the current prediction deviation dataset. These datasets will serve as an important basis for subsequent model selection and optimization. During the data organization process, it is necessary to ensure the accuracy and completeness of the data. The collected data must be cleaned and preprocessed as necessary to remove abnormal data and noise interference to improve the accuracy and reliability of the subsequent prediction model. For example, any obviously unreasonable values ​​in the environmental factor data need to be verified and corrected; any missing parts in the biological distribution data need to be supplemented through reasonable methods.

[0069] Example 2: This example provides a detailed description of step S2.

[0070] The historical statistical period is determined based on comprehensive considerations of data validity and comprehensiveness. Typically, a period encompassing multiple biological growth cycles, seasonal climate conditions, and varying operating states of nuclear power cooling systems is selected, such as the past three to five years. This ensures that the collected data reflects the impact of various possible environmental changes and system operating conditions on biological abundance and predictive models.

[0071] After setting the historical statistical period, data collection begins, along with the identified target species set and the set of influencing parameters for the prediction model. The target species set, derived from a previous survey and screening of biomes in the nuclear power plant's water intake area, includes key species that may impact water extraction efficiency. The influencing parameters for the prediction model include key parameters that influence biomass predictions, such as environmental factors like water temperature, salinity, and dissolved oxygen content, as well as biodiversity data and prediction deviation data.

[0072] The collection process specifically includes recording the application success rate data for each candidate prediction model during the historical statistical period. The application success rate is calculated based on the degree of agreement between the prediction model's predictions and the actual biomass counts. For example, the success or failure of each model in a specific time period can be determined by comparing the error range between the predicted values ​​and the measured values. These results are then statistically compiled. Simultaneously, the environmental factor data corresponding to each candidate prediction model during the corresponding time period is collected. This data is collected in the same manner as the environmental factor data in step S1, measured and recorded at different time points using specialized monitoring equipment.

[0073] Biodiversity distribution data and prediction deviation data for different time periods within the historical statistical period also need to be collected. Biodiversity distribution data is also collected using methods such as field sampling and video monitoring, detailing the distribution location and abundance changes of each target species over different time periods. Prediction deviation data is obtained from historical prediction records. The prediction results of each candidate prediction model during the historical statistical period are compared with the actual monitoring data to calculate the prediction deviation value for each time period.

[0074] Through this collection process, we can obtain a historical candidate model success rate dataset, a historical environmental factor dataset, a historical biological distribution dataset, and a historical prediction deviation dataset. These datasets contain a wealth of information about the performance of candidate prediction models over the historical statistical period, as well as the corresponding environmental and biological data. This provides rich material for the subsequent construction of the final model selection correlation equation and the final environmental factor weight dataset.

[0075] These datasets are used to construct the final model selection equations and the final dataset of environmental factor weights. This construction utilizes swarm intelligence optimization techniques, which simulate the intelligent behavior of biological populations to find the optimal solution. For example, the particle swarm optimization algorithm simulates the foraging behavior of bird flocks. Each particle represents a possible solution. Through mutual collaboration and information sharing among particles, they continuously adjust their positions in the solution space to find the optimal model selection equation parameters and environmental factor weights.

[0076] During implementation, the relevant parameters of swarm intelligence optimization technology, such as swarm size, number of iterations, and learning factor, must be determined. These parameters affect the effectiveness and efficiency of optimization and need to be adjusted appropriately based on the actual data size and problem complexity. Then, historical datasets on candidate model success rates, environmental factors, biodiversity distribution, and prediction deviations are fed into the optimization algorithm as a basis for optimization.

[0077] Based on this data, the optimization algorithm continuously iterates and adjusts the structure and parameters of the model selection equation, as well as the weights of various environmental factors. In each iteration, the algorithm evaluates how well the current combination of model selection equations and environmental factor weights fits the historical data. In other words, it calculates the accuracy of the prediction of the success rate of historical candidate models under the current combination. Through continuous optimization, the model selection equation better reflects the relationship between the success rate of candidate prediction models and environmental factors, biodistribution data, and prediction deviation data, while also determining the weight of each environmental factor in model selection.

[0078] When the optimization algorithm reaches a preset number of iterations or meets other termination conditions, iterations cease, resulting in the final model selection correlation equation and the final environmental factor weight dataset. The final model selection correlation equation describes the mathematical relationship between the success rate of the candidate prediction model and various data types, while the final environmental factor weight dataset clarifies the importance of each environmental factor in the model selection process.

[0079] Throughout the construction process, data must be properly preprocessed and normalized to ensure the stability and accuracy of the optimization algorithm. For example, environmental factor data of different dimensions must be standardized to keep them within the same numerical range to avoid deviations in the optimization results due to dimensional differences. Furthermore, outliers in the data must be identified and addressed to ensure that the constructed model selection correlation equations and environmental factor weight datasets accurately reflect the actual situation.

[0080] Example 3: This example describes the implementation process of step S3 in detail.

[0081] The initial prediction model must be set. This model can be selected based on historical application or preliminary screening. For example, if an LSTM model has been used in previous biomass predictions at nuclear power water intakes under similar environmental conditions, and its structure and parameter settings are consistent with the current prediction requirements, then the LSTM model can be set as the initial prediction model.

[0082] Substitute each data point from the current environmental factor dataset, the current biomass distribution dataset, the current prediction deviation dataset, and the final environmental factor weight dataset into the final model selection correlation equation for correlation calculation. For example, assuming the water temperature data from the current environmental factor dataset is 25°C, the salinity data is 30‰, and the dissolved oxygen data is 5 mg / L, these data must be formatted and normalized according to the input requirements of the correlation equation. For example, convert the water temperature data to a value in the interval [0, 1] using a linear transformation, such as (25-15) / (35-15)=0.5 (assuming the historical water temperature range is 15°C to 35°C).

[0083] The current biological distribution dataset shows a target species density of 100 individuals / m³ for this time period. This density needs to be normalized according to pre-set rules, for example, by dividing it by the historical maximum density of 200 individuals / m³ to obtain a normalized value of 0.5. The historical prediction deviation for this time period, recorded in the current prediction deviation dataset, is 5%, which also needs to be converted to a corresponding numerical form. In the final environmental factor weight dataset, assume that water temperature has a weight of 0.3, salinity has a weight of 0.2, dissolved oxygen has a weight of 0.3, and other factors have a weight of 0.2. These weights are used to adjust the influence of each environmental factor in the correlation calculation.

[0084] When substituting data, the operational logic of the final model selection correlation equation must be strictly followed. For example, the correlation equation may be a weighted summation expression, in which environmental factor data are multiplied by the corresponding weights and then accumulated. This is then combined with biological distribution data and prediction deviation data for a comprehensive calculation. Assume that the correlation equation is of the form: success rate = a × (water temperature × 0.3 + salinity × 0.2 + dissolved oxygen × 0.3 + biomass × 0.1 + prediction deviation × 0.1) + b, where a and b are equation parameters. In this case, substituting the standardized water temperature of 0.5, salinity of 0.6 (assuming 30‰ is 0.6 after treatment), dissolved oxygen of 0.5 (5 mg / L corresponds to 0.5), biomass density of 0.5 (100 / 200), and prediction deviation of 0.05 (5% is converted to 0.05) into the equation, the success rate of the candidate model is calculated for the current data.

[0085] The above data substitution and calculation process is performed for each candidate prediction model. For example, if the candidate models also include CNN models and BP neural network models, their success rates are calculated for the current environmental factor dataset, biological distribution dataset, and prediction deviation dataset, respectively. This yields the current candidate model success rate dataset. Assuming the calculated success rates for the LSTM model are 0.75, the CNN model 0.82, and the BP model 0.68, then the current candidate model success rate dataset contains the success rate values ​​for these three models.

[0086] Compare the values ​​in the success rate dataset for the current candidate model to identify the candidate prediction model with the highest success rate. In the above example, the CNN model has the highest success rate of 0.82. At this point, determine whether the candidate model corresponding to this highest success rate is consistent with the current initial prediction model. If the current initial prediction model is an LSTM model, and the model with the highest success rate is a CNN model, replace the current prediction model with a CNN model. If the model with the highest success rate is the same as the initial model, such as if the initial model is a CNN model and its calculated success rate is the highest, do not replace the model.

[0087] In practice, data collection and processing must strictly adhere to pre-set standards. For example, the frequency of environmental factor data collection must match the temporal resolution of the prediction model. If the model predicts on a daily basis, environmental factor data must be collected daily. During data processing, if water temperature data for a certain period is found to have an abnormal value due to a sensor failure, data from adjacent periods must be interpolated to ensure the accuracy and reliability of the data input into the correlation equation.

[0088] Furthermore, the parameters of the correlation equations used in the final model selection may need to be regularly updated based on historical data. For example, as new historical statistical data accumulates, swarm intelligence optimization techniques can be used to optimize the correlation equation parameters again to adapt to environmental changes and evolving model performance. In this example, it is assumed that the correlation equation parameters have already been determined through a previous optimization process and do not need to be adjusted in this step.

[0089] The entire model selection process must ensure data integrity and calculation accuracy. For example, when calculating the success rate, it is necessary to clearly define the success rate, whether it is based on the mean square error between the predicted and measured values, the mean absolute error, or other indicators. In this example, it is assumed that the success rate is calculated based on the proportion of prediction errors within a ±10% range. That is, if a model has an error within ±10% for 80 of the past 100 predictions, then its success rate is 80%.

[0090] Through the above steps, the most suitable current prediction model can be selected from candidate prediction models based on the current actual data, providing a foundation for subsequent biomass prediction and model optimization. During implementation, attention should be paid to the details of data processing and the rigor of calculation logic to avoid model selection errors caused by data errors or calculation errors.

[0091] Example 4: This example further details step S4. A second current statistical period is set. This period should be determined based on the characteristics of the organism's population changes and the model validation requirements. For example, assuming the target organism is a migratory fish, whose population is significantly affected by water temperature and currents between April and June each year, the second current statistical period can be set to April 1, 2025, to June 30, 2025, a duration of three months, to cover the organism's primary activity period.

[0092] Set several time points within the second current statistical period to form a current time point set. For example, using a 10-day interval, set nine time points within three months: April 10, April 20, and finally June 20. Also, set a set of model validation parameters, including water temperature, salinity, dissolved oxygen, and biomass density. These parameters should clearly reflect the effectiveness of the prediction model. For example, water temperature can be collected hourly using an underwater temperature sensor, and the daily average value can be used as the daily data.

[0093] Combine the model validation parameter type set with the current time node set to collect actual biomass data using the current prediction model. Assuming the current prediction model is the CNN model selected in step S3, on April 10th, using a combination of underwater camera monitoring and manual sampling, the target organism density data for that day was 85 individuals / m³. The average water temperature for that day was also recorded as 22°C, salinity as 28‰, and dissolved oxygen as 6.2 mg / L, forming a row of data in the current validation parameter matrix. This process continues until data collection for all time nodes is complete.

[0094] Collect average biomass data for several time points prior to the current prediction model to form a historical average biomass dataset. For example, select data from the three preceding years (April to June 2024) corresponding to the current time point and calculate the average biomass density at each time point. Assuming the biomass density on April 10, 2022, 2023, and 2024 was 80 individuals / m³, 78 individuals / m³, and 82 individuals / m³, respectively, the average biomass density at that time point in the historical average biomass dataset would be (80 + 78 + 82) / 3 = 80 individuals / m³.

[0095] Calculate the difference between the historical average population dataset and each row of data in the current validation parameter matrix to obtain the current effect difference dataset. For example, if the current biomass density on April 10th is 85 individuals / m³ and the historical average is 80 individuals / m³, the difference is 5 individuals / m³. If the current biomass density at another node is 70 individuals / m³ and the historical average is 80 individuals / m³, the difference is -10 individuals / m³.

[0096] Set the first and second effect difference thresholds. Assume the first threshold is 15 units / m³ and the second threshold is 5 units / m³. If the difference value in the current effect difference dataset is less than 5 units / m³, such as a node with a difference of 3 units / m³, the current prediction model is ineffective and you need to switch back to the initial prediction model (such as the LSTM model). If the difference value is greater than or equal to 5 units / m³ and less than 15 units / m³, such as 8 units / m³, proceed to step S5 to optimize the model structure.

[0097] If all values ​​in the current effect difference dataset are greater than or equal to 15 individuals / m³, for example, if the difference at a particular node is 20 individuals / m³, then further predictions of biomass at future time points are required. A feedforward neural network model is used to predict biomass data for future time points (e.g., July 10th and July 20th). Assuming the environmental factor data and biomass density data in the current validation parameter matrix are input, a future validation parameter matrix is ​​generated. For example, the biomass density on July 10th is predicted to be 90 individuals / m³, the water temperature to be 28°C, the salinity to be 29‰, and the dissolved oxygen to be 5.8 mg / L.

[0098] Calculate the difference between the historical average population dataset and each row of data in the future validation parameter matrix to obtain the future effect difference dataset. Assume that the historical average biomass density for the same period (July) was 85 individuals / m³ and the future predicted value was 90 individuals / m³, resulting in a difference of 5 individuals / m³. If any value in the future effect difference dataset is less than 5 individuals / m³, such as 4 individuals / m³, proceed to step S5. If all difference values ​​are greater than or equal to 5 individuals / m³, no processing is performed.

[0099] During data collection, equipment accuracy must be ensured. For example, temperature sensors require regular calibration, and underwater cameras must capture the primary area of ​​the water intake to avoid data bias caused by sampling blind spots. If dissolved oxygen data at a particular time point is missing due to equipment failure, linear interpolation using values ​​from adjacent time points is required to ensure data integrity.

[0100] Furthermore, the time span of the historical average population dataset must be appropriately selected. If the selected period is too short, it may not reflect long-term changes in biomass; if it is too long, environmental changes may reduce the data's reference value. In this example, data from the first three years was selected to balance data timeliness and regularity.

[0101] The entire judgment process must be strictly executed according to the preset critical values. For example, if the current effect difference dataset contains values ​​greater than 15 / m³ and less than 5 / m³, the case with a difference less than 5 / m³ must be prioritized, that is, switching back to the initial model to ensure prediction accuracy.

[0102] Through the above steps, based on the actual biomass data collected, it is possible to determine whether the current prediction model needs to be adjusted, providing a basis for subsequent model optimization. During implementation, attention should be paid to the standardization of data collection, the rigor of calculation logic, and the rationality of critical value setting to ensure the reliability of the judgment results.

[0103] Example 5: This example describes the implementation process of step S5 in detail. Assume that in step S4, it is determined that the current prediction model needs to be structurally optimized. For example, the current prediction model is a CNN model, and the effect difference value in the second current statistical period is 8 / m³, which is between the second effect difference critical value of 5 / m³ and the first effect difference critical value of 15 / m³. At this time, step S5 is entered.

[0104] Set the initial model structure optimization parameters, including the number of layers, number of neurons, and convolution kernel size. For example, initially set the CNN model to have three convolution layers, 64, 128, and 256 neurons in each layer, a 3×3 convolution kernel size, two pooling layers, and 100 neurons in the fully connected layer. Use these initial model structure optimization parameters to optimize the current CNN model, adjust the model's network structure, and store the optimized model.

[0105] After completing the structural optimization process, collect current biomass data based on the model validation parameter set. For example, after the optimized model is operational, collect target biomass density data for May 10, 2025, at 90 individuals / m³. Simultaneously, record the water temperature as 24°C, salinity as 29‰, and dissolved oxygen as 6.0 mg / L to form the current optimized validation parameter set. Calculate the difference between the biomass density data in the current optimized validation parameter set and the average biomass at the corresponding time in the historical average biomass dataset. Assuming the average biomass density on May 10 in the historical average biomass dataset was 80 individuals / m³, the current optimized difference is 10 individuals / m³.

[0106] Compare the current optimized difference value with the first effect difference threshold of 15 / m³. Since 10 / m³ is less than 15 / m³, the initial model structure optimization parameters need to be adjusted. First, set the value range of the initial model structure optimization parameters. For example, the number of convolutional layers ranges from 2 to 5 layers, the number of neurons per layer ranges from 32 to 512, the convolution kernel size ranges from 3×3 to 5×5, the number of pooling layers ranges from 1 to 3 layers, and the number of neurons in the fully connected layer ranges from 50 to 200.

[0107] Construct a model structure optimization swarm, consisting of multiple individuals, each representing a set of possible model structure optimization parameters. Suppose we construct an optimization swarm consisting of 50 individuals, each with a different parameter combination. Set the maximum number of iterations for the model structure optimization swarm to 100, and the current number of iterations to 1.

[0108] The initial position of each individual in the model structure optimization group is set according to the current parameter value range. For example, the first individual has 3 convolution layers, 64, 128, and 256 neurons, a 3×3 convolution kernel size, 2 pooling layers, and 100 neurons in the fully connected layer. The second individual has 4 convolution layers, 32, 64, 128, and 256 neurons, a 4×4 convolution kernel size, 1 pooling layer, and 150 neurons in the fully connected layer. And so on, to obtain the second initial position set.

[0109] Construct a fitness evaluation function for the model structure optimization group, which is used to evaluate the quality of each individual position. The design of the fitness evaluation function is based on the model's prediction error. For example, the inverse of the mean square error between the predicted value and the measured value is used as the fitness value. The smaller the mean square error, the greater the fitness value.

[0110] Begin the iterations, setting the current iteration count of the structure optimization to 1. During each iteration, a fitness evaluation function is used to calculate the fitness of each individual position in the optimized model structure population, updated in the previous iteration. For example, for each individual's model structure, a prediction is made using the current environmental factor dataset and the biological distribution dataset. The mean squared error between the predicted result and the actual biological population data is calculated, and the inverse of this error is taken as the fitness value.

[0111] The position of each individual in the model structure optimization swarm after the previous round of iterative updates is then updated. This update method uses the rules of swarm intelligence optimization algorithms. For example, in the particle swarm optimization algorithm, each individual adjusts its position based on its own historical optimal position and the global optimal position of the swarm to generate a new parameter combination.

[0112] When the current number of structural optimization iterations equals the maximum number of structural optimization iterations (100), the iteration stops and the second final global optimal fitness and the second final global optimal position are obtained. Assume that after the iteration, the optimal individual parameter combination is 4 convolutional layers, 64, 128, 256, and 512 neurons, a convolution kernel size of 3×3, 2 pooling layers, and 150 neurons in the fully connected layer, and the corresponding fitness value is 0.85.

[0113] The difference value corresponding to the second final global optimal fitness is used as the current optimized difference value after optimization. Assuming that according to the corresponding relationship between the fitness evaluation function and the difference value, the fitness of 0.85 corresponds to a difference value of 16 / m³, which is greater than the first effect difference critical value of 15 / m³, the parameter combination corresponding to the second final global optimal position is used to optimize the current prediction model structure and store it to obtain the final prediction model structure.

[0114] If the current post-optimization difference value after optimization does not reach the first effect difference critical value, for example, the fitness 0.7 corresponds to a difference value of 14 / m³, which is less than 15 / m³, then return to continue iterating until the current post-optimization difference value after optimization is greater than or equal to the first effect difference critical value.

[0115] During the iteration process, if the position of an individual in a round of iteration exceeds the parameter value range, it needs to be corrected, for example, by adjusting the number of convolutional layers to the maximum or minimum value within the range. At the same time, to avoid falling into local optimality, random factors can be introduced to increase the diversity of the population.

[0116] During data collection, if the dissolved oxygen data at a particular time point is abnormal, data from adjacent nodes must be interpolated to ensure the accuracy of the data input into the model. The historical average population dataset must be regularly updated to reflect the latest trends in biomass.

[0117] The entire optimization process requires strict recording of the parameter combinations and fitness values ​​for each iteration to facilitate analysis of the optimization results and adjustment of the optimization strategy. The final prediction model structure must be verified multiple times to ensure that it can maintain good prediction performance under different environmental conditions.

[0118] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0119] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A method for predicting the number of organisms at nuclear power water intakes based on deep learning, characterized in that: The following steps are involved: S1. Collect environmental factor data, biological species distribution data, and historical prediction deviation data of the nuclear power water intake area at different time periods to obtain the current environmental factor dataset, current biological distribution dataset, and current prediction deviation dataset; S2. Collect historical application success rate data of each candidate prediction model, environmental factor data of each candidate prediction model, biological distribution data at different time periods, and prediction deviation data, and construct a final model selection association equation and a final environmental factor weight data set, wherein the final model selection association equation describes the mathematical relationship between the success rate of the candidate prediction model and various types of data, and the final environmental factor weight data set clarifies the importance of each environmental factor in the model selection process; S3, selecting a current prediction model based on the final model selection correlation equation, the current environmental factor data set, the current biological distribution data set, and the current prediction deviation data set; S4, by collecting actual biological population data after using the current prediction model, determining whether the current prediction model needs to be parameter adjusted or structure optimized, and obtaining a determination result, said S4 includes the following steps: S41, setting a second current statistical period; setting a number of time nodes within the second current statistical period to obtain a current time node set; setting a number of types of verification parameters that can reflect the application effect of the prediction model to obtain a model verification parameter type set; S42. Collecting actual biological population data after the current prediction model is used, using the model validation parameter type set and the current time node set, to obtain a current validation parameter matrix; then collecting average biological population data at several time nodes before the current prediction model is used, to obtain a historical average population data set; and calculating the difference between the historical average population data set and each row of data in the current validation parameter matrix to obtain a current effect difference data set. S43, setting a first effect difference critical value and a second effect difference critical value; when the current effect difference value in the current effect difference data set is less than the second effect difference critical value, switching the current prediction model back to the current initial prediction model; when the current effect difference value in the current effect difference data set is greater than or equal to the second effect difference critical value and less than the first effect difference critical value, entering S5; otherwise, entering S44; S44. Predicting biological population data at future time points using the current verification parameter matrix and a feedforward neural network model to obtain a future verification parameter matrix; then calculating the difference between the historical average population data set and each row of data in the future verification parameter matrix to obtain a future effect difference data set; if the future effect difference value in the future effect difference data set is less than the second effect difference threshold, proceed to S5; otherwise, no processing is performed; S5. Optimize the structure of the current prediction model according to the determination result to obtain the final prediction model structure. S5 includes the following steps: S51, setting initial model structure optimization parameters; performing structural optimization processing on the current prediction model using the initial model structure optimization parameters and storing the result; after the structural optimization processing and storage of the current prediction model are completed, collecting current biological population data according to the model verification parameter type set to obtain a current optimized verification parameter set; calculating the difference value between the current optimized verification parameter set and the historical average population data set to obtain a current optimized difference value; S52. When the current optimized difference value is greater than or equal to the first effect difference critical value, the initial model structure optimization parameters are used as the final prediction model structure; otherwise, the initial model structure optimization parameters are adjusted until the current optimized difference value is greater than or equal to the first effect difference critical value.

2. The method for predicting the number of organisms at nuclear power water intakes based on deep learning according to claim 1 is characterized in that: Said S1 comprises the following steps: S11. Delineate the monitoring area of ​​the nuclear power water intake that needs to be predicted and the corresponding several types of biological species to obtain a target biological species set; then set several key parameter types that affect the selection of the prediction model to obtain a prediction model influencing parameter type set; S12. Set the first current statistical period; in conjunction with the first current statistical period and the prediction model influencing parameter type set, obtain the environmental factor data of each biological species in the target biological species set, the biological distribution data of different time periods and the prediction deviation data, and obtain the current environmental factor data set, the current biological distribution data set and the current prediction deviation data set.

3. The method for predicting the number of organisms at nuclear power water intakes based on deep learning according to claim 2 is characterized in that: The S2 comprises the following steps: S21. Setting a historical statistical period; in conjunction with the target biological species set and the prediction model influencing parameter type set, collecting application success rate data of each candidate prediction model within the historical statistical period, environmental factor data of each candidate prediction model, biological distribution data at different time periods, and prediction deviation data, to obtain a historical candidate model success rate dataset, a historical environmental factor dataset, a historical biological distribution dataset, and a historical prediction deviation dataset; S22. Use the historical candidate model success rate dataset, historical environmental factor dataset, historical biological distribution dataset, and historical prediction deviation dataset to construct the final model selection association equation and the final environmental factor weight dataset.

4. The method for predicting the number of organisms at nuclear power water intakes based on deep learning according to claim 3 is characterized in that: In S22, swarm intelligence optimization technology was used to construct the final model selection association equation and the final environmental factor weight data set.

5. The method for predicting the number of organisms at nuclear power water intakes based on deep learning according to claim 4 is characterized in that: The S3 includes the following steps: S31, setting the current initial prediction model; substituting each data in the current environmental factor dataset, the current biological distribution dataset, the current prediction deviation dataset, and the final environmental factor weight dataset into the final model selection correlation equation for correlation calculation to obtain the current candidate model success rate dataset; S32. When the candidate prediction model corresponding to the highest success rate data in the current candidate model success rate data set is the same as the current initial prediction model, the current prediction model is not replaced; otherwise, the candidate prediction model corresponding to the highest success rate data in the current candidate model success rate data set is used as the current prediction model, and is recorded as the current prediction model.

6. A method for predicting the number of organisms at nuclear power water intakes based on deep learning according to claim 5, characterized in that: Adjusting the initial model structure optimization parameters in S52 includes the following steps: S521, setting the value range of the initial model structure optimization parameter to obtain the current parameter value range; constructing a model structure optimization group; setting the maximum number of iterations and the current number of iterations of the model structure optimization group, which are recorded as the maximum number of iterations of structure optimization and the current number of iterations of structure optimization, respectively; S522. Setting the initial position of each individual in the model structure optimization group according to the current parameter value range to obtain a second initial position set; S523, constructing a fitness evaluation function for the model structure optimization group; S524, start iteration, and set the current iteration number of the structure optimization to 1 before the iteration; in each round of iteration, use the fitness evaluation function of the model structure optimization group to calculate the fitness value of each individual position in the model structure optimization group updated in the previous round of iteration, and update the position of each individual in the model structure optimization group updated in the previous round of iteration; S525. When the current number of iterations of the structural optimization is equal to the maximum number of iterations of the structural optimization, the iteration is stopped to obtain the second final global optimal fitness and the second final global optimal position; otherwise, the iteration is continued until the current number of iterations of the structural optimization is equal to the maximum number of iterations of the structural optimization; the second final global optimal fitness is used as the current post-optimization difference value after optimization; when the current post-optimization difference value after optimization is greater than or equal to the first effect difference critical value, the second final global optimal position is used to perform structural optimization processing on the current prediction model and store it to obtain the final prediction model structure; otherwise, return to S524 and continue to iterate until the current post-optimization difference value after optimization is greater than or equal to the first effect difference critical value.

7. The method for predicting the number of organisms at nuclear power water intakes based on deep learning according to claim 6 is characterized in that: The model validation parameter type set includes water temperature parameters, salinity parameters, dissolved oxygen parameters and biological density parameters.

8. The method for predicting the number of organisms at nuclear power water intakes based on deep learning according to claim 2 is characterized in that: The delineation of the target species set includes the following steps: S111. Conduct a biodiversity survey in the nuclear power plant water intake area, recording the species whose occurrence frequency exceeds the preset threshold within different time periods; S112. Based on the operating characteristics of the nuclear power cooling system, screen out key biological species that may affect water extraction efficiency and summarize them to form a target biological species set.

Citation Information

Patent Citations

  • Improved artificial ecosystem optimization method based on probability-dependent reverse regeneration of producer

    CN118154385A

  • Power prediction method, system and equipment based on group algorithm

    CN118643941A