A method for simulating phytoplankton group variation and analyzing driving mechanism
By acquiring growth data and environmental data of multiple phytoplankton groups, and using a target prediction model to generate growth prediction results and reference background prediction results, and calculating contribution values, the problem of low accuracy in multi-group simulation was solved, and the accuracy of simulation of synergistic changes in the growth of multiple groups and analysis of driving mechanisms was improved.
Patent Information
- Application Number
- CN202511273352.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2045-09-08
AI Technical Summary
Existing technologies only simulate the dynamic changes in the community structure of a single phytoplankton group, which makes it difficult to accurately reflect the changes in community structure among multiple groups during eutrophication, resulting in low simulation accuracy.
By acquiring phytoplankton growth data and environmental data for multiple groups in the target water body, predictions are made using a target prediction model to generate phytoplankton growth prediction results and reference background prediction results. The contribution of environmental data to the growth of each group is calculated, and the growth driving mechanism is analyzed.
This method enables the simulation of synergistic changes in the growth of multiple phytoplankton groups, improves the accuracy of growth-driving mechanism analysis, reduces the influence of the prediction model itself, and enhances the accuracy of prediction results.
Smart Images

Figure CN120764401B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to a phytoplankton group change simulation and driving mechanism analysis method. BACKGROUND
[0002] With the increasingly serious problem of water body eutrophication, seasonal algal bloom events occur frequently, which seriously threatens the improvement of water environmental quality and the safety of water ecological environment. Therefore, it is urgent to analyze the community structure change of phytoplankton groups in the process of water body eutrophication, and to provide reference for the prevention and control of water body eutrophication.
[0003] In related technologies, a time series deep learning model is used to simulate the dynamic change of the community structure of phytoplankton of a single group, and to analyze the influence of phytoplankton of a single group on the water quality of the water body. However, in actual application scenarios, the same water body often contains phytoplankton of multiple groups at the same time, and different phytoplankton groups produce complex and different response characteristics when the water quality of the water body changes, and react on the water quality, resulting in cross-influence between phytoplankton of each group. Therefore, only simulating the dynamic change of the community structure of phytoplankton of a single group cannot accurately reflect the community structure change rule between phytoplankton groups in the process of water body eutrophication. SUMMARY
[0004] Therefore, the purpose of the present application is to provide a phytoplankton group change simulation and driving mechanism analysis method to solve the problem of low accuracy in related technologies by only simulating the dynamic change of the community structure of phytoplankton of a single group.
[0005] In a first aspect, the present application provides a phytoplankton group change simulation and driving mechanism analysis method, which comprises:
[0006] obtaining a phytoplankton dataset of a target water body in a target period; the phytoplankton dataset comprises phytoplankton growth data of multiple groups and multiple types of environmental data corresponding to the phytoplankton growth data;
[0007] obtaining reference background data with the same data dimension as the phytoplankton dataset;
[0008] processing the phytoplankton dataset and the reference background data respectively by a target prediction model to generate phytoplankton growth prediction results and reference background prediction results of the target water body in a target future period; the target prediction model is pre-optimized by a historical period phytoplankton sample dataset;
[0009] comparing the phytoplankton growth prediction result and the reference background prediction result, to generate a contribution value of the plurality of types of environmental data to the phytoplankton growth prediction result of the plurality of groups respectively;
[0010] based on the contribution value, analyzing the driving mechanism of the plurality of types of environmental data to the growth of the phytoplankton of the plurality of groups.
[0011] In an optional implementation, the environmental data includes water quality data and meteorological data; and the target prediction model respectively processes the phytoplankton data set and the reference background data to generate the phytoplankton growth prediction result of the target water body in the target future period and the reference background prediction result, comprising:
[0012] The dimension segmentation embedding module is used for performing block and first feature extraction on the phytoplankton data set in time sequence to obtain first phytoplankton growth feature data, first water quality feature data and first meteorological feature data;
[0013] The first decoder is used for performing feature fusion on the first phytoplankton growth feature data, the first water quality feature data and the first meteorological feature data to generate first phytoplankton growth decoding data, first water quality decoding data and first meteorological decoding data; wherein the first phytoplankton growth feature data, the first water quality feature data and the first meteorological feature data are subjected to time attention feature fusion and then subjected to dimension attention feature fusion together;
[0014] The first encoder is used for performing feature extraction on the first phytoplankton growth feature data, the first water quality feature data and the first meteorological feature data to generate first phytoplankton growth encoding data, first water quality encoding data and first meteorological encoding data; wherein the first phytoplankton growth feature data, the first water quality feature data and the first meteorological feature data are subjected to time attention feature extraction and then subjected to dimension attention feature extraction together;
[0015] The second decoder is used for performing feature fusion on the first phytoplankton growth encoding data, the first water quality encoding data and the first meteorological encoding data to generate second phytoplankton growth decoding data, second water quality decoding data and second meteorological decoding data;
[0016] The second encoder is used for performing feature extraction on the first phytoplankton growth encoding data, the first water quality encoding data and the first meteorological encoding data to generate second phytoplankton growth encoding data, second water quality encoding data and second meteorological encoding data;
[0017] The third decoder is used for performing feature fusion on the second phytoplankton growth encoding data, the second water quality encoding data and the second meteorological encoding data to generate the phytoplankton growth prediction result of the target water body in the target future period.
[0018] In one optional implementation, comparing the phytoplankton growth prediction results with the reference background prediction results to generate contribution values of the various types of environmental data to the phytoplankton growth prediction results of the multiple groups includes:
[0019] When the target prediction model processes the phytoplankton dataset, it obtains the phytoplankton growth data of each group and the environmental data of each type, as well as the corresponding input and output data of each network layer.
[0020] Obtain the input and output data of each network layer when the target prediction model processes the reference background data;
[0021] Based on the target prediction model processing the phytoplankton dataset, the input and output data of each group of phytoplankton growth data and each type of environmental data at each network layer, as well as the input and output data of each network layer when the target prediction model processes the reference background data, calculate the input and output changes of each group of phytoplankton growth data and each type of environmental data at each network layer in the target prediction model.
[0022] Based on the phytoplankton growth data of each group and the input and output changes of each type of environmental data in each network layer of the target prediction model, calculate the unit contribution rate of each type of environmental data to the phytoplankton growth data of each group in each network layer.
[0023] Based on the unit contribution rate of each type of environmental data in each network layer to the phytoplankton growth data of each group, the contribution value of each type of environmental data to the phytoplankton growth prediction results of each group is calculated.
[0024] In one optional implementation, the step of analyzing the growth-driving mechanisms of the various types of environmental data on the multiple groups of phytoplankton based on the contribution value includes:
[0025] Based on the contribution of each type of environmental data to the phytoplankton growth prediction results for each group, the key driving factors of each group of phytoplankton at each stage of the growth cycle are identified, so as to analyze the growth driving mechanism of the multiple types of environmental data on the multiple groups of phytoplankton.
[0026] In one optional implementation, the method further includes:
[0027] Each data point in the phytoplankton sample dataset is processed to follow a normal distribution, resulting in a normally processed phytoplankton sample dataset.
[0028] divide the phytoplankton sample data set processed by the normal distribution into training set data and test set data;
[0029] generate a phytoplankton growth prediction result corresponding to the time period of the test set data by processing the training set data by the to-be-optimized target prediction model;
[0030] set a plurality of hyperparameter search spaces of the to-be-optimized target prediction model by using a grid search algorithm, and calculate root mean square errors of the test set data and the phytoplankton growth prediction result corresponding to the time period of the test set data under the plurality of hyperparameter search spaces;
[0031] select a hyperparameter corresponding to a maximum value of the root mean square errors under the plurality of hyperparameter search spaces, set the to-be-optimized target prediction model, and obtain the target prediction model.
[0032] In an optional implementation, before the phytoplankton data set and the reference background data are respectively processed by the target prediction model, the method further includes:
[0033] fit each parameter in the phytoplankton data set respectively to generate a prediction value corresponding to each parameter;
[0034] calculate a residual error between each parameter and the corresponding prediction value to generate a residual error sequence corresponding to each parameter;
[0035] perform anomaly recognition on the residual error sequence corresponding to each parameter to remove an abnormal value in the residual error sequence corresponding to each parameter;
[0036] perform interpolation on missing data of the residual error sequence corresponding to each parameter after anomaly recognition.
[0037] In a second aspect, the present application provides a phytoplankton group change simulation and driving mechanism analysis device, the device comprises:
[0038] a plant data acquisition module for acquiring a phytoplankton data set of a target water body in a target time period; the phytoplankton data set comprises phytoplankton growth data of a plurality of groups and a plurality of types of environmental data corresponding to the phytoplankton growth data;
[0039] a reference data acquisition module for acquiring reference background data with the same data dimension as the phytoplankton data set;
[0040] a prediction module for processing the phytoplankton data set and the reference background data by a target prediction model respectively to generate a phytoplankton growth prediction result of a target water body in a target future time period and a reference background prediction result; the target prediction model is pre-optimized by a hyperparameter through a phytoplankton sample data set of a historical time period.
[0041] a contribution calculation module configured to compare the phytoplankton growth prediction result and a reference background prediction result, and generate a contribution value of the plurality of types of environmental data to the phytoplankton growth prediction result of the plurality of groups respectively;
[0042] an analysis module configured to analyze a growth driving mechanism of the plurality of groups of phytoplankton based on the contribution values.
[0043] In a third aspect, the present application provides a computer device, comprising a memory and a processor, which are communicatively connected with each other, and the memory stores computer instructions, and the processor executes the computer instructions to perform the phytoplankton group change simulation and driving mechanism analysis method of the first aspect or any of the corresponding embodiments thereof.
[0044] In a fourth aspect, the present application provides a computer readable storage medium, which stores computer instructions, and the computer instructions are used to make a computer execute the phytoplankton group change simulation and driving mechanism analysis method of the first aspect or any of the corresponding embodiments thereof.
[0045] In a fifth aspect, the present application provides a computer program product, which comprises computer instructions, and the computer instructions are used to make a computer execute the phytoplankton group change simulation and driving mechanism analysis method of the first aspect or any of the corresponding embodiments thereof.
[0046] The technical solution provided by the present application can include the following beneficial effects:
[0047] The phytoplankton group change simulation and driving mechanism analysis method provided by the present application first acquires a phytoplankton data set of a target water body in a target period; the phytoplankton data set includes phytoplankton growth data of multiple groups and multiple types of environmental data corresponding to the phytoplankton growth data; then, reference background data with the same data dimension as the phytoplankton data set is acquired; then, a target prediction model is used to process the phytoplankton data set and the reference background data respectively, to generate phytoplankton growth prediction results of the target water body in a target future period and reference background prediction results; the target prediction model is pre-optimized by a phytoplankton sample data set of a historical period; then, the phytoplankton growth prediction results and the reference background prediction results are compared to generate contribution values of the multiple types of environmental data to the phytoplankton growth prediction results of the multiple groups; finally, based on the contribution values, the driving mechanism of the multiple types of environmental data to the growth of the multiple groups of phytoplankton is analyzed. The above scheme can simultaneously predict the phytoplankton growth data of multiple groups and the multiple types of environmental data corresponding to the phytoplankton growth data, realize coordinated change simulation of the growth of multiple groups of phytoplankton, and calculate the contribution values of the multiple types of environmental data to the phytoplankton growth prediction results of the multiple groups by taking the prediction of the reference background data as a control group, thereby reducing the influence of the prediction mechanism of the target prediction model itself on the prediction results, improving the accuracy of the contribution values, and further improving the accuracy of the growth driving mechanism analysis. BRIEF DESCRIPTION OF DRAWINGS
[0048] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the drawings needed in the specific embodiments or prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0049] Figure 1 is a flowchart of a phytoplankton group change simulation and driving mechanism analysis method according to an embodiment of the present application;
[0050] Figure 2 is a flowchart of another phytoplankton group change simulation and driving mechanism analysis method according to an embodiment of the present application;
[0051] Figure 3 is a structural diagram of a target prediction model according to an embodiment of the present application;
[0052] Figure 4 is a structural block diagram of a phytoplankton group change simulation and driving mechanism analysis device according to an embodiment of the present application;
[0053] Figure 5Fig. 1 is a schematic diagram of a hardware structure of a computer device according to an embodiment of the present application. DETAILED DESCRIPTION
[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort should fall within the protection scope of the present application.
[0055] With the increasingly serious problem of water eutrophication, seasonal algal bloom events occur frequently, which seriously threatens the improvement of water environmental quality and the safety of water ecological environment. Therefore, it is urgent to analyze the community structure changes of phytoplankton groups in the process of water eutrophication to provide a reference for the prevention and control of water eutrophication.
[0056] The changes of phytoplankton community structure and environmental factors often present a complex nonlinear relationship, combined with the huge amount of monitoring data and high variable dimension, the simulation of the dynamic changes of community structure and the identification of driving mechanism of phytoplankton community structure are challenged. In the related technology, a time series deep learning model is used to simulate the dynamic changes of the community structure of a single group of phytoplankton, and the influence of a single group of phytoplankton on the water quality of the water body is analyzed. However, in actual application scenarios, the same water body often contains multiple groups of phytoplankton, and different groups of phytoplankton produce complex and different response characteristics when the water quality of the water body changes, and react on the water quality, resulting in cross-influence between different groups of phytoplankton. Therefore, only simulating the dynamic changes of the community structure of a single group of phytoplankton cannot accurately reflect the community structure change rule between different groups of phytoplankton in the process of water eutrophication.
[0057] According to the embodiments of the present application, a phytoplankton group change simulation and driving mechanism analysis method embodiment is provided. It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a group of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0058] In this embodiment, a phytoplankton group change simulation and driving mechanism analysis method is provided, which can be used in desktop computers, notebook computers, servers, etc. Figure 1 Fig. 1 is a flowchart of a phytoplankton group change simulation and driving mechanism analysis method according to an embodiment of the present application, as shown in Figure 1 The flowchart includes the following steps:
[0059] Step S101, obtain a phytoplankton dataset of a target water body in a target period.
[0060] The target water body is a water body that needs to be simulated for phytoplankton group change and driving mechanism analysis. The target period is the period corresponding to the phytoplankton dataset obtained by data collection of the target water body. In the target period, data is collected multiple times according to a preset frequency to obtain phytoplankton data containing time sequence characteristics, and then a phytoplankton dataset is obtained. The phytoplankton dataset includes phytoplankton growth data of multiple groups and multiple types of environmental data corresponding to the phytoplankton growth data. The phytoplankton growth data is used to indicate the growth of phytoplankton, such as the concentration, biomass proportion, physiological activity index (such as photosynthetic efficiency), morphological characteristics (such as average cell volume), etc. The environmental data is used to indicate the growth environment data of the phytoplankton, such as air temperature, air pressure, oxygen concentration of the target water body, etc. The phytoplankton growth data is a data sequence containing time sequence characteristics composed of the single phytoplankton growth data collected at each time in chronological order, and the environmental data is a data sequence containing time sequence characteristics composed of the single environmental data collected at each time in chronological order.
[0061] Because different groups of phytoplankton have complex competition relationships in the growth process, and different groups of phytoplankton have different suitable growth environment conditions, such as different growth rates of different groups of phytoplankton under different temperature conditions, each group of phytoplankton exhibits synergistic or antagonistic characteristics. By collecting phytoplankton growth data of multiple groups and multiple types of environmental data corresponding to the phytoplankton growth data, sufficient information can be provided for subsequent simulation of phytoplankton growth of each group by a target prediction model, capturing the mutual influence of each group of phytoplankton in the growth process, and achieving cooperative simulation of each group of phytoplankton.
[0062] Optionally, an automatic monitoring device is used to collect relevant data according to a preset frequency to obtain a phytoplankton dataset. The automatic monitoring device can use aquatic ecology and particulate organic carbon technology to collect growth data of different groups of phytoplankton by combining fluorescence differentiation method. The preset frequency can be set according to actual needs, such as once a day, once a week, etc.
[0063] Step S102, obtain reference background data with the same data dimension as the phytoplankton dataset.
[0064] The reference background data is used to indicate background values that do not contain any information, serving as a control group for the phytoplankton dataset.
[0065] Step S103, the phytoplankton dataset and the reference background data are respectively processed by the target prediction model to generate a phytoplankton growth prediction result of the target water body in a target future period and a reference background prediction result.
[0066] The target prediction model is pre-optimized by a historical period phytoplankton sample dataset, which is different from the target period phytoplankton sample dataset in that the collection period is earlier. The historical period phytoplankton sample dataset is used to optimize the hyperparameters of the target prediction model, so that the target prediction model can simulate the group change of the plurality of phytoplankton groups.
[0067] The target prediction model processes the phytoplankton dataset to generate a phytoplankton growth prediction result of the target water body in a target future period. The target future period can be selected according to requirements, for example, within 24 hours in the future. The target prediction model processes the reference background data to generate a reference background prediction result of the target water body in the target future period, that is, a prediction result obtained by the target prediction model predicting the growth of phytoplankton without any information. As a control group of the phytoplankton growth prediction result obtained by the target prediction model processing the phytoplankton dataset, it reflects the influence of the phytoplankton dataset on the prediction result of the target prediction model.
[0068] Step S104, comparing the phytoplankton growth prediction result and the reference background prediction result, generating a contribution value of each type of environmental data to the phytoplankton growth prediction result of each group.
[0069] The contribution value is used to indicate the influence degree of environmental data on the phytoplankton growth prediction result. By calculating the contribution value of each type of environmental data to the phytoplankton growth prediction result of each group, the influence degree of each type of environmental data on the phytoplankton growth prediction result of each group can be obtained, so as to analyze the growth driving mechanism of the plurality of types of environmental data to the plurality of groups of phytoplankton.
[0070] Step S105, based on the contribution value, analyzing the growth driving mechanism of the plurality of types of environmental data to the plurality of groups of phytoplankton.
[0071] The growth driving mechanism is used to indicate the influence of environmental data on the growth of phytoplankton, such as the degree and order of influence of each type of environmental data on each growth stage of phytoplankton, so as to analyze the corresponding influence degree of each type of environmental data on each growth stage of each group of phytoplankton, and determine which types of environmental data drive the growth of each group of phytoplankton in each growth stage.
[0072] The phytoplankton group change simulation and driving mechanism analysis method provided in the embodiment can simultaneously predict phytoplankton growth data of multiple groups and multiple types of environmental data corresponding to the phytoplankton growth data, realize coordinated change simulation of the growth of phytoplankton of multiple groups, calculate the contribution value of multiple types of environmental data to the prediction result of the growth of phytoplankton of multiple groups by taking the background data as a control group, reduce the influence of the prediction mechanism of the target prediction model itself on the prediction result, improve the accuracy of the contribution value, and further improve the accuracy of the growth driving mechanism analysis.
[0073] In the embodiment, a phytoplankton group change simulation and driving mechanism analysis method is provided, which can be used for desktop computers, notebook computers, servers, and the like. Figure 2 FIG. 1 is a flowchart of a phytoplankton group change simulation and driving mechanism analysis method according to an embodiment of the present application, as shown in the figure, the flow includes the following steps: Figure 2
[0074] In step S201, a target prediction model is obtained.
[0075] A pre-trained model with time series prediction function can be selected as a base model (referred to as a to-be-optimized target prediction model), which is optimized through a historical period phytoplankton data set (referred to as a phytoplankton sample data set) to obtain the target prediction model.
[0076] Specifically, first, the phytoplankton sample data set is collected, then each data in the phytoplankton sample data set is processed in a normal distribution to obtain a normal distribution processed phytoplankton sample data set; then, the normal distribution processed phytoplankton sample data set is divided into training set data and test set data; then, the to-be-optimized target prediction model processes the training set data to generate phytoplankton growth prediction results corresponding to the time period of the test set data; then, a grid search algorithm is used to set a plurality of hyperparameter search spaces of the to-be-optimized target prediction model, and the root mean square error of the test set data and the phytoplankton growth prediction results corresponding to the time period of the test set data in the plurality of hyperparameter search spaces is calculated; finally, the hyperparameters corresponding to the maximum value of the root mean square error in the plurality of hyperparameter search spaces are selected to set the to-be-optimized target prediction model to obtain the target prediction model.
[0077] For example, the phytoplankton sample data set is divided into training set, validation set and test set according to the proportions of 70%, 10% and 20%.
[0078] Optionally, the Z-score standardization method is used to process each data in the phytoplankton sample data set in a normal distribution, which is converted into a standard normal distribution with a mean of 0 and a standard deviation of 1, and the calculation formula is as follows:
[0079]
[0080] in, It is the first in the phytoplankton sample dataset Each sample value It is the mean of the phytoplankton sample dataset. It is the standard deviation of the phytoplankton sample dataset. It is the processed first Each sample value.
[0081] Optionally, a hyperparameter search space for the target prediction model to be optimized is set according to the succession cycle and diurnal variation cycle of phytoplankton groups. This hyperparameter search space includes the model input window length and the feature embedding dimension. Next, the target prediction model to be optimized is trained using training set data under different hyperparameter search spaces. The phytoplankton growth prediction results for the multiple phytoplankton groups are calculated using test set data, and the root mean square error (RMSE) between the phytoplankton growth prediction results and the measured phytoplankton growth values is calculated. Finally, the hyperparameter search space corresponding to the maximum RMSE is used as the final hyperparameters of the target prediction model to be optimized, and the final hyperparameters are used to optimize the target prediction model to obtain the target prediction model. The formula for calculating the root mean square error is as follows:
[0082]
[0083] in, These are observed values. It is a predicted value. It is the mean of the observed values. It refers to the number of observations.
[0084] Step S202: Obtain the phytoplankton dataset of the target water body during the target time period.
[0085] This phytoplankton dataset includes phytoplankton growth data from multiple groups, as well as various types of environmental data corresponding to this phytoplankton growth data.
[0086] Optionally, the phytoplankton growth data corresponds to taxa including cyanobacteria, green algae, diatoms, and cryptophytes. This phytoplankton growth data includes the relative proportion of taxa (the percentage of each taxa in the total biomass), photosynthetic efficiency, pigment composition (e.g., phycocyanin levels), cell density, average cell volume, and aggregation state (single cell / population ratio). The environmental data includes water quality data and meteorological data. The water quality data includes dissolved oxygen, pH, conductivity, turbidity, temperature, and chlorophyll a concentration. The meteorological data includes air temperature, wind speed, air pressure, and precipitation.
[0087] Step S203, obtaining reference background data with the same data dimension as the phytoplankton data set.
[0088] Optionally, a random array with the same data dimension as the phytoplankton data set is obtained as the reference background data. The random array can be generated by using a random array generation method in the related art.
[0089] Step S204, processing the phytoplankton data set and the reference background data by a target prediction model respectively to generate a phytoplankton growth prediction result of the target water body in a target future period and a reference background prediction result.
[0090] The target prediction model is pre-optimized by a historical period phytoplankton sample data set. For details, please refer to step S201, which will not be repeated here.
[0091] Optionally, before processing the phytoplankton data set and the reference background data by the target prediction model, the data in the phytoplankton data set is pre-processed. Specifically, first, each parameter in the phytoplankton data set is fitted to generate a prediction value corresponding to each parameter; then, the residual between each parameter and the corresponding prediction value is calculated to generate a residual sequence corresponding to each parameter; then, the residual sequence corresponding to each parameter is identified for abnormality to remove abnormal values in the residual sequence corresponding to each parameter; finally, the missing data in the residual sequence corresponding to each parameter after abnormality identification is interpolated.
[0092] Optionally, Holt-Winters seasonal exponential smoothing method is used to fit each parameter in the phytoplankton data set to generate a prediction value corresponding to each parameter.
[0093] Optionally, the Generalized Extreme Studentized Deviate (GESD) algorithm is used to identify the abnormality of the residual sequence corresponding to each parameter to remove abnormal values in the residual sequence corresponding to each parameter.
[0094] Optionally, a threshold is set. When the number of missing data does not exceed the threshold, the residual sequence is determined to be a short-term missing sequence, and linear interpolation is used for interpolation. The interpolation value is calculated according to the following formula:
[0095]
[0096] wherein, x 0 is the interpolation value, t 1 is the time before the start of missing data, t 2 is the time after the end of missing data, x 1 ist Data corresponding to time 1 x 2 is t Data corresponding to time 2.
[0097] When the number of missing data exceeds the imputation threshold, the residual sequence is determined to be a long-term missing sequence, and a two-stage Kalman filter is used for imputation. The imputation value is calculated using the following formula:
[0098]
[0099] in, For this Kalman filter in The predicted state estimation vector at time t. for The interpolation value at time, for State transition parameters at time t, for The random state noise term at time t. for Measurement parameters at time, for The measurement error term for time.
[0100] The target prediction model includes a Dimension-Segment-Wise Embedding (DSW) module, multiple encoders, and multiple decoders. The DSW module is used to segment the phytoplankton dataset into blocks and perform preliminary feature extraction in chronological order. The encoders are used to extract features from the input data, and the decoders are used to fuse features from the input data. The number of encoders and decoders can be set according to requirements, with one more decoder than encoder.
[0101] With the target prediction model including three decoders and two encoders as an example, when the phytoplankton dataset is processed by the target prediction model respectively, first, the dimension segmentation embedding module blocks and extracts first phytoplankton growth feature data, first water quality feature data and first meteorological feature data from the phytoplankton dataset in chronological order; then, the first decoder is used to perform feature fusion on the first phytoplankton growth feature data, the first water quality feature data and the first meteorological feature data to generate first phytoplankton growth decoding data, first water quality decoding data and first meteorological decoding data; wherein, after time attention feature fusion of the first phytoplankton growth feature data, the first water quality feature data and the first meteorological feature data respectively, dimension attention feature fusion is performed together; then, the first encoder is used to extract features from the first phytoplankton growth feature data, the first water quality feature data and the first meteorological feature data to generate first phytoplankton growth encoding data, first water quality encoding data and first meteorological encoding data; wherein, after time attention feature extraction of the first phytoplankton growth feature data, the first water quality feature data and the first meteorological feature data respectively, dimension attention feature extraction is performed together; then, the second decoder is used to perform feature fusion on the first phytoplankton growth encoding data, the first water quality encoding data and the first meteorological encoding data to generate second phytoplankton growth decoding data, second water quality decoding data and second meteorological decoding data; then, the second encoder is used to extract features from the first phytoplankton growth encoding data, the first water quality encoding data and the first meteorological encoding data to generate second phytoplankton growth encoding data, second water quality encoding data and second meteorological encoding data; finally, the third decoder is used to perform feature fusion on the second phytoplankton growth encoding data, the second water quality encoding data and the second meteorological encoding data to generate a phytoplankton growth prediction result of the target water body in a target future period.
[0102] Optionally, the encoder of the target prediction model comprises a segment fusion module (Seg Merge) and a time-space attention module (TSA Layer, Time-Space Attention Layer) in sequence, the time-space attention module comprises a plurality of parallel time attention feature extraction modules (Time Attention) and a plurality of serial dimension attention feature extraction modules (Dim Attention) in sequence, and the plurality of parallel time attention feature extraction modules process phytoplankton growth data, water quality data and meteorological data respectively. The decoder of the target prediction model comprises a plurality of parallel time attention feature fusion modules (Time Attention), a plurality of serial dimension attention feature fusion modules (Dim Attention), a multi-head self-attention module (MSA, Multi-Head Attention) and a fully connected layer (FC, Fully Connected Layer) in sequence. The target prediction model further comprises a position embedding module (Position Embedding) for identifying position information of each data corresponding to a time step in the phytoplankton growth data, the water quality data and the meteorological data, the position information of the time step comprising time order between each data in each data block (data sequence) of the segmented phytoplankton data set and time order between each data block. Wherein, the time step refers to a time interval between one observation point and the next observation point in the sequence data.
[0103] Exemplarily, Figure 3 is a structural schematic diagram of a target prediction model according to an embodiment of the present application. As Figure 3 shown, the target prediction model comprises n encoders and n+1 decoders. enc is the encoder output, z dec is the decoder output.
[0104] Exemplarily, the target prediction model is a Crossformer model.
[0105] In step S205, the phytoplankton growth prediction result and the reference background prediction result are compared to generate a contribution value of each type of environmental data to the phytoplankton growth prediction result of each group.
[0106] Specifically, input data and output data corresponding to each network layer of each class group of phytoplankton growth data and each type of environmental data in the phytoplankton dataset processed by the target prediction model are obtained; then, input data and output data of each network layer when the target prediction model processes the reference background data are obtained; then, based on the input data and output data corresponding to each network layer of each class group of phytoplankton growth data and each type of environmental data in the phytoplankton dataset processed by the target prediction model and the input data and output data of each network layer when the target prediction model processes the reference background data, the input change and output change of each class group of phytoplankton growth data and each type of environmental data in each network layer in the target prediction model are calculated; then, according to the input change and output change of each class group of phytoplankton growth data and each type of environmental data in each network layer in the target prediction model, the unit contribution rate of each type of environmental data to the phytoplankton growth data of each class group in each network layer is calculated; finally, according to the unit contribution rate of each type of environmental data to the phytoplankton growth data of each class group in each network layer, the contribution value of each type of environmental data to the phytoplankton growth prediction result of each class group is calculated.
[0107] For example, assuming that the phytoplankton dataset is , and the reference background data is , the formula for calculating the output change and the input change is as follows:
[0108]
[0109]
[0110] , wherein, is the output change, indicating the difference between the phytoplankton growth prediction result and the reference background prediction result generated by the target prediction model under the conditions of actual parameter (phytoplankton dataset) input and information-free background data (reference background data) input; is the input change of the i-th feature (used to indicate the data of each type in the phytoplankton dataset, such as temperature), which is used to indicate the difference between the i-th feature in the phytoplankton dataset and the reference background data; i is the value of the i-th feature in the phytoplankton dataset i , and is the value of the i-th feature in the reference background data . i i
[0111] Next, the unit contribution rate of the input to the output of each network layer in the Crossformer model (including any module in the target prediction model that processes the input data and outputs the processed data) is calculated layer by layer. The formula is as follows:
[0112]
[0113] in, For the first j The unit contribution rate of the input to the output of each network layer is used to indicate the degree of influence of changes in the input data of the corresponding network layer on the output result; For the first j The local analytical solution of the contribution (Shapely value) of the input of each network layer to the output.
[0114] Then, following the order from the output layer to the input layer of the target prediction model, the unit contribution rates of each network layer are multiplied together to obtain the input of the model input layer. For the model output layer The unit contribution rate is calculated using the following formula:
[0115]
[0116] in, Input to the model input layer For the model output layer unit contribution rate Input to the model input layer For the intermediate layer of the model unit contribution rate Input for the intermediate layer of the model For the model output layer The unit contribution rate.
[0117] Finally, the final contribution of each input feature of the target prediction model to the output result is calculated using the following formula:
[0118]
[0119] in, For the first i The Shapely value of each feature.
[0120] Step S206: Based on the contribution value, analyze the growth driving mechanism of the various types of environmental data for the multiple groups of phytoplankton.
[0121] Specifically, based on the contribution value of each type of environmental data to the phytoplankton growth prediction result of each class group, the key driving factors of each class group of phytoplankton at each stage in the growth cycle are identified to analyze the driving mechanism of the plurality of types of environmental data on the growth of the plurality of classes of phytoplankton.
[0122] Optionally, the absolute value of the contribution value of each type of environmental data to the phytoplankton growth prediction result of each class group at each time step corresponding to the input time series data of the target prediction model is calculated, and the absolute values of the contribution values of each type of environmental data to the phytoplankton growth prediction result of each class group at different time steps are averaged as the driving importance indicators of each type of environmental data to the phytoplankton growth prediction result of each class group. The driving importance indicators corresponding to each type of environmental data are sorted according to size and time. The earlier the sorting, the more critical the impact on the growth of phytoplankton. The environmental data sorted at the most front position in the target stage of the growth cycle of the target class group of phytoplankton is the key driving factor corresponding to the target stage of the growth cycle of the target class group of phytoplankton.
[0123] Optionally, the contribution value of each type of environmental data to the phytoplankton growth prediction result of each class group and the key driving factors of each class group of phytoplankton at each stage in the growth cycle are introduced into different lag times (for example, the lag time is set to 1, 2, 3, …, 24 in hours) for time lag effect analysis. The lag time is realized by uniformly shifting the corresponding time of the environmental data in the original input sequence forward by τ hours (τ = 0, 1, 2, …, 24). Based on the time lag effect analysis result, the importance change trend of each type of environmental data with different lag times is obtained, and the time lag effect of the driving factor on the growth of phytoplankton is revealed. Specifically, the relationship among the lag time, the contribution value of each type of environmental data to the phytoplankton growth prediction result of each class group, and the key driving factors of each class group of phytoplankton at each stage in the growth cycle is fitted by local weighted regression, and the formula is as follows:
[0124]
[0125] wherein, represents the position of the selected prediction point, represents the th data point, represents the total number of data points, represents the regression estimation prediction value at ; is the distance weight assigned to the data point centered on ; is the data point corresponding to the measured value, represents the intercept of the local linear model, represents the slope of the local linear model, is the regression coefficient fitted within the local window.
[0126] Exemplarily, the various stages of the phytoplankton in the growth cycle include the growth period, the outbreak period and the extinction period. In analyzing the growth driving mechanism of the plurality of types of environmental data on the plurality of groups of phytoplankton, the Shapley values corresponding to each type of environmental data of the four groups of phytoplankton are sorted according to feature importance, the key driving factors corresponding to the changes in the community structure of the four groups of phytoplankton in different stages are identified, the relationship between the Shapley values corresponding to the four groups of phytoplankton and the key driving factors is fitted by local weighted regression respectively, and the differential effects of the driving factors corresponding to each type of environmental data on the changes in the community structure of the phytoplankton in the various stages of the growth cycle of the phytoplankton are revealed.
[0127] Exemplarily, in analyzing the growth driving mechanism of the plurality of types of environmental data on the plurality of groups of phytoplankton, in the day and night periods, the Shapley values corresponding to each type of environmental data of the four groups of phytoplankton are sorted according to feature importance, the key driving factors corresponding to the changes in the community structure of the four groups of phytoplankton in the day and night periods are identified, the relationship between the Shapley values corresponding to the four groups of phytoplankton and the key driving factors is fitted by local weighted regression respectively, and the differential effects of the driving factors corresponding to each type of environmental data on the changes in the community structure of the phytoplankton in the different periods of the day and night are revealed.
[0128] The phytoplankton group change simulation and driving mechanism analysis method provided in the embodiment can simultaneously predict the growth data of the plurality of groups of phytoplankton and the plurality of types of environmental data corresponding to the growth data of the phytoplankton, realize the coordinated change simulation of the growth of the plurality of groups of phytoplankton, calculate the contribution values of the plurality of types of environmental data to the growth prediction results of the plurality of groups of phytoplankton by taking the prediction based on the background data as a control group, reduce the influence of the prediction mechanism of the target prediction model itself on the prediction results, improve the accuracy of the contribution values, and further improve the accuracy of the growth driving mechanism analysis.
[0129] In the embodiment, a phytoplankton group variation simulation and driving mechanism analysis device is also provided, which is used to implement the above-mentioned embodiments and preferred embodiments, and details are not repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware, or a combination of software and hardware can also be implemented and conceived.
[0130] The embodiment provides a phytoplankton group variation simulation and driving mechanism analysis device, as shown in the accompanying drawings, comprising: Figure 4
[0131] A phytoplankton data acquisition module 401 is configured to acquire a phytoplankton dataset of a target water body in a target period; the phytoplankton dataset comprises phytoplankton growth data of a plurality of groups and a plurality of types of environmental data corresponding to the phytoplankton growth data;
[0132] A reference data acquisition module 402 is configured to acquire reference background data with the same data dimension as the phytoplankton dataset;
[0133] A prediction module 403 is configured to process the phytoplankton dataset and the reference background data respectively by a target prediction model to generate a phytoplankton growth prediction result of the target water body in a target future period and a reference background prediction result; the target prediction model is pre-optimized by a phytoplankton sample dataset of a historical period;
[0134] A contribution calculation module 404 is configured to compare the phytoplankton growth prediction result and the reference background prediction result to generate a contribution value of the plurality of types of environmental data to the phytoplankton growth prediction result of the plurality of groups;
[0135] An analysis module 405 is configured to analyze a growth driving mechanism of the plurality of types of environmental data to the plurality of groups of phytoplankton based on the contribution value.
[0136] In an optional implementation, the environmental data includes water quality data and meteorological data; the prediction module is further configured to: segment and perform first feature extraction on the phytoplankton data set in chronological order by a dimensional segmentation embedding module to obtain first phytoplankton growth feature data, first water quality feature data, and first meteorological feature data; perform feature fusion on the first phytoplankton growth feature data, the first water quality feature data, and the first meteorological feature data by a first decoder to generate first phytoplankton growth decoding data, first water quality decoding data, and first meteorological decoding data; wherein the first phytoplankton growth feature data, the first water quality feature data, and the first meteorological feature data are subjected to time attention feature fusion and then to dimensional attention feature fusion; perform feature extraction on the first phytoplankton growth feature data, the first water quality feature data, and the first meteorological feature data by a first encoder to generate first phytoplankton growth encoding data, first water quality encoding data, and first meteorological encoding data; wherein the first phytoplankton growth feature data, the first water quality feature data, and the first meteorological feature data are subjected to time attention feature extraction and then to dimensional attention feature extraction; perform feature fusion on the first phytoplankton growth encoding data, the first water quality encoding data, and the first meteorological encoding data by a second decoder to generate second phytoplankton growth decoding data, second water quality decoding data, and second meteorological decoding data; perform feature extraction on the first phytoplankton growth encoding data, the first water quality encoding data, and the first meteorological encoding data by a second encoder to generate second phytoplankton growth encoding data, second water quality encoding data, and second meteorological encoding data; and perform feature fusion on the second phytoplankton growth encoding data, the second water quality encoding data, and the second meteorological encoding data by a third decoder to generate a phytoplankton growth prediction result of the target water body in a target future period.
[0137] In an optional implementation, the contribution calculation module is further configured to: obtain the input data and the output data of each network layer when the target prediction model processes the reference background data; and calculate the input change and the output change of each class group of phytoplankton growth data and each type of environmental data at each network layer in the target prediction model based on the input data and the output data of each network layer when the target prediction model processes the reference background data, the input data and the output data of each class group of phytoplankton growth data and each type of environmental data corresponding to each network layer when the target prediction model processes the phytoplankton data set, and the input change and the output change of each class group of phytoplankton growth data and each type of environmental data at each network layer in the target prediction model; calculate the unit contribution rate of each type of environmental data to the phytoplankton growth data of each class group in each network layer based on the input change and the output change of each class group of phytoplankton growth data and each type of environmental data at each network layer in the target prediction model; and calculate the contribution value of each type of environmental data to the phytoplankton growth prediction result of each class group based on the unit contribution rate of each type of environmental data to the phytoplankton growth data of each class group in each network layer.
[0138] In an optional implementation, the analysis module is further configured to: identify the key driving factors of each class group of phytoplankton at each stage in the growth cycle based on the contribution value of each type of environmental data to the phytoplankton growth prediction result of each class group, so as to analyze the growth driving mechanism of the plurality of types of environmental data to the plurality of class groups of phytoplankton.
[0139] In an optional implementation, the device further comprises a hyperparameter optimization module configured to: perform normal distribution processing on each data in the phytoplankton sample data set to obtain a normal distribution processed phytoplankton sample data set; divide the normal distribution processed phytoplankton sample data set into training set data and test set data; process the training set data by a to-be-optimized target prediction model to generate a phytoplankton growth prediction result corresponding to the time period of the test set data; set a plurality of hyperparameter search spaces of the to-be-optimized target prediction model by using a grid search algorithm, and calculate the root mean square error of the test set data and the phytoplankton growth prediction result corresponding to the time period of the test set data under the plurality of hyperparameter search spaces; and select the hyperparameters corresponding to the maximum root mean square error in the plurality of hyperparameter search spaces to set the to-be-optimized target prediction model, thereby obtaining the target prediction model.
[0140] In one optional embodiment, the device further includes a preprocessing module, configured to: fit each parameter in the phytoplankton dataset to generate a predicted value for each parameter; calculate the residual between each parameter and the corresponding predicted value to generate a residual sequence for each parameter; perform anomaly identification on the residual sequence for each parameter and remove outliers from the residual sequence for each parameter; and impute missing data in the residual sequence for each parameter after anomaly identification.
[0141] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.
[0142] The phytoplankton community change simulation and driving mechanism analysis device in this embodiment is presented in the form of a functional unit. Here, a unit refers to an ASIC (Application Specific Integrated Circuit), a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0143] This invention also provides a computer device having the above-described features. Figure 4 The device shown is for simulating and analyzing the driving mechanism of phytoplankton group changes.
[0144] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 5 As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 5 Take a processor 10 as an example.
[0145] The processor 10 can be a central processing unit, a network processing unit, or a combination thereof. The processor 10 can further include a hardware chip. The hardware chip can be an application specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device can be a complex programmable logic device, a field programmable logic device, a generic array logic, or any combination thereof.
[0146] The memory 20 stores instructions executable by the at least one processor 10 to cause the at least one processor 10 to perform the methods illustrated in the above embodiments.
[0147] The memory 20 can include a program storage area and a data storage area. The program storage area can store an operating system and application programs required by at least one function. The data storage area can store data created according to the use of the computer device, and the like. In addition, the memory 20 can include a high-speed random access memory, and can further include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some alternative embodiments, the memory 20 can optionally include a memory disposed remotely from the processor 10, and these remote memories can be connected to the computer device through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0148] The memory 20 can include a volatile memory, such as a random access memory, and can also include a non-volatile memory, such as a flash memory, a hard disk, or a solid state disk. The memory 20 can further include a combination of the above-mentioned types of memories.
[0149] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30, and the output device 40 can be connected by a bus or other means, Figure 5 For example, by a bus.
[0150] The input device 30 can receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, and the like. The output device 40 can include a display device, an auxiliary lighting device (e.g., an LED), a tactile feedback device (e.g., a vibration motor), and the like. The display device includes, but is not limited to, a liquid crystal display, a light emitting diode, a display, and a plasma display. In some alternative embodiments, the display device can be a touch screen.
[0151] The embodiments of the present application further provide a computer readable storage medium, and the method according to the embodiments of the present application can be implemented in hardware, firmware, or recorded in a storage medium, or stored in a remote storage medium or a non-transitory machine readable storage medium and downloaded to a local storage medium through network, so that the method described herein can be processed by such software on a storage medium using a general purpose computer, a special purpose processor, or programmable or special hardware. The storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid state disk, etc. Further, the storage medium can also include a combination of the above-mentioned memories. It can be understood that the computer, the processor, the microprocessor controller, or the programmable hardware includes a storage component that can store or receive software or computer code, when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.
[0152] Part of the present application can be applied as a computer program product, for example, computer program instructions, when executed by a computer, through the operation of the computer, the method and / or technical solutions according to the present application can be called or provided. Those skilled in the art should understand that the form of computer program instructions in a computer readable medium includes but is not limited to source files, executable files, installation package files, etc. Correspondingly, the way of executing computer program instructions by computer includes but is not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Here, the computer readable medium can be any available computer readable storage medium or communication medium accessible to the computer.
[0153] Although the embodiments of the present application are described in conjunction with the accompanying drawings, various modifications and changes can be made by those skilled in the art without departing from the spirit and scope of the present application, and such modifications and changes fall within the scope of protection of the present application.
Claims
1. A method for simulating phytoplankton group change and analyzing driving mechanism, characterized in that, The method comprises: obtaining a phytoplankton dataset of a target water body in a target period; the phytoplankton dataset comprises phytoplankton growth data of a plurality of groups and a plurality of types of environmental data corresponding to the phytoplankton growth data; obtaining reference background data with the same data dimension as the phytoplankton dataset; processing the phytoplankton dataset and the reference background data respectively by a target prediction model to generate a phytoplankton growth prediction result of the target water body in a target future period and a reference background prediction result respectively; the target prediction model is pre-optimized by a historical phytoplankton sample dataset; comparing the phytoplankton growth prediction result and the reference background prediction result to generate a contribution value of the plurality of types of environmental data to the phytoplankton growth prediction result of the plurality of groups respectively; based on the contribution value, analyzing the growth driving mechanism of the plurality of types of environmental data to the plurality of groups of phytoplankton.
2. The method of claim 1, wherein, The environmental data comprises water quality data and meteorological data; the processing of the phytoplankton dataset and the reference background data respectively by the target prediction model to generate the phytoplankton growth prediction result of the target water body in the target future period and the reference background prediction result respectively comprises: performing block and first feature extraction on the phytoplankton dataset in time sequence by a dimension segmentation embedding module to obtain first phytoplankton growth feature data, first water quality feature data and first meteorological feature data; performing feature fusion on the first phytoplankton growth feature data, the first water quality feature data and the first meteorological feature data by a first decoder to generate first phytoplankton growth decoding data, first water quality decoding data and first meteorological decoding data; wherein the first phytoplankton growth feature data, the first water quality feature data and the first meteorological feature data are subjected to time attention feature fusion and then dimension attention feature fusion together; performing feature extraction on the first phytoplankton growth feature data, the first water quality feature data and the first meteorological feature data by a first encoder to generate first phytoplankton growth encoding data, first water quality encoding data and first meteorological encoding data; wherein the first phytoplankton growth feature data, the first water quality feature data and the first meteorological feature data are subjected to time attention feature extraction and then dimension attention feature extraction together; performing feature fusion on the first phytoplankton growth encoding data, the first water quality encoding data and the first meteorological encoding data by a second decoder to generate second phytoplankton growth decoding data, second water quality decoding data and second meteorological decoding data; performing feature extraction on the first phytoplankton growth encoding data, the first water quality encoding data and the first meteorological encoding data by a second encoder to generate second phytoplankton growth encoding data, second water quality encoding data and second meteorological encoding data; performing feature fusion on the second phytoplankton growth encoding data, the second water quality encoding data and the second meteorological encoding data by a third decoder to generate the phytoplankton growth prediction result of the target water body in the target future period.
3. The method of claim 1, wherein, The comparing the phytoplankton growth prediction result with the reference background prediction result generates a contribution value of each type of environmental data to the phytoplankton growth prediction result of each group. Obtaining input data and output data corresponding to each network layer of each group of phytoplankton growth data and each type of environmental data when the target prediction model processes the phytoplankton data set; Obtaining input data and output data of each network layer when the target prediction model processes the reference background data; Based on the input data and output data corresponding to each network layer of each group of phytoplankton growth data and each type of environmental data when the target prediction model processes the phytoplankton data set, and the input data and output data of each network layer when the target prediction model processes the reference background data, calculating the input change and output change of each network layer of each group of phytoplankton growth data and each type of environmental data in the target prediction model; According to the input change and output change of each network layer of each group of phytoplankton growth data and each type of environmental data in the target prediction model, calculating the unit contribution rate of each type of environmental data to the phytoplankton growth data of each group in each network layer; According to the unit contribution rate of each type of environmental data to the phytoplankton growth data of each group in each network layer, calculating the contribution value of each type of environmental data to the phytoplankton growth prediction result of each group.
4. The method of claim 3, wherein, Based on the contribution value, the method further comprises: Based on the contribution value of each type of environmental data to the phytoplankton growth prediction result of each group, identifying the key driving factor of each group of phytoplankton at each stage in the growth cycle to analyze the growth driving mechanism of the plurality of types of environmental data to the plurality of groups of phytoplankton.
5. The method according to any one of claims 1 to 4, characterized in that, The method further comprises: Normal distribution processing is performed on each data in the phytoplankton sample data set to obtain a normal distribution processed phytoplankton sample data set; The normal distribution processed phytoplankton sample data set is divided into training set data and test set data; A target prediction model is used to process the training set data to generate a phytoplankton growth prediction result corresponding to the time period of the test set data; A grid search algorithm is used to set a plurality of hyperparameter search spaces of the target prediction model, and the root mean square error of the test set data and the phytoplankton growth prediction result corresponding to the time period of the test set data in the plurality of hyperparameter search spaces is calculated; The hyperparameters corresponding to the maximum root mean square error in the plurality of hyperparameter search spaces are selected to set the target prediction model, and the target prediction model is obtained.
6. The method according to any one of claims 1 to 4, characterized in that, Before the target prediction model processes the phytoplankton data set and the reference background data respectively, the method further comprises: Each parameter in the phytoplankton data set is fitted to generate a prediction value corresponding to each parameter; Calculate the residual between each parameter and the corresponding predicted value, generate the residual sequence corresponding to each parameter; Abnormal identification is performed on the residual sequence corresponding to each parameter, and abnormal values in the residual sequence corresponding to each parameter are removed; The missing data of the residual sequence corresponding to each parameter after abnormal identification is interpolated.
7. An apparatus for simulating phytoplankton group change and analyzing driving mechanism, characterized in that, The device comprises: A plant data acquisition module is configured to acquire a phytoplankton data set of a target water body in a target period; the phytoplankton data set comprises phytoplankton growth data of a plurality of groups and a plurality of types of environmental data corresponding to the phytoplankton growth data; A reference data acquisition module is configured to acquire reference background data with the same data dimension as the phytoplankton data set; A prediction module is configured to process the phytoplankton data set and the reference background data respectively by a target prediction model to generate a phytoplankton growth prediction result of the target water body in a target future period and a reference background prediction result; the target prediction model is pre-optimized by a phytoplankton sample data set of a historical period; A contribution calculation module is configured to compare the phytoplankton growth prediction result and the reference background prediction result to generate a contribution value of the plurality of types of environmental data to the phytoplankton growth prediction result of the plurality of groups; An analysis module is configured to analyze a growth driving mechanism of the plurality of types of environmental data to the plurality of groups of phytoplankton based on the contribution value.
8. A computer device, comprising: It comprises: A memory and a processor, which are communicatively connected, the memory stores computer instructions, and the processor executes the computer instructions to perform the phytoplankton group change simulation and driving mechanism analysis method in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions, which are used to make the computer execute the phytoplankton group change simulation and driving mechanism analysis method in any one of claims 1 to 6.
10. A computer program product, characterised in that, It comprises computer instructions, which are used to make the computer execute the phytoplankton group change simulation and driving mechanism analysis method in any one of claims 1 to 6.
Citation Information
Patent Citations
Method for judging habitat threshold values of different ecological restoration types in drainage basin
CN113283743A
Nearshore phytoplankton community structure prediction method based on attention mechanism
CN118261203A