A method for predicting the acidity and alkalinity in the water quality for the growth of sea cucumbers
Through the method of variational modal decomposition and feature fusion, each modal component of sea cucumber growth water quality data is extracted, and fused with relevant factors. The prediction is performed using a recurrent neural network, which solves the problem of inaccurate prediction of the pH of sea cucumber growth water quality in the existing technology, improves the prediction accuracy and supports intelligent breeding.
Patent Information
- Application Number
- CN202210366226.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-08
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-04-08
AI Technical Summary
The prior art is difficult to accurately predict the pH in the water quality of sea cucumbers, especially in the presence of nonlinear and non-mechanical structures in the water quality of sea water.
Using a prediction method based on variational modal decomposition and feature fusion, we use multi-modal decomposition of sea cucumber growth water quality data, extract each modal component, and fuse it with water quality factors and meteorological factors to form extended modal component data, normalized standardization processing and segmentation, and use recurrent neural network for training and prediction.
It improves the prediction accuracy of pH in the water quality of sea cucumber growth, avoids overfitting, provides a scientific decision-making basis, and lays the foundation for the sea cucumber aquaculture industry to achieve intelligent breeding.
Smart Images

Figure CN114841412B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sea cucumber growth water quality management, and particularly relates to a method for predicting the pH value in the sea cucumber growth water quality based on variational mode decomposition and feature fusion. Background Art
[0002] Sea cucumbers belong to bioremediation species and have extremely high requirements for water quality. If the water quality is polluted, sea cucumbers will also show autolysis or even death. With the popularization of informatization, intelligent or even smart systems have been basically realized in various fields. As one of the important parameters affecting the sea cucumber growth water quality, making a reasonable prediction of pH in advance is a strong scientific decision-making basis for realizing intelligent sea cucumber farming. It is found in the research that the complex characteristics of hydrology can be summarized as follows: (1) water quality data belongs to time series data and has strong periodic characteristics; (2) due to large fluctuations in weather factors, the data shows strong instability; (3) there is complex correlation among various water quality data. If the original data is simply used to analyze and predict the pH in the sea cucumber growth water quality, the results are not accurate.
[0003] At present, the water quality prediction models for seawater can be divided into three types. The first is the hydrodynamic simulation model, such as S-P, EFDC, etc. This model describes the change rules of multi-dimensional water quality indicators, but has high requirements for regional data and it is difficult to master the accuracy. The second is the regression analysis model based on machine learning, such as statistical regression, SVR, etc. This model is simple but not applicable to the application of non-linear problems. The third is the neural network model, which does not reflect the time sequence relationship. Summary of the Invention
[0004] Aiming at the above-mentioned defects existing in the prior art, the present invention proposes a method for predicting the pH value in the sea cucumber growth water quality. By highly decomposing to explore the characteristics of the target factor pH, and fusing the relationship between multiple associated influencing factors and water quality prediction factors, a "decomposition-fusion-prediction-reconstruction" model is formed to improve the accuracy of water quality prediction.
[0005] To achieve the above object, the technical solution of the present application is: a method for predicting the pH value in the sea cucumber growth water quality, including:
[0006] Cleaning and repairing the collected multiple water quality data;
[0007] Screening out the water quality factors and meteorological factors related to the pH value for later model experiments;
[0008] Performing multi-modal decomposition on the repaired water quality factor data to obtain each modal component IMF;
[0009] Performing difference verification on each modal component IMF and the original data to obtain the residual verification data R;
[0010] Each of the modal components IMF is fused with water quality factors, meteorological factors, and residual verification data to form extended modal component data new-IMFs;
[0011] Perform normalization processing on the extended modal component data new-IMFs;
[0012] Segment the extended modal component data after normalization processing to obtain a training set, a test set, and a validation set;
[0013] The segmented extended modal component data is subjected to format conversion and fitting training and learning; and the data of the validation set and the test set of each extended modal component are evaluated for performance using the mean absolute error MAE, the mean absolute percentage error MAPE, and the root mean square error RMSE.
[0014] Perform reconstruction inverse normalization on the prediction result, automatically restore the original data column, and obtain the error of the prediction data again.
[0015] Furthermore, clean and repair the collected multiple water quality data, specifically:
[0016] Eliminate redundant data in the multiple water quality data;
[0017] Specifically, in the data cleaning work, since the self-built wireless communication data acquisition system may have unstable signals and network delays, resulting in non-periodic data acquisition, it is necessary to eliminate redundant data to ensure data synchronization;
[0018] Perform interpolation operations on the multiple water quality data after elimination using the linear interpolation algorithm. The formula used in the algorithm is:
[0019] f(x t )=(1 - x t - x t-1 / x t+1 - x t )f(x t-1 ) + x t - x t-1 / x t+1 - x t f(x t )
[0020] where f(x) is the value to be filled, x t is the target factor value at a certain moment, and x t-1 , x t+1 are the target factor values at the previous moment and the next moment respectively.
[0021] Specifically, in case of data loss, the linear interpolation algorithm is used to ensure data integrity.
[0022] Furthermore, water quality factors and meteorological factors related to pH are screened out. Specifically, the Spearman rank method is used to perform correlation analysis on the repaired data to screen out water quality factors and meteorological factors that have a strong correlation with the dissolved oxygen factor.
[0023] Furthermore, multi-modal decomposition is performed on the repaired water quality factor data. Specifically, first, the repaired water quality factor data is subjected to Hilbert transform to obtain a single-sided spectrum signal, and then the single-sided spectrum signal is translated to the baseband, and then bandwidth estimation is performed.
[0024] Furthermore, each modal component IMF is obtained. Specifically:
[0025] Initialize the parameters: the number of components k = 5, the double rise time step tau = 0, and the central constraint strength alpha = 7000;
[0026] Update all component lists u k and the central frequency value ω k according to the following formula;
[0027]
[0028]
[0029] where f refers to the Fourier transform formula;
[0030] According to the given evaluation criteria, continuously iterate until iterating K times and the filter iteration accuracy is less than ε, then stop the iteration; where the evaluation criteria are:
[0031]
[0032] Even further, the difference check is performed between each modal component IMF and the original data to obtain the residual check data R. Specifically:
[0033]
[0034] Perform multi-modal decomposition on the target factor Y to form k IMF components. The shape of Y is changed from (n,1) to (n,k), and at the same time, a difference component R is generated. The data format conversion is as follows:
[0035]
[0036] Even further, each modal component IMF is fused with water quality factors, meteorological factors, and residual check data to form extended modal component data new-IMFs. Specifically:
[0037] Add Gaussian white noise \(r\) to each modal component IMF, and after fusion, form matrix \(D\), as shown in the formula:
[0038]
[0039] In the formula, \(u\) i is the formed single modal component, \(\zeta\) i is the two-dimensional matrix of water quality factors, \(s\) 2 is the calculated variance of the modal component, \(N\) is the length of the data column of the modal component, and randn represents random generation;
[0040] Specifically, to prevent overfitting of the target factor prediction, add Gaussian white noise \(r\) to each modal component IMF. The number of components \(k\) is set to 6. Component 1 represents the trend component, components 2 - 3 represent the periodic components, and components 4 - 6 represent the generated noise components. To reduce the greater impact of adding Gaussian noise to the noise components, no Gaussian noise is added to components 4 - 6.
[0041] Change matrix \(D\) into a three-dimensional matrix \(T\) with the shape of \((12, q, m)\), and its data format is as follows:
[0042]
[0043] In the formula, \(m\) is the number of samples and \(q\) is the sample step size.
[0044] Furthermore, perform normalization processing on the extended modal component data new - IMFs, specifically:
[0045]
[0046] Change the water quality data with different dimensions into the same - dimension data between \([0, 1]\). \(x\) t is the measured value at time \(t\), \(min[x]\) is the minimum value in a certain measurement data column, and \(max[x]\) is the maximum value in a certain measurement data column.
[0047] As a further step, the evaluation functions of mean absolute error MAE, mean absolute percentage error MAPE, and root mean square error RMSE are respectively:
[0048]
[0049]
[0050]
[0051] Among them, \(y\) i is the true value of the target factor, is the predicted value of the target factor, and the performance of the constructed model is evaluated through different evaluation functions.
[0052] As a further step, the prediction result pre_Y is obtained using the following formula:
[0053]
[0054] IMF preyk_n is the predicted value of the k-th component at the n-th moment, and IMF yk is the column of the original values of each component.
[0055] The above technical solutions adopted by the present invention, compared with the prior art, have the following advantages:
[0056] 1. Since the strong corrosiveness of seawater will cause unstable sensor data transmission, the present invention realizes the repair processing of data.
[0057] 2. Aiming at the non-linear and non-mechanistic structure of seawater quality, the variational mode decomposition technology is used to extract the surface layer characteristics of the pH value in the water quality for sea cucumber growth.
[0058] 3. To prevent overfitting in the prediction of sea cucumber growth water quality data, Gaussian white noise is added to some surface layer characteristics, and other water quality influencing factors are further fused, so as to perform cyclic prediction of components and improve the accuracy of sea cucumber growth water quality prediction.
[0059] 4. The present invention can provide a scientific decision-making basis for breeders in this field and lay a good foundation for the realization of intelligent aquaculture in the sea cucumber aquaculture industry. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 is the flowchart of water quality prediction based on variational mode decomposition and multi-feature fusion long short-term memory network (VLSTM);
[0061] Figure 2 is the multi-feature correlation analysis diagram;
[0062] Figure 3 is the multi-modal decomposition (VMD) component diagram and the central frequency diagram;
[0063] Figure 4 is the comparison diagram of the PH prediction of a single long short-term memory network (LSTM) and VLSTM;
[0064] Figure 5 is the change curve diagram of the loss values of a single LSTM and VLSTM;
[0065] Figure 6 is the error curve diagram of the predicted value and the original data. DETAILED DESCRIPTION OF THE INVENTION
[0066] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application, that is, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Usually, the components of the embodiments of the present application described and shown in the accompanying drawings herein can be arranged and designed in various different configurations.
[0067] Therefore, the detailed description of the embodiments of the present application provided in the accompanying drawings below is not intended to limit the scope of the present application claimed, but only represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0068] Embodiment 1
[0069] As Figure 1 shown, this embodiment provides a method for predicting the acidity and alkalinity in the water quality for sea cucumber growth. First, the historical water quality data is decomposed into three variational mode components IMF, namely the trend component, the periodic component and the noise component. Secondly, the modal components are subjected to feature fusion processing. Then, the processed time series data is analyzed by multi-modal regression modeling through a recurrent neural network LSTM. Finally, the predicted values of each modal component are spliced and reconstructed, and at the same time, the residual value is input into the model for cyclic prediction, and finally the point prediction of the sea cucumber growth water quality index data is completed.
[0070] Using the intelligent aquaculture monitoring platform built by the research team, the water quality data of a sea cucumber farming circle in Jinzhou District, Dalian City was collected. The collection period was 30 minutes, and the collected data included twelve-dimensional related data such as water temperature, salinity, PH, air temperature, wind speed, air pressure, rainfall, air humidity and dissolved oxygen. The experimental data selected the data from June 22, 2021 to July 5, 2021. This time period belongs to the period of strong sea cucumber growth, and its data has a certain representativeness, and the data length is 645. Before model training, the entire data was divided into a training set, a test set and a validation set, and the splitting ratio was 7:2:1. Finally, the validation set and three evaluation functions (mae, mape, rmse) were used to determine the performance. This method has certain reference and practicality for short-term real-time water quality prediction.
[0071] First, the data was cleaned and analyzed, and it was found that there were abnormal missing and data redundancy in the data. In response to this phenomenon, redundant data in various water quality data was removed, and data repair work was carried out using the linear interpolation method, as shown in the following table.
[0072] Table 1 pH data repair values
[0073]
[0074]
[0075] After repairing the data, the Spearman rank correlation analysis method is used to screen out the water quality factors and meteorological factors that have a greater impact on pH. The analysis results are as Figure 2 shown. According to the water quality correlation evaluation standard, if the correlation |r| is less than 0.2, it is considered that the factor is not related to pH, otherwise they are related to each other. It can be observed from the figure that air humidity, indoor humidity, dew point, and total rainfall are negatively correlated with pH, and their |r| is greater than 0.2. It can be analyzed that as the air humidity and rainfall increase, the pH will gradually decrease. Relative air pressure, dissolved oxygen, and salinity are positively correlated with pH, and their |r| is greater than 0.2, which can prove that as the air pressure, dissolved oxygen, and salinity increase, the pH in the water quality for sea cucumber growth will gradually increase. Through analysis, seven-dimensional data of air and indoor humidity, relative air pressure, dew point, total rainfall, dissolved oxygen, and salinity can be screened out.
[0076] After performing the correlation analysis, pH is decomposed by variational mode decomposition (VMD) to form IMF components and central frequencies, as Figure 3 shown. The IMF components are reconstructed and differenced from the original data to obtain a residue R. To prevent overfitting in data prediction, Gaussian white noise r is added to the IMF components as shown in the following table. Then, it is fused with the water quality factors (salt content, dissolved oxygen) and meteorological factors (air and indoor humidity, relative air pressure, dew point, total rainfall) that have a greater correlation with pH to form extended mode component data new-IMFs.
[0077] Table 2 Gaussian white noise r added corresponding to IMF components
[0078]
[0079] Furthermore, preprocessing operations are performed on the extended mode component data new-IMFs, including normalization and data segmentation. The sliding window mechanism is used to form T: 7×3×431 and input it into the LSTM network for training. 7×3×431 means that the model uses 7 water quality factors, with a step size of 3, and inputs m = 431 samples for training. Each sample contains an array with a shape of (3,7).
[0080]
[0081] Among them, x in T nRefers to seven-dimensional data of air and indoor humidity, relative air pressure, dew point, total rainfall, dissolved oxygen, and salinity, i.e., n = 7. The final predicted result of pH is:
[0082]
[0083] IMF preyk_n is the predicted value at the nth moment in the future for the kth component, and IMF yk is the column of the original values of each component.
[0084] The prediction model constructed by the present invention includes an input layer, two LSTM layers, a Dropout layer added separately to prevent overfitting, a Flatten layer, a Dense layer, and an output layer. The number of neuron nodes in the two LSTM layers is set to 34 and 30 respectively. ADMA is used as the optimizer, and the ratio in the Dropout layer is set to 0.4. To prove that this model can reduce the prediction error of pH, a single LSTM network is used for effect comparison. The single LSTM network only includes an input layer, an LSTM layer, and an output layer, and the number of neural nodes in its LSTM layer is set to 40. Figure 4 Represents the prediction effects of the single LSTM network and the VLSTM model on pH. After 200 epochs of 20 - time fitting training, the optimal fitting effect diagram is manually selected; it can be found from the left figure that the prediction effect of the VLSTM model has been greatly improved, and it is found from the right figure that the three performances of the model of the present invention in the prediction of pH are reduced by 3, 3.9, and 0.4 percentage points respectively.
[0085] Figure 5 Includes the training loss diagrams of using the single LSTM and VLSTM models. It can be seen from the figure that before the 5th round, the loss of the VLSTM model has started to converge and has small fluctuations, and the single LSTM model starts to converge between the 7th and 10th rounds, with relatively larger fluctuations. Thus, it can be seen that the prediction model of the present invention has a faster convergence speed and a more robust model than the previous two prediction models, and the loss of the validation set is more stable than the performance of the LSTM model.
[0086] Finally, after inverse scaling the predicted results, the formula operation (E = Y - pre_Y) is performed using the predicted results and the original data to obtain the final error curve. Figure 6 It can be seen from [the figure] that the error values of the actual data are mostly between [-0.015, 0.015], and a small part of the errors are outside this interval.
[0087] The foregoing description of specific exemplary embodiments of the invention is for purposes of illustration and exemplification. These descriptions are not intended to limit the invention to the precise forms disclosed, and it is apparent that, according to the above teachings, many modifications and variations are possible. The purpose of selecting and describing the exemplary embodiments is to explain the specific principles of the invention and its practical applications, so that those skilled in the art can implement and utilize the various different exemplary embodiments of the invention, as well as various different selections and modifications. The scope of the invention is intended to be defined by the claims and their equivalents.
Claims
1. A method for predicting the acidity and alkalinity in the water quality for sea cucumber growth, characterized in that, Including: Clean and repair various collected water quality data; Screen out water quality factors and meteorological factors related to pH; Perform multi-modal decomposition on the repaired water quality factor data, decomposing it into three variational mode components: trend component, periodic component, and noise component, to obtain each modal component IMF; Perform difference verification on the each modal component IMF and the original data to obtain residual verification data R; Add Gaussian white noise to the trend component and periodic component in the variational mode component, and then fuse each modal component IMF with water quality factors and meteorological factors to form extended modal component data new-IMFs; Perform normalization processing on the extended modal component data new-IMFs; Segment the normalized extended modal component data to obtain a training set, a test set, and a validation set; Convert the format of the segmented extended modal component data, and then input it into a long short-term memory (LSTM) neural network for fitting training and learning; And evaluate the performance of the validation set and test set data of each extended modal component using mean absolute error mae, mean absolute percentage error mape, and root mean square error rmse; Stitch and reconstruct the predicted values of each modal component, and at the same time input the residual values into the model for cyclic prediction; Perform reconstruction inverse normalization on the prediction results, automatically restore the original data column, and obtain the error of the prediction data again.
2. The prediction method for the pH value in the water quality for sea cucumber growth according to claim 1, characterized in that, Clean and repair various collected water quality data, specifically: Eliminate redundant data in various water quality data; Perform interpolation operations on the various water quality data after elimination using a linear interpolation algorithm, and the formula used in the algorithm is: ; where f(x) is the value to be filled, and x t is the target factor value at a certain moment, and x t-1 , x t+1 are the target factor values at the previous moment and the next moment respectively.
3. The prediction method for the pH value in the water quality for sea cucumber growth according to claim 1, wherein Screen out water quality factors and meteorological factors related to pH, specifically: Use the Spearman rank method to perform correlation analysis on the repaired data, and screen out water quality factors and meteorological factors that have a strong correlation with the dissolved oxygen factor.
4. The prediction method for the pH value in the water quality for sea cucumber growth according to claim 1, wherein Perform multi-modal decomposition on the repaired water quality factor data, specifically: First, obtain the single-sided spectrum signal of the repaired water quality factor data through Hilbert transform, then shift the single-sided spectrum signal to the baseband, and then perform bandwidth estimation.
5. The prediction method for the acidity and alkalinity in the water quality for the growth of sea cucumbers according to claim 1 or 4, characterized in that, Obtain each modal component IMF, specifically: Initialize parameters: the number of components k = 5, the double rise time step tau = 0, and the central constraint strength alpha = 7000; Update all component lists \(u\) in the decomposition according to the following formula k and the center frequency value \(\omega\) k as follows; ; In the formula, f refers to the Fourier transform formula; According to the given evaluation criteria, continuously iterate until iterating K times and the filter iteration accuracy is less than ε, then stop iterating; where the evaluation criteria are:
6. The prediction method for the pH value in the water quality for sea cucumber growth according to claim 1, wherein, Perform difference verification on the each modal component IMF and the original data to obtain residual verification data R, specifically: ; Perform multi-modal decomposition on the target factor Y to form k IMF components, the shape of Y is converted from (n,1) to (n,k), and at the same time a difference component R is generated; the data format conversion is as follows:
7. The prediction method for the pH value in the water quality for sea cucumber growth according to claim 1, characterized in that, Add Gaussian white noise to the trend component and periodic component in the variational mode component, and then fuse each modal component IMF with water quality factors and meteorological factors to form extended modal component data new-IMFs, specifically: ; where \(u\) i is the formed single modal component, and \(\zeta\) i is the two-dimensional matrix of water quality factors, \(s\) 2 is the calculated variance of the modal component, \(N\) is the length of the data column of the modal component, and randn represents random generation; Convert the matrix D into a three-dimensional matrix T with a shape of (12, q, m), and its data format is as follows: ; In the formula, m is the number of samples and q is the sample step size.
8. The prediction method for the pH value in the water quality for sea cucumber growth according to claim 1, wherein Perform normalization processing on the extended modal component data new-IMFs, specifically: ; Convert water quality data with different dimensions into dimensionless data between [0, 1], where x t is the measured value at time t, min[x] is the minimum value in a certain measurement data series, and max[x] is the maximum value in a certain measurement data series.
9. The prediction method for the acidity and alkalinity in the water quality for the growth of sea cucumbers according to claim 1, characterized in that, The evaluation functions of the mean absolute error MAE, the mean absolute percentage error MAPE, and the root mean square error RMSE are respectively: ; Among them, y i is the true value of the target factor, is the predicted value of the target factor. The performance of the constructed model is evaluated through different evaluation functions.
10. The prediction method for the acidity and alkalinity in the water quality for the growth of sea cucumbers according to claim 1, wherein The prediction result pre_Y is obtained using the following formula: ; IMF preyk_n is the predicted value of the k-th component at the n-th moment, and IMF yk is the column of the original values of each component.
Citation Information
Patent Citations
Sea cucumber culture water temperature prediction method based on multi-dimensional data
CN114662790A
Wind power prediction method, system and device based on VMD and LSTM fusion model and medium
CN116127833A