Method for predicting dissolved oxygen based on improved hybrid neural network

By improving the hybrid neural network model, combining adaptive time convolution network, optimized Transformer encoder and enhanced GRU module, the existing dissolved oxygen prediction model has been solved in terms of accuracy and adaptability, and higher prediction accuracy and stability have been achieved.

CN119943211APending Publication Date: 2025-05-06HAINAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411900651.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing dissolved oxygen prediction models have insufficient accuracy in both short-term and long-term predictions, especially in dealing with complex multivariable environments and capturing long-term dependencies.

Method used

The improved hybrid neural network model is adopted, including adaptive time convolution network, optimized Transformer encoder, enhanced GRU module and linear regression error correction module. Through the combination and optimization of these modules, the feature extraction capability and prediction accuracy of the model are improved.

Benefits of technology

The accuracy is significantly improved in the short-term and long-term dissolved oxygen prediction tasks, enhanced the adaptability and stability of the model, and better handle complex environmental conditions and reduce prediction errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943211A_ABST
    Figure CN119943211A_ABST
Patent Text Reader

Abstract

The invention discloses a dissolved oxygen prediction method based on an improved hybrid neural network. The method is specifically implemented according to the following steps: step 1, collecting and arranging a dissolved oxygen related data set; step 2, constructing a hybrid neural network model comprising an adaptive time convolution network, an optimized Transform encoder, an enhanced GRU module and a linear regression error correction module; 3, training a hybrid neural network model and optimizing parameters; 4, verifying the hybrid neural network model and hyper-parameter tuning; 5, testing the hybrid neural network model and performing performance evaluation; and step 6, predicting dissolved oxygen by using the mixed neural network model. According to the dissolved oxygen prediction method based on the improved hybrid neural network, the problem that in the prior art, a dissolved oxygen prediction model is insufficient in precision in short-term and long-term prediction is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of big data and environmental information processing, and in particular relates to a method for predicting dissolved oxygen based on an improved hybrid neural network. Background Art

[0002] Dissolved oxygen is a critical parameter in the grouper aquaculture environment, which directly affects the growth and health of fish. Sufficient dissolved oxygen levels are the basis for normal metabolism and efficient aquaculture of fish. Fluctuations in dissolved oxygen may cause stress responses or even death in fish. At present, traditional dissolved oxygen prediction methods mostly rely on physical and chemical models, but these methods are complex in calculation and have poor real-time performance. They are difficult to cope with the rapid changes in the aquaculture environment, especially in a complex environment where multiple factors such as temperature, salinity, and light act together. The limitations of traditional methods are particularly obvious.

[0003] With the continuous advancement of technology, neural networks have gradually been applied to water quality prediction and have shown obvious advantages. However, a single neural network still faces the challenges of insufficient accuracy and adaptability when dealing with dissolved oxygen prediction in complex multivariate environments, especially in capturing long-term and short-term dependencies, multi-scale features, and local features. In addition, traditional neural networks perform poorly with non-stationarity and sudden changes and require a large amount of training data. The main problems include low prediction accuracy. Existing dissolved oxygen prediction models often have large errors and insufficient accuracy in short-term and long-term predictions. Secondly, the model has poor adaptability to complex data: the current dissolved oxygen prediction model does not take into account the complex impact of various environmental factors, resulting in unstable performance in practical applications. Summary of the invention

[0004] The purpose of the present invention is to provide a prediction method for dissolved oxygen based on an improved hybrid neural network, which solves the problem of insufficient accuracy of the dissolved oxygen prediction model in the prior art in short-term and long-term predictions.

[0005] The technical solution adopted by the present invention is to improve the prediction method of dissolved oxygen based on the hybrid neural network, which is specifically implemented in the following steps: Step 1, collect and organize dissolved oxygen related data sets; Step 2, construct a hybrid neural network model including an adaptive temporal convolutional network, an optimized Transformer encoder, an enhanced GRU module, and a linear regression error correction module; Step 3: training the hybrid neural network model and parameter optimization; Step 4: Verify the hybrid neural network model and hyperparameter tuning; Step 5, testing the hybrid neural network model and performance evaluation; Step 6: Predict dissolved oxygen using the hybrid neural network model.

[0006] The technical solution of the present invention is also characterized in that: Step 1 is as follows: Collect data sets, detect and remove outliers in the data sets, use mean filtering and median filtering to smooth the data to remove noise in the data, and set a threshold for each feature in the data set to remove outliers. The specific formula for the threshold is as follows: (1) In the formula, and are the lower and upper limits of the feature, respectively. represents the i-th data point; For data with missing values, linear interpolation is used to fill them. The calculation formula of linear interpolation is: (2) In the formula, and are the time of the two non-missing data points before and after, and are the corresponding values ​​respectively, and t represents the time point to be supplemented.

[0007] The data sets include dissolved oxygen, temperature, salinity, redox potential, pH, and electrical conductivity.

[0008] Step 2 is as follows: Step 2.1, construct an adaptive temporal convolutional network; Step 2.2, build an optimized Transformer model; Step 2.3, optimize the gated recurrent unit model.

[0009] Step 2.1 is as follows: Adaptive temporal convolutional network introduces adaptive weight vector , dynamically adjust the weight of the convolution kernel. For the case where the convolution kernel size is 𝑘, the learned weight vector Transformed into probability distribution through softmax function , specifically expressed as follows: (3) The adaptive convolution kernel W can be expressed as: (4) In the formula, Represents the weight of the convolution kernel at different positions, represents the corresponding probability distribution; When the dimensions of the input and output do not match, a 1x1 convolutional layer is used for downsampling; when there is inconsistency in the length of the time series, padding is used to adjust the length to ensure a smooth transition of the residual connection. The specific implementation is as follows: (5) (6) In the formula, and Represent the number of input and output channels respectively, and Represents the length of the output and residual branches.

[0010] Step 2.2 is as follows: The core of the encoder consists of a multi-head attention mechanism and a feedforward neural network. The feature extraction process can be summarized as follows: For the input sequence X, the feature extraction process is as follows: (7) Where Q, K, and V are the query, key, and value matrices obtained from the input X through linear transformation, respectively, which are used to calculate the correlation within the sequence in the self-attention mechanism. The multi-head attention mechanism (Q, K, V) extracts global features by capturing the relationship between different parts of the sequence. Subsequently, the feedforward neural network further performs nonlinear transformation and information compression, combined with residual connections and layer normalization to stabilize the training process and retain input information. In order to improve the performance of the Transformer model in processing complex time series data, optimization is performed. The first optimization is to use a multi-head attention mechanism. The query, key, and value are projected into the same embedding dimension space through linear transformation, and their dot product is calculated as the attention score. The score is normalized by the softmax function, and then the value vector is weighted to obtain the final attention output. The process can be expressed by the following formula: (8) Where Q, K, and V represent query, key, and value matrices, respectively. Represents the dimension of the key; The second optimization is to introduce relative position encoding, using sine and cosine functions to encode the position information and add it to the input sequence to retain the time information of the sequence. The process can be expressed by the following formula: (9) (10) In the formula, pos represents the position, i represents the dimension index, and d represents the embedding dimension.

[0011] Step 2.3 is as follows: The traditional Transformer decoder is replaced by an improved gated recurrent unit model. The key improvements include introducing adjustable gating factors and optimizing the internal structure and weight initialization of the GRU unit to capture dependencies. First, the adjustable gating factor is optimized, and the adjustable gating factors α and β are introduced into the reset gate and update gate. The specific formula is as follows: (11) (12) In the formula, and are the outputs of the reset gate and the update gate, respectively. and is the corresponding weight matrix, is the hidden state at the previous time step, is the current input. Secondly, the internal structure of the gated recurrent unit model is redesigned. The improved GRU unit resets the previous hidden state under the influence of the reset gate and calculates the new hidden state using the current input: (13) The final hidden state is updated using the following weighted formula: (14).

[0012] Step 3 is as follows: the data set is divided into training set, validation set and test set in a ratio of 8:1:1. The mean square error is used as the loss function during training, and the Adam optimizer is used for parameter update. The initial learning rate is set to 0.0001, and the scheduler is used to gradually reduce the learning rate when the validation error no longer decreases. The batch size is set to 128 and the number of training rounds is 200.

[0013] Step 4 is as follows: After the hybrid neural network model is trained, an independent test set is used to evaluate its performance, and the accuracy and stability of the model are verified using multiple indicators such as mean absolute error, root mean square error, mean absolute percentage error, and determination coefficient.

[0014] Compared with the prior art, the present invention has the following beneficial effects: (1) The prediction method of dissolved oxygen based on the improved hybrid neural network provided by the present invention can show higher accuracy in both short-term and long-term dissolved oxygen prediction tasks by combining an adaptive convolutional network, an optimized Transformer encoder and an improved GRU module. The test results on multiple objective indicators (such as MAE, RMSE, MAPE) show that it is far superior to traditional prediction methods in terms of accuracy.

[0015] (2) The prediction method for dissolved oxygen based on the improved hybrid neural network provided by the present invention enhances the feature extraction capability of the model, enabling it to better handle complex and changeable environmental conditions, and improves the stability and robustness of dissolved oxygen prediction. The model can maintain stable prediction performance under different environments by effectively extracting multi-factor features.

[0016] (3) The dissolved oxygen prediction method based on the improved hybrid neural network provided by the present invention has broad application prospects and significant advantages in the dissolved oxygen prediction tasks in the fields of aquaculture and environmental monitoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is the overall framework diagram of the hybrid neural network model of the present invention; Figure 2 It is a schematic diagram of the process of the prediction method of dissolved oxygen based on the improved hybrid neural network of the present invention; Figure 3 This is a comparison chart of the predicted value and the actual value in the initial stage of Example 6 of the present invention; Figure 4 This is a comparison chart of the predicted value and the true value at the final stage of Example 6 of the present invention. DETAILED DESCRIPTION

[0018] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0019] Example 1 The present invention provides a prediction method for dissolved oxygen based on an improved hybrid neural network, such as Figure 1 and Figure 2 As shown, the specific implementation steps are as follows: Step 1, collect and organize dissolved oxygen related data sets; Step 2, construct a hybrid neural network model including an adaptive temporal convolutional network, an optimized Transformer encoder, an enhanced GRU module, and a linear regression error correction module; Step 3: training the hybrid neural network model and parameter optimization; Step 4: Verify the hybrid neural network model and hyperparameter tuning; Step 5, testing the hybrid neural network model and performance evaluation; Step 6: Predict dissolved oxygen using the hybrid neural network model.

[0020] like Figure 1As shown in the figure, the overall model architecture of the present invention is designed. The hybrid model first performs data preprocessing and feature extraction through an improved adaptive temporal convolutional network (ATCN). The ATCN module combines residual connections, adaptive downsampling, and dilated convolutions to effectively capture key temporal patterns, preserve temporal dependencies, and reduce data dimensions, providing high-quality feature inputs for subsequent model training. Subsequently, the Transformer encoder is optimized on the infrastructure and a linear attention mechanism is introduced, which significantly reduces computational complexity while enhancing parallel processing capabilities. In addition, relative position encoding further enhances the model's ability to capture long-term dependencies, thereby providing a richer representation of time series features.

[0021] Next, the enhanced GRU module replaces the traditional Transformer decoder to capture short-term and long-term temporal dependencies to improve the prediction performance. Finally, the linear regression model is used to correct the errors of the preliminary prediction results, thereby significantly improving the accuracy and stability of the overall prediction. In summary, the proposed algorithm model integrates multiple advanced technologies such as ATCN, optimized Transformer encoder, enhanced GRU module, and error correction mechanism to improve the accuracy of dissolved oxygen prediction.

[0022] Example 2 On the basis of Example 1, Figure 1 and Figure 2 As shown, step 1 specifically involves that the data set collected by individuals is collected from the grouper farming environment through on-site sensor equipment, covering multiple environmental parameters such as dissolved oxygen (DO), temperature (WT), salinity, oxidation-reduction potential (ORP), pH value, conductivity, etc. The data is collected at intervals of 2 minutes to capture the dynamic changes in the farming environment.

[0023] In order to ensure the quality of the data required for model training, data cleaning is required after the data set is collected. The goal of data cleaning is to remove noise and fill in missing values, thereby providing more accurate and consistent data input for the model. The specific cleaning steps are as follows: First, outliers are detected and removed from the data set. Mean and median filters are used to smooth the data to remove noise. For each feature, a reasonable threshold is set to remove outliers. The formula is as follows: (1); in, and are the lower and upper limits of the feature, respectively, obtained through empirical data or statistical analysis. represents the ith data point.

[0024] For data with missing values, linear interpolation is used to fill in the missing values ​​to ensure the continuity and integrity of the data. The calculation formula for linear interpolation is: (2); in, and are the time of the two non-missing data points before and after, and are the corresponding values ​​respectively, and t represents the time point to be supplemented.

[0025] Step 2, construct a hybrid neural network model including an adaptive temporal convolutional network, an optimized Transformer encoder, an enhanced GRU module, and a linear regression error correction module; After data preprocessing, a hybrid neural network model including an adaptive temporal convolutional network (ATCN), an optimized Transformer encoder, an enhanced GRU module, and a linear regression error correction module is constructed. Each module has been specifically improved on the basis of the traditional model to improve the accuracy of dissolved oxygen prediction and the adaptability of the model. Among them, the adaptive temporal convolutional network is used to improve the ability to extract initial features; the improved Transformer encoder is used to capture long-term and short-term dependency features in time series; the enhanced GRU module is used to improve the model's ability to represent time series information; finally, error correction is performed through the linear regression module to further improve the accuracy of the prediction results. The following is a detailed description of the improved design of each module.

[0026] Example 3 Based on Example 2, step 2 is specifically as follows: Step 2.1, construct an adaptive temporal convolutional network; Step 2.2, build an optimized Transformer model; Step 2.3, optimize the gated recurrent unit model.

[0027] Among them, step 2.1 is specifically as follows: Temporal Convolutional Network (TCN) is an architecture based on Convolutional Neural Network (CNN) specifically used to process time series data. This study proposes an improved version based on the classic TCN model, called Adaptive Temporal Convolutional Network (ATCN). This model aims to further enhance the ability of preliminary feature extraction by introducing adaptive convolution kernels and optimized residual connection mechanism.

[0028] Adaptive temporal convolutional network introduces adaptive weight vector , dynamically adjust the weight of the convolution kernel. For the case where the convolution kernel size is 𝑘, the learned weight vector Transformed into probability distribution through softmax function , defined as follows: (3); The adaptive convolution kernel W can be expressed as: (4); in, Represents the weight of the convolution kernel at different positions, Represents the corresponding probability distribution. This adaptive mechanism enables the model to more flexibly capture local and global features in time series data, thereby enhancing the effect of feature extraction.

[0029] Residual connection optimization. To facilitate effective information transfer in deep networks, the adaptive temporal convolutional network incorporates an optimized residual connection mechanism. When the dimensions of the input and output do not match, a 1x1 convolutional layer is used for downsampling; when there is an inconsistency in the length of the time series, padding is used to adjust the length to ensure a smooth transition of the residual connection. The specific implementation is as follows: (5); (6); in, and Represent the number of input and output channels respectively, and Represents the length of the output and residual branches. This optimization ensures the smooth transmission of information flow in the network, prevents gradient vanishing, and improves the training efficiency and generalization ability of deep models.

[0030] Step 2.2 is as follows: The Transformer model performs well in natural language processing (NLP) and various sequence data processing tasks due to its unique encoder structure. The core of the encoder consists of a multi-head attention mechanism and a feedforward neural network (FFN). The feature extraction process can be summarized as follows: For the input sequence X, the feature extraction process is as follows: (7); Here, Q, K, and V are the query, key, and value matrices obtained from the input X through linear transformation, respectively, and are used to calculate the correlation within the sequence in the self-attention mechanism. The multi-head attention mechanism (Q, K, V) extracts global features by capturing the relationship between different parts of the sequence. Subsequently, the feedforward neural network further performs nonlinear transformation and information compression, combined with residual connections and layer normalization to stabilize the training process and retain the input information.

[0031] To improve the performance of the Transformer model in processing complex time series data, we first performed two important optimizations on the encoder. The first optimization is to introduce a linear attention mechanism, which aims to reduce the computational complexity and enhance the parallel processing capability of the model.

[0032] In the specific implementation, a multi-head attention mechanism is used to project the query, key, and value into the same embedding dimension space through linear transformation, and then calculate their dot product as the attention score. The score is normalized by the softmax function, and then the value vector is weighted to obtain the final attention output. The process can be expressed by the following formula: (8); Among them, Q, K, and V represent query, key, and value matrices respectively. Indicates the dimension of the key.

[0033] The second optimization is to introduce relative position encoding to further enhance the model's performance in capturing long-distance dependencies and complex nonlinear features. We use sine and cosine functions to encode the position information and add it to the input sequence to preserve the temporal information of the sequence. The process can be expressed by the following formula: (9); (10); Among them, pos represents the position, i represents the dimension index, and d represents the embedding dimension.

[0034] Step 2.3 specifically replaces the traditional Transformer decoder with an improved gated recurrent unit model to improve the model's performance in processing complex time series data. Key improvements include the introduction of adjustable gating factors and the optimization of the internal structure and weight initialization of the GRU unit to better capture these dependencies.

[0035] First, we optimized the adjustable gating factors. To improve the flexibility of GRU, we introduced adjustable gating factors α and β in the reset gate and update gate. These factors control the weight of each gate, allowing the model to fine-tune the impact of different time steps. The formula is as follows: (11); (12); in, and are the outputs of the reset gate and the update gate, respectively. and is the corresponding weight matrix, is the hidden state at the previous time step, is the current input. By adjusting α and β, the model can better adapt to different types of time series data.

[0036] Next, the internal structure of the gated recurrent unit model was redesigned to more effectively integrate new inputs and previous hidden states. Specifically, the improved GRU unit resets the previous hidden state under the influence of the reset gate and calculates the new hidden state using the current input: (13); The final hidden state is updated using the following weighted formula: (14); This optimization allows to better capture long-term dependencies in the sequence, thereby improving the performance of the model on long sequence data.

[0037] Example 4 Based on Example 3, step 3 is specifically as follows: after the model is constructed, the model is trained using the preprocessed data.

[0038] Training set partitioning and optimization algorithm: The data set is divided into training set, validation set and test set in a ratio of 8:1:1. The mean square error (MSE) is used as the loss function during training, and the Adam optimizer is used for parameter update. The initial learning rate is set to 0.0001, and the scheduler is used to gradually reduce the learning rate when the validation error no longer decreases. The batch size is set to 128, and the number of training rounds is 200.

[0039] Hyperparameter tuning: Monitor model performance through validation sets, and perform hyperparameter tuning on the model based on validation results to ensure good performance on different data sets. This includes adjusting the learning rate, number of hidden units, etc. to improve the accuracy and generalization ability of the model.

[0040] After training is completed, the model performance is evaluated using the validation set to check the generalization ability and prediction accuracy of the model. According to the feedback from the validation set, the model's hyperparameters are tuned, including the adjustment of the learning rate and regularization parameters, as well as the optimization of the neural network structure (such as the number of layers and neurons). Use methods such as grid search or Bayesian optimization to find the optimal hyperparameter combination to improve the performance of the model. At the same time, the model structure is appropriately modified to enhance the feature extraction capability or simplify the model complexity to ensure good performance in different data sets and environments. In addition, by setting an early stopping mechanism, when the loss on the validation set is no longer continuously reduced, the training is stopped in time to prevent overfitting and save computing resources.

[0041] Example 5 Based on Example 4, step 5 is specifically as follows: after the model training is completed, an independent test set is used to evaluate its performance, and multiple indicators such as mean absolute error (MAE), root mean square error (RMSE), mean absolute percentage error (MAPE) and determination coefficient (R²) are used to verify the accuracy and stability of the model. Through these comprehensive evaluations, it is ensured that the model performs well in the task of dissolved oxygen prediction and has application value in actual aquaculture and environmental monitoring.

[0042] Example 6 Using the prediction method of dissolved oxygen based on the improved hybrid neural network provided in the above embodiment, as Figure 3 and Figure 4 As shown, the prediction results of different time periods are compared with the true values. In the experiment, two different typical prediction intervals are selected from the data set to compare the fitting of the predicted values ​​of the model of the present invention with the actual values. The results show that the model of the present invention performs well in trend tracking ability. Whether in areas of stable changes or violent fluctuations, the prediction curve can closely follow the changing trend of the actual value. In addition, at the turning point of the curve or the interval of violent fluctuations, the model of the present invention also shows good capture ability. The prediction curve has no obvious lag at the turning point, reducing the possibility of large deviations. In long-term series prediction, the prediction curve of the model of the present invention maintains a small deviation from the actual curve, showing its stability and reliability in long-term prediction tasks.

[0043] Comparative Example 1 The dataset contains the parameter features: Dissolved oxygen (DO): unit is mg / L; pH: No unit; Water temperature (Temperature): unit is °C; Oxidation reduction potential (ORP): unit is mV; Salinity: The unit is ppt; Conductivity: The unit is mS / cm; Time: the unit is minute; According to the division ratio of the data set (80% training set, 10% test set and 10% validation set), the feature data samples in the data set numbers 48330 to 53558 (validation set) are used as the input for model reasoning dissolved oxygen prediction. In the model evaluation stage, the feature data in the validation set is input by the trained model in a sliding window manner. The model performs reasoning based on these input features to generate predicted dissolved oxygen concentration results. Subsequently, the predicted dissolved oxygen results are compared with the corresponding true dissolved oxygen values ​​in the validation set, and the performance indicators under different step lengths are calculated, including the mean absolute error (MAE), root mean square error (RMSE), mean absolute percentage error (MAPE) and determination coefficient (R²). This step ensures an objective evaluation of the model prediction accuracy and stability. Table 1 shows the dissolved oxygen prediction results of Transformer, Autoformer, Longformer, Reformer and the model proposed in this application when the step length is set to 72, and gives the real results for comparison. These data can intuitively reflect the prediction performance of each model under different samples.

[0044] Table 1 Partial display of the specific results of each model predicting dissolved oxygen (the step size is set to 72)

[0045] Based on the prediction results and true values ​​of each model in Table 1, according to the calculation formulas of various evaluation indicators such as MAE, RMSE, MAPE and R², the objective index performance results of each model are shown in Table 2, and the prediction performance comparison results of each model are given. In general, the model proposed in the present invention shows significant accuracy advantages and stability in both short-term (step length = 18) and long-term (step length = 72) predictions, and is superior to other comparison models in various evaluation indicators such as MAE, RMSE, MAPE and R², especially in the processing ability of complex time series data. It shows outstanding performance. The model of the present invention can effectively capture the dynamic change characteristics of dissolved oxygen, reduce prediction errors, and significantly improve prediction accuracy and reliability, proving that it has significant technical advantages in time series prediction tasks, and is particularly suitable for practical application scenarios with high requirements for accuracy and stability, such as environmental monitoring and aquaculture.

[0046] Table 2 Comparison of model prediction performance

Claims

1. A method for predicting dissolved oxygen based on an improved hybrid neural network, characterized in that: Please follow the steps below to implement: Step 1, collect and organize dissolved oxygen related data sets; Step 2, construct a hybrid neural network model including an adaptive temporal convolutional network, an optimized Transformer encoder, an enhanced GRU module, and a linear regression error correction module; Step 3: training the hybrid neural network model and parameter optimization; Step 4: Verify the hybrid neural network model and hyperparameter tuning; Step 5, testing the hybrid neural network model and performance evaluation; Step 6: Predict dissolved oxygen using the hybrid neural network model.

2. The method for predicting dissolved oxygen based on an improved hybrid neural network according to claim 1, characterized in that: The step 1 is specifically as follows: Collect data sets, detect and remove outliers in the data sets, use mean filtering and median filtering to smooth the data to remove noise in the data, and set a threshold for each feature in the data set to remove outliers. The specific formula for the threshold is as follows: (1) In the formula, and are the lower and upper limits of the feature, respectively. represents the i-th data point; For data with missing values, linear interpolation is used to fill them. The calculation formula of linear interpolation is: (2) In the formula, and are the time of the two non-missing data points before and after, and are the corresponding values ​​respectively, and t represents the time point to be supplemented.

3. The method for predicting dissolved oxygen based on an improved hybrid neural network according to claim 2, characterized in that: The data set includes dissolved oxygen, temperature, salinity, redox potential, pH, and conductivity.

4. The method for predicting dissolved oxygen based on an improved hybrid neural network according to claim 1, characterized in that: The step 2 is specifically as follows: Step 2.1, construct an adaptive temporal convolutional network; Step 2.2, build an optimized Transformer model; Step 2.3, optimize the gated recurrent unit model.

5. The method for predicting dissolved oxygen based on improved hybrid neural network according to claim 4, characterized in that: The step 2.1 is specifically as follows: Adaptive temporal convolutional network introduces adaptive weight vector , dynamically adjust the weight of the convolution kernel. For the case where the convolution kernel size is 𝑘, the learned weight vector Transformed into probability distribution through softmax function , specifically expressed as follows: (3) The adaptive convolution kernel W can be expressed as: (4) In the formula, Represents the weight of the convolution kernel at different positions, represents the corresponding probability distribution; When the dimensions of the input and output do not match, a 1x1 convolutional layer is used for downsampling; when there is inconsistency in the length of the time series, padding is used to adjust the length to ensure a smooth transition of the residual connection. The specific implementation is as follows: (5) (6) In the formula, and Represent the number of input and output channels respectively, and Represents the length of the output and residual branches.

6. The method for predicting dissolved oxygen based on improved hybrid neural network according to claim 4, characterized in that: The step 2.2 is specifically as follows: The core of the encoder consists of a multi-head attention mechanism and a feedforward neural network. The feature extraction process is summarized as follows: For the input sequence X, the feature extraction process is as follows: (7) Where Q, K, and V are the query, key, and value matrices obtained from the input X through linear transformation, respectively, which are used to calculate the correlation within the sequence in the self-attention mechanism. The multi-head attention mechanism extracts global features by capturing the relationship between different parts of the sequence; Subsequently, the feedforward neural network further performs nonlinear transformation and information compression, combined with residual connections and layer normalization to stabilize the training process and retain input information; To improve the performance of the Transformer model in processing complex time series data, we need to optimize it. The first optimization is to use a multi-head attention mechanism. Through linear transformation, we project the query, key, and value into the same embedding dimension space, calculate their dot product, and use it as the attention score. The score is normalized by the softmax function, and then the value vector is weighted to get the final attention output. The process can be expressed by the following formula: (8) Where Q, K, and V represent query, key, and value matrices, respectively. Represents the dimension of the key; The second optimization is to introduce relative position encoding, using sine and cosine functions to encode the position information and add it to the input sequence to retain the time information of the sequence. The process can be expressed by the following formula: (9) (10) In the formula, pos represents the position, i represents the dimension index, and d represents the embedding dimension.

7. The method for predicting dissolved oxygen based on improved hybrid neural network according to claim 4, characterized in that: The step 2.3 is specifically as follows: The traditional Transformer decoder is replaced by an improved gated recurrent unit model. The gated recurrent unit model includes the introduction of an adjustable gating factor and the optimization of the internal structure and weight initialization of the GRU unit to capture dependencies. First, the adjustable gating factor is optimized, and the adjustable gating factors α and β are introduced into the reset gate and the update gate. The specific formula is as follows: (11) (12) In the formula, and are the outputs of the reset gate and the update gate, respectively. and is the corresponding weight matrix, is the hidden state at the previous time step, is the current input; Secondly, the internal structure of the gated recurrent unit model is redesigned. Specifically, the improved GRU unit resets the previous hidden state under the influence of the reset gate and calculates the new hidden state using the current input: (13) The final hidden state is updated using the following weighted formula: (14)。 8. The method for predicting dissolved oxygen based on improved hybrid neural network according to claim 1, characterized in that: The step 3 is specifically as follows: the data set is divided into a training set, a validation set and a test set in a ratio of 8:1:1, the mean square error is used as the loss function during training, the Adam optimizer is used for parameter update, the initial learning rate is set to 0.0001, and the scheduler is used to gradually reduce the learning rate when the validation error no longer decreases, the batch size is set to 128, and the number of training rounds is 200.

9. The method for predicting dissolved oxygen based on improved hybrid neural network according to claim 1, characterized in that: The step 4 is specifically as follows: after the hybrid neural network model is trained, an independent test set is used to perform performance evaluation, and the accuracy and stability of the model are verified using multiple indicators such as mean absolute error, root mean square error, mean absolute percentage error, and determination coefficient.

Citation Information

Cited By

  • Method for predicting marine carbon dioxide blowout plume parameters based on attention mechanism and support vector regression

    CN122548702A