Lake permanganate index inversion method based on mixed density network
Through a hybrid density network-based method, combined with deep learning and mixed probability model, the existing permanganate index inversion method has been solved, and higher inversion accuracy and stronger model generalization capabilities have been achieved.
Patent Information
- Application Number
- CN202510141313.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-06-17
AI Technical Summary
The existing permanganate index inversion methods have problems such as insufficient simulation accuracy and limited multimodal data processing capabilities, which are difficult to meet the needs of rapid monitoring of water quality and large-area assessment.
Using a lake permanganate index inversion method based on a hybrid density network, combining the feature extraction ability of deep learning technology and the advantages of mixed probability models in multimodal data processing, a convolutional neural network model is constructed and the output layer is replaced with a mixed probability model, and the optimal hyperparameters are found through negative log-likelihood loss function and Bayesian optimization.
It significantly improves the inversion accuracy and generalization ability of the model, enhances the processing ability of data heterogeneity under complex water conditions, and can more accurately invert the lake's permanganate index.
Smart Images

Figure CN120164091A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of quantitative remote sensing inversion, and particularly relates to a method for inverting the permanganate index of lakes based on a mixture density network. Background Technique
[0002] The permanganate index COD Mn is an important indicator for evaluating the degree of organic pollution in water bodies, and is widely used in lake water quality monitoring, pollutant discharge monitoring and water ecological environment management. COD Mn refers to the amount of oxygen consumed when the organic matter in a given volume of water is oxidized by permanganate to CO2 and H2O. Since the determination method of the permanganate index is relatively complex and usually relies on laboratory wet chemical analysis methods, although the traditional methods have high accuracy, they have the problems of long monitoring period and high cost, and it is difficult to meet the needs of rapid water quality monitoring and large-area assessment.
[0003] Water quality parameter inversion is a method for obtaining the concentration of water quality parameters based on remote sensing reflectance, which has the advantages of high monitoring efficiency, wide range and low cost. In the optical remote sensing band range, the water body reflectance spectrum is significantly affected by the absorption and scattering of organic substances in the water, and there is a certain correlation between the remote sensing reflectance at different wavelengths and the permanganate index. Existing inversion methods usually rely on linear regression or empirical models, and have limited ability to handle non-linear relationships and complex characteristics, and the inversion accuracy and generalization ability of the models are insufficient.
[0004] In order to further improve the inversion accuracy and generalization ability, deep learning technology has been introduced into the research of water quality parameter inversion in recent years. The deep learning model can extract multi-scale spectral features and has significant advantages in dealing with non-linear relationships. However, when dealing with small and medium-sized sample sets, the overly complex structure of the deep learning model may cause model degradation problems such as overfitting, and it is difficult to handle specific problems existing in water quality parameter inversion, such as the same spectrum may correspond to multiple water quality parameter composition ratios.
[0005] Based on this, the present invention proposes a method for inverting the permanganate index of lakes based on a mixture density network, so as to make full use of the feature extraction ability of deep learning technology and the advantages of the mixture probability model in multi-modal data processing, which can not only effectively improve the inversion accuracy, but also enhance the model's ability to simulate data heterogeneity under complex water body conditions, and has important application value. Summary of the Invention
[0006] Aiming at the problems of insufficient simulation accuracy of existing inversion methods and the processing of multi-modal data, the present invention combines the feature extraction ability of deep learning technology and the advantages of the mixture probability model in multi-modal data processing, and proposes a method for inverting the permanganate index of lakes based on a mixture density network. For the target lake, the following steps are performed:
[0007] (1) Data collection and processing
[0008] Obtain the images of the study area where the transit date coincides with the measured water quality data, and preprocess the images, including splicing, resampling, and atmospheric correction, to obtain the surface reflectance images.
[0009] Extract the values of each band of the target pixels according to the longitude and latitude of the sampling points, use the ratio of two bands and the normalized value for feature enhancement, add the longitude, latitude, water temperature, and PH corresponding to the point and time for feature expansion, form a feature vector, and use the permanganate index value as the label to form a water body apparent reflectance dataset. Randomly divide the dataset into 5 equal parts for 5-fold cross-validation.
[0010] (2) Construction of convolutional neural network model
[0011] Define the input feature tensor and the label tensor, and use the Min-Max normalization method to scale each feature value to [0,1] to unify all feature dimensions and facilitate the comparison of weights between features.
[0012] The construction of the convolutional neural network includes defining a one-dimensional convolutional layer, a residual block, a one-dimensional average pooling layer, an activation function, a fully connected layer, and an output layer. Among them, the one-dimensional convolutional layer is set with 2 layers of 1*3 convolution. The output of the one-dimensional convolutional layer is stacked with two residual modules after average pooling to deepen the network depth. The construction of the residual block includes defining a skip connection, a one-dimensional convolutional layer, a batch normalization layer, and an activation function. The outputs of the convolutional layer and the residual module both need to go through the one-dimensional average pooling layer for feature simplification to reduce redundant parameters. The model has 4 fully connected layers. Before inputting into the fully connected layer, the dimension of the feature tensor needs to be transformed, and then the total number of neurons contained in the existing feature tensor is calculated. The number of neurons gradually increases in the fully connected layer, and a non-linear transformation value is performed through the activation function. A ReLU activation function is added after each convolutional layer and fully connected layer.
[0013] (3) Construction and optimization of the mixture probability model
[0014] (3.1) Replace the output layer with the mixture probability model
[0015] Replace part of the output layer with the mixture probability model, that is, replace the output layer with three linear layers, each of which outputs K neurons (K is the number of Gaussian distributions). The outputs of the three linear layers represent the mean, standard deviation, and weight of each Gaussian distribution respectively. The standard deviation ensures non-negative attributes through nested exponential operations, and the weight value ensures its non-negative and sum-to-1 attributes through the Softmax function operation.
[0016] (3.2) Define the negative log-likelihood loss function
[0017] Construct the negative log-likelihood loss function, and its expression is as follows:
[0018]
[0019] where: loss is the value of the loss function; K is the number of Gaussian distributions; α k is the weight of each Gaussian component; μ k is the mean of each Gaussian component; σ k is the standard deviation of each Gaussian component; y is the measured value of permanganate index in the training set part; N(y; μ k is the probability density of y on the k-th Gaussian component;
[0020] In each iteration, after obtaining the loss value of that iteration, the Gaussian distribution parameters are updated using the gradient descent method, and the optimizer is the Adam optimizer. The goal is to minimize the loss value, that is, to maximize the probability density value of y on the Gaussian mixture distribution.
[0021] (3.3) Define the sampling function of the high-score mixture distribution
[0022] The model output result is the Gaussian mixture distribution of the permanganate index prediction value. It is necessary to define a sampling function for component selection and distribution sampling. Select a component according to the weight of the mixture distribution, extract the corresponding mean and standard deviation according to the component index, define a Gaussian distribution and sample it to obtain the simulated value of the permanganate index.
[0023] (3.4) Model optimization
[0024] Given the types and value ranges of the hyperparameters to be optimized, including the convolutional kernel size, fully connected layer parameters, learning rate, number of iterations, and batch size, use Bayesian optimization to find the optimal combination of hyperparameters; use the 5-fold cross-validation method for training. Each time, select 1 subset as the validation set, and the remaining 4 subsets as the training set. Repeat the above process until each subset has been used as a validation set once. Finally, calculate the average value of the 5 results as the evaluation index of the model performance, and save the model parameters after training.
[0025] (4) Large-area inversion of permanganate index
[0026] Extract the water body boundary of the target lake image, read and process the pixel values of each band pixel by pixel to form a target vector, input the vector into the trained model, and process the simulated values of each pixel point into an image with geographical coordinates. Description of the drawings
[0027] Figure 1 is the flow chart of the method for inverting the permanganate index of lakes based on the mixture density network according to the embodiment of the present invention;
[0028] Figure 2The structural diagram of a convolutional neural network model combining a residual module and a mixture probability model provided according to an embodiment of the present invention;
[0029] Figure 3 The comparison chart of the simulated value and the true value of the permanganate index provided according to an example of the present invention;
[0030] Figure 4 The inversion result chart of the permanganate index of Taihu Lake provided according to an example of the present invention. Detailed implementation manners
[0031] The present invention will be further described in detail below with reference to the accompanying drawings:
[0032] As Figure 1 shown in the process, a method for inverting the permanganate index of lakes based on a mixture density network includes the following steps:
[0033] (1) Data collection and processing
[0034] Obtain the image of the study area where the transit date coincides with the measured water quality data, and preprocess the image, including mosaicking, resampling and atmospheric correction, to obtain the surface reflectance image.
[0035] Extract the values of each band of the target pixel according to the longitude and latitude of the sampling point, use the ratio of two bands and the normalized value for feature enhancement, add the longitude, latitude, water temperature, and PH of the corresponding point and time for feature expansion to form a feature vector, and use the permanganate index value as a label to form a water body apparent reflectance dataset. Randomly divide the dataset into 5 equal parts for 5-fold cross-validation.
[0036] (2) Construction of the convolutional neural network model
[0037] Define the input feature tensor and the label tensor, and use the Min-Max normalization method to scale each feature value to [0,1] to unify all feature dimensions and facilitate the comparison of weights between features.
[0038] As Figure 2As shown, the construction of the convolutional neural network includes defining a one-dimensional convolutional layer, a residual block, a one-dimensional average pooling layer, an activation function, a fully connected layer, and an output layer. Among them, the one-dimensional convolutional layer is set with 2 layers of 1*3 convolutions, the number of output channels is 16 / 32, the moving step is 1, the edge padding is 1, the padding value is 0, and the length N of the input and output feature vectors remains unchanged. The output of the one-dimensional convolutional layer undergoes average pooling, and then two residual modules are stacked to deepen the network depth. The residual block includes defining a skip connection, a one-dimensional convolutional layer, a batch normalization layer, and an activation function; among them, the skip connection means that through a cross-layer data path, the defined convolutional layer and normalization layer are skipped, and the input is directly added before the activation function; the one-dimensional convolutional layer has 3 layers for local feature extraction and increasing the feature dimension, where the first two layers are 1*3 convolutions, the number of output channels is the number of input channels * 2, the moving step is 1, the edge padding is 1, the padding value is 0, and the length N of the input and output feature vectors remains unchanged. The last layer is a 1*1 convolution for adjusting the number of input channels of the skip connection and adjusting the shape of the input tensor to the output shape after the first two layers of convolutions so that addition operations can be performed.
[0039] The outputs of both the convolutional layer and the residual module need to go through a one-dimensional average pooling layer for feature simplification to reduce redundant parameters. The sliding window and moving step of the pooling layer are both 2, and the length of the output feature vector is half of the input length. If the input length is odd, it is rounded up. The model has 4 fully connected layers. Before the input fully connected layer, the feature tensor needs to be dimensionally transformed, and then the total number of neurons contained in the existing feature tensor is calculated. The number of neurons is gradually increased in the fully connected layer, and non-linear transformation values are obtained through the activation function. A ReLU activation function is added after each convolutional layer and fully connected layer.
[0040] (3) Construction and optimization of the mixture probability model (3.1) Replace the output layer with the mixture probability model
[0041] Part of the output layer is replaced with a mixture probability model, that is, the output layer is replaced with three linear layers, each of which outputs K neurons (K is the number of Gaussian distributions). The outputs of the three linear layers represent the mean, standard deviation, and weight of each Gaussian distribution respectively. The standard deviation ensures non-negative attributes through nested exponential operations, and the weight values ensure their non-negative and sum-to-1 attributes through Softmax function operations.
[0042] (3.2) Define the negative log-likelihood loss function
[0043] Construct the negative log-likelihood loss function, and its expression is as follows:
[0044]
[0045] In the formula: loss is the value of the loss function; K is the number of Gaussian distributions; α k is the weight of each Gaussian component; μk is the mean of each Gaussian component; σ k is the standard deviation of each Gaussian component; y is the measured value of permanganate index in the training set part; N(y; μ k is the probability density of y on the k-th Gaussian component;
[0046] In each iteration, after obtaining the loss value of that iteration, the Gaussian distribution parameters are updated using the gradient descent method. The optimizer is the Adam optimizer, and the goal is to minimize the loss value, that is, to maximize the probability density value of y on the Gaussian mixture distribution.
[0047] (3.3) Define the sampling function for the high-score mixture distribution
[0048] The model output result is the Gaussian mixture distribution of the predicted value of permanganate index. It is necessary to define a sampling function for component selection and distribution sampling. Select a component according to the weight of the mixture distribution, extract the corresponding mean and standard deviation according to the component index, define a Gaussian distribution and sample it to obtain the simulated value of permanganate index. The comparison chart between the simulated value and the true value of permanganate index is as Figure 3 shown.
[0049] (3.4) Model optimization
[0050] Given the types and numerical ranges of hyperparameters to be optimized, including the convolution kernel size, fully connected layer parameters, learning rate, number of iterations, and batch size, use Bayesian optimization to find the optimal combination of hyperparameters; use the 5-fold cross-validation method for training. Each time, select 1 subset as the validation set, and the remaining 4 subsets as the training set. Repeat the above process until each subset has been used as a validation set once. Finally, calculate the average value of the 5 results as the evaluation index of the model performance, and save the model parameters after training.
[0051] (4) Large-area inversion of permanganate index
[0052] Extract the water body boundary of the target lake image, read and process the pixel values of each band pixel by pixel to form a target vector, input the vector into the saved trained model, and make the predicted simulated value of each pixel point an image with geographical coordinates. The image is as Figure 4 shown.
Claims
1. A lake permanganate index inversion method based on a mixed density network, characterized in that: The following steps are involved: S1: Construct a water body apparent reflectance dataset based on the sampling time and geographic coordinates of the measured water quality parameters; S2: Construct a convolutional neural network containing a residual module, replace part of the output layer with a mixed probability model, design a negative log-likelihood loss function, define a sampling function, and output the permanganate index parameter value; S3: Determine the optimal hyperparameters of the network through Bayesian optimization. Based on the optimal hyperparameters, use the 5-fold cross-validation method to train and validate the model and save the training model. S4: Extract the water boundary of the target lake image, read and process the pixel values of each band pixel by pixel to form a target vector, input the vector into the trained mixed density network to obtain the simulated value of each pixel point, and process it into an image with geographic coordinates.
2. A lake permanganate index inversion method based on a mixed density network according to claim 1, characterized in that: The specific steps of S1 are as follows: S1.1: Obtain images of the study area whose transit dates coincide with measured water quality data, and preprocess the images, including stitching, resampling, and atmospheric correction, to obtain surface reflectance images; S1.2: Extract the band values of the target pixel according to the longitude and latitude of the sampling point, use the pairwise band ratio and normalized value for feature enhancement, add the longitude, latitude, water temperature, and pH of the corresponding point and time for feature expansion to form a feature vector, and use the permanganate index value as a label to form a water body apparent reflectance data set; S1.3: The dataset was randomly divided into 5 equal parts for 5-fold cross validation.
3. The lake permanganate index inversion method based on a mixed density network according to claim 1, characterized in that: The specific steps of S2 are as follows: S2.1: Define the input feature tensor and label tensor, and use the Min-Max normalization method to scale each feature value to [0,1]; S2.2: Construct a residual block, which includes a defined skip connection, a one-dimensional convolution layer, a batch normalization layer, and an activation function; wherein the skip connection refers to skipping the defined convolution layer and normalization layer through a cross-layer data path, and adding the input directly before the activation function; the one-dimensional convolution layer has 3 layers for local feature extraction and increasing feature dimensions, wherein the first two layers are 1*3 convolutions, the number of output channels is the number of input channels*2, the moving step is 1, the edge padding is 1, the padding value is 0, the input and output feature vector length N remains unchanged, and the last layer is a 1*1 convolution, which is used to adjust the number of input channels of the skip connection, and adjust the shape of the input tensor to the output shape after the first two layers of convolution, so that it can be added; S2.3: Construct a convolutional neural network, including defining a one-dimensional convolutional layer, a residual block, a one-dimensional average pooling layer, an activation function, a fully connected layer, and an output layer; The one-dimensional convolution layer sets 2 layers of 1*3 convolution, the number of output channels is 16 / 32, the moving step is 1, the edge filling is 1, the filling value is 0, and the input and output feature vector length N remains unchanged; the output of the one-dimensional convolution layer is average pooled, and then two residual modules are stacked to deepen the network depth; the outputs of the convolution layer and the residual module need to pass through the one-dimensional average pooling layer to simplify the features and reduce redundant parameters. The sliding window and moving step of the pooling layer are both 2, and the output feature vector length is half of the input length. If the input length is an odd number, it is rounded up; the model has 4 fully connected layers. Before entering the fully connected layer, the feature tensor needs to be transformed in dimension, and then the total number of neurons contained in the existing feature tensor is calculated. The number of neurons is gradually increased in the fully connected layer, and the nonlinear transformation value is performed through the activation function; ReLU activation function is added after each convolution layer and fully connected layer; S2.4: Part of the output layer is replaced with a mixed probability model, that is, the output layer is replaced with three linear layers, each of which outputs K neurons, where K is the number of Gaussian distributions. The outputs of the three linear layers represent the mean, standard deviation, and weight of each Gaussian distribution, respectively. The standard deviation is guaranteed to be non-negative through nested exponential operations, and the weight value is guaranteed to be non-negative and sum to 1 through the Softmax function operation. S2.5: Construct the negative log-likelihood loss function, which is expressed as follows: Where: loss is the loss function value; K is the number of Gaussian distributions; α k is the weight of each Gaussian component; μ k is the mean of each Gaussian component; σ k is the standard deviation of each Gaussian component; y is the measured value of the permanganate index in the training set; N(y; μ k , σ k ) is the probability density of y on the kth Gaussian component; In each round of iteration, after obtaining the loss value of that round, the Gaussian distribution parameters are updated using the gradient descent method, and the optimizer is the Adam optimizer; S2.6: The model output result is a Gaussian mixture distribution of the predicted value of the permanganate index. It is necessary to define a sampling function for component selection and distribution sampling, select a component according to the weight of the mixture distribution, extract the corresponding mean and standard deviation according to the component index, define a Gaussian distribution and sample, and obtain the simulated value of the permanganate index.
4. The lake permanganate index inversion method based on a mixed density network according to claim 1, characterized in that: The specific steps of S3 are as follows: Given the types and value ranges of hyperparameters that need to be optimized, including convolution kernel size, fully connected layer parameters, learning rate, number of iterations and batch size, Bayesian optimization is used to find the optimal hyperparameter combination; the 5-fold cross-validation method is used for training, and one subset is selected as the validation set each time, and the remaining 4 subsets are used as training sets. The above process is repeated until each subset is used as a validation set once, and finally the average of the 5 results is calculated as the evaluation indicator of model performance.
Citation Information
Patent Citations
Multi-topological-node hyperspectral water quality parameter joint inversion method and related equipment
CN116754499A
Method, recording medium and system for detecting concentration of chlorophyll in water body
CN118427669A
Cited By
River and lake index ammonia nitrogen monitoring method, device and system and storage medium
CN121933447A