Online monitoring method for emerging marine pollutants

Through intelligent sensor networks and deep learning technology, high-dimensional features of emerging marine pollutants are automatically extracted and classified in combination with SVM and decision trees, which solves the problems of inaccuracy and low efficiency of monitoring in the existing technology, and achieves efficient online monitoring of emerging marine pollutants.

CN120102469APending Publication Date: 2025-06-06WUHAN UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510189363.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently and accurately monitor emerging marine pollutants online, especially when facing high-dimensional and complex pollutant data. Feature extraction depends on expert experience and the classification algorithm performs poorly, resulting in low classification accuracy.

Method used

Intelligent sensor network is used for multi-dimensional data acquisition, combined with deep learning technology for signal preprocessing and feature extraction, multi-layer convolutional neural network CNN is used to automatically extract high-dimensional features, and pollutant identification and classification is performed through the combined model SVM+ decision tree. Finally, XGBoost and random forest are used for data fusion and anomaly detection.

Benefits of technology

It has achieved efficient and accurate online monitoring of emerging marine pollutants, improved the efficiency and accuracy of feature extraction and classification, and can effectively deal with complex and changeable marine environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120102469A_ABST
    Figure CN120102469A_ABST
Patent Text Reader

Abstract

The invention discloses an online monitoring method for emerging marine pollutants, and relates to the technical field of pollutant monitoring, a plurality of sensors are adopted to work at the same time, multi-dimensional data are collected in real time, and comprehensive pollutant information is provided; a multi-layer convolutional neural network CNN is utilized to perform automatic feature extraction, automatic learning and extraction of high-dimensional complex features in data, a combined model of a support vector machine SVM and a decision tree is combined, preliminary classification is performed by the SVM, refined classification is performed by the decision tree, and the classification accuracy is improved; a deep learning-based denoising automatic encoder DAE is adopted to carry out denoising processing on sensor data, the DAE effectively removes background noise in nonlinear and complex environments through a learning noise mode, in addition, a self-adaptive filter is used, filtering parameters are dynamically adjusted according to environmental changes, the denoising effect is further optimized, and the signal quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pollutant monitoring, and in particular to an online monitoring method for emerging marine pollutants. Background Art

[0002] In recent years, with the increasingly serious problem of environmental pollution, traditional pollutants such as sulfur dioxide, nitrogen oxides, PM2.5, etc. have received extensive attention and monitoring; however, with the advancement of industrialization and modernization, the emergence of emerging pollutants has gradually become an important environmental issue; these emerging pollutants include persistent organic pollutants, endocrine disruptors, antibiotics, etc., which have a wide range of sources, a wide variety, and complex chemical properties, posing a potential threat to the environment and human health.

[0003] The monitoring and control of emerging pollutants have become a hot area of ​​international research. Compared with traditional pollutants, emerging pollutants are persistent, concealed and cumulative, which poses new challenges to existing monitoring technologies. Traditional monitoring methods usually show problems such as insufficient sensitivity and limited processing capacity when dealing with these complex, organic and emerging pollutants, and cannot meet the needs of real-time, comprehensive and efficient monitoring. Therefore, it is particularly important to develop a method that can efficiently and accurately monitor emerging marine pollutants online. Emerging pollutants refer to pollutants that cannot be effectively treated by traditional pollution control measures, such as new organic compounds, pharmaceutical residues and microplastics. These pollutants have complex chemical properties, wide distribution and strong environmental persistence, which poses great challenges to traditional monitoring and control methods. Existing detection technologies often cannot meet the needs of high-sensitivity monitoring of low-concentration, multi-component pollutants.

[0004] Currently, traditional methods show obvious limitations when dealing with high-dimensional and complex pollutant data. The feature extraction in traditional methods relies on expert experience for manual operations, which is time-consuming and labor-intensive and difficult to capture the complex features in high-dimensional data. At the same time, the classification algorithms commonly used in traditional methods, such as linear regression or basic decision trees, perform poorly when faced with complex multidimensional features, resulting in low classification accuracy and difficulty in coping with the complex and changeable marine environment. Therefore, an online monitoring method for emerging marine pollutants is urgently needed to solve such problems. Summary of the invention

[0005] In view of the shortcomings of the prior art, the present invention provides an online monitoring method for emerging marine pollutants, which solves the problem that feature extraction in the prior art relies on manual operation based on expert experience, is time-consuming and labor-intensive, and is difficult to capture complex features in high-dimensional data. At the same time, classification algorithms commonly used in traditional methods, such as linear regression or basic decision trees, perform poorly when faced with complex multidimensional features, resulting in low classification accuracy and difficulty in coping with the complex and changeable marine environment.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions:

[0007] The present invention provides an online monitoring method for emerging marine pollutants, comprising:

[0008] Step 1. Multi-sensor data collection, deployment of intelligent sensor network, deployment of intelligent sensor network with self-organizing network and adaptive adjustment functions in the target sea area, multi-dimensional data collected by sensors in real time, including spectral signals, electrochemical signals and SPR signals;

[0009] Step 2: Signal preprocessing: Use a deep learning-based denoising autoencoder (DAE) to denoise the sensor data, remove background noise, and retain the characteristic signals of pollutants. The denoised high signal-to-noise ratio data is passed to the signal enhancement and feature extraction steps.

[0010] Adaptive filters are used to dynamically adjust filter parameters according to environmental changes. The high-quality data processed here contains low-concentration pollutant signals;

[0011] Step 3. Feature extraction based on multi-layer convolutional neural network (CNN): use multi-layer CNN to extract features from the preprocessed data, automatically extract high-dimensional features, and pass the extracted high-dimensional feature signals to the pollutant identification and classification step;

[0012] Step 4. Pollutant identification and classification. The classification model adopts a combined model, namely, SVM + decision tree model. Combining SVM and decision tree, SVM first performs preliminary classification, and then the decision tree performs detailed classification to obtain the classified pollutant types and corresponding concentration information; then it is passed to the data fusion and ensemble learning step;

[0013] Step 5. Data fusion: Use XGBoost and random forest to perform fusion analysis on the classified data from different sensors, and the comprehensive detection results after fusion analysis to improve the accuracy and sensitivity of detection;

[0014] Step 6. Anomaly detection: Use the autoencoder model to learn the normal environment data, establish the background noise model, identify abnormal signals by calculating the reconstruction error, and pass the identified abnormal signals and the detection results of low-concentration pollutants to the result output and feedback step;

[0015] Step 7. Output the results. Use data visualization technology to display pollutant concentrations and distribution in real time, establish an intelligent early warning system, automatically detect anomalies and send notifications; monitor data and early warning information in real time, and generate detailed monitoring reports.

[0016] The present invention is further configured such that the self-organizing network nodes of the intelligent sensor network in step 1 dynamically adjust the network topology structure according to network requirements;

[0017] Each sensor node is equipped with a wireless communication module to automatically scan surrounding nodes and establish connections;

[0018] The routing algorithm uses AODV and DSR algorithms; it provides stable guarantee for data network transmission; the network topology is automatically adjusted according to the addition and removal of nodes. The adaptive adjustment function here refers to the sensor node adjusting the working parameters according to environmental changes and network status, including sampling frequency and transmission power. The sensor monitors changes in the surrounding environment, including water temperature and salinity, and adjusts the sensor monitoring parameters according to environmental changes;

[0019] The node dynamically adjusts the transmission power according to the communication distance and channel quality; and dynamically adjusts the sampling frequency of the sensor according to the change of pollutant concentration:

[0020] Increase the sampling frequency when the pollutant concentration is high to obtain more data; reduce the sampling frequency when the concentration is low to save energy;

[0021] The sensor collects data in multiple dimensions. The sensor network collects data in multiple dimensions at the same time, including spectral signals, electrochemical signals, and surface plasmon resonance (SPR) signals.

[0022] The present invention is further configured such that in step 2, the sensor signal is denoised using the deep learning denoising autoencoder DAE in the following manner:

[0023] Assume that the collected noisy sensor signal data is x noisy , the noise-free target signal is x clean , divide the data into training set and test set;

[0024] Build an autoencoder model, which consists of an encoder and a decoder:

[0025] The encoder maps the input data x to a latent representation h, h = f(x) = σ(W e x+b e ), where x is the input noisy signal, W e is the weight matrix of the encoder, b e is the bias vector of the encoder, σ is the activation function;

[0026] The decoder maps the latent representation h back to the original data space and reconstructs the noise-free signal Where W d is the weight matrix of the decoder, b d is the bias vector of the decoder;

[0027] Use mean square error MSE as the loss function to calculate the reconstructed output With the noise-free signal x clean The error between Where n is the number of samples, represents the noise-free signal of the ith sample, is the reconstructed output signal of the i-th sample;

[0028] The present invention is further configured such that the denoising processing method in step 2 further includes:

[0029] Use backpropagation and gradient descent to optimize model weights and biases to minimize the loss function:

[0030] Where η represents the learning rate, Represents the loss function L on the encoder weight matrix W e The partial derivative of Represents the loss function L for the encoder bias vector b e The partial derivative of The loss function L is the decoder weight matrix W d The partial derivative of The loss function L is used to bias the decoder vector b d The partial derivative of

[0031] Then, the noisy sensor signal x is denoised. noisy Input the trained autoencoder, h = f(x noisy )=σ(W e x noisy +b e ), where x noisy is the noisy input signal, is the output signal after denoising;

[0032] Compare the denoised signals With the actual noise-free signal x clean , calculate the denoising effect indicators: signal-to-noise ratio SNR and mean square error MSE: Among them, SNR is the signal-to-noise ratio, represents the noise-free signal of the ith sample, is the denoised output signal of the i-th sample;

[0033] The present invention is further configured such that the high-dimensional feature extraction method in step 3 is:

[0034] Input data X, X is a denoised high signal-to-noise ratio sensor signal matrix with a size of m×n, where m is the length of the time series and n is the dimension of the signal, representing spectral signals and electrochemical signals;

[0035] Build a multi-layer convolutional neural network (CNN) model, including:

[0036] Convolution layer uses convolution kernel to perform convolution operation on input data to extract local features. The convolution operation formula is: H l,j =σ(W l,j *X l-1 +b l,j ), where H l,j represents the output feature map of the jth convolution kernel in the lth layer, W l,j is the weight matrix of the jth convolution kernel in the lth layer, X l-1 is the input data of the l-1th layer. For the first layer, the input data is the signal matrix X after denoising. l,j is the bias of the jth convolution kernel in the lth layer, σ is the activation function, and * represents the convolution operation; the pooling layer downsamples the output of the convolution layer to extract the main features and perform maximum pooling. The formula is: P l =max pool(H l ), where: P l is the output after pooling at layer l, H l Output feature map of the lth convolutional layer, max pool represents the maximum pooling operation;

[0037] The fully connected layer expands the output of the pooling layer into a one-dimensional vector and then performs feature extraction and classification. The calculation formula of the fully connected layer is: Z k =σ(W fc,k v l-1 +b fc,k ), where Z k It represents the output of the kth fully connected layer, W fc,k is the weight matrix of the kth fully connected layer, v l-1 is the input vector of the l-1th layer, b fc,k It represents the bias of the kth fully connected layer;

[0038] The present invention is further configured such that the high-dimensional feature extraction method in step 3 further includes:

[0039] Train the CNN model, use mean square error MSE and cross entropy loss as loss functions, calculate the error between the predicted output and the actual label, and use mean square error to predict pollutant concentration:

[0040] For the classification pollutant type problem, the cross entropy loss is used: Where N is the number of samples, is the predicted value of the i-th sample, y i is the true value of the i-th sample;

[0041] Use the back propagation algorithm and gradient descent method to optimize the model parameters and minimize the loss function. Where W is the weight matrix of the convolutional layer and the fully connected layer, b is the bias of the convolutional layer and the fully connected layer, and η represents the learning rate;

[0042] Through multi-layer convolution operations, local features and high-dimensional features of the input signal are extracted. Through pooling operations, the dimensions of the feature map are reduced and the main features are retained. Through the fully connected layer, the feature map is converted into a feature vector for subsequent pollutant identification and classification.

[0043] Output high-dimensional feature vector after multi-layer convolution and pooling processing;

[0044] The present invention is further configured such that the pollutant identification and classification method in step 4 is:

[0045] Input data X, X is a high-dimensional feature vector extracted by a multi-layer convolutional neural network CNN, with a size of m×n, where m is the number of samples and n is the feature dimension;

[0046] Support Vector Machine SVM preliminary classification classifies high-dimensional feature vectors into initial categories using SVM, and the hyperplane is defined as: w·x+b=0, where w is the weight vector, x is the input feature vector, and b is the bias;

[0047] Maximize the margin and minimize the loss function, where y i is the label of the i-th sample;

[0048] Convert the optimization problem into a dual problem and solve it using the Lagrange multiplier method:

[0049]

[0050] where α i is the Lagrange multiplier, C represents the regularization parameter;

[0051] For a new input sample x, the classification decision function is: f(x) = sign(w·x+b), where sign(·) is a sign function that determines the initial category to which the sample belongs;

[0052] The decision tree model is trained with the goal of minimizing the classification error, using information gain as the splitting criterion by recursively dividing the dataset into subsets;

[0053] For feature A, the information gain is defined as: Where H(D) is the entropy of the data set (D), D v It represents the subset divided according to the value v of feature A, and Values(A) is all possible values ​​of feature A;

[0054] The entropy of a data set (D) is defined as: where p i represents the proportion of the i-th category in the data set (D), and k represents the number of categories;

[0055] The present invention is further configured such that the pollutant identification and classification method in step 4 further includes:

[0056] Each node of the decision tree corresponds to a single feature test, and each node corresponds to a specific category and concentration range:

[0057] in represents the refined classification result of the sample, x is the input feature vector, and Tree(·) is the decision tree model;

[0058] Output the classified pollutant types and corresponding concentration information, including the preliminary classification categories and specific details of the refined classification;

[0059] The present invention is further configured that the data fusion method includes:

[0060] Integrate the data after SVM preliminary classification and decision tree refined classification into a comprehensive data set, which contains the classification results and corresponding feature vectors of each sample;

[0061] Perform feature engineering on the integrated data set, and then split the integrated data set into a training set and a test set. The training set is used to train the XGBoost and random forest models, and the test set is used to verify the performance of the model.

[0062] Train XGBoost model and random forest model;

[0063] The trained XGBoost model and random forest model are integrated, and the output of XGBoost and random forest is used as new features by stacking method to train the logistic regression model for final prediction;

[0064] The present invention is further configured such that the abnormality detection method is:

[0065] Based on the sensor data collected under normal conditions, it is divided into a training set and a validation set;

[0066] Construct an autoencoder model, input normal environment data into the autoencoder model, and train the model to minimize the error between the input data and the reconstructed data;

[0067] The model parameters are optimized through multiple iterations until the model converges, that is, the error between the input data and the reconstructed data is minimized;

[0068] After training is completed, save the autoencoder model and use the validation set data to evaluate the reconstruction performance of the autoencoder model;

[0069] The sensor data monitored in real time is input into the autoencoder model, and the real-time data is reconstructed through the autoencoder model to generate an output signal similar to the input data;

[0070] The error between the input data and the reconstructed data is calculated, and a threshold is set based on the distribution of the reconstruction error during training; if the reconstruction error exceeds the threshold, the signal is considered to be an abnormal signal.

[0071] Compared with the prior art, the present invention has the following beneficial effects:

[0072] In the aspect of high-dimensional and complex pollutant data processing, this application uses multiple sensors to work simultaneously, collect multi-dimensional data in real time, and provide comprehensive pollutant information; uses multi-layer convolutional neural network CNN for automatic feature extraction, automatically learns and extracts high-dimensional complex features in data, and improves the efficiency and accuracy of feature extraction; combines the combined model of support vector machine SVM and decision tree, first performs preliminary classification by SVM, and then refines classification by decision tree, so as to improve the accuracy of classification;

[0073] In this application, a denoising autoencoder DAE based on deep learning is used to denoise the sensor data. DAE effectively removes background noise in nonlinear and complex environments by learning noise patterns. In addition, an adaptive filter is used to dynamically adjust the filter parameters according to environmental changes to further optimize the denoising effect and improve signal quality.

[0074] In this application, an autoencoder model is used to learn normal environmental data, establish a background noise model, and accurately identify abnormal signals by calculating reconstruction errors, thereby improving the detection capabilities of low-concentration pollutants and abnormal signals;

[0075] It solves the problem that feature extraction in the existing technology relies on manual operation based on expert experience, which is time-consuming and labor-intensive and difficult to capture complex features in high-dimensional data. At the same time, the classification algorithms commonly used in traditional methods, such as linear regression or basic decision trees, perform poorly when faced with complex multi-dimensional features, resulting in low classification accuracy and difficulty in coping with complex and changeable marine environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] Figure 1 The figure is a flow chart of the online monitoring method for emerging marine pollutants of the present invention. DETAILED DESCRIPTION

[0077] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0078] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0079] The present invention is further described in detail below in conjunction with the accompanying drawings:

[0080] Example 1

[0081] See also Figure 1 The present invention provides an online monitoring method for emerging marine pollutants, comprising:

[0082] Step 1. Multi-sensor data collection, deployment of intelligent sensor network, deployment of intelligent sensor network with self-organizing network and adaptive adjustment functions in the target sea area, multi-dimensional data collected by sensors in real time, including spectral signals, electrochemical signals and SPR signals;

[0083] The self-organizing nodes of the intelligent sensor network dynamically adjust the network topology according to the network demand;

[0084] Each sensor node is equipped with a wireless communication module to automatically scan surrounding nodes and establish connections;

[0085] The routing algorithm uses AODV and DSR algorithms; it provides stable guarantee for data network transmission; the network topology is automatically adjusted according to the addition and removal of nodes. The adaptive adjustment function here refers to the sensor node adjusting the working parameters according to environmental changes and network status, including sampling frequency and transmission power. The sensor monitors changes in the surrounding environment, including water temperature and salinity, and adjusts the sensor monitoring parameters according to environmental changes;

[0086] The node dynamically adjusts the transmission power according to the communication distance and channel quality; and dynamically adjusts the sampling frequency of the sensor according to the change of pollutant concentration:

[0087] Increase the sampling frequency when the pollutant concentration is high to obtain more data; reduce the sampling frequency when the concentration is low to save energy;

[0088] The sensor collects data in multiple dimensions. The sensor network collects data in multiple dimensions at the same time, including spectral signals, electrochemical signals, and surface plasmon resonance (SPR) signals.

[0089] Step 2: Signal preprocessing: Use a deep learning-based denoising autoencoder (DAE) to denoise the sensor data, remove background noise, and retain the characteristic signals of pollutants. The denoised high signal-to-noise ratio data is passed to the signal enhancement and feature extraction steps.

[0090] Adaptive filters are used to dynamically adjust filter parameters according to environmental changes. The high-quality data processed here contains low-concentration pollutant signals;

[0091] The method of denoising the sensor signal using the deep learning denoising auto encoder DAE is as follows:

[0092] Assume that the collected noisy sensor signal data is x noisy , the noise-free target signal is x clean , divide the data into training set and test set;

[0093] Build an autoencoder model, which consists of an encoder and a decoder:

[0094] The encoder maps the input data x to a latent representation h, h = f(x) = σ(W e x+b e ), where x is the input noisy signal, W e is the weight matrix of the encoder, b e is the bias vector of the encoder, σ is the activation function;

[0095] The decoder maps the latent representation h back to the original data space and reconstructs the noise-free signal Where W d is the weight matrix of the decoder, b d is the bias vector of the decoder;

[0096] Use mean square error MSE as the loss function to calculate the reconstructed output With the noise-free signal x clean The error between Where n is the number of samples, represents the noise-free signal of the ith sample, is the reconstructed output signal of the i-th sample;

[0097] Use backpropagation and gradient descent to optimize model weights and biases to minimize the loss function:

[0098] Where η represents the learning rate, Represents the loss function L on the encoder weight matrix W e The partial derivative of Represents the loss function L for the encoder bias vector b e The partial derivative of The loss function L is the decoder weight matrix W d The partial derivative of The loss function L is used to bias the decoder vector b d The partial derivative of

[0099] Then, the noisy sensor signal x is denoised. noisy Input the trained autoencoder, h = f(x noisy )=σ(W e x noisy +b e ), where x noisy is the noisy input signal, is the output signal after denoising;

[0100] Compare the denoised signals Compared with the actual noise-free signal xclean, calculate the denoising effect indicators: signal-to-noise ratio SNR and mean square error MSE: Among them, SNR is the signal-to-noise ratio, represents the noise-free signal of the ith sample, is the denoised output signal of the i-th sample; by removing the background noise in step 2 and retaining the characteristic signal of the pollutant, the signal-to-noise ratio and detection accuracy are significantly improved;

[0101] Step 3. Feature extraction based on multi-layer convolutional neural network (CNN): use multi-layer CNN to extract features from the preprocessed data, automatically extract high-dimensional features, and pass the extracted high-dimensional feature signals to the pollutant identification and classification step;

[0102] The high-dimensional feature extraction method is:

[0103] Input data X, X is a denoised high signal-to-noise ratio sensor signal matrix with a size of m×n, where m is the length of the time series and n is the dimension of the signal, representing spectral signals and electrochemical signals;

[0104] Build a multi-layer convolutional neural network (CNN) model, including:

[0105] Convolution layer uses convolution kernel to perform convolution operation on input data to extract local features. The convolution operation formula is: H l,j =σ(W l,j *X l-1 +b l,j ), where H l,j represents the output feature map of the jth convolution kernel in the lth layer, W l,j is the weight matrix of the jth convolution kernel in the lth layer, X l-1 is the input data of the l-1th layer. For the first layer, the input data is the signal matrix X after denoising. l,j is the bias of the jth convolution kernel in the lth layer, σ is the activation function, and * represents the convolution operation; the pooling layer downsamples the output of the convolution layer to extract the main features and perform maximum pooling. The formula is: P l =max pool(H l ), where: P l is the output after pooling at layer l, H l Output feature map of the lth convolutional layer, max pool represents the maximum pooling operation;

[0106] The fully connected layer expands the output of the pooling layer into a one-dimensional vector and then performs feature extraction and classification. The calculation formula of the fully connected layer is: Z k =σ(W fc,k v l-1 +b fc,k ), where Z k It represents the output of the kth fully connected layer, W fc,k is the weight matrix of the kth fully connected layer, v l-1 is the input vector of the l-1th layer, b fc,k It represents the bias of the kth fully connected layer;

[0107] Train the CNN model, use mean square error MSE and cross entropy loss as loss functions, calculate the error between the predicted output and the actual label, and use mean square error to predict pollutant concentration:

[0108]

[0109] For the classification pollutant type problem, the cross entropy loss is used: Where N is the number of samples, is the predicted value of the i-th sample, y i is the true value of the i-th sample;

[0110] Use the back propagation algorithm and gradient descent method to optimize the model parameters and minimize the loss function. Where W is the weight matrix of the convolutional layer and the fully connected layer, b is the bias of the convolutional layer and the fully connected layer, and η represents the learning rate;

[0111] Through multi-layer convolution operations, local features and high-dimensional features of the input signal are extracted. Through pooling operations, the dimensions of the feature map are reduced and the main features are retained. Through the fully connected layer, the feature map is converted into a feature vector for subsequent pollutant identification and classification.

[0112] Output high-dimensional feature vector after multi-layer convolution and pooling processing;

[0113] Step 4. Pollutant identification and classification. The classification model adopts a combined model, namely, SVM + decision tree model. Combining SVM and decision tree, SVM first performs preliminary classification, and then the decision tree performs detailed classification to obtain the classified pollutant types and corresponding concentration information, which are then passed to the data fusion and ensemble learning step; the pollutant identification and classification method is as follows:

[0114] Input data X, X is a high-dimensional feature vector extracted by a multi-layer convolutional neural network CNN, with a size of m×n, where m is the number of samples and n is the feature dimension;

[0115] Support Vector Machine SVM preliminary classification classifies high-dimensional feature vectors into initial categories using SVM, and the hyperplane is defined as: w·x+b=0, where w is the weight vector, x is the input feature vector, and b is the bias;

[0116] Maximize the margin and minimize the loss function, where y i is the label of the i-th sample;

[0117] Convert the optimization problem into a dual problem and solve it using the Lagrange multiplier method:

[0118]

[0119] where α i is the Lagrange multiplier, C represents the regularization parameter;

[0120] For a new input sample x, the classification decision function is: f(x) = sign(w·x+b), where sign(·) is a sign function that determines the initial category to which the sample belongs;

[0121] The decision tree model is trained with the goal of minimizing the classification error, using information gain as the splitting criterion by recursively dividing the dataset into subsets;

[0122] For feature A, information gain is defined as: Where H(D) is the entropy of the data set (D), D vIt represents the subset divided according to the value v of feature A, and Values(A) is all possible values ​​of feature A;

[0123] The entropy of a data set (D) is defined as: where p i represents the proportion of the i-th category in the data set (D), and k represents the number of categories;

[0124] Each node of the decision tree corresponds to a single feature test, and each node corresponds to a specific category and concentration range:

[0125] in represents the refined classification result of the sample, x is the input feature vector, and Tree(·) is the decision tree model;

[0126] Output the classified pollutant types and corresponding concentration information, including the preliminary classification categories and specific details of the refined classification;

[0127] Step 5. Data fusion: Use XGBoost and random forest to perform fusion analysis on the classified data from different sensors, and the comprehensive detection results after fusion analysis can improve the accuracy and sensitivity of detection; data fusion methods include:

[0128] Integrate the data after SVM preliminary classification and decision tree refined classification into a comprehensive data set, which contains the classification results and corresponding feature vectors of each sample;

[0129] Perform feature engineering on the integrated data set, and then split the integrated data set into a training set and a test set. The training set is used to train the XGBoost and random forest models, and the test set is used to verify the performance of the model.

[0130] Train XGBoost model and random forest model;

[0131] Use the training set to train the XGBoost model, adjust the model parameters through cross-validation, build multiple decision trees, and train each tree based on the residual of the previous tree to continuously reduce the prediction error;

[0132] Use the training set to train the random forest model, adjust the model parameters through cross-validation, build multiple decision trees, each tree is trained on a different subset of random samples, and finally make predictions by averaging;

[0133] The trained XGBoost model and random forest model are integrated, and the stacking method is used to use the output of XGBoost and random forest as new features to train the logistic regression model for final prediction; the comprehensive detection results are output in real time to display the type and concentration information of pollutants;

[0134] Step 6. Anomaly detection: Use the autoencoder model to learn the normal environment data, establish the background noise model, identify abnormal signals by calculating the reconstruction error, and pass the identified abnormal signals and the detection results of low-concentration pollutants to the result output and feedback step; the anomaly detection method is:

[0135] Based on the sensor data collected under normal conditions, it is divided into a training set and a validation set;

[0136] Construct an autoencoder model, input normal environment data into the autoencoder model, and train the model to minimize the error between the input data and the reconstructed data;

[0137] The model parameters are optimized through multiple iterations until the model converges, that is, the error between the input data and the reconstructed data is minimized;

[0138] After training is completed, save the autoencoder model and use the validation set data to evaluate the reconstruction performance of the autoencoder model;

[0139] The sensor data monitored in real time is input into the autoencoder model, and the real-time data is reconstructed through the autoencoder model to generate an output signal similar to the input data;

[0140] Calculate the error between the input data and the reconstructed data, and set a threshold based on the distribution of the reconstruction error during training; if the reconstruction error exceeds the threshold, the signal is considered to be an abnormal signal;

[0141] Step 7. Output the results. Use data visualization technology to display pollutant concentrations and distribution in real time, establish an intelligent early warning system, automatically detect anomalies and send notifications; monitor data and early warning information in real time, and generate detailed monitoring reports.

[0142] The present invention arranges sensor nodes in the target sea area, and the nodes are equipped with wireless communication modules, which can automatically scan and connect to surrounding nodes to form a stable network topology structure; AODV and DSR routing algorithms are used to provide stability of data transmission, and the sensor nodes can dynamically adjust working parameters according to environmental changes and network status, including sampling frequency and transmission power; the sensors monitor environmental changes of water temperature and salinity, and adjust parameters according to changes in environmental parameters. When the concentration of pollutants is high, the sensors increase the sampling frequency to obtain more data; when the concentration is low, the sampling frequency is reduced to save energy; and the sensor network can simultaneously collect multi-dimensional data of spectral signals, electrochemical signals and surface plasmon resonance (SPR) signals;

[0143] In the signal preprocessing stage, a denoising autoencoder DAE based on deep learning is used to denoise the sensor data, remove background noise, and retain the characteristic signals of pollutants;

[0144] In the high-dimensional feature extraction stage, a multi-layer convolutional neural network (CNN) is used to extract features from the preprocessed data. The input data is a denoised high signal-to-noise ratio sensor signal matrix. A multi-layer CNN model is constructed. The pooling layer downsamples the output of the convolution layer to retain the main features. The fully connected layer expands the output of the pooling layer into a one-dimensional vector for feature extraction and classification. The high-dimensional features of the input signal are extracted through multi-layer convolution operations.

[0145] In the pollutant identification and classification stage, a combined model SVM+decision tree model is used for classification. First, SVM performs preliminary classification and finds the optimal hyperplane to classify high-dimensional feature vectors into initial categories. Then, a decision tree is used to refine the results of the SVM preliminary classification, and the data set is recursively divided into subsets through information gain, and finally the classified pollutant types and corresponding concentration information are output.

[0146] In the data fusion stage, XGBoost and random forest are used to perform fusion analysis on the classified data from different sensors. First, the data that has undergone preliminary SVM classification and decision tree refinement classification are integrated into a comprehensive data set. After feature engineering, the comprehensive data set is divided into a training set and a test set. The training set is used to train the XGBoost model and the random forest model respectively. The trained XGBoost model and the random forest model are fused, and the stacking method is used to use their outputs as new features to train the logistic regression model for final prediction. The comprehensive detection results are output in real time to display the type and concentration information of pollutants.

[0147] The above contents are only for explaining the technical idea of ​​the present invention and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution in accordance with the technical idea proposed by the present invention shall fall within the protection scope of the claims of the present invention.

Claims

1. An online monitoring method for emerging marine pollutants, characterized in that: include: Step 1. Multi-sensor data collection, deployment of intelligent sensor network, deployment of intelligent sensor network with self-organizing network and adaptive adjustment functions in the target sea area, multi-dimensional data collected by sensors in real time, including spectral signals, electrochemical signals and SPR signals; Step 2: Signal preprocessing: Use a deep learning-based denoising autoencoder (DAE) to denoise the sensor data, remove background noise, and retain the characteristic signals of pollutants. The denoised high signal-to-noise ratio data is passed to the signal enhancement and feature extraction steps. Adopt adaptive filter to dynamically adjust filter parameters according to environmental changes; Step 3. Feature extraction based on multi-layer convolutional neural network (CNN): use multi-layer CNN to extract features from the preprocessed data, automatically extract high-dimensional features, and pass the extracted high-dimensional feature signals to the pollutant identification and classification step; Step 4. Pollutant identification and classification: The classification model adopts a combined model, namely, SVM + decision tree model, which combines SVM and decision tree. SVM first performs preliminary classification, and then the decision tree performs detailed classification to obtain the classified pollutant types and corresponding concentration information; Step 5. Data fusion: Use XGBoost and random forest to perform fusion analysis on the classified data from different sensors, and obtain the comprehensive detection results after fusion analysis; Step 6. Anomaly detection: Use the autoencoder model to learn the normal environment data, establish a background noise model, and identify abnormal signals by calculating the reconstruction error; Step 7. Output the results, using data visualization technology to display pollutant concentration and distribution in real time, establish an intelligent early warning system, automatically detect anomalies and send notifications.

2. The online monitoring method for emerging marine pollutants according to claim 1, characterized in that: In step 1, the self-organizing nodes of the intelligent sensor network dynamically adjust the network topology according to network requirements; Each sensor node is equipped with a wireless communication module to automatically scan surrounding nodes and establish connections; The routing algorithm uses AODV and DSR algorithms; the network topology is automatically adjusted according to the addition and removal of nodes. The adaptive adjustment function here refers to the sensor node adjusting the working parameters according to environmental changes and network status, including sampling frequency and transmission power. The sensor monitors changes in the surrounding environment, including water temperature and salinity, and adjusts the sensor monitoring parameters according to environmental changes; The node dynamically adjusts the transmission power according to the communication distance and channel quality; and dynamically adjusts the sampling frequency of the sensor according to the change of pollutant concentration: Increase the sampling frequency when the pollutant concentration is high; decrease the sampling frequency when the concentration is low; The sensors collect data in multiple dimensions, and the sensor network collects data in multiple dimensions simultaneously, including spectral signals, electrochemical signals, and surface plasmon resonance (SPR) signals.

3. The online monitoring method for emerging marine pollutants according to claim 2, characterized in that: In step 2, the deep learning denoising autoencoder DAE is used to denoise the sensor signal as follows: Assume that the collected noisy sensor signal data is x noisy , the noise-free target signal is x clean , divide the data into training set and test set; Build an autoencoder model, which consists of an encoder and a decoder: The encoder maps the input data x to a latent representation h, h = f(x) = σ(W e x+b e ), where x is the input noisy signal, W e is the weight matrix of the encoder, b e is the bias vector of the encoder, σ is the activation function; The decoder maps the latent representation h back to the original data space and reconstructs the noise-free signal Where W d is the weight matrix of the decoder, b d is the bias vector of the decoder; Use mean square error MSE as the loss function to calculate the reconstructed output With the noise-free signal x clean The error between Where n is the number of samples, represents the noise-free signal of the ith sample, is the reconstructed output signal of the i-th sample.

4. The online monitoring method for emerging marine pollutants according to claim 3 is characterized in that: The denoising process in step 2 also includes: Use backpropagation and gradient descent to optimize model weights and biases to minimize the loss function: Where η represents the learning rate, Represents the loss function L on the encoder weight matrix W e The partial derivative of Represents the loss function L for the encoder bias vector b e The partial derivative of The loss function L is the decoder weight matrix W d The partial derivative of The loss function L is used to bias the decoder vector b d The partial derivative of Then, the noisy sensor signal x is denoised. noisy Input the trained autoencoder, h = f(x noisy )=σ(W e x noisy +b e ), where x noisy is the noisy input signal, is the output signal after denoising; Compare the denoised signals With the actual noise-free signal x clean , calculate the denoising effect indicators: signal-to-noise ratio SNR and mean square error MSE: Among them, SNR is the signal-to-noise ratio, represents the noise-free signal of the ith sample, is the denoised output signal of the i-th sample.

5. The online monitoring method for emerging marine pollutants according to claim 4 is characterized in that: The high-dimensional feature extraction method in step 3 is: Input data X, X is a denoised high signal-to-noise ratio sensor signal matrix with a size of m×n, where m is the length of the time series and n is the dimension of the signal, representing spectral signals and electrochemical signals; Build a multi-layer convolutional neural network (CNN) model, including: Convolution layer uses convolution kernel to perform convolution operation on input data to extract local features. The convolution operation formula is: H l,j =σ(W l,j *X l-1 +b l,j ), where H l,j represents the output feature map of the jth convolution kernel in the lth layer, W l,j is the weight matrix of the jth convolution kernel in the lth layer, X l-1 is the input data of the l-1th layer. For the first layer, the input data is the signal matrix X after denoising. l,j is the bias of the jth convolution kernel in the lth layer, σ is the activation function, and * represents the convolution operation; the pooling layer downsamples the output of the convolution layer to extract the main features and perform maximum pooling. The formula is: P l =max pool(H l ), where: P l is the output after pooling at layer l, H l Output feature map of the lth convolutional layer, max pool represents the maximum pooling operation; The fully connected layer expands the output of the pooling layer into a one-dimensional vector and then performs feature extraction and classification. The calculation formula of the fully connected layer is: Z k =σ(W fc,k v l-1 +b fc,k ), where Z k It represents the output of the kth fully connected layer, W fc,k is the weight matrix of the kth fully connected layer, v l-1 is the input vector of the l-1th layer, b fc,k It represents the bias of the kth fully connected layer.

6. The online monitoring method for emerging marine pollutants according to claim 5, characterized in that: The high-dimensional feature extraction method in step 3 also includes: Train the CNN model, use mean square error MSE and cross entropy loss as loss functions, calculate the error between the predicted output and the actual label, and use mean square error to predict pollutant concentration: For the classification pollutant type problem, the cross entropy loss is used: Where N is the number of samples, is the predicted value of the i-th sample, y i is the true value of the i-th sample; Use the back propagation algorithm and gradient descent method to optimize the model parameters and minimize the loss function. Where W is the weight matrix of the convolutional layer and the fully connected layer, b is the bias of the convolutional layer and the fully connected layer, and η represents the learning rate; Output a high-dimensional feature vector after multi-layer convolution and pooling processing.

7. The online monitoring method for emerging marine pollutants according to claim 6, characterized in that: The pollutant identification and classification method in step 4 is: Input data X, X is a high-dimensional feature vector extracted by a multi-layer convolutional neural network CNN, with a size of m×n, where m is the number of samples and n is the feature dimension; Support Vector Machine SVM preliminary classification classifies high-dimensional feature vectors into initial categories using SVM, and the hyperplane is defined as: w·x+b=0, where w is the weight vector, x is the input feature vector, and b is the bias; Maximize the margin and minimize the loss function, where y i is the label of the i-th sample; Convert the optimization problem into a dual problem and solve it using the Lagrange multiplier method: where α i is the Lagrange multiplier, C represents the regularization parameter; For a new input sample x, the classification decision function is: f(x) = sign(w·x+b), where sign(·) is a sign function that determines the initial category to which the sample belongs; The decision tree model is trained with the goal of minimizing the classification error, using information gain as the splitting criterion by recursively dividing the dataset into subsets; For feature A, information gain is defined as: Where H(D) is the entropy of the data set (D), D v It represents the subset divided according to the value v of feature A, and Values(A) is all possible values ​​of feature A; The entropy of a data set (D) is defined as: where p i represents the proportion of the i-th category in the data set (D), and k represents the number of categories.

8. The online monitoring method for emerging marine pollutants according to claim 7 is characterized in that: The pollutant identification and classification methods in step 4 also include: Each node of the decision tree corresponds to a single feature test, and each node corresponds to a specific category and concentration range: in represents the refined classification result of the sample, x is the input feature vector, and Tree(·) is the decision tree model; Output the classified pollutant types and corresponding concentration information, including the preliminary classification categories and specific details of the refined classification.

9. The online monitoring method for emerging marine pollutants according to claim 8, characterized in that: Data fusion methods include: Integrate the data after SVM preliminary classification and decision tree refined classification into a comprehensive data set, which contains the classification results and corresponding feature vectors of each sample; Perform feature engineering on the integrated data set, and then split the integrated data set into a training set and a test set. The training set is used to train the XGBoost and random forest models, and the test set is used to verify the performance of the model. Train XGBoost model and random forest model; The trained XGBoost model and random forest model are fused, and the stacking method is used to use the output of XGBoost and random forest as new features to train the logistic regression model for the final prediction.

10. The online monitoring method for emerging marine pollutants according to claim 9, characterized in that: The anomaly detection methods are: Based on the sensor data collected under normal conditions, it is divided into a training set and a validation set; Construct an autoencoder model, input normal environment data into the autoencoder model, and train the model to minimize the error between the input data and the reconstructed data; The model parameters are optimized through multiple iterations until the model converges, that is, the error between the input data and the reconstructed data is minimized; After training is completed, save the autoencoder model and use the validation set data to evaluate the reconstruction performance of the autoencoder model; The sensor data monitored in real time is input into the autoencoder model, and the real-time data is reconstructed through the autoencoder model to generate an output signal similar to the input data; Calculate the error between the input data and the reconstructed data, and set a threshold based on the distribution of reconstruction errors during training; If the reconstruction error exceeds the threshold, the signal is considered to be an abnormal signal.

Citation Information

Cited By

  • Water quality detection method and device based on electrochemical sensor

    CN120490266A

  • Leakage detection method based on multi-sensor time shift correction and deep learning

    CN120492986A