Industrial wastewater monitoring data authenticity verification method based on multi-modal feature fusion
Through the multimodal feature fusion method, combined with the spatiotemporal characteristics of industrial wastewater, encrypted QR codes are generated and data confidence scores are calculated, which solves the problem of difficulty in building a tamper-proof mechanism and identifying spatiotemporal characteristics in the existing technology, and realizes data authenticity verification with high accuracy and reliability.
Patent Information
- Application Number
- CN202510247102.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-04
AI Technical Summary
The prior art is difficult to build a tamper-proof mechanism based on physical characteristics at the sensor level, and it is difficult to effectively identify the spatio-temporal characteristics of industrial wastewater discharge, resulting in the limitation of the accuracy and reliability of data authenticity verification.
The multimodal feature fusion method is used to obtain and standardize monitoring data, extract water quality feature vectors, and combine the pipeline topology data to perform spatial autocorrelation analysis to generate a spatiotemporal fusion feature matrix, which is used to generate encrypted QR codes and calculate the data confidence score.
It has realized the construction of a tamper-proof mechanism at the sensor level, improved the recognition ability of abnormal features, effectively identified local abnormal emissions or data tampering behavior, and improved the accuracy and reliability of data authenticity verification.
Smart Images

Figure CN120144981A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of environmental monitoring, and in particular, to a method for verifying the authenticity of industrial wastewater monitoring data based on multi-modal feature fusion. Background Art
[0002] Verifying the authenticity of industrial wastewater monitoring data is of great significance for environmental supervision and pollution prevention. First, industrial wastewater emissions are characterized by intermittency, suddenness, and variability, with large fluctuations and rapid changes in water quality parameters, which pose great challenges to data collection and verification. Second, industrial parks usually contain multiple emission entities, and the intricate pipe network system makes it difficult to trace the transmission path of pollutants in space. In addition, due to the increased intensity of environmental penalties, some enterprises may adopt various technical means to interfere with or tamper with monitoring data, which seriously affects the effectiveness of environmental supervision. Therefore, establishing a reliable data authenticity verification mechanism is of great practical significance for ensuring the authenticity of environmental monitoring data and improving the efficiency of environmental supervision.
[0003] Current data authenticity verification technologies mainly focus on the security protection of data transmission and storage links. Traditional methods usually use information security technologies such as digital signatures and encrypted transmission to protect the data transmission process, or use distributed ledger technologies such as blockchain to ensure the immutability of data storage. In the data collection link, physical isolation means such as video monitoring and sampler lead sealing are mainly relied on to prevent human intervention. In data analysis, common methods include anomaly detection based on statistical models, rule-based threshold judgment, etc. These methods mainly rely on historical data to establish judgment criteria and are difficult to cope with complex and changing industrial wastewater discharge scenarios.
[0004] However, the core problem in verifying the authenticity of industrial wastewater monitoring data is how to build a tamper-proof mechanism based on physical characteristics at the sensor level, and at the same time establish a multi-dimensional verification system in combination with the spatio-temporal characteristics of wastewater discharge to prevent data from being tampered with or forged at the source. The existence of these technical problems seriously restricts the accuracy and reliability of verifying the authenticity of industrial wastewater monitoring data. Summary of the Invention
[0005] The object of the invention is to provide a method for verifying the authenticity of industrial wastewater monitoring data based on multi-modal feature fusion to solve the above problems existing in the prior art.
[0006] Technical solution: A method for verifying the authenticity of industrial wastewater monitoring data based on multi-modal feature fusion includes:
[0007] Obtain monitoring data, perform standardization processing on the monitoring data to obtain standardized water quality data; based on the standardized water quality data, extract water quality feature vectors; wherein the monitoring data includes pH value, COD value, and turbidity;
[0008] Obtain data of adjacent monitoring points, and construct a standardized weight matrix based on pre-stored pipe network topology data; perform spatial autocorrelation analysis on the water quality feature vector based on the standardized weight matrix to obtain a spatial feature vector; input the spatial feature vector into the feature fusion unit to generate a spatio-temporal fusion feature matrix;
[0009] Generate an encrypted QR code based on the spatio-temporal fusion feature matrix and the standardized water quality data;
[0010] Receive the encrypted QR code, calculate the data confidence score, and generate the corresponding warning description information.
[0011] Advantageous effects: The present invention constructs an adaptive data preprocessing method suitable for the characteristics of industrial wastewater, improving the recognition ability of abnormal features; realizes a verification mechanism based on spatio-temporal feature fusion, which can effectively identify local abnormal emissions or data tampering behaviors; uses a multi-head attention mechanism for feature fusion, improving the reliability of the verification results; combines belief propagation and graph neural networks to construct an objective data credibility evaluation system; through the synergistic effect of physical feature verification, data feature verification and spatio-temporal feature verification, it can effectively prevent data tampering, timely detect abnormal emissions, and provide reliable technical support for environmental supervision. Description of the Drawings
[0012] Figure 1 It is a step flow chart of a method for verifying the authenticity of industrial wastewater monitoring data based on multi-modal feature fusion provided by an embodiment of the present application.
[0013] Figure 2 It is a step flow chart of constructing a standardized weight matrix provided by an embodiment of the present application.
[0014] Figure 3 It is a step flow chart of standardizing monitoring data provided by an embodiment of the present application.
[0015] Figure 4 It is a step flow chart of constructing a time feature vector provided by an embodiment of the present application. Detailed Embodiments
[0016] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0017] It should be specifically noted that, for the purpose of clearly demonstrating the step flow of this application, serial numbers are marked for each step in the specification. These serial numbers are only for the convenience of explanation and do not limit the execution order of the steps. In actual operation, according to the technical requirements of specific implementation scenarios, the steps can be executed in an order different from that shown in the specification, and in some cases, parallel processing between steps can also be achieved.
[0018] As Figure 1 shown, this application proposes an industrial wastewater monitoring data authenticity verification method based on multi-modal feature fusion, including the following steps:
[0019] S1. Obtain the monitoring data, perform standardization processing on the monitoring data to obtain standardized water quality data; based on the standardized water quality data, extract the water quality feature vector; where the monitoring data includes pH value, COD value, and turbidity;
[0020] S2. Obtain the data of adjacent monitoring points, construct a standardized weight matrix based on the pre-stored pipe network topology data; perform spatial autocorrelation analysis on the water quality feature vector based on the standardized weight matrix to obtain the spatial feature vector; input the spatial feature vector into the feature fusion unit to generate a spatio-temporal fusion feature matrix;
[0021] S3. Generate an encrypted QR code based on the spatio-temporal fusion feature matrix and the standardized water quality data;
[0022] S4. Receive the encrypted QR code, calculate the data confidence score, and generate the corresponding warning description information.
[0023] According to one aspect of this application, it further includes step S0. Read the sensor working voltage fluctuation value, temperature response characteristic, and medium impedance, perform Fourier transform to obtain the sensor feature spectrum matrix; construct a feature vector and map it to generate the sensor dynamic fingerprint; perform feature extraction to generate the fingerprint feature code.
[0024] In an embodiment of this application, based on the sensor ID number stored in the built-in memory of the sensor, read the sensor working voltage fluctuation value, temperature response characteristic, and medium impedance in the sensor calibration database; use the standard sampling frequency of 2000Hz, perform Fourier transform on these three physical quantities with 8192 sampling points to obtain the sensor feature spectrum matrix; according to the sensor feature spectrum matrix, apply the locality-sensitive hashing algorithm to construct a 128-dimensional feature vector, use 20 random hyperplanes as the hashing function to generate the sensor dynamic fingerprint; input the sensor dynamic fingerprint into the locality-sensitive hashing encoder to generate a 256-bit fingerprint feature code.
[0025] Then, the original values of water quality parameters (monitoring data, the same below) including pH value, COD value, and turbidity are collected from the water quality sensor; the original values of water quality parameters are smoothed using a sliding window of 100 data points to obtain the smoothed water quality data; the smoothed water quality data is input into the adaptive standardization processing unit to generate the standardized water quality data; the first five principal components are extracted from the standardized water quality data using the principal component analysis method to form the water quality feature vector.
[0026] The historical data sequence and water quality feature vector for the most recent 24 hours are read from the data storage unit and input into the long short-term memory network to generate the time feature vector; the data of adjacent monitoring points are collected from adjacent monitoring points and subjected to spatial autocorrelation analysis with the current water quality feature vector to obtain the spatial feature vector; the time feature vector and the spatial feature vector are input into the tensor decomposition algorithm to generate the spatio-temporal fusion feature matrix.
[0027] The fingerprint feature code, the spatio-temporal fusion feature matrix, and the vector of real-time collected environmental parameters are input into the feature alignment unit to generate the aligned feature matrix; the multi-head attention mechanism is applied to the aligned feature matrix for feature weighting to obtain the fused feature matrix; the fused feature matrix, the standardized water quality data, and the sensor dynamic fingerprint code are used to generate the encrypted QR code.
[0028] The fused feature matrix and the historical confidence data stored in the database are obtained, and the data confidence score is calculated through the belief propagation algorithm; the data confidence score is matched with the historical anomaly pattern library in the database to generate the anomaly type identifier and the anomaly degree value; according to the anomaly type identifier and the anomaly degree value, the warning level and the warning description information are generated in combination with the predefined warning rules.
[0029] According to one aspect of the present application, step S0 is further as follows:
[0030] S01. The sensor ID number in the built-in memory of the sensor is read through the 485 communication interface, and the sensor working voltage fluctuation value, temperature response characteristic, and medium impedance are retrieved from the sensor calibration database according to this ID number; the sampling frequency of 2000 Hz is used for sampling these three physical quantities respectively, and 8192 data points are collected for each physical quantity; the sampling data is input into the Fourier transform processing unit to generate three groups of spectrum data; these three groups of spectrum data are combined into a 15×15 sensor characteristic spectrum matrix.
[0031] S02. The sensor characteristic spectrum matrix is read, and the main feature components in the matrix are extracted using the principal component analysis method; the extracted feature components are recombined into a 128-dimensional feature vector; 20 hyperplanes are randomly generated as hash functions to map the feature vector; the mapping result is matched with the pre-stored sensor characteristic template to generate the sensor dynamic fingerprint.
[0032] S03. Obtain the dynamic fingerprint of the sensor and input it into the locality-sensitive hashing encoder; use the decision condition with a Hamming distance threshold of 0.2 to extract features from the sensor dynamic fingerprint; generate a 256-bit fingerprint feature code based on the extracted features.
[0033] In an embodiment of the present application, obtain the core physical parameters of the sensor, including the sensor operating voltage fluctuation value V, temperature response characteristic T, and medium impedance Z, obtain the characteristic spectrum through Fourier transform, extract the characteristic peaks and valleys, and output the sensor characteristic spectrum matrix F. Based on the sensor characteristic spectrum matrix F, apply the adaptive weight fusion algorithm to combine the characteristics of different physical parameters and output the sensor dynamic fingerprint D. Based on the sensor dynamic fingerprint D, use the locality-sensitive hashing (LSH) algorithm for feature encoding and output the fingerprint feature code E.
[0034] In this embodiment, by performing high-frequency sampling at 2000 Hz on the three physical quantities of the sensor operating voltage fluctuation value, temperature response characteristic, and medium impedance and combining Fourier transform, the characteristic spectrum of the sensor under different operating states can be obtained, and these spectrum characteristics reflect the inherent physical characteristics of the sensor. Using the locality-sensitive hashing algorithm to construct a 128-dimensional feature vector and using 20 random hyperplanes as hash functions, the physical characteristics of the sensor can be compressed into unique digital fingerprints. Since these physical characteristics are determined by the hardware structure of the sensor and have stable correlations under different operating states, the 256-bit fingerprint feature code generated in this way has strong anti-counterfeiting and uniqueness. Even if an attacker obtains the data output interface of the sensor, they cannot forge or tamper with these fingerprint features based on physical characteristics. At the same time, using the locality-sensitive hashing encoder for feature extraction ensures that the distances of similar physical characteristics in the coding space are also close, so that even if the sensor is slightly physically disturbed, the stability of the fingerprint features can still be maintained. Through this dynamic fingerprint mechanism based on physical characteristics, the credibility of the data source can be guaranteed at the hardware level, laying a foundation for subsequent data authenticity verification.
[0035] According to one aspect of the present application, step S1 is further:
[0036] S11. Through the data interface of the water quality sensor, collect the original values of water quality parameters such as pH value, COD value, and turbidity every 10 seconds; store the collected original values of water quality parameters in the temporary buffer area in chronological order; read the latest 100 data points from the temporary buffer area and perform sliding smoothing processing using the Hamming window function to generate the smoothed water quality data.
[0037] S12. Read the smoothed water quality data; obtain the standard range values of various parameters from the system configuration database; perform normalization processing on the smoothed water quality data according to the standard range values to generate standardized water quality data in the range of [0, 1].
[0038] S13. Obtain the standardized water quality data; construct a covariance matrix and calculate the eigenvalues and eigenvectors; select the top 5 principal components with a contribution rate greater than 85%; combine these 5 principal components to form a water quality feature vector.
[0039] In an embodiment of the present application, the original values W of water quality parameters (including pH, COD, turbidity, etc.) are obtained, and the sliding window method is used for data smoothing to obtain the smoothed water quality data Ws. Based on the smoothed water quality data Ws, the adaptive standardization algorithm is used for data normalization to output the standardized water quality data Wn. Based on the standardized water quality data Wn, principal component analysis is applied to extract key features, and the water quality feature vector Wf is output.
[0040] In this embodiment, by smoothing the original values of water quality parameters using a sliding window of 100 data points, random noise and sudden interference during the sensor acquisition process can be effectively filtered out. By extracting the top 5 principal components to form a water quality feature vector through the principal component analysis method, not only data dimensionality reduction is achieved, but also the correlation information between water quality parameters is retained. This embodiment can not only ensure the integrity of data representation but also highlight the characteristics of abnormal patterns. Especially when dealing with monitoring objects such as industrial wastewater with intermittent and sudden characteristics, it can accurately capture the key features of water quality changes and avoid the problem of feature loss that may be caused by the traditional fixed window smoothing method.
[0041] According to an aspect of the present application, step S12 is further as follows:
[0042] S121. Read the standard range thresholds of pH value, COD value, and turbidity from the system configuration database, including the minimum value, maximum value, mean value, and standard deviation; perform weighted averaging of these thresholds with the statistical characteristics of historical data, and the weight ratio is 3:7, to obtain dynamic reference thresholds; generate a standardized interval for each parameter according to the dynamic reference thresholds.
[0043] S122. Obtain the smoothed water quality data; apply the Tukey criterion to calculate the outlier limits of each parameter, set the inner limit coefficient to 1.5 and the outer limit coefficient to 3.0; identify the data points exceeding the limit and mark them as outliers; correct the outliers using the local linear interpolation method to generate outlier corrected data.
[0044] S123. Read the outlier correction data and the normalization interval; construct an adaptive normalization function, where the slope of the function is dynamically adjusted according to the data distribution density; set the normalization curve in the data-intensive area to an S shape and the sparse area to a linear shape; map all parameters to the [0, 1] interval through this function to generate preliminary normalized water quality data.
[0045] S124. Obtain the preliminary normalized water quality data; calculate the skewness and kurtosis of the data distribution of each parameter; when the absolute value of the skewness is greater than 0.5 or the kurtosis value exceeds the range of [2, 4], apply the Box-Cox transformation for correction; perform variance normalization on the corrected data to generate the final normalized water quality data.
[0046] In this embodiment, by reading the standard range thresholds of each water quality parameter from the system configuration database and performing a weighted average of 3:7 on its statistical characteristics with historical data, a dynamic reference threshold is generated, which solves the problem that fixed thresholds are difficult to adapt to the fluctuations in industrial wastewater quality. The Tukey criterion is used to identify outliers in the data, with an inner limit coefficient of 1.5 and an outer limit coefficient of 3.0, and combined with the local linear interpolation method for outlier correction, which not only ensures the continuity of the data but also avoids the influence of outliers on subsequent analysis. When constructing the adaptive normalization function, the slope of the function is dynamically adjusted according to the data distribution density. In the data-intensive area, an S-shaped curve is used to improve the resolution, and in the sparse area, a linear mapping is used to maintain the original distribution characteristics. This differential normalization strategy is particularly suitable for the uneven distribution characteristics of industrial wastewater water quality parameters. By calculating the skewness and kurtosis of the data distribution, when the absolute value of the skewness is greater than 0.5 or the kurtosis value exceeds the range of [2, 4], the Box-Cox transformation is applied for correction to ensure the statistical characteristics of the normalization result. This multi-level data normalization processing scheme not only solves the problems of frequent outliers and uneven distribution in industrial wastewater monitoring data but also provides a reliable data basis for subsequent feature extraction and pattern recognition.
[0047] According to one aspect of the present application, step S123 is further as follows:
[0048] S1231. Read the outlier correction data; classify according to the physicochemical properties of the water quality parameters; use piecewise linear mapping for the acidity and alkalinity; use logarithmic mapping for the concentration parameters; generate a parameter feature classification table.
[0049] S1232. Obtain the parameter feature classification table; calculate the data distribution characteristics of each type of parameter; determine the type of transformation function according to the skewness; generate a set of normalization functions.
[0050] S1233. Read the set of normalization functions and the outlier correction data; apply the corresponding normalization function to each type of parameter; perform interval mapping; generate preliminary normalized water quality data.
[0051] In this embodiment, the water quality parameters are classified according to their physical and chemical properties. For the pH value, a piecewise linear mapping is adopted, and for the concentration parameters, a logarithmic mapping is used to establish a special parameter feature classification table. According to the data distribution characteristics of each type of parameter, the type of transformation function is determined. In particular, the shape of the normalization function is adjusted according to the skewness to ensure the rationality of the normalization result. This differential normalization strategy fully considers the physical and chemical properties and numerical distribution laws of different water quality parameters, avoiding the information loss problem caused by the traditional unified normalization method. Through this refined data processing scheme, not only the physical meaning of the data is maintained, but also the accuracy of subsequent analysis is improved, which is particularly suitable for the monitoring scenario of industrial wastewater with diverse parameter types and uneven numerical distributions.
[0052] According to one aspect of the present application, step S13 is further as follows:
[0053] S131. Read the normalized water quality data; calculate the pairwise covariance between each parameter to generate an initial covariance matrix; perform a condition number analysis on the covariance matrix, and when the condition number is greater than 1000, perform regularization processing with a regularization coefficient of 0.01; generate an optimized feature covariance matrix.
[0054] S132. Obtain the feature covariance matrix; use an improved QR decomposition algorithm to calculate the eigenvalues, set the convergence threshold to 1e-6; arrange the eigenvalues in descending order; calculate the cumulative contribution rate to generate an eigenvalue sequence and the corresponding eigenvector set.
[0055] S133. Read the eigenvalue sequence and the eigenvector set; select the first K principal components (K ≤ 5) according to the 85% cumulative contribution rate threshold; orthogonalize the selected eigenvectors; generate a water quality feature vector with a dimension of K × 3 through linear combination.
[0056] In this embodiment, an initial covariance matrix is constructed by calculating the pairwise covariance between each water quality parameter, and a condition number analysis is performed on it. When the condition number is greater than 1000, a regularization coefficient of 0.01 is used for processing, effectively avoiding the numerical instability problem in the feature extraction process. An improved QR decomposition algorithm is used to calculate the eigenvalues, and a convergence threshold of 1e-6 is set to ensure the accuracy of the eigenvalue decomposition result. The first K principal components (K ≤ 5) are selected according to the 85% cumulative contribution rate threshold, and the selected eigenvectors are orthogonalized to generate a water quality feature vector with a dimension of K × 3, realizing the dimensionality reduction of the data while retaining the main feature information. This principal component analysis method based on regularization is particularly suitable for the monitoring data of industrial wastewater with strong correlations between parameters. Through the accurate calculation of eigenvalues and eigenvectors, the internal relationships between water quality parameters can be accurately captured, providing a reliable feature expression for the subsequent verification of data authenticity.
[0057] According to one aspect of the present application, as Figure 3 shown, the steps for standardizing monitoring data include:
[0058] Dynamically adjust the slope of a pre-configured normalization function according to the data distribution density of the monitoring data;
[0059] Based on the adjusted slope, perform normalization on the data-intensive area using an S-shaped curve; based on the adjusted slope, perform normalization on the data-sparse area using linear mapping; generate preliminary standardized water quality data;
[0060] Based on the preliminary standardized water quality data, calculate the skewness and kurtosis of the data distribution, and perform correction using the Box-Cox transformation to generate standardized water quality data.
[0061] According to one aspect of the present application, the steps for extracting water quality feature vectors based on the standardized water quality data include:
[0062] Based on the standardized water quality data, construct a covariance matrix and perform condition number analysis; when the condition number is less than a preset threshold, directly use the covariance matrix as the optimized feature covariance matrix; when the condition number is greater than the preset threshold, perform regularization on the covariance matrix to obtain the optimized feature covariance matrix;
[0063] Based on the optimized feature covariance matrix, use an improved QR decomposition algorithm to calculate eigenvalues, generating an eigenvalue sequence and a corresponding set of eigenvectors;
[0064] Based on the eigenvalue sequence and the corresponding set of eigenvectors, select principal components according to a preset cumulative contribution rate threshold to obtain the selected eigenvectors;
[0065] Orthogonalize the selected eigenvectors to generate water quality feature vectors.
[0066] According to one aspect of the present application, the steps for calculating eigenvalues using an improved QR decomposition algorithm include:
[0067] Set a weight factor according to the importance degree of pollutant parameters in the monitoring data;
[0068] Decompose the optimized feature covariance matrix into an upper triangular matrix and an orthogonal matrix;
[0069] Iteratively optimize the decomposition process based on the weight factor;
[0070] Calculate the change rates of eigenvalues and eigenvectors during the iteration process, and dynamically adjust the calculation step size;
[0071] When the change rates of both the eigenvalue and the eigenvector are less than the corresponding preset thresholds, determine the final eigenvalue;
[0072] Arrange the final eigenvalues in descending order, calculate the cumulative contribution rate of the final eigenvalues, and generate an eigenvalue sequence and the corresponding eigenvector set.
[0073] For the eigenvalue calculation scenario of industrial wastewater monitoring data, this implementation uses an improved QR decomposition algorithm. By introducing a weight adjustment mechanism based on pollutant characteristics, the QR decomposition process better adapts to the characteristics of water quality data. Specifically, for the pollutant parameters under key monitoring (such as COD, ammonia nitrogen, etc.), higher weight factors are assigned. These weight factors act on the decomposition process of the eigen covariance matrix, making the decomposition result better reflect the change characteristics of key pollutants. At the same time, the algorithm also designs a dynamic step size adjustment mechanism, that is, during the iteration process, the calculation step size is adaptively adjusted according to the change rate of the eigenvalue, and a dual-index convergence judgment strategy for eigenvalues and eigenvectors is adopted. This improvement ensures that while maintaining numerical stability, it can more accurately capture the mutation information in water quality data. On the one hand, this embodiment improves the sensitivity to the changes in key pollutant parameters, making the eigenvalue calculation results better reflect water quality anomalies; on the other hand, through the dynamic step size and dual-index convergence judgment, it optimizes the algorithm efficiency while ensuring the calculation accuracy, and is especially suitable for scenarios that require rapid response in industrial wastewater monitoring.
[0074] According to one aspect of the present application, step S2 is further as follows:
[0075] S21. Sequentially read the historical data sequence of the most recent 24 hours from the data storage unit; obtain the water quality eigenvector; input the historical data sequence and the water quality eigenvector into a long short-term memory network with 64 hidden layer nodes; extract the time series features through the network to generate a time feature vector with a dimension of 32.
[0076] S22. Obtain the adjacent monitoring point data of the adjacent 5 monitoring points from the communication module; read the water quality eigenvector; calculate the Moran index and Geary index between the adjacent monitoring point data and the water quality eigenvector; generate a spatial feature vector with a dimension of 32 according to the calculation results.
[0077] S23. Obtain the time feature vector and the spatial feature vector; construct a third-order tensor and perform CP decomposition; select the decomposition result with a rank of 10; reorganize the decomposition result into a 32×32 spatio-temporal fusion feature matrix.
[0078] In one embodiment of the present application, based on the water quality feature vector Wf and the historical data sequence H, a long short-term memory network is applied to extract temporal features, and a temporal feature vector Tf is obtained. Based on the water quality feature vector Wf and the data N of adjacent monitoring points, a spatial autocorrelation algorithm is used to analyze the spatial consistency of the data, and a spatial feature vector Sf is output. Based on the temporal feature vector Tf and the spatial feature vector Sf, a tensor decomposition algorithm with adaptive weights is used for feature fusion, and a spatio-temporal fusion feature matrix M is output.
[0079] This embodiment analyzes the 24-hour historical data sequence and combines it with the network structure of 64 hidden layer nodes, which can effectively capture the long-term change trend and short-term fluctuation characteristics of water quality parameters. Especially when dealing with monitoring data with obvious temporal characteristics such as industrial wastewater, by calculating the Moran index and Geary index of adjacent monitoring points, the spatial autocorrelation of the data can be accurately evaluated, so as to identify abnormal discharge behaviors. The tensor decomposition algorithm is used to fuse the temporal feature vector and the spatial feature vector, and the decomposition result with a rank of 10 is selected to reconstruct a 32×32 feature matrix, which not only retains the spatio-temporal correlation characteristics of the data but also realizes the dimensionality reduction and compression of the features. This verification mechanism based on spatio-temporal feature fusion can consider both the temporal continuity and spatial consistency of the monitoring data, and is particularly suitable for monitoring scenarios with complex pipe network structures in industrial parks. Through the spatial autocorrelation analysis of the data of adjacent monitoring points, local abnormal discharges or data tampering behaviors can be effectively identified, improving the reliability and accuracy of the verification method.
[0080] According to one aspect of the present application, step S21 is further as follows:
[0081] S211. Read the historical data sequence of the most recent 24 hours in chronological order; segment the sequence with a sliding window, the window size is set to 30 minutes, and the step size is 5 minutes; calculate the statistical features within each window, including mean, variance, kurtosis, and skewness; generate temporal statistical features.
[0082] S212. Obtain the water quality feature vector and the temporal statistical features; align and merge the two sets of data in time; perform feature extraction through a 64-node LSTM network, and the forgetting gate threshold is set to 0.3; generate the original temporal features.
[0083] S213. Read the original temporal features; apply the backpropagation algorithm to optimize the feature weights, and the learning rate is set to 0.01; reduce the dimensionality of the features to 32 dimensions; generate the temporal feature vector through normalization processing.
[0084] In this embodiment, by adaptively segmenting the 24-hour historical data sequence, combining with an LSTM network with 64 nodes for feature extraction, and setting a forgetting gate threshold of 0.3, the long-term dependence relationship of the data sequence can be accurately captured. Through backpropagation optimization with a learning rate of 0.01, the features are reduced to 32 dimensions and normalized, which not only retains the key features of the time-series data but also realizes the compression of features. This time-series feature extraction method based on an adaptive window is particularly suitable for monitoring objects such as industrial wastewater with obvious time-varying characteristics, can accurately identify changes in the discharge conditions, and provides an important basis for the detection of abnormal discharges.
[0085] According to one aspect of the present application, step S211 is further as follows:
[0086] S2111. Read the historical data sequence; calculate the first-order difference value of the data, and determine the data change rate; mark the inflection points of the data change according to the size of the difference value; generate a data change feature sequence.
[0087] S2112. Obtain the data change feature sequence; calculate the time interval between adjacent inflection points; determine the reference window size according to the distribution characteristics of the intervals; generate window reference parameters.
[0088] S2113. Read the window reference parameters and the data change feature sequence; when the data change rate exceeds the threshold, reduce the window size to 1 / 2 of the reference value; when the data is stable, increase the window size to 2 times the reference value; generate an adaptive window sequence.
[0089] S2114. Obtain the historical data sequence and the adaptive window sequence; calculate the statistical features within each window; generate time-series statistical features.
[0090] In this embodiment, by calculating the first-order difference value of the historical data sequence and marking the inflection points of the data change, and determining the reference window size in combination with the distribution characteristics of the time intervals between adjacent inflection points, the adaptive adjustment of the window size is realized. When the data change rate exceeds the threshold, the window size is automatically reduced to 1 / 2 of the reference value, and when the data is stable, it is increased to 2 times the reference value. This dynamic adjustment mechanism is particularly suitable for the intermittent discharge characteristics of industrial wastewater. By calculating statistical features such as mean, variance, kurtosis, and skewness within each adaptive window, not only the main features of the data are retained but also the noise interference can be effectively filtered. This embodiment solves the problem that traditional fixed windows are prone to splitting the discharge process, and improves the accuracy and reliability of time-series feature extraction.
[0091] According to one aspect of the present application, step S22 is further as follows:
[0092] S221. Obtain the adjacent monitoring point data of the adjacent 5 monitoring points from the communication module in real time; construct a spatial weight matrix according to the physical distance between the monitoring points, and the weight decays exponentially with the distance; perform row normalization on the weight matrix; generate a normalized weight matrix.
[0093] S222. Obtain the water quality feature vector and the normalized weight matrix; calculate the local Moran's I index with the spatial lag order set to 2; calculate the local Geary's C index with the neighborhood range set to 1000 meters; generate spatial correlation indicators.
[0094] S223. Read the spatial correlation indicators; construct a 32-dimensional spatial feature mapping function; map the spatial correlation indicators to the feature space; optimize the mapping parameters by the least squares method to generate a spatial feature vector.
[0095] In this embodiment, an exponential decay function is used to construct the spatial weight matrix. Compared with the traditional fixed threshold method, it can more precisely reflect the non-linear distance effect between adjacent monitoring points, especially capture the strong correlation at medium and short distances. Through the dual processing of physical distance weights and row normalization, not only the magnitude information of the original spatial relationship is retained, but also the interference of inter-row differences on subsequent analysis is eliminated, making different monitoring units comparable. At the same time, the local Moran's I index and the local Geary's C index are calculated. The former detects spatial aggregation characteristics (such as pollution diffusion hotspots), and the latter identifies spatial outliers (such as sudden pollution sources), forming a three-dimensional analysis framework for mutual verification. Through the dual definition of the 2nd-order spatial lag and the 1000-meter neighborhood, both the topological adjacent relationship (network propagation path) and the actual geographical range are considered, and a multi-dimensional spatial correlation network is constructed. The 32-dimensional spatial feature mapping function breaks through the dimensionality limitation of traditional spatial statistical indicators, can transform discrete spatial correlation indicators into continuously distributed high-dimensional representations, and provides an adapted input structure for deep learning models. The application of the least squares method in mapping parameter optimization ensures the physical relevance between the feature space and the target variable (such as water quality level) by minimizing the loss function, avoiding feature distortion caused by simple mathematical transformations. Compared with the traditional water quality monitoring scheme, this embodiment reduces the spatial analysis error by about 37% (simulated test data), and at the same time improves the feature representation efficiency by 2.8 times, providing a more accurate spatial decision-making basis for water quality anomaly warning and pollution source tracing.
[0096] According to one aspect of the present application, step S221 is further as follows:
[0097] S2211. Read the pipe network topology data from the industrial park pipe network database; extract the upstream contribution coefficient and downstream influence coefficient of each monitoring point; calculate the flow capacity coefficient according to the pipe diameter, flow velocity and hydraulic gradient; generate a pipe network feature vector.
[0098] S2212. Read the network feature vector; construct the pipe segment flow direction matrix, where the upstream node takes a positive value of 1, the downstream node takes a negative value of 1, and the value is 0 for no direct connection; calculate the shortest network path between nodes; generate the flow direction correlation matrix.
[0099] S2213. Obtain the network feature vector and the flow direction correlation matrix; calculate the diffusion attenuation coefficient of pollutants at different flow rates; calculate the pollutant transfer time between nodes based on the diffusion equation; generate the spatio-temporal transfer matrix.
[0100] S2214. Read the spatio-temporal transfer matrix; according to the network flow direction, assign a higher weight (0.6 - 0.8) to the upstream monitoring points, a medium weight (0.3 - 0.5) to the same-level monitoring points, and a lower weight (0.1 - 0.3) to the downstream monitoring points; perform a non-linear adjustment on the weights, where the adjustment coefficient is proportional to the flow rate and pipe diameter; generate the initial weight matrix.
[0101] S2215. Obtain the initial weight matrix; apply row normalization; dynamically correct the weights according to the real-time flow rate; generate the final standardized weight matrix.
[0102] In this embodiment, by reading the network topology data from the network database, extracting the upstream contribution coefficient and downstream influence coefficient of each monitoring point, and combining the pipe diameter size, flow rate, and hydraulic gradient to calculate the flow capacity coefficient, a complete network feature vector is constructed. According to the pipe segment flow direction, a flow direction matrix is constructed, with the upstream node taking a positive value of 1 and the downstream node taking a negative value of 1, and the shortest network path between nodes is calculated, accurately reflecting the pollutant transmission path. Based on the diffusion equation, the diffusion attenuation coefficient of pollutants at different flow rates and the transfer time between nodes are calculated to generate the spatio-temporal transfer matrix. When allocating weights, according to the characteristics of the network flow direction, a high weight of 0.6 - 0.8 is assigned to the upstream monitoring points, a medium weight of 0.3 - 0.5 to the same-level monitoring points, and a low weight of 0.1 - 0.3 to the downstream monitoring points, and dynamic correction is performed according to the real-time flow rate. This spatial correlation analysis method based on the network topology fully considers the characteristics of the complex network structure in the industrial park. By establishing a pollutant transmission model and a dynamic weight system, it can accurately evaluate the spatial consistency of monitoring data and effectively identify illegal emissions or data tampering behaviors.
[0103] According to one aspect of the present application, step S23 is further as follows:
[0104] S231. Read the time feature vector and the spatial feature vector; construct a third-order tensor with a dimension of 32×32×10; perform normalization processing on the tensor; generate the initial feature tensor.
[0105] S232. Obtain the initial feature tensor; apply the CP decomposition algorithm, set the maximum number of iterations to 200; select the decomposition result with a rank of 10; generate a group of decomposed feature matrices.
[0106] S233. Read the decomposed feature matrix group; evaluate the reconstruction error through the Tucker norm; select the feature combinations with a reconstruction error less than 0.01; and recombine them into a 32×32 spatio-temporal fusion feature matrix.
[0107] In this embodiment, a third-order tensor with a dimension of 32×32×10 is constructed to fuse the time feature vector and the space feature vector. The CP decomposition algorithm is used for tensor decomposition, with a maximum number of iterations set to 200, and the decomposition result with a rank of 10 is selected. The Tucker norm is used to evaluate the reconstruction error, and the feature combinations with a reconstruction error less than 0.01 are selected and recombined into a 32×32 spatio-temporal fusion feature matrix. This spatio-temporal feature fusion method based on tensor decomposition not only maintains the independence of time features and space features but also realizes the organic fusion of the two features. Especially in the monitoring scenario of industrial wastewater with obvious spatio-temporal correlation, the coupling relationship of spatio-temporal features can be effectively extracted through tensor decomposition, providing a more comprehensive feature expression for data authenticity verification. At the same time, through the control of the reconstruction error, the accuracy and reliability of the feature fusion process are ensured.
[0108] As Figure 2 shown, according to one aspect of the present application, the steps of constructing a standardized weight matrix based on pre-stored pipe network topology data include:
[0109] Read the pipe network topology data from the industrial park pipe network database, and extract the upstream contribution coefficient and downstream influence coefficient of each monitoring point based on the pipe network topology data;
[0110] Calculate the flow capacity coefficient according to the pre-stored pipe diameter, flow velocity, and hydraulic gradient, and generate a pipe network feature vector based on the upstream contribution coefficient, downstream influence coefficient, and flow capacity coefficient;
[0111] Based on the pipe network feature vector, calculate the shortest pipe network path between nodes, and construct a flow direction correlation matrix, where the upstream node takes a positive value of 1, the downstream node takes a negative value of 1, and the value is 0 for no direct connection;
[0112] Based on the pipe network feature vector and the flow direction correlation matrix, calculate the diffusion attenuation coefficient of pollutants at different flow velocities;
[0113] Based on the diffusion attenuation coefficient, calculate the pollutant transfer time between nodes through the diffusion equation to generate a spatio-temporal transfer matrix;
[0114] Assign weights to the monitoring points according to the spatio-temporal transfer matrix and perform non-linear adjustment. The coefficient of non-linear adjustment is proportional to the flow velocity and pipe diameter to obtain a standardized weight matrix.
[0115] As Figure 4As shown, according to one aspect of the present application, when constructing the spatial feature vector, the time feature vector is constructed as follows:
[0116] Read the historical data sequence of a predetermined time from the data storage unit;
[0117] Based on the historical data sequence, calculate the data change rate and determine the data change inflection point;
[0118] Determine the reference window size according to the time interval between adjacent data change inflection points;
[0119] When the data change rate exceeds the preset threshold, reduce the reference window size to 1 / a of the original; when the data change rate does not exceed the preset threshold, increase the reference window size to a times the original; obtain the adjusted window size;
[0120] Based on the historical data sequence and the adjusted window size, calculate the statistical features within each window to generate the time series statistical features;
[0121] Based on the time series statistical features and the water quality feature vector, generate the time feature vector; where a is a natural number greater than 0.
[0122] According to one aspect of the present application, the process of generating the spatio-temporal fusion feature matrix further includes:
[0123] Perform tensor decomposition on the time feature vector and the spatial feature vector to generate the spatio-temporal fusion feature matrix;
[0124] The steps of tensor decomposition include:
[0125] Construct a third-order tensor with dimensions M×N×K;
[0126] Perform CP decomposition on the third-order tensor and select the decomposition result with rank P;
[0127] Based on the decomposition result, evaluate the reconstruction error through the Tucker norm and select the feature combination with the reconstruction error less than the preset threshold;
[0128] Recombine the feature combination into the spatio-temporal fusion feature matrix, where M, N, K, and P are preset values.
[0129] According to one aspect of the present application, step S3 is further as follows:
[0130] S31. Read the fingerprint feature code and the spatio-temporal fusion feature matrix; collect parameters such as temperature, humidity, and air pressure from the environmental sensor to form the environmental parameter vector; perform timestamp alignment and dimension unification processing on these three groups of data to generate a 64×64 aligned feature matrix.
[0131] S32. Obtain the aligned feature matrix; construct 8 attention heads, each with a dimension of 8; calculate the feature weights through multi-head attention; perform weighted combination of the features according to the weights to generate a 128-dimensional fused feature matrix.
[0132] S33. Read the fused feature matrix, normalized water quality data, and sensor dynamic fingerprints; perform data encoding using the Reed-Solomon coding algorithm; convert the encoding result into a 29×29 encrypted QR code.
[0133] In an embodiment of the present application, based on the fingerprint feature code E, spatio-temporal fusion feature matrix M, and environmental parameter vector P, a multi-modal feature alignment algorithm is used for feature mapping to obtain the aligned feature matrix A; based on the aligned feature matrix A, an attention mechanism is adopted for feature weighted fusion to output the fused feature matrix R. Based on the fused feature matrix R, normalized water quality data Wn, and sensor dynamic fingerprints D, a high-density QR code is generated using the Reed-Solomon coding algorithm to obtain the encrypted QR code Q.
[0134] This embodiment unifies the fingerprint feature code, spatio-temporal fusion feature matrix, and environmental parameter vector into the same feature space to generate a 64×64 aligned feature matrix, solving the problem of multi-source heterogeneous data fusion. 8 attention heads are used for feature weighting, each with a dimension of 8. Through the multi-head attention mechanism, feature weights can be assigned from different angles, improving the robustness of feature fusion. The Reed-Solomon coding algorithm is used to generate a 29×29 encrypted QR code, which not only realizes high-density storage of data but also provides 30% error correction ability, ensuring the recognition reliability in an industrial environment. This multi-modal feature fusion scheme constructs a multi-level data authenticity verification system by making full use of sensor physical features, monitoring data features, and environmental parameter features. Especially in scenarios such as industrial wastewater monitoring that require high reliability, through cross-verification of multi-source data, single-point data tampering can be effectively prevented, improving the security and credibility of the entire monitoring system.
[0135] According to one aspect of the present application, step S32 is further as follows:
[0136] S321. Read the aligned feature matrix; divide the matrix into 8 subspaces, each with a dimension of 8; calculate the query matrix Q, key matrix K, and value matrix V for each subspace; generate a group of attention calculation matrices.
[0137] S322. Obtain the group of attention calculation matrices; calculate the scaled dot-product attention, with the scaling factor set to the square root of 8; apply the Softmax function to calculate the weight distribution; generate the multi-head feature matrix.
[0138] S323. Read the multi-head feature matrix; convert the feature dimension to 128 through a linear transformation layer; apply residual connection and layer normalization; generate a fused feature matrix.
[0139] In this embodiment, the aligned feature matrix is divided into 8 subspaces with each subspace dimension being 8, and the query matrix Q, key matrix K, and value matrix V are calculated to construct a complete multi-head attention calculation system. The square root of 8 is used as the scaling factor to calculate the scaled dot-product attention, and the Softmax function is applied to calculate the weight distribution, ensuring the reasonable distribution of attention weights. The feature dimension is converted to 128 through a linear transformation layer, and residual connection and layer normalization are applied to generate the final fused feature vector. This feature fusion method based on the multi-head attention mechanism can capture the correlation relationships between features from different perspectives, and is particularly suitable for monitoring data of industrial wastewater with multi-dimensional features. Through the design of residual connection, both the original feature information is retained, and the expression ability of the model is enhanced, improving the accuracy and stability of feature fusion.
[0140] According to one aspect of the present application, step S33 is further as follows:
[0141] S331. Obtain three groups of input data; compress the fused feature matrix to 64 dimensions; merge all data into an ordered byte stream; generate encoded preprocessing data.
[0142] S332. Read the encoded preprocessing data; perform encoding using the Reed-Solomon algorithm with the error correction level set to level H (30%); add parity bits and synchronization codes; generate an encoded data stream.
[0143] S333. Obtain the encoded data stream; perform data layout according to the 29×29 specification; add positioning patterns and calibration patterns; generate an encrypted QR code.
[0144] In this embodiment, the fused feature vector is compressed to 64 dimensions, and combined with sensor fingerprints and water quality data to generate an encoded preprocessing data stream. Encoding is performed using the Reed-Solomon algorithm, with the error correction level set to level H (30%), and parity bits and synchronization codes are added, finally generating an encrypted QR code with a 29×29 specification. This encoding scheme based on parameter characteristics improves the encoding efficiency and is particularly suitable for monitoring data of industrial wastewater with multiple parameters and different value ranges. Through the design of a high error correction level, the recognition reliability in a harsh industrial environment is ensured, providing a strong guarantee for the secure transmission and verification of data.
[0145] According to one aspect of the present application, step S332 is further as follows:
[0146] S3321. Read the preprocessed encoded data; map the pH value (0 - 14) to a 4-bit code; divide other parameters into different code lengths according to their value ranges; generate a parameter code mapping table.
[0147] S3322. Obtain the parameter code mapping table; allocate the error correction bits of the Reed - Solomon code according to the importance of different parameters; perform double verification on key parameters; generate an error correction coding rule.
[0148] S3323. Read the preprocessed encoded data and the error correction coding rule; perform hierarchical coding; add a parameter type identifier; generate an encoded data stream.
[0149] In this embodiment, by establishing a targeted parameter code mapping table, mapping the pH value (0 - 14) to a 4-bit code, and dividing other parameters into different code lengths according to their value ranges, the coding efficiency is improved. According to the importance of different parameters, dynamically allocate the error correction bits of the Reed - Solomon code, and adopt a double verification mechanism for key parameters to ensure the reliability of data transmission. When performing hierarchical coding, by adding a parameter type identifier, the rapid identification and parsing of parameters are realized. This embodiment not only improves the information density of the two-dimensional code but also enhances the anti-interference ability through reasonable allocation of error correction bits, and is especially suitable for use in complex industrial environments.
[0150] According to one aspect of the present application, the steps for generating an encrypted two-dimensional code include:
[0151] Based on the spatio-temporal fusion feature matrix and the real-time collected environmental parameter vector, generate a fusion feature matrix; combine the fusion feature matrix and the standardized water quality data to generate preprocessed encoded data;
[0152] Divide different code lengths according to the preprocessed encoded data to generate a parameter code mapping table;
[0153] Based on the importance of different parameters in the parameter code mapping table, allocate different Reed - Solomon code error correction bits to each parameter to obtain an allocation result; perform double verification coding on the key parameters in the allocation result to generate an error correction coding rule;
[0154] Based on the preprocessed encoded data and the error correction coding rule, perform hierarchical coding and add a parameter type identifier to generate an encoded data stream;
[0155] Layout the encoded data stream according to a predetermined specification to generate an encrypted two-dimensional code.
[0156] According to one aspect of the present application, step S4 is further:
[0157] S41. Read the fused feature matrix; extract the historical confidence data for the most recent 30 days from the database; construct a factor graph network, and set the number of message propagation iterations to 50; calculate the node marginal probability through message propagation; generate a data confidence score ranging from [0, 1] based on the marginal probability value.
[0158] S42. Obtain the data confidence score; read the historical anomaly pattern library containing 100 typical anomaly patterns from the system database; construct a graph neural network with 3 hidden layers (the number of nodes is 64, 32, and 16 respectively); obtain the anomaly type identifier through network matching analysis; calculate the anomaly degree value ranging from [0, 1] based on the matching similarity.
[0159] S43. Read the anomaly type identifier and the anomaly degree value; load 5 levels of warning rule templates from the rule library; map the anomaly type identifier to a warning level between 1 and 5 according to the rule template; generate warning description information including the anomaly type, degree, and recommended measures from the predefined information template library based on the warning level and the anomaly degree value.
[0160] In an embodiment of the present application, based on the fused feature matrix R and the historical confidence data C, the belief propagation algorithm is applied to calculate the data credibility, and a data confidence score S is obtained. Based on the data confidence score S and the historical anomaly pattern library P, a graph neural network based on the attention mechanism is used for anomaly pattern matching, and the anomaly type identifier L and the anomaly degree value G are output. Receive the anomaly type identifier L and the anomaly degree value G, and generate warning information based on the predefined rule template to obtain the warning level Y and the warning description information I.
[0161] This embodiment analyzes the fused feature vector and the historical confidence data, constructs a factor graph network, calculates the node marginal probability through 50 iterations of message propagation, and obtains a data confidence score ranging from [0, 1]. A three-layer graph neural network (the number of nodes is 64, 32, and 16 respectively) is used to perform anomaly pattern matching on the data. Combining with the library of 100 typical anomaly patterns, the anomaly type can be accurately identified and the anomaly degree can be quantified. According to the predefined 5-level warning rule template, combining the anomaly type identifier and the anomaly degree value, warning information including specific measure suggestions is generated. This evaluation method based on belief propagation and graph neural network can not only evaluate the credibility of data in real time, but also detect and warn of abnormal situations in a timely manner. Especially in the scenario of industrial wastewater monitoring involving multiple stakeholders, through objective confidence scoring and a standardized warning mechanism, it can provide a reliable decision-making basis for regulatory authorities and also provide clear rectification guidance for enterprises.
[0162] According to one aspect of the present application, step S41 is further as follows:
[0163] S411. Obtain the fused feature matrix and the historical confidence data for the most recent 30 days; construct a factor graph with a bipartite graph structure; define variable nodes and factor nodes; generate an initial factor graph.
[0164] S412. Read the initial factor graph; design message update rules, including message passing equations from variables to factors and from factors to variables; initialize node messages; generate a message propagation rule set.
[0165] S413. Obtain the message propagation rule set; perform message propagation iterations, with a maximum of 50 iterations; calculate node belief values; generate a node confidence distribution.
[0166] S414. Read the node confidence distribution; calculate marginal probabilities; apply the sigmoid function to map to the [0, 1] interval; generate a data confidence score.
[0167] In this embodiment, by constructing a factor graph with a bipartite graph structure, taking the fused feature vector and 30-day historical confidence data as inputs, defining the message passing rules between variable nodes and factor nodes, and achieving the full propagation of messages through 50 iterations of optimization. During the message update process, the message passing equations from variables to factors and from factors to variables are designed respectively to ensure the effective flow of information in the network. By calculating node belief values and using the sigmoid function to map marginal probabilities to the [0, 1] interval, a final data confidence score is generated. This belief propagation method based on the factor graph makes full use of the statistical laws contained in historical data and can accurately evaluate the credibility of current data. Especially in the monitoring scenario of industrial wastewater with large data fluctuations, through the modeling of the probabilistic graph model, normal fluctuations and abnormal changes can be effectively distinguished, providing a reliable evaluation mechanism for data authenticity verification.
[0168] According to one aspect of the present application, step S42 is further as follows:
[0169] S421. Read the data confidence score and 100 typical abnormal patterns; construct a three-layer graph neural network with 64, 32, and 16 nodes; initialize network parameters; generate a graph network model.
[0170] S422. Obtain the graph network model; perform graph convolution operations, using the ReLU activation function; extract graph feature representations; generate abnormal pattern features.
[0171] S423. Read the abnormal pattern features; calculate the cosine similarity with the typical patterns; select the type with the highest similarity as the abnormal type identifier; map the maximum similarity value to the abnormal degree value.
[0172] In this embodiment, a three-layer graph neural network (with 64, 32, and 16 nodes respectively) is constructed to process the confidence data, and it is matched with an anomaly library containing 100 typical patterns. The ReLU activation function is used during the graph convolution operation, and the extracted graph feature representation can effectively capture the structural features of data anomaly patterns. By calculating the cosine similarity with the typical patterns, the type with the highest similarity is selected as the anomaly type identifier, and the maximum similarity value is mapped to the anomaly degree value. This anomaly recognition method based on graph neural network not only considers the temporal features of the monitoring data but also incorporates spatial topological information, and is particularly suitable for monitoring scenarios with complex pipe network structures in industrial parks. By comparing with a rich anomaly pattern library, it can not only accurately identify the anomaly type but also quantify the anomaly degree, providing a reliable basis for subsequent early warning decisions.
[0173] According to one aspect of the present application, the steps of calculating the data confidence score include:
[0174] Obtain the fusion feature matrix in the encrypted QR code and the pre-stored historical confidence data; construct a factor graph with a bipartite graph structure, define variable nodes and factor nodes, and generate an initial factor graph;
[0175] Based on the initial factor graph, construct message update rules, including message passing equations from variables to factors and from factors to variables;
[0176] Based on the message update rules, perform message propagation iterations and calculate the node belief values to obtain the node confidence distribution; where the maximum number of iterations for performing message propagation iterations is c; c is a preset value;
[0177] Calculate the marginal probability based on the node confidence distribution and map it to the [0, 1] interval through the sigmoid function to obtain the data confidence score.
[0178] According to one aspect of the present application, the steps of generating corresponding early warning description information include:
[0179] Construct a graph neural network with three hidden layers;
[0180] Input the data confidence score into the graph neural network, perform graph convolution operations, and generate anomaly pattern features;
[0181] Calculate the cosine similarity between the anomaly pattern features and D pre-stored typical anomaly patterns; where D is a preset value;
[0182] Select the type with the highest cosine similarity as the anomaly type identifier, and map the maximum similarity value in the cosine similarity to the anomaly degree value;
[0183] Based on the anomaly type identifier and the anomaly degree value, generate an early warning level from a predefined rule template;
[0184] Generate warning description information including abnormal type, abnormal degree, and recommended measures from a predefined information template library according to the warning level.
[0185] Existing data authenticity verification technologies lack a tamper-proof mechanism based on physical characteristics at the sensor level, and existing digital encryption schemes cannot prevent attackers from fabricating data by simulating sensor signals; in the data preprocessing link, the traditional fixed-window smoothing method is prone to artificially splitting the emission process, resulting in the loss of key features; in the feature extraction process, the physical and chemical characteristics of industrial wastewater quality parameters are not fully considered, and the unified standardization processing method is prone to weakening or losing the feature information of some parameters; in the spatial correlation analysis, only the physical distance between monitoring points is considered, ignoring the impact of the pipe network structure on pollutant transmission, resulting in inaccurate extraction of spatial features; in the multi-modal feature fusion link, existing methods often use simple feature splicing or weighted averaging, making it difficult to capture the complex correlation relationships between different features; in the time-series data analysis, the fixed feature extraction window is difficult to adapt to the intermittent emission characteristics of industrial wastewater and is prone to missing key change features.
[0186] Regarding the physical anti-counterfeiting problem of sensors, by collecting three physical quantities: the working voltage fluctuation value, temperature response characteristics, and medium impedance of the sensor, and generating a dynamic fingerprint through Fourier transform, anti-counterfeiting at the hardware level is achieved. For the data segmentation problem, an adaptive window technology based on data change characteristics is adopted, and the window size is dynamically adjusted according to the data change rate, effectively avoiding the loss of key features. For the problem of parameter feature preservation, a differentiated standardization strategy is adopted according to the physical and chemical characteristics of different water quality parameters. For example, piecewise linear mapping is used for acidity and alkalinity, and logarithmic mapping is used for concentration parameters. Regarding the problem of pipe network influence, a flow matrix is constructed based on the pipe network topology, and the pollutant transfer time is calculated considering factors such as pipe diameter and flow velocity, achieving accurate extraction of spatial features. For the feature fusion problem, dynamic weight allocation of features is achieved through the multi-head attention mechanism, effectively capturing the complex correlations between different features.
[0187] This application collects the physical characteristics of sensors and generates dynamic fingerprints; performs adaptive preprocessing and feature extraction on water quality data; conducts integrity verification based on spatio-temporal feature sequences; generates encrypted QR codes through multi-modal feature fusion; evaluates data consistency and generates warning information. This method for verifying the authenticity of industrial wastewater monitoring data based on multi-modal feature fusion realizes the all-round guarantee of the authenticity of industrial wastewater monitoring data by constructing a complete technical chain from sensor physical feature extraction, data preprocessing, spatio-temporal feature verification, multi-modal feature fusion to confidence evaluation. Specifically, it is reflected in: firstly, a dynamic fingerprint mechanism based on sensor physical characteristics is proposed to ensure the credibility of data sources at the hardware level; an adaptive data preprocessing method suitable for the characteristics of industrial wastewater is designed to improve the recognition ability of abnormal features; a verification mechanism based on spatio-temporal feature fusion is realized, which can effectively identify local abnormal emissions or data tampering behaviors; a multi-head attention mechanism is used for feature fusion to improve the reliability of verification results; combined with belief propagation and graph neural networks, an objective data credibility evaluation system is constructed. The present invention constructs a multi-level and highly reliable data authenticity verification system, which is particularly suitable for complex monitoring scenarios such as industrial parks. Through the synergistic effect of physical feature verification, data feature verification and spatio-temporal feature verification, data tampering can be effectively prevented, abnormal emissions can be timely detected, and reliable technical support can be provided for environmental supervision.
[0188] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all belong to the protection scope of the present invention.
Claims
1. The authenticity verification method of industrial wastewater monitoring data based on multimodal feature fusion is characterized by: include: Acquire monitoring data, perform standardization on the monitoring data, and obtain standardized water quality data; extract water quality feature vectors based on the standardized water quality data; The monitoring data include pH value, COD value and turbidity; Acquire the data of adjacent monitoring points and construct a standardized weight matrix based on the pre-stored pipe network topology data; perform spatial autocorrelation analysis on the water quality feature vector based on the standardized weight matrix to obtain the spatial feature vector; input the spatial feature vector into the feature fusion unit to generate a spatiotemporal fusion feature matrix; Generate encrypted QR codes based on spatiotemporal fusion feature matrix and standardized water quality data; Receive the encrypted QR code, calculate the data confidence score, and generate the corresponding warning description information.
2. The method for verifying the authenticity of industrial wastewater monitoring data based on multimodal feature fusion according to claim 1 is characterized in that: The steps of constructing a standardized weight matrix based on pre-stored pipe network topology data include: Read the pipe network topology data from the industrial park pipe network database, and extract the upstream contribution coefficient and downstream influence coefficient of each monitoring point based on the pipe network topology data; Calculate the flow capacity coefficient according to the pre-stored pipe diameter, flow velocity and hydraulic slope, and generate the pipe network characteristic vector based on the upstream contribution coefficient, downstream influence coefficient and flow capacity coefficient; Based on the network feature vector, the shortest network path between nodes is calculated and the flow direction association matrix is constructed, where the upstream node takes a positive value of 1, the downstream node takes a negative value of 1, and no direct connection takes a value of 0; Based on the pipe network characteristic vector and flow direction correlation matrix, the diffusion attenuation coefficient of pollutants at different flow rates is calculated; Based on the diffusion attenuation coefficient, the pollutant transfer time between nodes is calculated through the diffusion equation to generate a space-time transfer matrix; The monitoring points are weighted according to the space-time transfer matrix and nonlinear adjustment is performed. The coefficient of the nonlinear adjustment is proportional to the flow velocity and the pipe diameter, and a standardized weight matrix is obtained.
3. The method for verifying the authenticity of industrial wastewater monitoring data based on multimodal feature fusion according to claim 1 is characterized in that: The steps for standardizing the monitoring data and obtaining standardized water quality data include: Dynamically adjust the slope of a pre-configured normalization function according to the data distribution density of the monitoring data; Based on the adjusted slope, the data-intensive area is normalized using an S-shaped curve; based on the adjusted slope, the data-sparse area is normalized using a linear mapping; preliminary standardized water quality data are generated; Based on the preliminary standardized water quality data, the skewness and kurtosis of the data distribution were calculated, and the Box-Cox transformation was used to correct them to generate standardized water quality data.
4. The method for verifying the authenticity of industrial wastewater monitoring data based on multimodal feature fusion according to claim 1 is characterized in that: When constructing the spatial feature vector, the steps to construct the temporal feature vector are as follows: Reading a predetermined time series of historical data from a data storage unit; Based on the historical data series, calculate the data change rate and determine the turning point of data change; Determine the size of the benchmark window according to the time interval between adjacent data change inflection points; When the data change rate exceeds the preset threshold, the benchmark window size is reduced to 1 / a of the original; when the data change rate does not exceed the preset threshold, the benchmark window size is increased to a times of the original; and the adjusted window size is obtained; Calculate the statistical features in each window based on the historical data sequence and the adjusted window size to generate time series statistical features; Based on the time series statistical characteristics and the water quality characteristic vector, a time characteristic vector is generated, where a is a natural number greater than 0.
5. The method for verifying the authenticity of industrial wastewater monitoring data based on multimodal feature fusion according to claim 4 is characterized in that: The process of generating the spatiotemporal fusion feature matrix also includes: Perform tensor decomposition on the temporal feature vector and the spatial feature vector to generate a temporal and spatial fusion feature matrix; The steps of tensor decomposition include: Construct a third-order tensor of dimension M×N×K; Perform CP decomposition on the third-order tensor and select the decomposition result with rank P; Based on the decomposition results, the reconstruction error is evaluated by Tucker norm, and the feature combination with a reconstruction error less than the preset threshold is selected; The feature combinations are reorganized into a spatiotemporal fusion feature matrix, where M, N, K, and P are preset values.
6. The method for verifying the authenticity of industrial wastewater monitoring data based on multimodal feature fusion according to claim 1 is characterized in that: The steps to generate an encrypted QR code include: Based on the spatiotemporal fusion feature matrix and the environmental parameter vector collected in real time, a fusion feature matrix is generated; the fusion feature matrix and the standardized water quality data are merged to generate coded preprocessed data; Divide different encoding lengths according to the encoding preprocessing data and generate a parameter encoding mapping table; Based on the importance of different parameters in the parameter coding mapping table, different Reed-Solomon coding error correction bits are allocated to each parameter to obtain an allocation result; double-check coding is performed on the key parameters in the allocation result to generate error correction coding rules; Performing layered encoding based on encoding preprocessing data and error correction encoding rules, and adding parameter type identifiers to generate an encoded data stream; The encoded data stream is laid out according to predetermined specifications to generate an encrypted QR code.
7. The method for verifying the authenticity of industrial wastewater monitoring data based on multimodal feature fusion according to claim 6 is characterized in that: The steps to calculate the data confidence score include: Obtain the fused feature matrix and pre-stored historical confidence data in the encrypted QR code; construct a factor graph of a bipartite graph structure, define variable nodes and factor nodes, and generate an initial factor graph; Based on the initial factor graph, construct the message update rules, including the message passing equations from variables to factors and factors to variables; Based on the message update rule, the message propagation iteration is performed, and the node belief value is calculated to obtain the node confidence distribution; the maximum number of iterations for executing the message propagation iteration is c; c is a preset value; The marginal probability is calculated based on the node confidence distribution and mapped to the [0,1] interval through the sigmoid function to obtain the data confidence score.
8. The method for verifying the authenticity of industrial wastewater monitoring data based on multimodal feature fusion according to claim 1 is characterized in that: The steps for generating the corresponding warning description information include: Build a graph neural network with three hidden layers; Input the data confidence score into the graph neural network, perform graph convolution operations, and generate abnormal pattern features; Calculate the cosine similarity between the abnormal pattern feature and the pre-stored D typical abnormal patterns; where D is a preset value; The type with the highest cosine similarity is selected as the abnormal type identifier, and the maximum similarity value in the cosine similarity is mapped to the abnormality degree value; Generate warning levels from predefined rule templates based on anomaly type identification and anomaly severity value; According to the warning level, warning description information including abnormality type, abnormality degree and recommended measures is generated from the predefined information template library.
9. The method for verifying the authenticity of industrial wastewater monitoring data based on multimodal feature fusion according to claim 1 is characterized in that: Based on the standardized water quality data, the steps of extracting the water quality feature vector include: Based on the standardized water quality data, the covariance matrix is constructed and the condition number analysis is performed; when the condition number is less than the preset threshold, the covariance matrix is directly used as the optimized characteristic covariance matrix; when the condition number is greater than the preset threshold, the covariance matrix is regularized to obtain the optimized characteristic covariance matrix; Based on the optimized eigencovariance matrix, the eigenvalues are calculated using the improved QR decomposition algorithm to generate eigenvalue sequences and corresponding eigenvector sets; Based on the eigenvalue sequence and the corresponding eigenvector set, the principal component is selected according to a preset cumulative contribution rate threshold to obtain the selected eigenvector; The selected eigenvectors are orthogonalized to generate water quality eigenvectors.
10. The method for verifying the authenticity of industrial wastewater monitoring data based on multimodal feature fusion according to claim 9 is characterized in that: The steps of calculating eigenvalues using the improved QR decomposition algorithm and generating eigenvalue sequences and corresponding eigenvector sets include: Set weight factors according to the importance of pollutant parameters in monitoring data; Decompose the optimized feature covariance matrix into an upper triangular matrix and an orthogonal matrix; Iteratively optimize the decomposition process based on weight factors; Calculate the rate of change of eigenvalues and eigenvectors during the iteration process, and dynamically adjust the calculation step size; When the rates of change of the eigenvalue and the eigenvector are both less than the corresponding preset thresholds, the final eigenvalue is determined; Arrange the final eigenvalues in descending order, calculate the cumulative contribution rate of the final eigenvalues, and generate an eigenvalue sequence and a corresponding eigenvector set.
Citation Information
Patent Citations
Intelligent design system based on AI
CN117407556A
Vision and language-based multi-modal mixed fusion fine-grained recognition method
CN118094172A
Intelligent automobile remote unlocking method and device based on multi-modal biological characteristics
CN118470836A
Indoor positioning method, system and device based on cross-modal fusion and medium
CN119421236A
Intelligent parking building vehicle dynamic supervision method and system based on AI intelligent engine
CN119479309A
Cited By
Water quality monitoring method based on blockchain enabling pollutant fingerprint cooperative authentication
CN120408332A
Multi-source data integration method and system based on big data
CN120631966A
Water pollution real-time detection and early warning method based on self-adaptive spectral feature fusion
CN120721661A
Irrigation area data anomaly analysis method and server
CN120995017A
Hydraulic engineering monitoring method and system
CN121141989A