Industrial wastewater monitoring data authenticity verification method based on multi-modal feature fusion
By employing a multimodal feature fusion method, the problems of data tampering and spatiotemporal feature verification at the sensor level were solved, enabling adaptive preprocessing and reliability verification of industrial wastewater monitoring data, and improving the accuracy of anomaly identification and data authenticity verification.
Patent Information
- Application Number
- CN202510247102.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-03-04
AI Technical Summary
Existing technologies struggle to build tamper-proof mechanisms based on physical characteristics at the sensor level, and also find it difficult to establish a multi-dimensional verification system that incorporates the spatiotemporal characteristics of wastewater discharge, resulting in insufficient accuracy and reliability in verifying the authenticity of industrial wastewater monitoring data.
By using a multimodal feature fusion method, monitoring data is acquired, standardized, and water quality feature vectors are extracted. A standardized weight matrix is constructed based on pipeline topology data, spatial autocorrelation analysis is performed, a spatiotemporal fusion feature matrix is generated, and an encrypted QR code is generated by combining the sensor feature spectrum matrix. Data confidence scores are calculated, and early warning description information is generated.
It achieves adaptive preprocessing of industrial wastewater monitoring data, improves the ability to identify abnormal features, identifies local abnormal emissions or data tampering, improves the reliability of verification results, constructs an objective data credibility assessment system, and prevents data tampering.
Smart Images

Figure CN120144981B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of environmental monitoring, and in particular, it is a method for verifying the authenticity of industrial wastewater monitoring data based on multimodal feature fusion. Background Technology
[0002] Verifying the authenticity of industrial wastewater monitoring data is crucial for environmental regulation and pollution control. First, industrial wastewater discharge is intermittent, sudden, and highly variable, with water quality parameters fluctuating widely and rapidly, posing significant challenges to data collection and verification. Second, industrial parks typically contain multiple discharge entities, and the complex pipeline systems make it difficult to trace the spatial transport paths of pollutants. Furthermore, due to increased environmental penalties, some enterprises may employ various technical means to interfere with or tamper with monitoring data, severely impacting the effectiveness of environmental regulation. Therefore, establishing a reliable data authenticity verification mechanism is of significant practical importance for ensuring the authenticity of environmental monitoring data and improving the effectiveness of environmental regulation.
[0003] Current data authenticity verification technologies primarily focus on security protection during data transmission and storage. Traditional methods typically employ information security technologies such as digital signatures and encrypted transmission to protect the data transmission process, or use distributed ledger technologies such as blockchain to ensure the immutability of data storage. In the data acquisition stage, physical isolation measures such as video surveillance and sampler seals are mainly relied upon to prevent human intervention. In data analysis, commonly used methods include anomaly detection based on statistical models and rule-based threshold judgment. These methods primarily rely on historical data to establish judgment criteria and are ill-suited to the complex and ever-changing scenarios of industrial wastewater discharge.
[0004] However, the core challenge in verifying the authenticity of industrial wastewater monitoring data lies in how to construct a tamper-proof mechanism based on physical characteristics at the sensor level, and simultaneously establish a multi-dimensional verification system that incorporates the spatiotemporal characteristics of wastewater discharge to prevent data from being tampered with or falsified at the source. These technical challenges severely restrict the accuracy and reliability of verifying the authenticity of industrial wastewater monitoring data. Summary of the Invention
[0005] The purpose of this invention is to provide a method for verifying the authenticity of industrial wastewater monitoring data based on multimodal feature fusion, so as to solve the above-mentioned problems existing in the prior art.
[0006] The technical solution, a method for verifying the authenticity of industrial wastewater monitoring data based on multimodal feature fusion, includes:
[0007] Acquire monitoring data, standardize the monitoring data to obtain standardized water quality data; extract water quality feature vectors based on standardized water quality data; the monitoring data includes pH value, COD value and turbidity;
[0008] Data from adjacent monitoring points is acquired, and a standardized weight matrix is constructed based on pre-stored pipeline topology data. Spatial autocorrelation analysis is performed on water quality feature vectors based on the standardized weight matrix to obtain spatial feature vectors. The spatial feature vectors are then input into a feature fusion unit to generate a spatiotemporal fusion feature matrix.
[0009] An encrypted QR code is generated based on a spatiotemporal fusion feature matrix and standardized water quality data.
[0010] Receive the encrypted QR code, calculate the data confidence score, and generate the corresponding warning description information.
[0011] Beneficial effects: This invention constructs an adaptive data preprocessing method adapted to the characteristics of industrial wastewater, improving the ability to identify abnormal features; it implements a verification mechanism based on spatiotemporal feature fusion, which can effectively identify local abnormal emissions or data tampering; it adopts a multi-head attention mechanism for feature fusion, improving the reliability of verification results; it combines confidence propagation and graph neural networks to construct an objective data credibility assessment system; through the synergistic effect of physical feature verification, data feature verification, and spatiotemporal feature verification, it can effectively prevent data tampering, promptly detect abnormal emissions, and provide reliable technical support for environmental supervision. Attached Figure Description
[0012] Figure 1 A flowchart illustrating the steps of a method for verifying the authenticity of industrial wastewater monitoring data based on multimodal feature fusion, provided in this application embodiment.
[0013] Figure 2 A flowchart illustrating the steps for constructing a standardized weight matrix provided in this application embodiment.
[0014] Figure 3 This is a flowchart illustrating the steps for standardizing monitoring data as provided in an embodiment of this application.
[0015] Figure 4 A flowchart illustrating the steps for constructing a time feature vector as provided in this application embodiment. Detailed Implementation
[0016] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0017] It should be noted that, for the purpose of clearly demonstrating the steps of this application, each step has been numbered in the specification. These numbers are for ease of explanation only and do not limit the execution order of the steps. In actual operation, depending on the technical requirements of the specific implementation scenario, the steps may be executed in a different order than that shown in the specification, and in some cases, parallel processing between steps can also be achieved.
[0018] like Figure 1 As shown, this application proposes a method for verifying the authenticity of industrial wastewater monitoring data based on multimodal feature fusion, comprising the following steps:
[0019] S1. Acquire monitoring data, standardize the monitoring data to obtain standardized water quality data; extract water quality feature vectors based on the standardized water quality data; the monitoring data includes pH value, COD value and turbidity;
[0020] S2. Acquire data from adjacent monitoring points and construct a standardized weight matrix based on pre-stored pipeline topology data; perform spatial autocorrelation analysis on water quality feature vectors based on the standardized weight matrix to obtain spatial feature vectors; input the spatial feature vectors into the feature fusion unit to generate a spatiotemporal fusion feature matrix.
[0021] S3. Generate encrypted QR codes based on spatiotemporal fusion feature matrices and standardized water quality data;
[0022] S4. Receive the encrypted QR code, calculate the data confidence score, and generate the corresponding warning description information.
[0023] According to one aspect of this application, the method further includes step S0: reading the sensor operating voltage fluctuation value, temperature response characteristics and dielectric impedance, performing Fourier transform to obtain the sensor feature spectrum matrix; constructing feature vectors and mapping them to generate a sensor dynamic fingerprint; and performing feature extraction to generate a fingerprint feature code.
[0024] In one embodiment of this application, based on the sensor ID number stored in the sensor's built-in memory, the sensor operating voltage fluctuation value, temperature response characteristics, and dielectric impedance are read from the sensor calibration database; using a standard sampling frequency of 2000Hz, Fourier transforms are performed on these three physical quantities with 8192 sampling points to obtain the sensor feature spectrum matrix; based on the sensor feature spectrum matrix, a 128-dimensional feature vector is constructed using a locality-sensitive hash algorithm, and 20 random hyperplanes are used as hash functions to generate a sensor dynamic fingerprint; the sensor dynamic fingerprint is input into a locality-sensitive hash encoder to generate a 256-bit fingerprint feature code.
[0025] Then, raw values of water quality parameters, including pH, COD, and turbidity, are collected from the water quality sensor (monitoring data, the same below); the raw values of water quality parameters are smoothed using a sliding window of 100 data points to obtain smoothed water quality data; the smoothed water quality data is input into the adaptive standardization processing unit to generate standardized water quality data; the first 5 principal components are extracted from the standardized water quality data using principal component analysis to form a water quality feature vector.
[0026] The system reads the historical data sequence and water quality feature vector of the most recent 24 hours from the data storage unit, inputs them into the long short-term memory network to generate a time feature vector; it collects data from adjacent monitoring points and performs spatial autocorrelation analysis with the current water quality feature vector to obtain a spatial feature vector; it inputs the time feature vector and spatial feature vector into the tensor decomposition algorithm to generate a spatiotemporal fusion feature matrix.
[0027] The fingerprint feature code, the spatiotemporal fusion feature matrix, and the real-time collected environmental parameter vector are input into the feature alignment unit to generate the alignment feature matrix; the alignment feature matrix is weighted by a multi-head attention mechanism to obtain the fusion feature matrix; the fusion feature matrix, standardized water quality data, and sensor dynamic fingerprint encoding are used to generate an encrypted QR code.
[0028] The system acquires the fusion feature matrix and historical confidence data stored in the database, and calculates the data confidence score using the confidence propagation algorithm. The data confidence score is then matched with the historical anomaly pattern library in the database to generate anomaly type identifiers and anomaly severity values. Based on the anomaly type identifiers and anomaly severity values, and combined with predefined early warning rules, an early warning level and early warning description information are generated.
[0029] According to one aspect of this application, step S0 further comprises:
[0030] S01. Read the sensor ID number from the sensor's built-in memory via the 485 communication interface. Based on this ID number, retrieve the sensor's operating voltage fluctuation value, temperature response characteristics, and dielectric impedance from the sensor calibration database. Sample these three physical quantities using a sampling frequency of 2000Hz, collecting 8192 data points for each physical quantity. Input the sampled data into the Fourier transform processing unit to generate three sets of spectrum data. Combine these three sets of spectrum data into a 15×15 sensor characteristic spectrum matrix.
[0031] S02. Read the sensor feature spectrum matrix and extract the main feature components in the matrix using principal component analysis. Reorganize the extracted feature components into a 128-dimensional feature vector. Randomly generate 20 hyperplanes as hash functions to map the feature vectors. Match the mapping results with the pre-stored sensor feature templates to generate the sensor dynamic fingerprint.
[0032] S03. Acquire the dynamic fingerprint of the sensor and input it into the local sensitive hash encoder; use the Hamming distance threshold of 0.2 as a criterion to extract features from the dynamic fingerprint of the sensor; generate a 256-bit fingerprint feature code based on the extracted features.
[0033] In one embodiment of this application, core physical parameters of the sensor are obtained, including the sensor operating voltage fluctuation value V, temperature response characteristic T, and dielectric impedance Z. A feature spectrum is obtained through Fourier transform, and characteristic peaks and valleys are extracted to output a sensor feature spectrum matrix F. Based on the sensor feature spectrum matrix F, an adaptive weighted fusion algorithm is applied to combine the features of different physical parameters to output a sensor dynamic fingerprint D. Based on the sensor dynamic fingerprint D, a Locality Sensitive Hash (LSH) algorithm is used for feature encoding to output a fingerprint feature code E.
[0034] This embodiment obtains the sensor's characteristic spectrum under different operating conditions by sampling three physical quantities—sensor operating voltage fluctuation, temperature response characteristics, and dielectric impedance—at a high frequency of 2000Hz and combining this with Fourier transform. These spectral features reflect the sensor's inherent physical characteristics. A 128-dimensional feature vector is constructed using a Locality Sensitive Hashing (LSH) algorithm, and 20 random hyperplanes are used as hash functions to compress the sensor's physical features into a unique digital fingerprint. Since these physical features are determined by the sensor's hardware structure and exhibit stable correlation under different operating conditions, the 256-bit fingerprint generated in this way possesses strong anti-counterfeiting and uniqueness. Even if an attacker gains access to the sensor's data output interface, they cannot forge or tamper with these physically based fingerprint features. Simultaneously, the use of an LSH encoder for feature extraction ensures that similar physical features are also close in distance within the encoding space, thus maintaining the stability of the fingerprint features even when the sensor is subjected to slight physical interference. This dynamic fingerprint mechanism based on physical characteristics guarantees the credibility of the data source at the hardware level, laying the foundation for subsequent data authenticity verification.
[0035] According to one aspect of this application, step S1 further comprises:
[0036] S11. Collect raw values of water quality parameters such as pH, COD, and turbidity every 10 seconds through the data interface of the water quality sensor; store the collected raw values of water quality parameters in a temporary buffer in chronological order; read the most recent 100 data points from the temporary buffer, apply the Hanning window function for sliding smoothing, and generate smoothed water quality data.
[0037] S12. Read the smoothed water quality data; obtain the standard range values of each parameter from the system configuration database; normalize the smoothed water quality data according to the standard range values to generate standardized water quality data in the range of [0,1].
[0038] S13. Obtain standardized water quality data; construct the covariance matrix and calculate the eigenvalues and eigenvectors; select the top 5 principal components with a contribution rate greater than 85%; combine these 5 principal components to form a water quality feature vector.
[0039] In one embodiment of this application, raw values W of water quality parameters (including pH, COD, turbidity, etc.) are obtained, and the data is smoothed using a sliding window method to obtain smoothed water quality data Ws. Based on the smoothed water quality data Ws, an adaptive standardization algorithm is used to normalize the data, outputting standardized water quality data Wn. Based on the standardized water quality data Wn, principal component analysis is applied to extract key features, outputting a water quality feature vector Wf.
[0040] This embodiment uses a sliding window of 100 data points to smooth the raw water quality parameter values, effectively filtering out random noise and sudden interference during sensor acquisition. Principal component analysis is used to extract the first five principal components to form a water quality feature vector, achieving not only dimensionality reduction but also preserving the correlation information between water quality parameters. This embodiment ensures the integrity of data representation while highlighting the characteristics of abnormal patterns. Especially when dealing with intermittent and sudden monitoring targets such as industrial wastewater, it can accurately capture key features of water quality changes, avoiding the feature loss problems that may occur with traditional fixed-window smoothing methods.
[0041] According to one aspect of this application, step S12 further comprises:
[0042] S121. Read the standard range thresholds for pH, COD, and turbidity from the system configuration database, including the minimum, maximum, mean, and standard deviation; calculate a weighted average of these thresholds with the statistical characteristics of historical data, with a weight ratio of 3:7, to obtain the dynamic reference thresholds; generate a standardized interval for each parameter based on the dynamic reference thresholds.
[0043] S122. Obtain smoothed water quality data; apply the Tukey criterion to calculate the outlier limits for each parameter, set the inner limit coefficient to 1.5 and the outer limit coefficient to 3.0; identify data points that exceed the limits and mark them as outliers; correct the outliers using a local linear interpolation method to generate outlier corrected data.
[0044] S123. Read outlier correction data and standardized intervals; construct an adaptive normalization function, with the function slope dynamically adjusted according to the data distribution density; set the normalization curve of the dense data region to S-shape and the sparse region to linear; map all parameters to the [0,1] interval through this function to generate preliminary standardized water quality data.
[0045] S124. Obtain preliminary standardized water quality data; calculate the data distribution skewness and kurtosis of each parameter; when the absolute value of skewness is greater than 0.5 or the kurtosis value exceeds the range of [2,4], apply Box-Cox transformation for correction; perform variance normalization on the corrected data to generate the final standardized water quality data.
[0046] This embodiment generates dynamic reference thresholds by reading the standard range thresholds of each water quality parameter from the system configuration database and performing a 3:7 weighted average with the statistical characteristics of historical data. This solves the problem that fixed thresholds are difficult to adapt to fluctuations in industrial wastewater quality. The Tukey criterion is used to identify outliers, with an inner limit coefficient of 1.5 and an outer limit coefficient of 3.0. Local linear interpolation is used to correct outliers, ensuring data continuity while avoiding the impact of outliers on subsequent analysis. When constructing the adaptive normalization function, the slope is dynamically adjusted according to the data distribution density. An S-curve is used to improve resolution in dense data areas, while linear mapping is used in sparse areas to maintain the original distribution characteristics. This differentiated normalization strategy is particularly suitable for the uneven distribution of industrial wastewater quality parameters. By calculating the skewness and kurtosis of the data distribution, when the absolute value of skewness is greater than 0.5 or the kurtosis value exceeds the range of [2,4], a Box-Cox transformation is applied for correction, ensuring the statistical properties of the normalization results. This multi-level data standardization processing scheme not only solves the problems of frequent outliers and uneven distribution in industrial wastewater monitoring data, but also provides a reliable data foundation for subsequent feature extraction and pattern recognition.
[0047] According to one aspect of this application, step S123 further comprises:
[0048] S1231. Read outlier correction data; classify water quality parameters according to their physicochemical properties; apply piecewise linear mapping to pH; apply logarithmic mapping to concentration parameters; generate a parameter feature classification table.
[0049] S1232. Obtain the parameter feature classification table; calculate the data distribution characteristics of each type of parameter; determine the transformation function type based on the skewness; generate a normalized function set.
[0050] S1233. Read the normalization function set and outlier correction data; apply the corresponding normalization function to each type of parameter; perform interval mapping; generate preliminary standardized water quality data.
[0051] This embodiment classifies water quality parameters according to their physicochemical properties, employing piecewise linear mapping for pH and logarithmic mapping for concentration parameters, establishing a specialized parameter characteristic classification table. The transformation function type is determined based on the data distribution characteristics of each parameter category, and the form of the normalization function is adjusted according to skewness to ensure the rationality of the standardization results. This differentiated normalization strategy fully considers the physicochemical characteristics and numerical distribution patterns of different water quality parameters, avoiding the information loss problems caused by traditional uniform standardization methods. This refined data processing scheme preserves the physical meaning of the data while improving the accuracy of subsequent analysis, making it particularly suitable for monitoring scenarios such as industrial wastewater, where parameter types are diverse and numerical distributions are uneven.
[0052] According to one aspect of this application, step S13 further comprises:
[0053] S131. Read standardized water quality data; calculate the pairwise covariance between each parameter to generate an initial covariance matrix; perform condition number analysis on the covariance matrix, and perform regularization when the condition number is greater than 1000, with the regularization coefficient set to 0.01; generate the optimized feature covariance matrix.
[0054] S132. Obtain the feature covariance matrix; use the improved QR decomposition algorithm to calculate eigenvalues, and set the convergence threshold to 1e-6; sort the eigenvalues in descending order; calculate the cumulative contribution rate, and generate the eigenvalue sequence and the corresponding eigenvector set.
[0055] S133. Read the feature value sequence and feature vector set; select the top K principal components (K≤5) based on the 85% cumulative contribution rate threshold; orthogonalize the selected feature vectors; generate a water quality feature vector with dimension K×3 through linear combination.
[0056] This embodiment constructs an initial covariance matrix by calculating the pairwise covariances between various water quality parameters and performs condition number analysis on it. When the condition number is greater than 1000, a regularization coefficient of 0.01 is applied to effectively avoid numerical instability during feature extraction. An improved QR decomposition algorithm is used to calculate eigenvalues, and a convergence threshold of 1e-6 is set to ensure the accuracy of the feature decomposition results. The top K principal components (K≤5) are selected based on the 85% cumulative contribution rate threshold, and the selected eigenvectors are orthogonalized to generate water quality feature vectors with a dimension of K×3, achieving dimensionality reduction of the data while retaining the main feature information. This regularization-based principal component analysis method is particularly suitable for monitoring data such as industrial wastewater where parameters are strongly correlated. Through accurate calculation of eigenvalues and eigenvectors, the intrinsic correlation between water quality parameters can be accurately captured, providing a reliable feature representation for subsequent data authenticity verification.
[0057] According to one aspect of this application, such as Figure 3 As shown, the steps for standardizing monitoring data include:
[0058] The slope of the pre-configured normalization function is dynamically adjusted based on the data distribution density of the monitoring data.
[0059] Based on the adjusted slope, S-curve normalization is performed on data-dense areas; based on the adjusted slope, linear mapping is performed on data-sparse areas; preliminary standardized water quality data are generated.
[0060] Based on preliminary standardized water quality data, the skewness and kurtosis of the data distribution are calculated and corrected using Box-Cox transformation to generate standardized water quality data.
[0061] According to one aspect of this application, the steps for extracting water quality feature vectors based on standardized water quality data include:
[0062] Based on standardized water quality data, a covariance matrix is constructed and condition number analysis is performed. When the condition number is less than a preset threshold, the covariance matrix is directly used as the optimized feature covariance matrix. When the condition number is greater than the preset threshold, the covariance matrix is regularized to obtain the optimized feature covariance matrix.
[0063] Based on the optimized feature covariance matrix, the improved QR decomposition algorithm is used to calculate the eigenvalues, generating an eigenvalue sequence and the corresponding eigenvector set;
[0064] Based on the feature value sequence and the corresponding feature vector set, principal components are selected according to a preset cumulative contribution rate threshold to obtain the selected feature vectors;
[0065] The selected feature vectors are orthogonalized to generate water quality feature vectors.
[0066] According to one aspect of this application, the step of calculating eigenvalues using an improved QR decomposition algorithm includes:
[0067] Weighting factors are set according to the importance of pollutant parameters in the monitoring data;
[0068] The optimized feature covariance matrix is decomposed into an upper triangular matrix and an orthogonal matrix;
[0069] The decomposition process is iteratively optimized based on weighting factors;
[0070] During the iteration process, the rate of change of eigenvalues and eigenvectors is calculated, and the calculation step size is dynamically adjusted;
[0071] When the rate of change of both the eigenvalue and the eigenvector is less than the corresponding preset threshold, the final eigenvalue is determined.
[0072] The final feature values are sorted in descending order, and the cumulative contribution rate of the final feature values is calculated to generate a feature value sequence and a corresponding feature vector set.
[0073] For the eigenvalue calculation scenario of industrial wastewater monitoring data, this implementation uses an improved QR decomposition algorithm. By introducing a weight adjustment mechanism based on pollutant characteristics, the QR decomposition process is better adapted to the characteristics of water quality data. Specifically, high weight factors are assigned to key monitored pollutant parameters (such as COD and ammonia nitrogen). These weight factors act on the decomposition process of the feature covariance matrix, making the decomposition results more reflective of the changing characteristics of key pollutants. Simultaneously, the algorithm also designs a dynamic step-size adjustment mechanism, that is, during the iteration process, the calculation step size is adaptively adjusted according to the rate of change of eigenvalues, and a dual-index convergence judgment strategy of eigenvalues and eigenvectors is adopted. This improvement ensures that while maintaining numerical stability, it can more accurately capture abrupt changes in water quality data. This embodiment improves the sensitivity to changes in key pollutant parameters, making the eigenvalue calculation results better reflect water quality anomalies; on the other hand, through dynamic step size and dual-index convergence judgment, it optimizes algorithm efficiency while ensuring calculation accuracy, making it particularly suitable for scenarios requiring rapid response in industrial wastewater monitoring.
[0074] According to one aspect of this application, step S2 further comprises:
[0075] S21. Read the historical data sequence of the most recent 24 hours sequentially from the data storage unit; obtain the water quality feature vector; input the historical data sequence and water quality feature vector into a long short-term memory network with 64 hidden layer nodes; extract the temporal features through the network to generate a time feature vector with a dimension of 32.
[0076] S22. Obtain data from adjacent monitoring points of 5 adjacent monitoring points from the communication module; read the water quality feature vector; calculate the Moran index and GEARY index between the adjacent monitoring point data and the water quality feature vector; generate a spatial feature vector with dimension 32 based on the calculation results.
[0077] S23. Obtain the temporal and spatial feature vectors; construct a third-order tensor and perform CP decomposition; select the decomposition result with a rank of 10; reorganize the decomposition result into a 32×32 spatiotemporal fusion feature matrix.
[0078] In one embodiment of this application, based on the water quality feature vector Wf and the historical data sequence H, a long short-term memory network is applied to extract temporal features, resulting in a time feature vector Tf. Based on the water quality feature vector Wf and data from adjacent monitoring points N, a spatial autocorrelation algorithm is used to analyze the spatial consistency of the data, outputting a spatial feature vector Sf. Based on the time feature vector Tf and the spatial feature vector Sf, an adaptive weighted tensor decomposition algorithm is used for feature fusion, outputting a spatiotemporal fusion feature matrix M.
[0079] This embodiment analyzes 24-hour historical data sequences and, combined with a network structure of 64 hidden nodes, effectively captures the long-term trends and short-term fluctuations of water quality parameters. Especially when processing monitoring data with significant temporal characteristics, such as industrial wastewater, the spatial autocorrelation of the data can be accurately assessed by calculating the Moran's index and the Gearry index of adjacent monitoring points, thereby identifying abnormal discharge behavior. A tensor decomposition algorithm is used to fuse temporal and spatial feature vectors, and the decomposition result with a rank of 10 is recombined into a 32×32 feature matrix. This preserves the spatiotemporal correlation characteristics of the data while achieving dimensionality reduction and compression of the features. This verification mechanism based on spatiotemporal feature fusion can simultaneously consider the temporal continuity and spatial consistency of monitoring data, making it particularly suitable for monitoring scenarios with complex pipe network structures, such as industrial parks. Spatial autocorrelation analysis of data from adjacent monitoring points can effectively identify localized abnormal discharges or data tampering, improving the reliability and accuracy of the verification method.
[0080] According to one aspect of this application, step S21 further comprises:
[0081] S211. Read the historical data sequence of the most recent 24 hours in order of timestamp; divide the sequence into sliding window segments with a window size of 30 minutes and a step size of 5 minutes; calculate the statistical characteristics within each window, including mean, variance, kurtosis and skewness; generate time series statistical characteristics.
[0082] S212. Obtain water quality feature vectors and time-series statistical features; align and merge the two sets of data according to time; extract features through a 64-node LSTM network, with the forget gate threshold set to 0.3; generate the original time features.
[0083] S213. Read the original time features; apply the backpropagation algorithm to optimize the feature weights, with the learning rate set to 0.01; reduce the feature dimension to 32; generate the time feature vector through normalization.
[0084] This embodiment employs adaptive window segmentation of a 24-hour historical data sequence, combined with feature extraction using a 64-node LSTM network, and sets a forget gate threshold of 0.3 to accurately capture long-term dependencies in the data sequence. Backpropagation optimization with a learning rate of 0.01 reduces the feature dimensionality to 32 dimensions and performs normalization, preserving key features of the time-series data while achieving feature compression. This adaptive window-based time-series feature extraction method is particularly suitable for monitoring industrial wastewater, which exhibits significant time-varying characteristics. It can accurately identify changes in discharge conditions and provides crucial evidence for detecting abnormal emissions.
[0085] According to one aspect of this application, step S211 further comprises:
[0086] S2111. Read the historical data sequence; calculate the first difference value of the data and determine the rate of change of the data; mark the inflection point of data change according to the magnitude of the difference value; generate the data change feature sequence.
[0087] S2112. Obtain the data change characteristic sequence; calculate the time interval between adjacent inflection points; determine the reference window size based on the distribution characteristics of the interval; generate window reference parameters.
[0088] S2113. Read the window baseline parameters and data change characteristic sequence; when the data change rate exceeds the threshold, reduce the window size to 1 / 2 of the baseline value; when the data is stable, increase the window size to twice the baseline value; generate an adaptive window sequence.
[0089] S2114. Obtain historical data sequences and adaptive window sequences; calculate statistical features within each window; generate time-series statistical features.
[0090] This embodiment calculates the first-order difference of historical data sequences and marks inflection points in data changes. It then determines the baseline window size based on the time interval distribution characteristics between adjacent inflection points, achieving adaptive adjustment of the window size. When the data change rate exceeds a threshold, the window size automatically decreases to half the baseline value, while it increases to twice the baseline value when the data is stable. This dynamic adjustment mechanism is particularly suitable for the intermittent discharge characteristics of industrial wastewater. By calculating statistical characteristics such as mean, variance, kurtosis, and skewness within each adaptive window, the main features of the data are preserved while effectively filtering out noise interference. This embodiment solves the problem of traditional fixed windows easily segmenting the discharge process, improving the accuracy and reliability of time-series feature extraction.
[0091] According to one aspect of this application, step S22 further comprises:
[0092] S221. Obtain data of adjacent monitoring points from the communication module in real time for the five adjacent monitoring points; construct a spatial weight matrix based on the physical distance between the monitoring points, with the weights decaying exponentially with distance; perform row standardization on the weight matrix; generate a standardized weight matrix.
[0093] S222. Obtain water quality feature vectors and standardized weight matrices; calculate the local Moran index with a spatial lag order of 2; calculate the local GEARY index with a neighborhood range of 1000 meters; generate spatial correlation indicators.
[0094] S223. Read the spatial correlation index; construct a 32-dimensional spatial feature mapping function; map the spatial correlation index to the feature space; optimize the mapping parameters using the least squares method to generate spatial feature vectors.
[0095] This embodiment employs an exponential decay function to construct a spatial weight matrix, which, compared to the traditional fixed threshold method, more precisely reflects the nonlinear distance effect between adjacent monitoring points, especially capturing strong correlations over short to medium distances. Through dual processing of physical distance weights and row standardization, the magnitude information of the original spatial relationships is preserved while eliminating the interference of inter-row differences on subsequent analysis, making different monitoring units comparable. Simultaneously, local Moran's index and local Geary index are calculated; the former detects spatial clustering characteristics (such as pollution diffusion hotspots), and the latter identifies spatial outliers (such as sudden pollution sources), forming a mutually verifying three-dimensional analysis framework. By using a dual definition of second-order spatial lag and a 1000-meter neighborhood, both topological adjacency relationships (network propagation paths) and actual geographical scope are considered, constructing a multi-dimensional spatial association network. The 32-dimensional spatial feature mapping function breaks through the dimensional limitations of traditional spatial statistical indicators, transforming discrete spatial association indicators into continuously distributed high-dimensional representations, providing a suitable input structure for deep learning models. The application of the least squares method in mapping parameter optimization ensures the physical correlation between the feature space and the target variable (such as water quality level) by minimizing the loss function, avoiding feature distortion caused by simple mathematical transformations. Compared with traditional water quality monitoring schemes, this embodiment reduces spatial analysis error by approximately 37% (simulated test data) while improving feature representation efficiency by 2.8 times, providing more accurate spatial decision-making basis for water quality anomaly early warning and pollution source tracing.
[0096] According to one aspect of this application, step S221 further comprises:
[0097] S2211. Read the pipeline topology data from the industrial park pipeline database; extract the upstream contribution coefficient and downstream influence coefficient of each monitoring point; calculate the flow capacity coefficient based on the pipe diameter, flow velocity and hydraulic gradient; generate the pipeline feature vector.
[0098] S2212. Read the pipeline feature vector; construct the pipeline flow direction matrix, with upstream nodes taking a positive value of 1, downstream nodes taking a negative value of 1, and no direct connection taking a value of 0; calculate the shortest pipeline path between nodes; generate the flow direction correlation matrix.
[0099] S2213. Obtain the feature vector and flow direction correlation matrix of the pipeline network; calculate the diffusion attenuation coefficient of pollutants at different flow velocities; calculate the pollutant transfer time between nodes based on the diffusion equation; generate the spatiotemporal transfer matrix.
[0100] S2214. Read the spatiotemporal transfer matrix; based on the pipeline flow direction, assign higher weights (0.6-0.8) to upstream monitoring points, medium weights (0.3-0.5) to monitoring points on the same level, and lower weights (0.1-0.3) to downstream monitoring points; perform nonlinear adjustment on the weights, with the adjustment coefficient being proportional to the flow velocity and pipe diameter; generate the initial weight matrix.
[0101] S2215. Obtain the initial weight matrix; apply row normalization; dynamically adjust the weights according to the real-time flow rate; generate the final standardized weight matrix.
[0102] This embodiment reads pipeline topology data from a pipeline database and extracts the upstream contribution coefficient and downstream influence coefficient for each monitoring point. Combined with pipe diameter, flow velocity, and hydraulic gradient, it calculates the flow capacity coefficient, constructing a complete pipeline feature vector. A flow direction matrix is constructed based on the flow direction of the pipe segments, with upstream nodes assigned a positive value of 1 and downstream nodes assigned a negative value of 1. The shortest pipeline path between nodes is calculated, accurately reflecting the pollutant transport path. Based on the diffusion equation, the diffusion attenuation coefficient of pollutants at different flow velocities and the transfer time between nodes are calculated, generating a spatiotemporal transport matrix. In weight allocation, based on the characteristics of the pipeline flow direction, upstream monitoring points are assigned a high weight of 0.6-0.8, monitoring points at the same level have a medium weight of 0.3-0.5, and downstream monitoring points have a low weight of 0.1-0.3, with dynamic adjustments based on real-time flow velocity. This spatial correlation analysis method based on pipeline topology fully considers the characteristics of the complex pipeline structure in industrial parks. By establishing a pollutant transport model and a dynamic weight system, it can accurately assess the spatial consistency of monitoring data and effectively identify illegal emissions or data tampering.
[0103] According to one aspect of this application, step S23 further comprises:
[0104] S231. Read the temporal and spatial feature vectors; construct a third-order tensor with dimensions of 32×32×10; normalize the tensor; generate the initial feature tensor.
[0105] S232. Obtain the initial feature tensor; apply the CP decomposition algorithm, setting the maximum number of iterations to 200; select the decomposition result with a rank of 10; generate the decomposed feature matrix group.
[0106] S233. Read the decomposed feature matrix group; evaluate the reconstruction error using the Tucker norm; select feature combinations with a reconstruction error less than 0.01; reconstruct the feature matrix into a 32×32 spatiotemporal fusion feature matrix.
[0107] This embodiment fuses temporal and spatial feature vectors by constructing a 32×32×10 third-order tensor. The CP decomposition algorithm is used for tensor decomposition, with a maximum of 200 iterations, and decomposition results with a rank of 10 are selected. The Tucker norm is used to evaluate the reconstruction error, and feature combinations with reconstruction errors less than 0.01 are recombined into a 32×32 spatiotemporal fusion feature matrix. This spatiotemporal feature fusion method based on tensor decomposition maintains the independence of temporal and spatial features while achieving their organic fusion. Especially in monitoring scenarios with significant spatiotemporal correlation, such as industrial wastewater, tensor decomposition can effectively extract the coupling relationship between spatiotemporal features, providing a more comprehensive feature representation for data authenticity verification. Simultaneously, by controlling the reconstruction error, the accuracy and reliability of the feature fusion process are ensured.
[0108] like Figure 2 As shown, according to one aspect of this application, the step of constructing a standardized weight matrix based on pre-stored pipeline topology data includes:
[0109] Read pipeline topology data from the industrial park's pipeline database, and extract the upstream contribution coefficient and downstream impact coefficient of each monitoring point based on the pipeline topology data;
[0110] The flow capacity coefficient is calculated based on the pre-stored pipe diameter, flow velocity, and hydraulic gradient. A network feature vector is generated based on the upstream contribution coefficient, downstream influence coefficient, and flow capacity coefficient.
[0111] Based on the pipeline feature vector, the shortest pipeline path between nodes is calculated, and a flow direction correlation matrix is constructed, where upstream nodes take a positive value of 1, downstream nodes take a negative value of 1, and no direct connection takes a value of 0.
[0112] Based on the pipeline network feature vector and the flow direction correlation matrix, the diffusion attenuation coefficient of pollutants at different flow velocities is calculated.
[0113] Based on the diffusion attenuation coefficient, the pollutant transfer time between nodes is calculated using the diffusion equation to generate a spatiotemporal transfer matrix.
[0114] Weights are assigned to monitoring points based on the spatiotemporal transfer matrix, and nonlinear adjustments are made. The coefficients of the nonlinear adjustments are proportional to the flow velocity and pipe diameter, resulting in a standardized weight matrix.
[0115] like Figure 4As shown, according to one aspect of this application, when constructing spatial feature vectors, temporal feature vectors are constructed, and the steps are as follows:
[0116] Read historical data sequences at a predetermined time from the data storage unit;
[0117] Based on historical data sequences, calculate the rate of change of data and determine the inflection point of data change;
[0118] The size of the baseline window is determined based on the time interval between adjacent data change inflection points;
[0119] When the data change rate exceeds a preset threshold, the baseline window size is reduced to 1 / a of the original size; when the data change rate does not exceed the preset threshold, the baseline window size is increased to a times the original size; thus obtaining the adjusted window size.
[0120] Statistical features within each window are calculated based on historical data sequences and adjusted window sizes to generate time-series statistical features.
[0121] Based on time-series statistical features and water quality feature vectors, a time feature vector is generated; where a is a natural number greater than 0.
[0122] According to one aspect of this application, the process of generating the spatiotemporal fusion feature matrix further includes:
[0123] Tensor decomposition is performed on the temporal and spatial feature vectors to generate a spatiotemporal fusion feature matrix.
[0124] The steps of tensor decomposition include:
[0125] Construct a third-order tensor with dimensions M×N×K;
[0126] Perform CP decomposition on the third-order tensor and select the decomposition result with rank P;
[0127] Based on the decomposition results, the reconstruction error is evaluated using the Tucker norm, and feature combinations with reconstruction errors less than a preset threshold are selected.
[0128] The features are combined and recombined into a spatiotemporal fusion feature matrix, where M, N, K and P are preset values.
[0129] According to one aspect of this application, step S3 further comprises:
[0130] S31. Read the fingerprint feature code and spatiotemporal fusion feature matrix; collect parameters such as temperature, humidity, and air pressure from environmental sensors to form an environmental parameter vector; perform timestamp alignment and dimension unification processing on these three sets of data to generate a 64×64 aligned feature matrix.
[0131] S32. Obtain the aligned feature matrix; construct 8 attention heads, each with a dimension of 8; calculate the feature weights through multi-head attention; weight the features according to the weights to generate a 128-dimensional fusion feature matrix.
[0132] S33. Read the fused feature matrix, standardized water quality data, and sensor dynamic fingerprint; encode the data using the Reed-Solomon encoding algorithm; convert the encoding result into a 29×29 encrypted QR code.
[0133] In one embodiment of this application, based on the fingerprint feature code E, the spatiotemporal fusion feature matrix M, and the environmental parameter vector P, a multimodal feature alignment algorithm is used for feature mapping to obtain an aligned feature matrix A. Based on the aligned feature matrix A, an attention mechanism is used for feature weighted fusion to output a fused feature matrix R. Based on the fused feature matrix R, standardized water quality data Wn, and the sensor dynamic fingerprint D, a high-density QR code is generated using the Reed-Solomon coding algorithm to obtain an encrypted QR code Q.
[0134] This embodiment unifies fingerprint feature codes, spatiotemporal fusion feature matrices, and environmental parameter vectors into a single feature space, generating a 64×64 aligned feature matrix, thus solving the problem of fusion of multi-source heterogeneous data. Eight attention heads are used for feature weighting, each with an 8-dimensional dimension. This multi-head attention mechanism allows for weight allocation of features from different perspectives, improving the robustness of feature fusion. The Reed-Solomon coding algorithm is used to generate a 29×29 encrypted QR code, achieving not only high-density data storage but also providing 30% error correction capability, ensuring reliable identification in industrial environments. This multimodal feature fusion scheme fully utilizes sensor physical features, monitoring data features, and environmental parameter features to construct a multi-layered data authenticity verification system. Especially in scenarios requiring high reliability, such as industrial wastewater monitoring, cross-validation of multi-source data effectively prevents single-point data tampering, improving the security and reliability of the entire monitoring system.
[0135] According to one aspect of this application, step S32 further comprises:
[0136] S321. Read the alignment feature matrix; divide the matrix into 8 subspaces, each with a dimension of 8; calculate the query matrix Q, key matrix K, and value matrix V for each subspace; generate the attention calculation matrix group.
[0137] S322. Obtain the attention calculation matrix group; calculate the scaled dot product attention, with the scaling factor set to the square root of 8; apply the Softmax function to calculate the weight distribution; generate the multi-head feature matrix.
[0138] S323. Read the multi-head feature matrix; convert the feature dimension to 128 through a linear transformation layer; apply residual connections and layer normalization; generate the fused feature matrix.
[0139] This embodiment constructs a complete multi-head attention computation system by dividing the aligned feature matrix into eight subspaces, each with a dimension of 8, and calculating the query matrix Q, key matrix K, and value matrix V. The square root of 8 is used as a scaling factor to calculate the scaled dot product attention, and the Softmax function is applied to calculate the weight distribution, ensuring a reasonable allocation of attention weights. A linear transformation layer converts the feature dimension to 128, and residual connections and layer normalization are applied to generate the final fused feature vector. This feature fusion method based on a multi-head attention mechanism can capture the correlation between features from different perspectives, making it particularly suitable for monitoring data with multi-dimensional features, such as industrial wastewater. The design of residual connections preserves the original feature information while enhancing the model's expressive power, improving the accuracy and stability of feature fusion.
[0140] According to one aspect of this application, step S33 further comprises:
[0141] S331. Obtain three sets of input data; compress the fusion feature matrix to 64 dimensions; merge all data into an ordered byte stream; generate encoded preprocessed data.
[0142] S332. Read the pre-processed encoding data; encode using the Reed-Solomon algorithm, setting the error correction level to H (30%); add check bits and synchronization codes; generate the encoded data stream.
[0143] S333. Obtain the encoded data stream; lay out the data according to the 29×29 specification; add positioning and calibration graphics; generate an encrypted QR code.
[0144] This embodiment compresses the fused feature vector to 64 dimensions and combines it with sensor fingerprints and water quality data to generate an encoded preprocessed data stream. Encoding is performed using the Reed-Solomon algorithm, with an H-level (30%) error correction level, added check bits and synchronization codes, ultimately generating a 29×29 encrypted QR code. This parameter-based encoding scheme improves encoding efficiency and is particularly suitable for monitoring data such as industrial wastewater, which has multiple parameters with varying value ranges. The high error correction level design ensures reliable identification even in harsh industrial environments, providing strong protection for secure data transmission and verification.
[0145] According to one aspect of this application, step S332 further comprises:
[0146] S3321. Read the pre-processed encoding data; map the pH value (0-14) to a 4-bit code; divide other parameters into different code lengths according to their numerical range; generate a parameter encoding mapping table.
[0147] S3322. Obtain the parameter encoding mapping table; allocate the number of error correction bits for Reed-Solomon encoding according to the importance of different parameters; apply double verification to key parameters; generate error correction encoding rules.
[0148] S3323: Read the preprocessed encoding data and error correction encoding rules; perform layered encoding; add parameter type identifiers; generate encoded data stream.
[0149] This embodiment improves encoding efficiency by establishing a targeted parameter encoding mapping table, mapping pH values (0-14) to 4-bit codes, and dividing other parameters into different code lengths according to their numerical ranges. The number of error correction bits in the Reed-Solomon encoding is dynamically allocated based on the importance of different parameters, and a double check mechanism is used for key parameters to ensure the reliability of data transmission. During hierarchical encoding, parameter type identifiers are added to achieve rapid parameter identification and parsing. This embodiment not only improves the information density of the QR code but also enhances anti-interference capabilities through reasonable error correction bit allocation, making it particularly suitable for use in complex industrial environments.
[0150] According to one aspect of this application, the steps for generating an encrypted QR code include:
[0151] Based on the spatiotemporal fusion feature matrix and the real-time collected environmental parameter vectors, a fusion feature matrix is generated; the fusion feature matrix and standardized water quality data are then merged to generate coded preprocessed data.
[0152] Based on the preprocessed encoding data, different encoding lengths are divided to generate a parameter encoding mapping table;
[0153] Based on the importance of different parameters in the parameter encoding mapping table, different Reed-Solomon encoding error correction bits are assigned to each parameter to obtain the allocation result; double check encoding is performed on the key parameters in the allocation result to generate error correction encoding rules.
[0154] Layered encoding is performed based on preprocessed encoding data and error correction encoding rules, and parameter type identifiers are added to generate encoded data streams;
[0155] The encoded data stream is laid out according to a predetermined specification to generate an encrypted QR code.
[0156] According to one aspect of this application, step S4 further comprises:
[0157] S41. Read the fusion feature matrix; extract the historical confidence data for the most recent 30 days from the database; construct a factor graph network and set the message propagation iteration number to 50; calculate the marginal probability of nodes through message propagation; generate data confidence scores in the range [0,1] based on the marginal probability values.
[0158] S42. Obtain the data confidence score; read the historical anomaly pattern library containing 100 typical anomaly patterns from the system database; construct a graph neural network with 3 hidden layers (64, 32, and 16 nodes respectively); obtain the anomaly type identifier through network matching analysis; calculate the anomaly degree value in the range [0,1] based on the matching similarity.
[0159] S43. Read the anomaly type identifier and anomaly severity value; load 5 levels of early warning rule templates from the rule base; map the anomaly type identifier to an early warning level between 1 and 5 according to the rule templates; generate early warning description information containing anomaly type, severity, and recommended measures from the predefined information template library based on the early warning level and anomaly severity value.
[0160] In one embodiment of this application, based on the fused feature matrix R and historical confidence data C, a confidence propagation algorithm is applied to calculate data confidence, resulting in a data confidence score S. Based on the data confidence score S and the historical anomaly pattern library P, an attention-based graph neural network is used for anomaly pattern matching, outputting anomaly type identifier L and anomaly severity value G. Upon receiving the anomaly type identifier L and anomaly severity value G, warning information is generated based on a predefined rule template, resulting in a warning level Y and warning description information I.
[0161] This embodiment analyzes fused feature vectors and historical confidence data, constructs a factor graph network, and calculates the marginal probability of nodes through 50 iterations of message propagation to obtain a data confidence score ranging from [0,1]. A three-layer graph neural network (with 64, 32, and 16 nodes respectively) is used to perform anomaly pattern matching on the data. Combined with a library of 100 typical anomaly patterns, it can accurately identify anomaly types and quantify the degree of anomaly. Based on a predefined 5-level early warning rule template, combined with anomaly type identifiers and anomaly degree values, early warning information containing specific action recommendations is generated. This evaluation method based on confidence propagation and graph neural networks can not only assess the credibility of data in real time but also promptly detect and warn of anomalies. Especially in scenarios involving multiple stakeholders, such as industrial wastewater monitoring, objective confidence scoring and standardized early warning mechanisms can provide reliable decision-making basis for regulatory authorities and clear rectification guidance for enterprises.
[0162] According to one aspect of this application, step S41 further comprises:
[0163] S411. Obtain the fused feature matrix and historical confidence data for the most recent 30 days; construct a factor graph with a bipartite graph structure; define variable nodes and factor nodes; generate the initial factor graph.
[0164] S412. Read the initial factor graph; design message update rules, including message passing equations from variables to factors and from factors to variables; initialize node messages; generate a set of message propagation rules.
[0165] S413. Obtain the message propagation rule set; execute message propagation iterations, with a maximum of 50 iterations; calculate node belief values; generate node confidence distributions.
[0166] S414. Read the node confidence distribution; calculate the marginal probability; apply the sigmoid function to map to the [0,1] interval; generate the data confidence score.
[0167] This embodiment constructs a bipartite graph-based factor graph, using fused feature vectors and 30-day historical confidence data as input. It defines message propagation rules between variable nodes and factor nodes, achieving full message propagation through 50 iterations. During message updates, transfer equations are designed for both variable-to-factor and factor-to-variable communication, ensuring effective information flow within the network. The final data confidence score is generated by calculating node belief values and mapping marginal probabilities to the [0,1] interval using the sigmoid function. This factor graph-based confidence propagation method fully utilizes the statistical regularities inherent in historical data, accurately assessing the credibility of current data. Especially in monitoring scenarios with significant data fluctuations, such as industrial wastewater, the probabilistic graphical model effectively distinguishes between normal fluctuations and abnormal changes, providing a reliable evaluation mechanism for data authenticity verification.
[0168] According to one aspect of this application, step S42 further comprises:
[0169] S421. Read the data confidence score and 100 typical anomaly patterns; construct a three-layer graph neural network with 64, 32, and 16 nodes; initialize the network parameters; generate the graph network model.
[0170] S422. Obtain the graph network model; perform graph convolution operations, using ReLU as the activation function; extract graph feature representations; generate abnormal pattern features.
[0171] S423. Read the features of the abnormal pattern; calculate the cosine similarity with the typical pattern; select the type with the highest similarity as the abnormal type identifier; map the maximum similarity value to the abnormality degree value.
[0172] This embodiment constructs a three-layer graph neural network (64, 32, and 16 nodes respectively) to process confidence data and matches it against an anomaly database containing 100 typical patterns. The ReLU activation function is used during graph convolution operations, and the extracted graph feature representations effectively capture the structural features of data anomaly patterns. By calculating the cosine similarity with typical patterns, the type with the highest similarity is selected as the anomaly type identifier, and the maximum similarity value is mapped to an anomaly severity value. This graph neural network-based anomaly identification method considers both the temporal characteristics of the monitoring data and incorporates spatial topological information, making it particularly suitable for monitoring scenarios with complex pipeline structures, such as industrial parks. By comparing with a rich anomaly pattern database, it can not only accurately identify anomaly types but also quantify the severity of anomalies, providing a reliable basis for subsequent early warning decisions.
[0173] According to one aspect of this application, the steps for calculating the data confidence score include:
[0174] Obtain the fusion feature matrix and pre-stored historical confidence data from the encrypted QR code; construct a bipartite graph factor graph, define variable nodes and factor nodes, and generate the initial factor graph.
[0175] Based on the initial factor graph, message update rules are constructed, including message passing equations from variables to factors and from factors to variables;
[0176] Based on the message update rules, message propagation iterations are performed, and node belief values are calculated to obtain the node confidence distribution; the maximum number of iterations for message propagation is c; c is a preset value.
[0177] Marginal probabilities are calculated based on the node confidence distribution and mapped to the [0,1] interval using the sigmoid function to obtain the data confidence score.
[0178] According to one aspect of this application, the steps for generating corresponding warning description information include:
[0179] Construct a graph neural network with three hidden layers;
[0180] Input the data confidence scores into a graph neural network, perform graph convolution operations, and generate abnormal pattern features;
[0181] Calculate the cosine similarity between the features of the anomaly pattern and D pre-stored typical anomaly patterns; where D is a preset value.
[0182] The type with the highest cosine similarity is selected as the anomaly type identifier, and the maximum similarity value in the cosine similarity is mapped to the anomaly degree value;
[0183] Based on the anomaly type identifier and anomaly severity value, an early warning level is generated from a predefined rule template;
[0184] Based on the warning level, a warning description message containing the anomaly type, anomaly severity, and recommended measures is generated from a predefined information template library.
[0185] Existing data authenticity verification technologies lack tamper-proof mechanisms based on physical characteristics at the sensor level, and existing digital encryption schemes cannot prevent attackers from falsifying data by simulating sensor signals. In the data preprocessing stage, traditional fixed-window smoothing methods are prone to artificially segmenting the emission process, resulting in the loss of key features. In the feature extraction process, the physicochemical characteristics of industrial wastewater quality parameters are not fully considered, and standardized processing methods can easily weaken or lose the feature information of certain parameters. In spatial correlation analysis, only the physical distance between monitoring points is considered, ignoring the impact of pipeline structure on pollutant transport, leading to inaccurate spatial feature extraction. In the multimodal feature fusion stage, existing methods often use simple feature splicing or weighted averaging, which is difficult to capture the complex correlations between different features. In time-series data analysis, fixed feature extraction windows are difficult to adapt to the intermittent discharge characteristics of industrial wastewater, and key changing features are easily missed.
[0186] For the physical anti-counterfeiting of sensors, hardware-level anti-counterfeiting is achieved by collecting three physical quantities: sensor operating voltage fluctuation, temperature response characteristics, and dielectric impedance, combined with Fourier transform to generate a dynamic fingerprint. For data segmentation, an adaptive windowing technique based on data change characteristics is employed, with the window size dynamically adjusted according to the data change rate, effectively preventing the loss of key features. For parameter feature preservation, differentiated standardization strategies are used for the physicochemical properties of different water quality parameters, such as piecewise linear mapping for pH and logarithmic mapping for concentration parameters. For the impact of pipe networks, a flow direction matrix is constructed based on the pipe network topology, and pollutant transport time is calculated considering factors such as pipe diameter and flow velocity, achieving accurate spatial feature extraction. For feature fusion, a multi-head attention mechanism is used to achieve dynamic weight allocation of features, effectively capturing the complex correlations between different features.
[0187] This application collects sensor physical features and generates dynamic fingerprints; performs adaptive preprocessing and feature extraction on water quality data; verifies integrity based on spatiotemporal feature sequences; generates encrypted QR codes through multimodal feature fusion; assesses data consistency and generates early warning information. This method for verifying the authenticity of industrial wastewater monitoring data based on multimodal feature fusion constructs a complete technology chain from sensor physical feature extraction, data preprocessing, spatiotemporal feature verification, multimodal feature fusion to confidence assessment, achieving comprehensive assurance of the authenticity of industrial wastewater monitoring data. Specifically, it proposes for the first time a dynamic fingerprint mechanism based on sensor physical features, ensuring the credibility of the data source from a hardware perspective; designs an adaptive data preprocessing method adapted to the characteristics of industrial wastewater, improving the ability to identify abnormal features; implements a verification mechanism based on spatiotemporal feature fusion, which can effectively identify local abnormal emissions or data tampering; uses a multi-head attention mechanism for feature fusion, improving the reliability of verification results; and constructs an objective data credibility assessment system by combining confidence propagation and graph neural networks. This invention constructs a multi-level, highly reliable data authenticity verification system, particularly suitable for complex monitoring scenarios such as industrial parks. By combining physical feature verification, data feature verification, and spatiotemporal feature verification, data tampering can be effectively prevented, abnormal emissions can be detected in a timely manner, and reliable technical support can be provided for environmental supervision.
[0188] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the protection scope of the present invention.
Claims
1. A method for verifying the authenticity of industrial wastewater monitoring data based on multimodal feature fusion, characterized in that, include: The sensor's operating voltage fluctuation value, temperature response characteristics, and dielectric impedance are read, and a Fourier transform is performed to obtain the sensor's characteristic spectrum matrix. Construct feature vectors and map them to generate dynamic fingerprints for the sensor; Perform feature extraction to generate fingerprint feature codes; Acquire monitoring data, standardize the monitoring data to obtain standardized water quality data; extract water quality feature vectors based on the standardized water quality data; The monitoring data includes pH value, COD value, and turbidity; Data from adjacent monitoring points is acquired, and a standardized weight matrix is constructed based on pre-stored pipeline topology data. Spatial autocorrelation analysis is performed on water quality feature vectors based on the standardized weight matrix to obtain spatial feature vectors. The spatial feature vectors are then input into a feature fusion unit to generate a spatiotemporal fusion feature matrix. An encrypted QR code is generated based on fingerprint feature code, spatiotemporal fusion feature matrix and standardized water quality data; Receive encrypted QR codes, calculate data confidence scores, and generate corresponding warning description information; In constructing the spatial feature vector, the temporal feature vector is constructed through the following steps: Read historical data sequences at a predetermined time from the data storage unit; Based on historical data sequences, calculate the rate of change of data and determine the inflection point of data change; The size of the baseline window is determined based on the time interval between adjacent data change inflection points; When the data change rate exceeds a preset threshold, the baseline window size is reduced to 1 / a of the original size; when the data change rate does not exceed the preset threshold, the baseline window size is increased to a times the original size; thus obtaining the adjusted window size. Statistical features within each window are calculated based on historical data sequences and adjusted window sizes to generate time-series statistical features. Based on time-series statistical features and water quality feature vectors, a time feature vector is generated; where a is a natural number greater than 0. The process of generating the spatiotemporal fusion feature matrix also includes: Tensor decomposition is performed on the temporal and spatial feature vectors to generate a spatiotemporal fusion feature matrix. The steps of tensor decomposition include: Construct a third-order tensor with dimensions M×N×K; Perform CP decomposition on the third-order tensor and select the decomposition result with rank P; Based on the decomposition results, the reconstruction error is evaluated using the Tucker norm, and feature combinations with reconstruction errors less than a preset threshold are selected. The features are combined and recombined into a spatiotemporal fusion feature matrix, where M, N, K and P are preset values.
2. The method for verifying the authenticity of industrial wastewater monitoring data based on multimodal feature fusion according to claim 1, characterized in that, The steps for constructing a standardized weight matrix based on pre-stored pipeline topology data include: Read pipeline topology data from the industrial park's pipeline database, and extract the upstream contribution coefficient and downstream impact coefficient of each monitoring point based on the pipeline topology data; The flow capacity coefficient is calculated based on the pre-stored pipe diameter, flow velocity, and hydraulic gradient. A network feature vector is generated based on the upstream contribution coefficient, downstream influence coefficient, and flow capacity coefficient. Based on the pipeline feature vector, the shortest pipeline path between nodes is calculated, and a flow direction correlation matrix is constructed, where upstream nodes take a positive value of 1, downstream nodes take a negative value of 1, and no direct connection takes a value of 0. Based on the pipeline network feature vector and the flow direction correlation matrix, the diffusion attenuation coefficient of pollutants at different flow velocities is calculated. Based on the diffusion attenuation coefficient, the pollutant transfer time between nodes is calculated using the diffusion equation to generate a spatiotemporal transfer matrix. Weights are assigned to monitoring points based on the spatiotemporal transfer matrix, and nonlinear adjustments are made. The coefficients of the nonlinear adjustments are proportional to the flow velocity and pipe diameter, resulting in a standardized weight matrix.
3. The method for verifying the authenticity of industrial wastewater monitoring data based on multimodal feature fusion according to claim 1, characterized in that, The steps for standardizing monitoring data to obtain standardized water quality data include: The slope of the pre-configured normalization function is dynamically adjusted based on the data distribution density of the monitoring data. Based on the adjusted slope, S-curve normalization is performed on data-dense areas; based on the adjusted slope, linear mapping is performed on data-sparse areas; preliminary standardized water quality data are generated. Based on preliminary standardized water quality data, the skewness and kurtosis of the data distribution are calculated and corrected using Box-Cox transformation to generate standardized water quality data.
4. The method for verifying the authenticity of industrial wastewater monitoring data based on multimodal feature fusion according to claim 1, characterized in that, The steps to generate an encrypted QR code include: Read the fingerprint feature code and spatiotemporal fusion feature matrix; collect parameters including temperature, humidity and air pressure from environmental sensors to form an environmental parameter vector; perform timestamp alignment and dimension unification processing on the fingerprint feature code, spatiotemporal fusion feature matrix and environmental parameter vector to generate a 64×64 aligned feature matrix; Obtain the aligned feature matrix; construct 8 attention heads, each with a dimension of 8; calculate the feature weights through multi-head attention; weight the features according to the feature weights to generate a 128-dimensional fusion feature matrix; Read the fused feature matrix, standardized water quality data, and sensor dynamic fingerprints; encode the data using the Reed-Solomon encoding algorithm; and convert the encoding results into a 29×29 encrypted QR code.
5. The method for verifying the authenticity of industrial wastewater monitoring data based on multimodal feature fusion according to claim 1, characterized in that, The steps for calculating the confidence score of the data include: Obtain the fusion feature matrix and pre-stored historical confidence data from the encrypted QR code; construct a bipartite graph factor graph, define variable nodes and factor nodes, and generate the initial factor graph. Based on the initial factor graph, message update rules are constructed, including message passing equations from variables to factors and from factors to variables; Based on the message update rules, message propagation iterations are performed, and node belief values are calculated to obtain the node confidence distribution; the maximum number of iterations for message propagation is c; c is a preset value. Marginal probabilities are calculated based on the node confidence distribution and mapped to the [0,1] interval using the sigmoid function to obtain the data confidence score.
6. The method for verifying the authenticity of industrial wastewater monitoring data based on multimodal feature fusion according to claim 1, characterized in that, The steps to generate the corresponding warning description information include: Construct a graph neural network with three hidden layers; Input the data confidence scores into a graph neural network, perform graph convolution operations, and generate abnormal pattern features; Calculate the cosine similarity between the features of the anomaly pattern and D pre-stored typical anomaly patterns; where D is a preset value. The type with the highest cosine similarity is selected as the anomaly type identifier, and the maximum similarity value in the cosine similarity is mapped to the anomaly degree value; Based on the anomaly type identifier and anomaly severity value, an early warning level is generated from a predefined rule template; Based on the warning level, a warning description message containing the anomaly type, anomaly severity, and recommended measures is generated from a predefined information template library.
7. The method for verifying the authenticity of industrial wastewater monitoring data based on multimodal feature fusion according to claim 1, characterized in that, The steps for extracting water quality feature vectors based on standardized water quality data include: Based on standardized water quality data, a covariance matrix is constructed and condition number analysis is performed. When the condition number is less than a preset threshold, the covariance matrix is directly used as the optimized feature covariance matrix. When the condition number is greater than the preset threshold, the covariance matrix is regularized to obtain the optimized feature covariance matrix. Based on the optimized feature covariance matrix, the improved QR decomposition algorithm is used to calculate the eigenvalues, generating an eigenvalue sequence and the corresponding eigenvector set; Based on the feature value sequence and the corresponding feature vector set, principal components are selected according to a preset cumulative contribution rate threshold to obtain the selected feature vectors; The selected feature vectors are orthogonalized to generate water quality feature vectors.
8. The method for verifying the authenticity of industrial wastewater monitoring data based on multimodal feature fusion according to claim 7, characterized in that, The steps for calculating eigenvalues and generating a sequence of eigenvalues and a corresponding set of eigenvectors using the improved QR decomposition algorithm include: Weighting factors are set according to the importance of pollutant parameters in the monitoring data; The optimized feature covariance matrix is decomposed into an upper triangular matrix and an orthogonal matrix; The decomposition process is iteratively optimized based on weighting factors; During the iteration process, the rate of change of eigenvalues and eigenvectors is calculated, and the calculation step size is dynamically adjusted; When the rate of change of both the eigenvalue and the eigenvector is less than the corresponding preset threshold, the final eigenvalue is determined. The final feature values are sorted in descending order, and the cumulative contribution rate of the final feature values is calculated to generate a feature value sequence and a corresponding feature vector set.
Citation Information
Patent Citations
Intelligent automobile remote unlocking method and device based on multi-modal biological characteristics
CN118470836A
Intelligent parking building vehicle dynamic supervision method and system based on AI intelligent engine
CN119479309A