Water pollution early warning method and system based on big data

By collecting and integrating hydropower usage data with water quality monitoring data, performing anomaly clustering and multi-dimensional parameter verification, and combining historical behavior pattern matching, the pollution source was identified and the diffusion path was simulated. This solved the problem of the inability to quickly identify the pollution source in existing technologies, and achieved accurate pollution source location and early warning notification.

CN121301831AActive Publication Date: 2026-01-09重庆知行数联智能科技有限责任公司

Patent Information

Application Number
CN202511872428.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-01-09
Estimated Expiration
2045-12-12

AI Technical Summary

Technical Problem

Existing technologies cannot quickly and accurately pinpoint the source of pollution after a pollution incident, resulting in the data value not being fully realized.

Method used

Data on hydropower usage and water quality monitoring are collected, multi-source data are standardized and fused, data attributes are extracted and anomaly clustering analysis is performed, multi-dimensional parameter correlation verification is conducted, historical behavior pattern matching and identity recognition are combined, benchmark coordinate retrieval and multi-source data cross-validation are performed, the specific location information of potential pollution sources is determined, and pollution diffusion simulation and impact range determination are carried out.

Benefits of technology

Through deep data fusion and multi-dimensional parameter correlation verification, it is possible to accurately identify the pollution source equipment or enterprise, dynamically deduce the pollutant diffusion path and affected area, generate accurate early warning notification data, and support environmental regulatory departments in formulating emergency strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121301831A_ABST
    Figure CN121301831A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of water pollution monitoring and big data analysis, and discloses a water pollution early warning method and system based on big data, and the method comprises the steps: collecting water and electricity use data and water quality monitoring data, and carrying out the standardized fusion, and obtaining an integrated data set; performing abnormal clustering analysis and multi-dimensional parameter relevance verification according to the integrated data set to obtain an abnormal linkage distribution result; performing historical behavior pattern matching and identity recognition based on the abnormal linkage distribution result to obtain hidden abnormal recognition data; performing reference coordinate retrieval and cross validation according to the hidden anomaly identification data, and determining specific position information of a potential pollution source; and carrying out pollution diffusion simulation on the position information, determining distribution information of an affected area and generating early warning notification data. According to the method, hidden pollution discharge behaviors can be effectively identified, and accurate positioning of pollution sources and dynamic evaluation of diffusion risks are realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of water pollution monitoring and big data analysis, and particularly relates to a water pollution early warning method and system based on big data. BACKGROUND

[0002] At present, with the acceleration of industrialization, water resource security is facing severe challenges. In order to protect water environment quality, real-time monitoring technology based on Internet of Things has been widely used, and regulatory authorities have increasingly paid attention to using big data analysis technology to improve the intelligent level of environmental governance, trying to obtain valuable regulatory information from massive monitoring data.

[0003] In one prior art, water pollution supervision mainly relies on real-time reading of water quality sensor data and over-limit alarm, or only calculates the total amount of pollution discharged by enterprises through simple industrial data statistics. However, this traditional monitoring mode often separates the production behavior of enterprises (such as power load, production conditions) and the end-of-pipe discharge behavior, and lacks deep integration of multi-dimensional industrial data analysis. Simply relying on water quality data is difficult to trace the specific emission source in a park where multiple enterprises coexist, and single energy consumption monitoring cannot directly prove the violation of pollution discharge, resulting in that the value of data cannot be fully released.

[0004] Therefore, in the prior art, there is a technical problem that the pollution source cannot be quickly and accurately locked after the pollution event occurs. SUMMARY

[0005] The present application provides a water pollution early warning method and system based on big data to solve the technical problem in the prior art that the pollution source cannot be quickly and accurately locked after the pollution event occurs.

[0006] In a first aspect, to solve the above technical problem, the present application provides a water pollution early warning method based on big data, comprising:

[0007] Collecting water and electricity usage data and water quality monitoring data, and performing multi-source data standardization fusion processing to obtain an integrated data set;

[0008] According to the integrated data set, data attribute extraction and abnormal clustering analysis are performed to obtain abnormal pattern data;

[0009] For the abnormal pattern data, multi-dimensional parameter correlation verification is performed to obtain abnormal linkage distribution results;

[0010] If the abnormal linkage distribution results meet the preset linkage determination condition, historical behavior pattern matching and identity recognition are performed to obtain implicit abnormal recognition data;

[0011] According to the implicit abnormality identification data, benchmark coordinate retrieval and multi-source data cross verification are performed to determine potential pollution source specific location information;

[0012] According to the potential pollution source specific location information, pollution diffusion simulation and influence range determination are performed to determine affected area distribution information;

[0013] According to the potential pollution source specific location information and the affected area distribution information, early warning generation and transmission processing are performed to obtain early warning notification data.

[0014] In a second aspect, the present application provides a water pollution early warning system based on big data, comprising:

[0015] A data processing module is configured to collect water and electricity usage data and water quality monitoring data, and perform multi-source data standardization and fusion processing to obtain an integrated data set;

[0016] An abnormality mining module is configured to perform data attribute extraction and abnormality clustering analysis according to the integrated data set to obtain abnormality pattern data;

[0017] An association verification module is configured to perform multi-dimensional parameter association verification on the abnormality pattern data to obtain abnormality linkage distribution results;

[0018] An identity recognition module is configured to perform historical behavior pattern matching and identity recognition to obtain implicit abnormality identification data if the abnormality linkage distribution results meet preset linkage determination conditions;

[0019] A positioning verification module is configured to perform benchmark coordinate retrieval and multi-source data cross verification according to the implicit abnormality identification data to determine potential pollution source specific location information;

[0020] A diffusion simulation module is configured to perform pollution diffusion simulation and influence range determination on the potential pollution source specific location information to determine affected area distribution information;

[0021] An early warning processing module is configured to perform early warning generation and transmission processing according to the potential pollution source specific location information and the affected area distribution information to obtain early warning notification data.

[0022] Compared with the prior art, the present application has the following advantages:

[0023] (1) The present application breaks the information island between industrial production data and environmental monitoring data through collecting water and electricity use data and water quality monitoring data and performing multi-source data standardization fusion processing, and uses data attribute extraction and abnormal clustering analysis technology to mine potential abnormal patterns. This deep data fusion mechanism can effectively identify the implicit correlation between production activities and pollution behavior from massive data, so as to find strong hidden and complex illegal discharge behavior, and significantly improve the active discovery ability of potential pollution risk.

[0024] (2) The present application verifies the multi-dimensional parameter correlation of abnormal pattern data, and combines the pre-set device characteristic fingerprint library to perform historical behavior pattern matching and identity recognition. This method not only quantifies the causal correlation strength between production activities and water quality abnormalities from a statistical point of view, effectively eliminates false positives caused by sensor errors or environmental fluctuations, but also accurately locates specific pollution source equipment or enterprise identity based on behavior characteristic fingerprints, solving the traceability problem of traditional monitoring methods that only know pollution but not who discharges.

[0025] (3) The present application verifies the multi-dimensional parameter correlation of abnormal pattern data, and combines the pre-set device characteristic fingerprint library to perform historical behavior pattern matching and identity recognition. This method not only quantifies the causal correlation strength between production activities and water quality abnormalities from a statistical point of view, effectively eliminates false positives caused by sensor errors or environmental fluctuations, but also accurately locates specific pollution source equipment or enterprise identity based on behavior characteristic fingerprints, solving the traceability problem of traditional monitoring methods that only know pollution but not who discharges. BRIEF DESCRIPTION OF DRAWINGS

[0026] Figure 1 is a water pollution early warning method flowchart provided by the first embodiment of the present application based on big data;

[0027] Figure 2 is a water pollution early warning system structure diagram provided by the second embodiment of the present application based on big data. DETAILED DESCRIPTION

[0028] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0029] Referring to Figure 1 , the first embodiment of the present application provides a water pollution early warning method based on big data, including the following steps:

[0030] S11, collect water and electricity use data and water quality monitoring data, and perform multi-source data standardization fusion processing to obtain an integrated data set;

[0031] S12, according to the integrated data set, perform data attribute extraction and abnormal clustering analysis to obtain abnormal pattern data;

[0032] S13, for the abnormal pattern data, perform multi-dimensional parameter correlation verification to obtain abnormal linkage distribution results;

[0033] S14, if the abnormal linkage distribution results meet the preset linkage determination condition, perform historical behavior pattern matching and identity recognition to obtain implicit abnormal recognition data;

[0034] S15, according to the implicit abnormal recognition data, perform reference coordinate retrieval and multi-source data cross verification to determine the specific location information of the potential pollution source;

[0035] S16, for the specific location information of the potential pollution source, perform pollution diffusion simulation and influence range determination to determine the affected area distribution information;

[0036] S17, according to the specific location information of the potential pollution source and the affected area distribution information, perform early warning generation and transmission processing to obtain early warning notification data.

[0037] In step S11, water and electricity use data and water quality monitoring data are collected, and multi-source data standardization fusion processing is performed to obtain an integrated data set, including:

[0038] The water and electricity use data and the water quality monitoring data are collected and time axis calibration is performed to obtain a first data set;

[0039] The first data set is subjected to multi-dimensional verification and abnormal data point correction to obtain an intermediate data set;

[0040] According to the intermediate data set, normalization processing is performed to obtain the integrated data set.

[0041] It should be noted that the collection of water and electricity use data and water quality monitoring data is achieved by deploying industrial-grade intelligent electric meters and multi-parameter water quality analyzers in the target monitoring area. The intelligent electric meter collects real-time power load data (unit: kilowatt) of the enterprise at a preset power sampling frequency (for example, every 15 minutes), and the water quality analyzer synchronously collects three key water quality index data of chemical oxygen demand (COD), ammonia nitrogen and pH value. Time axis calibration is achieved by using linear interpolation resampling method. This method first defines a standard time reference axis (for example, every 10 minutes is a scale starting from the whole point time), and then for the original discrete data points collected from different sources, the formula computing its value on a standard time scale, wherein and are the immediately preceding and following sample points in time in the raw data. By this operation, the heterogeneous data of different frequencies are mapped to a uniform time scale, obtaining the first data set.

[0042] It is worth noting that the determination of the preset power sampling frequency is based on the spectral analysis of the historical industrial production load fluctuation period. By performing fast Fourier transform (FFT) on the historical electricity consumption data, the main frequency component of the load change is identified, and according to the Nyquist sampling theorem, a frequency greater than twice the main frequency component is selected as the sampling frequency to ensure that the fluctuation details of the production activities can be restored without distortion.

[0043] It is worth noting that the multi-dimensional verification and abnormal data point correction of the first data set is achieved by using the three standard deviation rule (3-Sigma Rule) and the sliding average interpolation method based on statistics. This step aims to eliminate outliers caused by sensor failure or transmission error under the original physical dimension, preventing them from interfering with subsequent normalization calculations. The verification operation first calculates the local mean and the local standard deviation of each feature dimension in the sliding window If the raw data point at a certain time satisfies , then it is determined that the point is an outlier. The correction operation replaces it with the arithmetic mean of the adjacent valid data points before and after the abnormal point, obtaining the intermediate data set.

[0044] It is worth noting that the determination of the size of the sliding window (for example, 24 hours) is based on the autocorrelation analysis of the historical data smoothness. Specifically, the autocorrelation function of the historical data is calculated, and the calculation formula is

[0045]

[0046] where N is the total length of the data, is the lag time, is the mean. The autocorrelation function value The lag time at which the correlation decays to a preset correlation decay threshold (e.g., 0.5) is used as the sliding window size. This correlation decay threshold is set based on the definition of coherence time in signal processing, aiming to ensure that the data points within the window maintain statistically significant correlation and consistent distribution, avoiding the introduction of non-stationary trends due to an excessively large window or the instability of statistical characteristics due to an excessively small window. Specifically, the autocorrelation function of historical data under normal operating conditions is calculated, and the time point at which the autocorrelation value decays to 50% of the initial value is selected as the window size benchmark to ensure that the data within the window maintains significant correlation while avoiding interference from non-stationary trends.

[0047] It should be noted that the normalization process performed on the intermediate dataset aims to eliminate the order-of-magnitude differences between different physical units (such as kilowatts and milligrams per liter), ensuring the convergence of subsequent algorithms. This embodiment employs the Min-Max Normalization method. For each feature dimension X in the intermediate dataset (specifically, electricity consumption, COD, ammonia nitrogen, or pH value), the data at all time points are traversed to obtain the corrected maximum value. and minimum value And using the formula, All data are mapped to the dimensionless interval [0,1]. The processed data is the integrated dataset.

[0048] For example, electricity consumption data from a chemical plant was collected as [100kW, 102kW, 5000kW, 101kW] (where 5000kW is a spike caused by a transmission error), and COD data was collected as [450mg / L, 460mg / L, 455mg / L, 458mg / L]. After time axis calibration, the two data were aligned. First, multi-dimensional verification was performed. For the electricity consumption data, a local mean was calculated based on the neighborhood normal values ​​(100, 102, 101). Standard deviation Regarding the value of 5000kW, the degree of deviation is... far greater than (i.e., 3) is judged as abnormal. The correction operation replaces it with the average of the preceding and following values ​​(i.e., 101.5kW). Then, normalization is performed, at which point the maximum value of the electricity consumption data is obtained. minimum value The corrected value of 101.5 kW is mapped to... If anomalies are not removed first, 5000kW will be taken as the maximum value, causing normal fluctuations (100-102kW) to be compressed into a very small range of [0, 0.0004], losing their analytical value. The final dimensionless data matrix that retains the characteristics of normal fluctuations is the integrated data set.

[0049] In step S12, based on the integrated data set, data attribute extraction and anomaly clustering analysis are performed to obtain anomaly pattern data, including:

[0050] Based on the integrated data set, density values ​​are calculated and preliminary screening is performed to obtain a filtered data set.

[0051] The filtered data set is clustered and organized to obtain clustered grouped data and isolated data subsets;

[0052] Based on the clustered grouped data and the isolated data subset, a time-dimensional distribution pattern analysis is performed to obtain an anomaly marker set;

[0053] Based on the set of anomaly markers, a distribution pattern comparison analysis is performed to obtain the anomaly pattern data.

[0054] It should be noted that the density value calculation and preliminary screening based on the integrated dataset are implemented using a density estimation algorithm based on k-nearest neighbor distance. This operation first traverses each multidimensional data point in the integrated dataset (including normalized electricity consumption and water quality characteristics); then, it calculates the Euclidean distance between the data point and all other points in the dataset, and finds the λ nearest neighbor points; finally, it calculates the average distance from the data point to these λ neighbors. The reciprocal of the average distance is used as the local density evaluation value of the point. If the local density value of a data point is lower than the preset sparsity threshold (meaning that the point is far away from its neighbors and belongs to an outlier), it is marked as a potential anomaly candidate, and these marked data points and their original features are combined into the filtered data set.

[0055] It is worth noting that the determination of λ (e.g., 20) and the preset sparsity threshold is based on statistical distribution analysis of historical normal production data. Specifically, the average k-nearest neighbor distance of all points in the historical data is calculated, and a distance probability distribution histogram is constructed. The reciprocal of the distance value corresponding to the 95th quantile of this distribution is selected as the sparsity threshold to ensure that outlier data that are statistically significantly deviated from the normal cluster centers can be filtered out.

[0056] It should be noted that the clustering of the selected data set is implemented using a density-based spatial clustering algorithm (DBSCAN). This algorithm does not require pre-specifying the number of clusters and can automatically identify clusters of arbitrary shapes and separate noise. Data points in the selected data set are mapped to a multi-dimensional feature space; scanning is performed using preset neighborhood radius (Eps) and minimum number of points (MinPts) parameters. If a point has more than MinPts of neighbors within its Eps neighborhood, it is marked as a core point and forms a new cluster; if a point cannot reach any core point within its Eps neighborhood, it is identified as a noise point. All the formed clusters constitute the clustered grouped data (representing some regular abnormal behavior, such as continuous high-load sewage discharge), while the points identified as noise constitute the isolated data subset (representing sudden, irregular anomalies, such as instantaneous sensor fluctuations).

[0057] It is worth noting that the neighborhood radius (Eps) is determined based on the elbow rule analysis of the K-distance graph. Specifically, the distance from each data point to its λ-th nearest neighbor is calculated, and these distances are sorted from largest to smallest and plotted as a curve. The distance value corresponding to the inflection point where the curve drops sharply and then flattens out is selected as Eps. The minimum number of points (MinPts) is determined based on the dimension of the feature space (4 in this embodiment) and the minimum sample size of historical anomalies. According to the formula, the minimum number of points is greater than or equal to the dimension plus 1, and combined with the statistics of the minimum number of data points contained in historical short-term anomalies (such as anomalies lasting more than 15 minutes), the larger of the two values ​​(e.g., 5) is selected as the preset value.

[0058] It should be noted that the analysis of the temporal distribution patterns of the clustered grouped data and the isolated data subset is achieved using a sliding time window frequency statistics method. A fixed-length time window (e.g., 1 hour) and a step size (e.g., 10 minutes) are defined and slid along the time axis. For each window, the total number of sampling points in the integrated data set included within that window is first counted. Then, count the number of outlier data points within the window that belong to the clustered data or the isolated data subset. Finally, use the formula Calculate the anomaly density ratio. If the anomaly density ratio R of a window exceeds a preset anomaly proportion threshold, an anomaly time segment object is generated. The object generation process involves recording the start and end timestamps of the window; identifying the dominant anomaly type within the window (i.e., the most frequently occurring cluster label or isolated marker); and calculating the arithmetic mean of all anomalies within the window across each feature dimension. This information is then assigned to the object's Start_Time, End_Time, Type, and Avg_Features attributes, thus completing the object's instantiation. The collection of all these objects constitutes the anomaly marker set.

[0059] It is worth noting that the preset anomaly percentage threshold is determined based on statistical analysis of historical false alarm rates. Specifically, the time periods marked as normal fluctuations in historical monitoring records are reviewed, the distribution of anomaly percentages in these periods under a sliding window is calculated, and the 99th percentile of this distribution is selected as the threshold (e.g., 0.5) to ensure that only periods with a high concentration of anomalies are marked, thereby effectively reducing the false alarm rate caused by occasional noise.

[0060] For example, the total number of sampling points was counted within the time window from 10:00 to 11:00. There are 6 points (one point every 10 minutes). Among them, 5 belong to the high power consumption-high COD cluster (i.e., anomalies) identified by the DBSCAN algorithm. The anomaly density ratio is calculated as follows: The ratio exceeds the preset threshold of 0.5, therefore an anomaly flag is generated: {Start:"10:00",End:"11:00",Type:"Cluster-1",Avg_COD:0.8}.

[0061] It should be noted that the distribution pattern comparison analysis based on the aforementioned anomaly marker set aims to integrate scattered anomaly fragments into a pattern description with business semantics. This operation is achieved through feature vectorization and centroid extraction. First, a time segment merging operation is performed, sorting all anomaly time segments according to their start time; the sorted segments are then traversed, determining whether the end time of the current segment is later than or equal to the start time of the next segment; if so, the two segments are merged, and the end time of the merged segment is updated to the maximum of the two, and this process is repeated until there are no more segments to merge. Then, for each consecutive anomaly event after merging, all the original data points contained within it are extracted, and the arithmetic mean vector of these points in the multidimensional feature space is calculated. This vector is the geometric centroid, and it is defined as the feature fingerprint vector of the event (e.g., [average electricity consumption 0.9, average COD 0.8]). Finally, an encapsulation operation is performed to create a structured data entity (such as a JSON object), assign the merged start and end times to the Duration field, assign the calculated feature fingerprint vector to the Feature_Vector field, and assign the anomaly score (calculated based on density ratio) to the Anomaly_Score field, thus obtaining the anomaly pattern data.

[0062] For example, if two consecutive anomalous segments, 10:00-11:00 and 11:00-12:00, are detected and meet the merging criteria, they are merged into a single 2-hour event (10:00-12:00). The feature values ​​of all anomalous points within this time period are obtained, and their arithmetic mean is calculated to obtain the feature vector. (These correspond to normalized electricity consumption and COD, respectively). Output the final abnormal pattern data: {ID:"Pattern-20251126-01",Duration:"10:00-12:00",Feature_Vector:[0.85,0.78],Anomaly_Score:0.9}.

[0063] In step S13, multi-dimensional parameter correlation verification is performed on the abnormal pattern data to obtain the abnormal linkage distribution results, including:

[0064] Range data extraction is performed on the abnormal pattern data to obtain a range data set;

[0065] The set of range data is compared with a preset fluctuation threshold to obtain a set of linked anomaly points;

[0066] For the set of linked anomaly points, a distribution pattern matching is performed to obtain a set of linked feature values;

[0067] Based on the set of linkage feature values, the correlation strength is determined to obtain the linkage distribution feature set;

[0068] Based on the aforementioned linkage distribution feature set, correlation segmentation analysis is performed to obtain the abnormal linkage distribution results.

[0069] It should be noted that the range data extraction of the abnormal pattern data is a backtracking operation of the original data based on the abnormal time window output in step S12. This is based on the start time recorded in the abnormal pattern data. and end time From the integrated data set generated in step S11, normalized electricity consumption sequences for the specified time period are extracted respectively. and normalized water quality index series Then, the range of the two sequences is calculated, i.e. and The two scalar values ​​and their corresponding timestamps are combined into a record and stored in the range data set.

[0070] It should be noted that the preset fluctuation threshold comparison of the aforementioned range data set is implemented using a dual threshold filtering logic. The extracted... With the preset power fluctuation threshold Compare, With respect to the preset water quality fluctuation threshold Compare. Only if the condition is met. When this occurs, it indicates that a significant change in electricity load and a drastic fluctuation in water quality have occurred simultaneously within this time window. This time window is marked as a potential linkage event, and the relevant data is stored in the linkage anomaly point set.

[0071] It is worth noting that the preset power fluctuation threshold and water quality fluctuation threshold The determination of the range is based on statistical analysis of historical stable operation data. Specifically, historical data from non-peak discharge periods (such as equipment maintenance periods and low-load operation periods) are collected from enterprises, and range samples are extracted. The Kernel Density Estimation (KDE) method is used to smooth the sample data with a Gaussian kernel function, and the probability density function curve of the range distribution is calculated. The bandwidth is selected based on the Silverman rule. The range value corresponding to a cumulative probability of 95% under this curve is selected as the corresponding fluctuation threshold to ensure that normal process fluctuations and background noise can be effectively filtered out.

[0072] It should be noted that the distribution pattern matching of the aforementioned set of linked anomalies aims to quantify the temporal correlation between electricity consumption fluctuations and water quality changes by calculating the maximum cross-correlation coefficient. This calculation simultaneously covers the detection of synchronicity (no delay) and lagging correlation (with delay). This embodiment uses the Cross-Correlation Function (CCF) for calculation. For each set of linked anomaly sequences E and W, its time lag is calculated. Normalized cross-correlation coefficients under the following conditions The calculation formula is as follows:

[0073]

[0074] in, and Represent the arithmetic mean of sequences E and W within the given time window; iterate through the preset lag range (e.g. (minutes), searching maximum value and its corresponding optimal lag time These two parameters It objectively reflects the morphological similarity and time delay of the changes between the two, and constitutes the set of linkage feature values.

[0075] It is worth noting that the preset lag range (e.g., [0, 60] minutes) is determined based on the hydraulic residence time analysis of the industrial drainage system. Specifically, the pipe length and average flow velocity from the production equipment end to the end monitoring point are measured, the theoretical maximum flow time is calculated, and then multiplied by a safety factor (e.g., 1.5 times) to cover the longest possible physical delay.

[0076] It should be noted that the determination of correlation strength based on the aforementioned set of linkage feature values ​​is achieved by calculating a comprehensive linkage index. This index... Combining morphological correlation and volatility amplitude, its calculation formula is as follows:

[0077]

[0078] This formula combines morphological similarity through multiplication. With fluctuations The coupling is represented by a scalar value. A higher value indicates greater morphological similarity and more dramatic fluctuations between the two, suggesting a stronger physical correlation. Calculate the coupling for each outlier. The value is then stored as a core feature in the linked distribution feature set.

[0079] It should be noted that the correlation segmentation analysis based on the aforementioned linkage distribution feature set is implemented using a time-based statistical method for production shifts. First, the company's pre-set production shift schedule is obtained, and the 24 hours of a day are divided into several business time periods (e.g., morning shift 08:00-16:00, afternoon shift 16:00-24:00, night shift 00:00-08:00). Then, all events in the aforementioned linkage distribution feature set are traversed to determine which time period the event's start timestamp falls into, thus establishing a mapping relationship between events and time periods. Next, statistical calculations are performed for each time period, using formulas... Calculate the average linkage index using the formula Calculate the frequency of occurrence of high-intensity linkage events (where...) for The number of events exceeding the preset intensity benchmark, (This refers to the duration of the time period); Finally, perform data encapsulation to construct a JSON object containing the fields "Time_Slot" (time period name), "Avg_Index" (average index), and "Frequency" (occurrence frequency), which is the abnormal linkage distribution result.

[0080] It is worth noting that the predetermined intensity benchmark was determined based on statistical analysis of historically confirmed pollution events. Specifically, data on historically verified pollution events caused by production were collected, and the linkage index of these events was calculated. The distribution is used as a benchmark, and the 20th percentile of the distribution is selected to filter out weakly correlated events that coincidentally overlap.

[0081] For example, in one analysis, the period from 10:00 to 10:30 was extracted using electrode difference. The water quality is extremely poor. The values ​​exceeded preset thresholds for power fluctuations (e.g., 0.5) and water quality fluctuations (e.g., 0.4), respectively. Cross-correlation analysis showed that, in the lag... At minute 10, the correlation coefficient between the two reached its maximum value. (This means that water quality begins to deteriorate 15 minutes after the machine starts). The correlation index is calculated. The results were categorized under the early morning shift for statistical analysis. The final output of the abnormal linkage distribution results showed that during the early morning shift, the average linkage index was 0.65, and the frequency of high-intensity linkage was 2 times per hour, indicating that there is a regular production and sewage discharge correlation during this period.

[0082] It should be noted that the average linkage index is calculated as follows: for a given business time period (such as the morning shift), the comprehensive linkage index of all events marked as linkage anomalies within that time period is summed, and then divided by the total number of linkage anomalies within that time period; the calculation of the high-intensity linkage frequency depends on a preset intensity benchmark; first, the number of linkage anomalies whose comprehensive linkage index exceeds the intensity benchmark within the business time period is counted; then, this number is divided by the total duration of the time period (in hours) to obtain the high-intensity linkage frequency, which is expressed as times / hour.

[0083] In step S14, if the abnormal linkage distribution result meets the preset linkage judgment condition, then historical behavior pattern matching and identity recognition are performed to obtain latent anomaly identification data, including:

[0084] If the abnormal linkage distribution results meet the preset linkage judgment conditions, then the overlapping characteristics are extracted to obtain a preliminary abnormal behavior dataset.

[0085] Historical data records and a preset device feature fingerprint database are obtained, and the preliminary abnormal behavior dataset is matched and compared with the historical data records and the preset device feature fingerprint database to obtain a set of latent abnormal behavior features.

[0086] Feature mapping analysis is performed on the set of latent abnormal behavior features to determine the associated device identity identifier, thereby obtaining the latent abnormality identification data.

[0087] It should be noted that if the abnormal linkage distribution result meets the preset linkage judgment condition, it refers to the comprehensive linkage index calculated in step S13. The value exceeds the preset effective linkage threshold. When the condition is met, overlap characteristic extraction is performed. This extraction operation is based on the time window in which the anomaly occurred. Extract the normalized electricity consumption sequence for that period from the integrated dataset. and normalized water quality index series Subsequently, the statistical eigenvectors of these two sequences are calculated. ,in Peak electricity consumption The duration of the anomaly. The optimal lag time calculated for S13 This is the ratio of the ranges of the two. This eigenvector... This constitutes the preliminary abnormal behavior dataset.

[0088] It is worth noting that the determination of the preset effective linkage threshold is based on statistical analysis of historical false alarm data. Specifically, data on events historically marked as false alarms or accidental overlaps are collected, and their linkage index samples are extracted; a Gaussian kernel function is used to perform kernel density estimation (KDE) to construct a probability density function for the linkage index of false alarm events; the value at which the cumulative probability of this function reaches 99% is selected as the threshold to ensure that only events with significant causal strength enter the subsequent pattern matching stage.

[0089] Specifically, the probability density function is expressed as follows: ,in, It is at point The estimated probability density at the location; It is the amount of historical range sample data; It is a kernel function, specified in this embodiment as a Gaussian kernel function, specifically as follows: ; is a smoothing parameter called bandwidth, used to control the smoothness of the estimate. Its selection is based on the Silverman rule, which provides an empirically optimal bandwidth estimate for the Gaussian kernel. ;in It is the standard deviation of the sample data.

[0090] It should be noted that obtaining the pre-built device feature fingerprint database refers to loading a pre-built key-value pair database. The database is built by performing multi-source log correlation and source tracing analysis on historically confirmed pollution events. Specifically, for each historical pollution event's time window, instead of relying solely on potentially missing or tampered operation logs, an energy consumption-pollution time-series correlation analysis is performed. First, based on the historical pollution event's time window, current load data sequences of all key production equipment within the corresponding time period are retrieved from the historical power monitoring database. These are then resampled using linear interpolation to the same time frequency as the water quality data, thus constructing a real-time current load curve. Second, the Granger Causality Test algorithm is used to calculate the statistical causal significance between the load change sequence of each device and the water quality deterioration sequence. Specifically, the lag order of the Granger causality test is determined based on minimizing the Akaike Information Criterion (AIC): the AIC value is calculated within a preset order range (e.g., 1-10), and the order with the lowest AIC is selected. The accompanying probability threshold of 0.05 for the Granger causality test F-statistic is set according to the statistical significance level standard to control the false positive rate below 5%. Third, devices with a causality test P-value less than the preset significance threshold (e.g., 0.05) are selected and identified as the objective pollution source devices for that event. Next, extract the feature vector corresponding to the event (same as ID); (Definition of the fingerprint database); Finally, for multiple historical feature vectors associated with the same device, their arithmetic mean vector is calculated as the standard fingerprint vector for that device. The structure of the fingerprint database is {Device_ID:Standard_Feature_Vector}.

[0091] It should be noted that the matching and comparison between the preliminary abnormal behavior dataset and the pre-set device feature fingerprint database is achieved using Euclidean distance as the similarity metric. In this process, to eliminate weight biases caused by differences in units and orders of magnitude between different feature dimensions (such as power values ​​and time duration), a secondary standardization process is first performed on each component within the feature vector before distance calculation. This embodiment uses the Z-Score standardization method to calculate the mean of all standard fingerprint vectors in the fingerprint database for each feature dimension. and standard deviation Using the formula For the current event feature vector fingerprint database vector A uniform transformation is performed. Then, each standard fingerprint vector in the fingerprint database is traversed. Calculate its relationship with the current event characteristics Distance between The smaller the calculated distance, the more similar the current abnormal behavior is to the device's historical behavior patterns. The top N matches with the smallest distances and their corresponding distance values ​​are recorded to form the latent abnormal behavior feature set.

[0092] It is worth noting that the determination of the matching quantity N (e.g., 5) is based on the statistical analysis of the average number of similar functional devices (such as pumps and centrifuges) within the target area. Choosing an integer value slightly larger than the average number of similar devices as N aims to ensure that the candidate set covers all possible similar devices and prevents omissions.

[0093] It should be noted that the feature mapping analysis performed on the aforementioned set of latent abnormal behaviors to determine the associated device identity is implemented using a nearest neighbor classification strategy. The match with the smallest distance (Top-1 Match) is selected. If the match tolerance threshold is less than the preset threshold, the current anomaly is directly determined to be caused by the device corresponding to the match result, and the device ID is identified as the associated device identity. If the value exceeds the threshold, it is marked as an unknown source. The final generated result includes the device ID and the matching confidence score (defined as...). The data packets of the abnormal time window are the latent anomaly identification data.

[0094] It is worth noting that the preset matching tolerance threshold is determined based on intra-class divergence analysis of the feature space. Specifically, this involves calculating the average distance between different historical samples from the same device in the fingerprint database and their standard fingerprint vectors. Set the threshold to This is to cover the normal fluctuation range of equipment operating status.

[0095] For example, the linkage index output in step S13 is 0.75, which is greater than the preset effective linkage threshold of 0.6. The original feature vector of the current event is then extracted. The vector and all fingerprint vectors in the database are Z-score standardized using global statistical parameters (mean and standard deviation) from the fingerprint database. In the standardized feature space, the Euclidean distance between the current event and the standard fingerprint of device "Pump-03" is calculated to be 0.12, and the Euclidean distance with the standard fingerprint of device "Heater-01" is calculated to be 3.5. Since 0.12 is much smaller than 3.5 and smaller than the preset matching tolerance threshold of 0.25, the anomaly is determined to be caused by "Pump-03", and the latent anomaly identification data is output as {Device_ID:"Pump-03",Confidence:0.88,Time:"10:00-10:45"}.

[0096] In step S15, based on the latent anomaly identification data, reference coordinate retrieval and multi-source data cross-validation are performed to determine the specific location information of potential pollution sources, including:

[0097] Based on the device identification, database indexing and coordinate extraction are performed to obtain the baseline geographic coordinates;

[0098] Real-time monitoring data at the reference geographic coordinates is obtained and its consistency with the latent anomaly identification data is verified to obtain the verification result.

[0099] If the verification result is consistent, then the reference geographic coordinates are confirmed as the specific location information of the potential pollution source.

[0100] It should be noted that indexing and extracting coordinates based on the device identification refers to accessing a pre-built database of basic pollution source information. This database is built during the initialization phase. The construction process involves batch importing equipment ledger data from enterprises within the target area, extracting four key fields: unique device identification (Device ID), physical installation location (latitude and longitude coordinates), enterprise name, and discharge outlet type. These fields are then stored as key-value pairs using the Device ID as the index key. This operation uses key-value query technology to directly retrieve the static latitude and longitude corresponding to the device based on the input ID, which is the reference geographic coordinate.

[0101] It should be noted that acquiring real-time monitoring data at the aforementioned baseline geographic coordinates and performing state consistency verification through multi-source data cross-validation is a secondary confirmation mechanism for confirming the authenticity of the pollution source, aiming to eliminate false alarms caused by historical pattern matching errors. This verification process involves comparison in two dimensions: numerical dimension verification and state dimension verification. Numerical dimension verification is performed by accessing the device's current real-time COD concentration via the IoT interface. Simultaneously, the mean feature value corresponding to the anomaly pattern is extracted from the latent anomaly identification data output in step S14. Using formulas Calculate the relative deviation rate between the two. Status dimension verification checks the status bit of the real-time data to confirm whether the device is currently in an "operating," "faulty," or "maintained" state, thus verifying whether it possesses the physical conditions to generate pollution. If the relative deviation rate... If the deviation is less than the preset consistency deviation threshold and the device status shows "running," then the status is considered consistent, and the verification result is affirmative. This means that the suspect, deduced from historical big data, does indeed exhibit consistent abnormal characteristics at the current moment, thus physically confirming its identity as a source of pollution.

[0102] It is worth noting that the preset consistency deviation threshold (e.g., 15%) is determined based on a joint statistical analysis of sensor measurement errors and the fluctuation range of the production process. Specifically, the variance of numerical fluctuations of the same equipment during stable operation under historical normal operating conditions is analyzed, and combined with the accuracy error of the sensor's factory calibration (e.g., ±2%), the boundary of the normal fluctuation range with a confidence level of 95% is calculated and used as the critical threshold for judgment.

[0103] For example, step S14 identifies a hidden anomaly at a discharge outlet with ID "ChemPlant-B-Discharge01," with pattern characteristics indicating its COD should be around 120 mg / L. A database query yields its coordinates [114.23, 22.56]. Subsequently, the real-time reading from the sensor at this discharge outlet is read as 125 mg / L, with a relative deviation of... The percentage is less than the 15% threshold, and the device status is displayed as "Open" (running). The verification result indicates that the status is consistent. Based on this, the coordinates [114.23, 22.56] are confirmed as the exact location of the pollution source in this warning, i.e., the specific location information of the potential pollution source.

[0104] In step S16, based on the specific location information of the potential pollution source, pollution diffusion simulation and impact range determination are performed to determine the distribution information of the affected area, including:

[0105] Based on the specific location information of the potential pollution sources, environmental data is loaded and a fluid model is constructed to obtain the preliminary diffusion range;

[0106] Environmental variable data is acquired, and the preliminary diffusion range is dynamically boundary-corrected based on the environmental variable data to obtain diffusion distribution data;

[0107] From the diffusion distribution data, coverage analysis and gridding mapping are performed to determine the distribution information of the affected areas.

[0108] It should be noted that the loading of environmental data and the construction of the fluid model were achieved using a numerical solution method for the two-dimensional advection-diffusion equation. First, gridded topographic data (Bathymetry) and basic flow field data of the target water area were loaded. Then, using the specific location information of the potential pollution sources as the pollutant release points (Source Term), the finite difference method (FDM) was used to solve the governing equations. The solution is obtained through iterative steps. Here, C represents the pollutant concentration. Let be the velocity component, D be the diffusion coefficient, and k be the attenuation coefficient. By simulating the evolution of the concentration field over a predetermined time period (e.g., 24 hours), the spatial region where the concentration value exceeds the safety limit is obtained, which is the initial diffusion range.

[0109] It is worth noting that the diffusion coefficient D and attenuation coefficient k are determined based on parameter inversion of historical tracer test data for this water area. Specifically, a genetic algorithm is used for global optimization. First, a population containing N random (D,k) parameter pairs is initialized. Second, the fitness function is defined as the reciprocal of the root mean square error (RMSE) between the model simulation value and the historical measured value. Then, roulette wheel selection, single-point crossover (crossover probability 0.8), and Gaussian mutation (mutation probability 0.01) are performed to generate a new generation of population. The above process is repeated until the RMSE is less than a preset convergence threshold (e.g., 0.01), and the optimal parameter combination is output as the preset physical parameters. The preset convergence threshold is determined based on the measurement accuracy limit of the water quality monitoring equipment. Specifically, the minimum resolution of the tracer concentration detection instrument (e.g., 0.01 mg / L) is selected as the lower limit of error tolerance to ensure that the inversion results converge within the instrument's detectable accuracy range.

[0110] It should be noted that the dynamic boundary correction of the preliminary diffusion range based on the environmental variable data is achieved using a real-time flow field correction method. First, the real-time monitored water flow velocity is extracted from the environmental variable data obtained in this step. and wind speed Then, compare it with the base flow rate used in the model. Compare and calculate the flow velocity deviation factor. The factor is then used to perform an affine transformation on the boundary vector of the initial diffusion range. Specifically, the transformation is centered on the location of the pollution source. Construct a scaling matrix for each point on the boundary. Using the formula The diffusion distribution data is obtained by stretching or compressing along the streamline direction to obtain data that conforms to the current hydrological conditions.

[0111] It should be noted that the coverage resolution and gridding mapping are achieved using spatial topological overlay analysis. First, threshold segmentation and vectorization are performed, traversing the diffusion distribution data (i.e., a continuous concentration field raster matrix). Raster values ​​with concentrations greater than a preset threshold are marked as foreground, and the rest as background. A moving squares algorithm is used to trace the boundary between the foreground and background, extracting the closed vector contour of the highly polluted area. Subsequently, a Boolean intersection operation is performed between this contour and the water function zone geographic grid base map to identify all grid cells that spatially intersect with the highly polluted area contour (such as specific drinking water source protection zone IDs and water intake IDs). The spatial distribution list of these grid cells is then determined as the distribution information of the affected area.

[0112] It is worth noting that the determination of the preset concentration threshold is based on the relevant provisions of the "Surface Water Environmental Quality Standard" (GB3838-2002). Specifically, the standard limit value (e.g., 20 mg / L) of the target pollutant (such as COD) under the functional zone category (such as Class III water) of the water body is selected as the critical threshold.

[0113] For example, model calculations show that the pollutant will reach 3 kilometers downstream in 2 hours. However, real-time sensors show that the current flow rate is 20% faster than the historical average (i.e., The correction method then stretches the location vector of the diffusion front downstream to 3.6 km. Overlay analysis shows that the corrected extent covers the drinking water source protection area grid with ID "WaterSource-A", from which the distribution information of the affected area is output.

[0114] In step S17, based on the specific location information of the potential pollution source and the distribution information of the affected area, early warning generation and transmission processing are performed to obtain early warning notification data, including:

[0115] The specific location information of the potential pollution sources is associated and integrated with the distribution information of the affected areas to determine the alarm information content;

[0116] The alarm information content is encapsulated and transmitted, and transmission status feedback is obtained;

[0117] If the transmission status feedback includes a preset success identifier, the alarm information content is archived, stored, and its integrity is verified to obtain the early warning notification data.

[0118] It should be noted that the association and integration of the specific location information of the potential pollution sources with the distribution information of the affected areas is achieved using template data binding technology based on the CAP (Common Alerting Protocol) standard. First, a pre-built XML data template conforming to the CAP v1.2 standard is loaded. This template is constructed based on the OASIS standard schema during the initialization phase. The construction process involves defining static fields and... <status>Set to "Actual", <msgtype>Set to "Alert", <scope>Set to "Public"; define dynamic placeholders for <area> , <description>Equal field reserved variable interface. Then, the pollution source longitude and latitude coordinates confirmed in the extraction step S15 are filled into the template <area> under the element <circle>in the attribute (in the format "latitude, longitude, radius"); the extraction step S16 (diffusion simulation) determines the affected area boundary (i.e. the sequence of polygon vertices) which is filled into <area> under the element <polygon>In the attribute; at the same time, the associated pollutant type, the predicted concentration peak value and the recommended emergency measures (such as stopping water taking) are filled into <description>and <instruction>In the field, thereby generating complete alarm information content.

[0119] It should be noted that the data encapsulation and transmission of the alarm information content is achieved by using the RESTful API interface calling mode based on the HTTPS protocol. First, multi-level encryption encapsulation is performed. In the first step, the generated XML alarm information string is converted into Base64 encoding format to eliminate the influence of special characters on transmission. In the second step, a JSON object is constructed, and the Base64 string is assigned to the payload field. In the third step, the JSON object is encrypted as a whole by using the AES-256 symmetric encryption algorithm to generate the final binary encrypted data packet. Then, a POST request is constructed, and the data packet is sent to the API gateway address of the supervision system. Finally, the HTTP response message returned by the API gateway is received, the response body is deserialized by using a standard JSON parser, the response body is converted into a key-value pair object, and the values of the code (status code) and transaction_id (transaction ID) fields in the object are directly read to obtain the transmission status feedback.

[0120] It should be noted that the determination of the preset success identifier is based on the definition of the standard status code of the HTTP protocol. Specifically, the HTTP status code is 200 (OK) or 201 (Created), and the response body JSON includes the "status":"success" field and a non-empty transaction ID, which are defined as the unique identifier of successful transmission.

[0121] It should be noted that the archiving storage and integrity verification of the alarm information content are aimed at ensuring the non-tamperability of the data chain. First, the SHA-256 algorithm is used to calculate the hash value of the sent alarm data, which is used as the original digital fingerprint. Then, the alarm information content is written into a WORM (Write Once Read Many) storage device together with the digital fingerprint. After writing is completed, a read-back operation is immediately performed to read the data just written from the storage device, recalculate the SHA-256 hash value, and compare it with the original digital fingerprint. If they are completely consistent, it is marked as passed, and the pre-warning notification data containing the storage index (such as file path or block hash) and the verification result is generated. It should be noted that the WORM device is a once-write and multiple-read optical disc library. When archiving, the alarm information is written into an optical disc, and the SHA-256 hash value is immediately read back for verification to ensure that the data cannot be tampered with. If the integrity verification fails, the retransmission process is triggered.

[0122] For example, a CAP alert package containing the coordinates of the pollution source 32.15, 118.88 and the affected area (a polygon containing 10 vertices) is generated. After sending it to the regulatory platform through an HTTPS POST request, a status code 200 and transaction ID "TX-20251127-9988" are received. The SHA-256 value of the alert package is "a1b2c3d4...". After writing the data to the WORM optical disc library, the hash value is consistent when reading back, confirming that the early warning has been completed and safely archived.

[0123] In summary, the present application constructs a full-process regulatory system from the standardized fusion and abnormal clustering mining of multi-source heterogeneous data (water and electricity and water quality), to the precise tracing based on multi-dimensional parameter correlation verification and device feature fingerprint matching, to the dynamic diffusion simulation combined with real-time environmental flow field and closed-loop early warning generation. The present application deeply couples and analyzes industrial production behavior and environmental monitoring data. The present application innovatively introduces historical behavior pattern matching and dynamic boundary correction technology based on fluid model, effectively solves the technical problems of weak active discovery ability, difficulty in accurately locating the pollution source and lack of spatial dimension of early warning information caused by data silos in the prior art, and significantly improves the initiative of water pollution regulation, the accuracy of tracing and the scientificity of emergency response.

[0124] Reference Figure 2 The second embodiment of the present application provides a water pollution early warning system based on big data, comprising:

[0125] A data processing module is configured to collect water and electricity usage data and water quality monitoring data, and perform multi-source data standardized fusion processing to obtain an integrated data set;

[0126] An abnormal mining module is configured to perform data attribute extraction and abnormal clustering analysis based on the integrated data set to obtain abnormal pattern data;

[0127] An association verification module is configured to perform multi-dimensional parameter association verification on the abnormal pattern data to obtain abnormal linkage distribution results;

[0128] An identity recognition module is configured to perform historical behavior pattern matching and identity recognition to obtain implicit abnormal recognition data if the abnormal linkage distribution results meet the preset linkage determination conditions;

[0129] A positioning verification module is configured to perform benchmark coordinate retrieval and multi-source data cross verification based on the implicit abnormal recognition data to determine the specific location information of the potential pollution source;

[0130] A diffusion simulation module is configured to perform pollution diffusion simulation and impact range determination based on the specific location information of the potential pollution source to determine the distribution information of the affected area;

[0131] The early warning processing module is configured to generate and transmit early warning according to the specific location information of the potential pollution source and the distribution information of the affected area, and obtain early warning notification data.

[0132] It should be noted that the water pollution early warning system based on big data provided by the embodiments of the present application is used to execute all process steps of the water pollution early warning method based on big data provided by the above embodiments, and the working principles and beneficial effects of the two are one-to-one correspondence, so they will not be repeated.

[0133] The embodiments of the present application also provide an electronic device. The electronic device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a water pollution early warning program based on big data. The processor executes the computer program to implement the steps in each of the above water pollution early warning methods based on big data, such as step S11 shown in the figure. Alternatively, the processor executes the computer program to implement the functions of each module / unit in each of the above system embodiments, such as the data processing module. Figure 1

[0134] For example, the computer program can be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program in the electronic device.

[0135] The electronic device can be a desktop computer, a notebook, a palm computer, and a smart tablet, etc. The electronic device can include, but is not limited to, a processor, a memory. Those skilled in the art can understand that the above components are only examples of the electronic device and do not constitute a limitation on the electronic device, and can include more or fewer components than the above, or combine certain components, or different components, for example, the electronic device can also include an input / output device, a network access device, a bus, etc.

[0136] ​The processor can be a central processing unit (CPU), and can also be other general-purpose processors, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The processor is a control center of the electronic device, and connects various parts of the electronic device through various interfaces and lines.

[0137] The memory can be used to store the computer program and / or modules, and the processor realizes various functions of the electronic device by running or executing the computer program and / or modules stored in the memory, and calling data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.), etc.; and the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory can include a high-speed random access memory, and can also include a nonvolatile memory, for example, a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state memory device.

[0138] The modules / units integrated in the electronic device, if realized in the form of software function units and sold or used as independent products, can be stored in a computer readable storage medium. Based on such understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware, and the computer program can be stored in a computer readable storage medium. The computer program can implement the steps of each method embodiment when executed by a processor. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer readable medium can include any entity or system, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium, etc. that can carry the computer program code. It should be noted that the contents included in the computer readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, the computer readable medium does not include electric carrier signals and telecommunication signals.

[0139] It should be noted that the above-described system embodiments are only illustrative, and the units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. In addition, the connection relationship between the modules in the system embodiment provided by the present application indicates that there is a communication connection between them, which can be realized as one or more communication buses or signal lines. Those skilled in the art can understand and implement it without creative labor.

[0140] The above-described specific embodiments further detail the purpose, technical solutions, and beneficial effects of the present application. It should be understood that the above-described specific embodiments are only examples of the present application and are not intended to limit the scope of protection of the present application. It is particularly pointed out that any modification, equivalent replacement, improvement, etc. made by those skilled in the art within the spirit and principles of the present application should be included in the scope of protection of the present application.< / instruction> < / description> < / polygon> < / circle> < / description> < / scope> < / msgtype> < / status>

Claims

1. A water pollution early warning method based on big data, characterized in that, include: Collect hydropower usage data and water quality monitoring data, and perform multi-source data standardization and fusion processing to obtain an integrated data set; Based on the integrated dataset, data attribute extraction and anomaly clustering analysis are performed to obtain anomaly pattern data; For the abnormal pattern data, multidimensional parameter correlation verification is performed to obtain the abnormal linkage distribution results; If the abnormal linkage distribution results meet the preset linkage judgment conditions, then historical behavior pattern matching and identity recognition are performed to obtain hidden anomaly identification data. Based on the latent anomaly identification data, benchmark coordinate retrieval and multi-source data cross-validation are performed to determine the specific location information of potential pollution sources; Based on the specific location information of the potential pollution sources, pollution diffusion simulation and impact range determination are performed to identify the distribution information of the affected areas; Based on the specific location information of the potential pollution sources and the distribution information of the affected areas, early warnings are generated and transmitted to obtain early warning notification data.

2. The water pollution early warning method based on big data according to claim 1, characterized in that, The collected hydropower usage data and water quality monitoring data are then subjected to multi-source data standardization and fusion processing to obtain an integrated data set, including: Collect the hydropower usage data and the water quality monitoring data, and perform time axis calibration to obtain the first dataset; The first dataset is subjected to multi-dimensional verification and outlier correction to obtain an intermediate dataset; The intermediate dataset is normalized to obtain the integrated dataset.

3. The water pollution early warning method based on big data according to claim 1, characterized in that, The step of extracting data attributes and performing anomaly clustering analysis based on the integrated data set to obtain anomaly pattern data includes: Based on the integrated data set, density values ​​are calculated and preliminary screening is performed to obtain a filtered data set. The filtered data set is clustered and organized to obtain clustered grouped data and isolated data subsets; Based on the clustered grouped data and the isolated data subset, a time-dimensional distribution pattern analysis is performed to obtain an anomaly marker set; Based on the set of anomaly markers, a distribution pattern comparison analysis is performed to obtain the anomaly pattern data.

4. The water pollution early warning method based on big data according to claim 1, characterized in that, The process of performing multi-dimensional parameter correlation verification on the abnormal pattern data to obtain abnormal linkage distribution results includes: Range data extraction is performed on the abnormal pattern data to obtain a range data set; The set of range data is compared with a preset fluctuation threshold to obtain a set of linked anomaly points; For the set of linked anomaly points, a distribution pattern matching is performed to obtain a set of linked feature values; Based on the set of linkage feature values, the correlation strength is determined to obtain the linkage distribution feature set; Based on the aforementioned linkage distribution feature set, correlation segmentation analysis is performed to obtain the abnormal linkage distribution results.

5. The water pollution early warning method based on big data according to claim 1, characterized in that, If the abnormal linkage distribution result meets the preset linkage judgment condition, then historical behavior pattern matching and identity recognition are performed to obtain latent anomaly identification data, including: If the abnormal linkage distribution results meet the preset linkage judgment conditions, then the overlapping characteristics are extracted to obtain a preliminary abnormal behavior dataset. Historical data records and a preset device feature fingerprint database are obtained, and the preliminary abnormal behavior dataset is matched and compared with the historical data records and the preset device feature fingerprint database to obtain a set of latent abnormal behavior features. Feature mapping analysis is performed on the set of latent abnormal behavior features to determine the associated device identity identifier, thereby obtaining the latent abnormality identification data.

6. The water pollution early warning method based on big data according to claim 5, characterized in that, The step of determining the specific location information of potential pollution sources by performing benchmark coordinate retrieval and multi-source data cross-validation based on the latent anomaly identification data includes: Based on the device identification, database indexing and coordinate extraction are performed to obtain the baseline geographic coordinates; Real-time monitoring data at the reference geographic coordinates is obtained and its consistency with the latent anomaly identification data is verified to obtain the verification result. If the verification result is consistent, then the reference geographic coordinates are confirmed as the specific location information of the potential pollution source.

7. The water pollution early warning method based on big data according to claim 1, characterized in that, The process of simulating pollution diffusion and determining the impact range based on the specific location information of the potential pollution source, and identifying the distribution information of the affected area, includes: Based on the specific location information of the potential pollution sources, environmental data is loaded and a fluid model is constructed to obtain the preliminary diffusion range; Environmental variable data is acquired, and the preliminary diffusion range is dynamically boundary-corrected based on the environmental variable data to obtain diffusion distribution data; From the diffusion distribution data, coverage analysis and gridding mapping are performed to determine the distribution information of the affected areas.

8. The water pollution early warning method based on big data according to claim 1, characterized in that, The step of generating and transmitting early warning data based on the specific location information of the potential pollution source and the distribution information of the affected area, includes: The specific location information of the potential pollution sources is correlated and integrated with the distribution information of the affected areas to determine the alarm information content; The alarm information content is encapsulated and transmitted, and transmission status feedback is obtained; If the transmission status feedback includes a preset success identifier, the alarm information content is archived, stored, and its integrity is verified to obtain the early warning notification data.

9. A water pollution early warning system based on big data, characterized in that, include: The data processing module is used to collect hydropower usage data and water quality monitoring data, and to perform standardized fusion processing of multi-source data to obtain an integrated data set; The anomaly detection module is used to extract data attributes and perform anomaly clustering analysis based on the integrated data set to obtain anomaly pattern data. The correlation verification module is used to perform multi-dimensional parameter correlation verification on the abnormal pattern data to obtain the abnormal linkage distribution results. The identity recognition module is used to perform historical behavior pattern matching and identity recognition to obtain latent anomaly recognition data if the abnormal linkage distribution result meets the preset linkage judgment conditions. The location verification module is used to perform benchmark coordinate retrieval and multi-source data cross-verification based on the hidden anomaly identification data to determine the specific location information of potential pollution sources. The diffusion simulation module is used to simulate the diffusion of pollution and determine the scope of impact based on the specific location information of the potential pollution source, and to determine the distribution information of the affected area. The early warning processing module is used to generate and transmit early warnings based on the specific location information of the potential pollution sources and the distribution information of the affected areas, and obtain early warning notification data.

Citation Information

Patent Citations

  • Water quality change trend rapid prediction method based on multi-source data fusion and physical constraint

    CN120598102A

  • Water pollutant concentration detection system and method based on big data analysis

    CN120832537A

  • Water-pollution environmental-protection verification method and apparatus based on power grid and tax data fusion

    US20240232771A1

Cited By

  • Abnormal state early warning method and system applied to underground water online monitoring

    CN122245075A