Water quality monitoring method and system based on big data analysis

By standardizing the processing of multi-type sensor data and combining principal component analysis, random forest and convolutional neural networks, a real-time water quality monitoring system was constructed. This solves the problems of data processing lag and insufficient warning accuracy in traditional water quality monitoring methods, and realizes real-time assessment and early warning of water pollution.

CN120744703AInactive Publication Date: 2025-10-03湖南云河信息科技有限公司 +1
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202511187968.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-10-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional water quality monitoring methods are unable to cope with the efficient integration of multi-source heterogeneous data and the dynamic analysis of complex pollution scenarios, resulting in delayed data processing, insufficient early warning accuracy, and lack of targeted governance recommendations, making it difficult to achieve real-time and accurate water quality assessment and early warning.

Method used

By acquiring multi-type sensor data, using standardized algorithms to process and eliminate outliers, using principal component analysis to extract feature vectors, constructing a pollution assessment model based on the random forest algorithm, combining it with convolutional neural networks for time series analysis, and matching historical data through a distributed computing framework, a governance strategy is generated.

Benefits of technology

It has achieved real-time monitoring, assessment, early warning and traceability of water pollution, provided comprehensive technical support, and provided a scientific basis and decision-making basis for water environment protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744703A_ABST
    Figure CN120744703A_ABST
Patent Text Reader

Abstract

The invention discloses a water quality monitoring method and system based on big data analysis, and the method comprises the steps: obtaining water quality parameter data collected by multiple types of sensors, carrying out the format unification processing of the water quality parameter data through a standardization algorithm, and carrying out the elimination through a median filtering algorithm if abnormal data points are detected, thereby obtaining a standardized data set; aiming at the standardized data set, performing feature extraction on the chemical parameters, the spectral features and the biological indexes by adopting a principal component analysis algorithm to obtain feature vectors, and if the variance contribution rate of the feature vectors exceeds a preset threshold value, retaining the feature vectors and generating a dimension reduction feature data set; according to the dimension reduction feature data set, a random forest algorithm is adopted to construct a pollution evaluation model, a pollution index is calculated, if the pollution index exceeds a preset threshold value, a high pollution state mark is generated, and a pollution evaluation result is output. According to the invention, real-time monitoring, evaluation, early warning and traceability of water quality pollution are realized, and comprehensive technical support is provided for water environment protection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular discloses a water quality monitoring method and system based on big data analysis. Background Art

[0002] Water quality monitoring is a core area of ​​environmental governance and ecological protection, directly related to drinking water safety, industrial emission compliance, and agricultural non-point source pollution control. Its importance is self-evident. Traditional water quality monitoring methods rely on single-parameter detection or fixed threshold warnings, making them difficult to efficiently integrate multi-source heterogeneous data and dynamically analyze complex pollution scenarios. They are often plagued by data processing lags, inaccurate warnings, and weakly targeted remediation recommendations. These limitations make it difficult for environmental management departments to quickly respond to sudden pollution incidents or long-term pollution trends or develop precise remediation strategies.

[0003] In the field of water quality monitoring, the core challenge stems from the complexity of data collection and processing and the dual demands for real-time performance and accuracy. Diverse water quality parameters, such as chemical oxygen demand, ammonia nitrogen, and total phosphorus, need to be collected in real time by multiple types of sensors. However, the data formats and operating conditions of different devices vary significantly, easily generating outliers that interfere with analysis results. Data heterogeneity further exacerbates the difficulty of building comprehensive pollution models. Traditional methods struggle to integrate spectral features, biological indicators, and chemical parameters to form a unified pollution assessment system. The complexity of the model places higher demands on computing resources. Centralized cloud computing struggles to meet the needs of low-latency real-time warnings, while the computing power of terminal devices is insufficient to support complex analysis. These factors, layered together, collectively restrict the intelligence and efficiency of water quality monitoring systems.

[0004] Therefore, how to build a real-time and accurate water quality assessment and early warning system through efficient data association, fusion analysis and edge computing technology in a multi-source heterogeneous data environment has become a key issue that needs to be urgently solved in the field of water quality monitoring. Summary of the Invention

[0005] The present invention provides a water quality monitoring method and system based on big data analysis, aiming to solve at least one defect existing in the above-mentioned prior art.

[0006] One aspect of the present invention relates to a water quality monitoring method based on big data analysis, comprising the following steps: Acquire water quality parameter data collected by multiple types of sensors, use a standardized algorithm to unify the format of the water quality parameter data, and if any abnormal data points are detected, remove them through a median filter algorithm to obtain a standardized data set; water quality parameter data includes chemical parameters, spectral characteristics and biological indicators; For the standardized data set, the principal component analysis algorithm is used to extract the chemical parameters, spectral characteristics and biological indicators to obtain the feature vector. If the variance contribution rate of the feature vector exceeds the preset threshold, the feature vector is retained to generate a reduced-dimensional feature data set. Based on the dimensionality-reduced feature dataset, a pollution assessment model is constructed using the random forest algorithm to calculate the pollution index. If the pollution index exceeds the preset threshold, a high pollution state mark is generated and the pollution assessment result is output; Based on the pollution assessment results, a convolutional neural network algorithm is used to analyze the time series characteristics. If a continuous upward trend in the pollution index is detected, an early warning signal is generated and time series warning information is output; Based on the time-series warning information, a distributed computing framework is used to match and analyze the pollution assessment results with historical pollution data. If the matching results point to industrial emission areas, governance strategy data is generated.

[0007] Furthermore, water quality parameter data collected by multiple types of sensors are obtained, and the water quality parameter data are formatted uniformly using a standardization algorithm. If abnormal data points are detected, they are eliminated using a median filtering algorithm. The steps of obtaining a standardized data set include: Acquire water quality parameter data including chemical parameters, spectral characteristics and biological indicators from multiple types of sensors, classify and store the water quality parameter data through a preset acquisition protocol to obtain the original data set; The original data set is formatted uniformly using a standardized algorithm, and the data structure of chemical parameters, spectral characteristics, and biological indicators is adjusted based on a preset format template to obtain an initialized data set. The abnormal data points in the initialized data set are detected by statistical methods. If the abnormal data points deviate from the preset threshold, the median filtering algorithm is used to eliminate the abnormal data points to obtain a standardized data set.

[0008] Furthermore, for the standardized data set, a principal component analysis algorithm is used to extract features of chemical parameters, spectral features, and biological indicators to obtain a feature vector. If the variance contribution rate of the feature vector exceeds a preset threshold, the feature vector is retained. The steps of generating a dimensionality-reduced feature data set include: Chemical parameters, spectral characteristics and biological indicators are obtained from the standardized data set, and the principal component analysis algorithm is used to calculate the eigenvectors and corresponding variance contributions of each principal component to obtain a set of eigenvectors; For the set of eigenvectors, the variance analysis method is used to calculate the variance contribution rate of each eigenvector. If the variance contribution rate exceeds the preset threshold, the eigenvector is retained to obtain the filtered eigenvector set; Based on the set of screened feature vectors, the chemical parameters, spectral features and biological indicators in the standardized dataset are mapped to a low-dimensional space using the matrix projection method to generate a reduced-dimensional feature dataset.

[0009] Furthermore, based on the dimensionality-reduced feature dataset, a pollution assessment model is constructed using a random forest algorithm to calculate a pollution index. If the pollution index exceeds a preset threshold, a high pollution state flag is generated. The steps of outputting the pollution assessment result include: Chemical parameters, spectral features, and biological indicators are obtained from the dimensionality-reduced feature dataset, and the pollution index of each data point is calculated using the random forest algorithm to obtain a pollution index set; For the pollution index set, determine whether each pollution index exceeds the threshold when the preset threshold is exceeded. If it exceeds, generate a high pollution mark to obtain a mark set; According to the label set, the high-contamination labels are combined with the dimensionality-reduced feature data through the matrix projection method to generate low-dimensional space data containing labels, thus obtaining a labeled data set; High pollution marks and corresponding pollution indexes are obtained from the labeled data set. The marks are arranged according to the size of the pollution index using a sorting method to obtain the pollution assessment results.

[0010] Furthermore, based on the pollution assessment results, a convolutional neural network algorithm is used to perform time series feature analysis. If a continuous upward trend in the pollution index is detected, a warning signal is generated. The steps of outputting time series warning information include: Obtain time series data from pollution assessment results, use convolutional neural networks to extract features from the time series data, and obtain a time series feature set; For the time series feature set, the upward trend of each feature is analyzed by trend detection method to determine the continuously rising feature set; Determine whether at least one feature in the continuously rising feature set exceeds a preset threshold, and if so, generate a corresponding warning signal to obtain a warning signal set; According to the warning signal set, the warning signal is combined with the time series characteristics through the matrix mapping method to output the time series warning information.

[0011] Furthermore, based on the time-series warning information, a distributed computing framework is used to match and analyze the pollution assessment results with historical pollution data. If the matching result points to an industrial emission area, the steps for generating treatment strategy data include: Obtain a data set from time-series warning information, compare the pollution assessment results with historical pollution data through a distributed computing framework, and obtain a matching result set; For the set of matching results, a conditional judgment method is used. If at least one matching result points to an industrial emission area, the corresponding area identification data is extracted; Based on the regional identification data, the control data related to the emission area is searched through a pre-established mapping table to determine the control data set; The random forest algorithm is used to jointly process the governance data set and time series data to output the final governance data.

[0012] Another aspect of the present invention relates to a water quality monitoring system based on big data analysis, which is used to implement the above-mentioned water quality monitoring method based on big data analysis. The water quality monitoring system based on big data analysis includes: The acquisition module is used to obtain water quality parameter data collected by multiple types of sensors. The water quality parameter data is formatted uniformly using a standardized algorithm. If abnormal data points are detected, they are removed using a median filtering algorithm to obtain a standardized data set. The water quality parameter data includes chemical parameters, spectral characteristics, and biological indicators. The first generation module is used to extract features of chemical parameters, spectral features and biological indicators using the principal component analysis algorithm for the standardized data set to obtain feature vectors. If the variance contribution rate of the feature vector exceeds a preset threshold, the feature vector is retained to generate a reduced-dimensional feature data set; The first output module is used to build a pollution assessment model based on the dimensionality reduction feature data set using the random forest algorithm, calculate the pollution index, generate a high pollution state mark if the pollution index exceeds a preset threshold, and output the pollution assessment result; The second output module is used to perform time series feature analysis based on the pollution assessment results using a convolutional neural network algorithm. If a continuous upward trend in the pollution index is detected, an early warning signal is generated and time series warning information is output; The second generation module is used to match and analyze the pollution assessment results with historical pollution data based on the time-series warning information using a distributed computing framework. If the matching result points to an industrial emission area, governance strategy data is generated.

[0013] Furthermore, the acquisition module includes: The first acquisition unit is used to acquire water quality parameter data including chemical parameters, spectral characteristics and biological indicators from multiple types of sensors, classify and store the water quality parameter data through a preset acquisition protocol to obtain an original data set; The second acquisition unit is used to unify the format of the original data set using a standardized algorithm, and adjust the data structure of chemical parameters, spectral characteristics and biological indicators based on a preset format template to obtain an initialized data set; The third acquisition unit is used to detect abnormal data points in the initialized data set through a statistical method. If the abnormal data points deviate from a preset threshold, the median filtering algorithm is used to eliminate the abnormal data points to obtain a standardized data set.

[0014] Furthermore, the first generation module includes: a fourth acquisition unit, configured to acquire chemical parameters, spectral characteristics, and biological indicators from the standardized data set, and calculate the eigenvectors and corresponding variance contributions of each principal component using a principal component analysis algorithm to obtain a set of eigenvectors; a fifth acquiring unit, configured to calculate the variance contribution rate of each eigenvector using a variance analysis method for the eigenvector set, and retain the eigenvector if the variance contribution rate exceeds a preset threshold, thereby obtaining a screened eigenvector set; The generation unit is used to map the chemical parameters, spectral features and biological indicators in the standardized data set to a low-dimensional space based on the screening feature vector set using a matrix projection method to generate a reduced-dimensional feature data set.

[0015] Furthermore, the first output module includes: a sixth acquisition unit, configured to acquire chemical parameters, spectral features, and biological indicators from the dimensionally reduced feature data set, and calculate the pollution index of each data point using a random forest algorithm to obtain a pollution index set; a seventh obtaining unit, configured to determine, for the pollution index set, whether each pollution index exceeds a threshold value when the threshold value is set, and generate a high pollution mark if the pollution index exceeds the threshold value, thereby obtaining a mark set; an eighth acquisition unit, configured to combine the high-pollution markers with the dimensionality reduction feature data by a matrix projection method according to the marker set, to generate low-dimensional spatial data containing the markers, and obtain a labeled dataset; The ninth acquisition unit is used to obtain high pollution marks and corresponding pollution indexes from the labeled data set, and arrange the marks according to the size of the pollution index using a sorting method to obtain a pollution assessment result.

[0016] The beneficial effects achieved by the present invention are: The present invention provides a water quality monitoring method and system based on big data analysis. Water quality parameter data, including chemical parameters, spectral characteristics and biological indicators, are collected through multiple types of sensors. The data is standardized and outliers are eliminated. Principal component analysis is then used to extract features and reduce dimensionality. Based on the feature data set after dimensionality reduction, a pollution assessment model is constructed using a random forest algorithm to calculate the pollution index and generate an assessment result. Furthermore, a convolutional neural network is used to perform time series analysis on the pollution index, detect the upward trend and generate an early warning signal. Finally, the assessment results are matched and analyzed with historical data through a distributed computing framework to locate possible pollution sources and formulate a control strategy. The present invention realizes real-time monitoring, assessment, early warning and tracing of water pollution, providing comprehensive technical support for water environment protection. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 The figure is a flow chart of an embodiment of a water quality monitoring method based on big data analysis of the present invention. DETAILED DESCRIPTION

[0018] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0019] like Figure 1 As shown, the first embodiment of the present invention proposes a water quality monitoring method based on big data analysis, comprising the following steps: Step S100: Acquire water quality parameter data collected by multiple types of sensors, use a standardization algorithm to unify the format of the water quality parameter data, and if abnormal data points are detected, remove them through a median filtering algorithm to obtain a standardized data set; the water quality parameter data includes chemical parameters, spectral characteristics and biological indicators.

[0020] Multi-sensor technology integrates multiple sensors with different operating principles, functions, or measurement objectives into a single system or device. These sensors work together to achieve comprehensive sensing and data collection for a wide range of physical, chemical, or biological quantities. The core goal of multi-sensor technology is to enhance the system's sensing capabilities, accuracy, and environmental adaptability through multi-dimensional and multi-modal data fusion.

[0021] Water quality parameter data is a quantitative descriptive indicator of the physical, chemical and biological properties of water bodies. It is measured in real time or periodically through multiple types of sensors or detection equipment and is used to assess water quality status, pollution level and ecological health level.

[0022] Normalization algorithms are computational methods used to transform data into a uniform scale or standardized form. They aim to eliminate bias caused by differences in dimension, range, or distribution, ensuring comparability across data sources and types, thereby improving the reliability of data analysis, model training, and decision-making. The core of normalization algorithms lies in establishing a fair comparison benchmark across data through mathematical transformation.

[0023] Format unification refers to the standardization of data structure, style, or presentation through pre-set rules or technical means to ensure consistency, readability, and maintainability in specific scenarios. Format unification is widely used in software development, document editing, data analysis, and other fields. Its core goals include eliminating redundancy, improving collaboration efficiency, and reducing the complexity of subsequent processing.

[0024] Anomalous data points are observations in a dataset that significantly deviate from the normal distribution pattern or expected range. These may be caused by measurement errors, equipment failures, environmental changes, or real-world extreme events. They must be identified and addressed through statistical analysis and domain knowledge. In scenarios such as water quality monitoring, anomalous points may reflect pollution incidents, sensor failures, or extreme natural phenomena.

[0025] Median filtering is a nonlinear signal processing technique that replaces the original center value with the median value of the data within a sliding window. This effectively eliminates outliers such as impulse noise and salt-and-pepper noise while preserving signal edges and details to the greatest extent possible. Its core advantage lies in its strong resistance to outlier interference, making it suitable for image processing and sensor data denoising.

[0026] Normalizing a dataset involves uniformly adjusting the feature scale of raw data through mathematical transformations, eliminating dimensional differences and uneven distribution, and making different features comparable at the same scale, thereby adapting to the needs of machine learning model training and analysis. Its core goals are to improve data quality, optimize algorithm performance, and enhance model generalization.

[0027] Chemical parameters refer to a set of key indicators used to characterize the composition, reactivity, and environmental status of substances by measuring the type, concentration, or physicochemical properties of specific chemical components in a sample through quantitative or qualitative analysis methods. In water quality monitoring, chemical parameters are the core basis for assessing the degree of water pollution, ecological health, and safety. Spectral signatures refer to the wavelength-dependent absorption, reflection, or emission characteristics of a substance within a specific electromagnetic band (such as ultraviolet, visible light, and infrared). By quantifying the differences in energy response at different wavelengths, they characterize the substance's composition, structure, concentration, and environmental state. Essentially, they represent the "fingerprint" of the interaction between matter and electromagnetic radiation, serving as a key discriminant in fields such as remote sensing, chemical analysis, and environmental monitoring.

[0028] A biological indicator (Biomarker) is a biological entity or characteristic parameter that indirectly reflects environmental quality, pollutant exposure, or ecological health by monitoring changes in the composition, structure, function, or physiological response of individual organisms, populations, or communities. Its core is to leverage the sensitivity, accumulation, or adaptability of organisms to the environment to provide comprehensive assessment information on long-term pollutant exposure and its combined effects.

[0029] Step S200: For the standardized data set, the principal component analysis algorithm is used to extract the chemical parameters, spectral characteristics and biological indicators to obtain a feature vector. If the variance contribution rate of the feature vector exceeds a preset threshold, the feature vector is retained to generate a reduced-dimensional feature data set.

[0030] Principal Component Analysis (PCA) is an unsupervised machine learning algorithm and multivariate statistical method. Its core goal is to project high-dimensional, linearly correlated observational data into a low-dimensional space through an orthogonal transformation, generating a set of linearly independent variables called principal components (PCs) while maximizing the preservation of the variance information of the original data.

[0031] An eigenvector is a non-zero vector associated with a square matrix in linear algebra. When the matrix acts on the vector, it only changes its length (scales) without changing its direction. The scaling ratio is represented by the corresponding eigenvalue.

[0032] Variance Contribution Rate (VCR) is an indicator used in dimensionality reduction methods such as principal component analysis (PCA) to measure the proportion of the total variance of the original data explained by a single principal component.

[0033] A dimensionality-reduced feature dataset is a low-dimensional dataset created by reducing the feature dimensionality of the original data while preserving key information and data structure. Its core goal is to convert high-dimensional data into a more concise low-dimensional representation through mathematical mapping or feature filtering, thereby addressing issues such as high computational complexity and model overfitting associated with high-dimensional data.

[0034] Step S300: Based on the dimensionality reduction feature data set, a pollution assessment model is constructed using the random forest algorithm to calculate the pollution index. If the pollution index exceeds a preset threshold, a high pollution state mark is generated and the pollution assessment result is output.

[0035] Random Forest is a supervised machine learning algorithm based on ensemble learning. It builds multiple decision trees and integrates their predictions to improve the model's generalization and robustness. Its core concept is "swarm intelligence." Through bagging (bootstrap aggregating) and random feature subspace selection, it reduces the risk of overfitting in individual decision trees and synthesizes the voting (classification) or average (regression) results of multiple trees as the final output.

[0036] Pollution assessment models are quantitative tools based on environmental science principles and data analysis techniques. They integrate multi-source environmental data and pollution factors to systematically assess the pollution level, sources, and potential risks in specific areas or media. Their core goal is to provide a scientific basis for pollution control, health warnings, and policy development.

[0037] The pollution index converts complex environmental pollutant concentration data into a single numerical indicator through mathematical methods, which is used to comprehensively characterize the pollution degree or quality level of a specific environmental medium (such as air, water, soil) or polluted object. Its core function is to simplify the public's understanding of the pollution situation and guide risk response.

[0038] High pollution status marking is a process where pollutant-discharging units, during automated pollutant emission monitoring, use specific rules to identify monitoring data resulting from equipment failure, maintenance and commissioning, or abnormal operating conditions at production or pollution control facilities. This system aims to distinguish between normal and abnormal emission data and define its legal validity in environmental regulation. This marking mechanism standardizes data validity determinations and provides a basis for scientific regulation of pollution emissions.

[0039] Pollution assessment results are conclusive outputs obtained through a comprehensive analysis of the distribution, concentration, and potential impacts of pollutants in environmental media (such as air, water, and soil) or specific areas using scientific quantitative methods. They aim to reflect the degree of pollution, health risks, and environmental quality levels, and provide a decision-making basis for pollution control and policy making.

[0040] Step S400: Based on the pollution assessment results, a convolutional neural network algorithm is used to perform time series feature analysis. If it is detected that the pollution index is showing a continuous upward trend, an early warning signal is generated and time series warning information is output.

[0041] Convolutional Neural Networks (CNNs) are deep feedforward neural networks based on convolutional operations. They simulate the hierarchical feature extraction mechanism of biological visual systems to achieve representation learning and pattern recognition for grid-based data such as images and audio. Their core features include local perception, weight sharing, and spatial hierarchical abstraction, enabling them to automatically extract multi-scale features from raw data while significantly reducing the number of parameters.

[0042] Time series feature analysis is the process of systematically identifying and modeling the temporal dependencies, dynamic patterns, and key event characteristics implicit in time series data. It aims to extract the inherent patterns in the data through mathematical, statistical, and machine learning methods to support downstream tasks such as prediction, classification, and anomaly detection. Its core is to capture the local correlations, long-term trends, and cyclical fluctuations of the sequence to construct a feature set that reflects the data's evolutionary mechanisms.

[0043] A sustained upward trend refers to a state in which a variable (such as asset prices, economic indicators, pollution index, etc.) shows a continuous and stable growth trend over a period of time, manifested as prices or values ​​forming successively rising highs and lows, and demand in the supply and demand structure dominates the market balance for a long time.

[0044] Early warning signals are standardized warning signs issued by authoritative organizations (such as meteorological and emergency management departments). Using color grading (e.g., blue, yellow, orange, and red) and clear standards, they communicate the severity and potential risks of natural disasters or emergencies to the public and guide the implementation of targeted preventative measures. Their core function is to reduce blindness in disaster response and improve emergency response efficiency through quantitative indicators and intuitive identification.

[0045] Time series warning information is a real-time or near-real-time warning signal generated based on the dynamic evolution of time series data. By integrating historical patterns, current status, and predictive models, it identifies risk patterns that exceed normal thresholds (such as abnormal fluctuations, trend deviations, or cyclical mutations) and triggers a graded response mechanism at key time points to achieve preemptive risk intervention. Its core is to extract time series features and match them with risk patterns, transforming the continuous changes in data streams into actionable evidence for defensive decisions.

[0046] Step S500: Based on the time series warning information, a distributed computing framework is used to match and analyze the pollution assessment results with historical pollution data. If the matching result points to an industrial emission area, governance strategy data is generated.

[0047] A distributed computing framework is a system architecture that supports large-scale data processing and parallel computing. It achieves efficient resource utilization and improved computing performance by breaking down complex tasks into multiple subtasks and executing them collaboratively on multiple computer nodes.

[0048] An industrial emission area refers to a specific geographical or administrative boundary where pollutants are concentratedly generated and released into the environment during industrial production activities. Its definition requires a comprehensive consideration of multiple attributes such as geographical distribution, emission types, and legal regulations.

[0049] Governance strategy data is a dynamic collection of information supporting industrial pollution control decision-making. It covers core elements such as pollutant emission characteristics, technical parameters, policy standards, and implementation effectiveness. Through systematic collection, modeling, and analysis, it provides a quantitative basis for the formulation, optimization, and evaluation of pollution control measures. Its core function is to connect environmental governance goals with practical operational pathways, enabling data-driven management across the entire pollution control chain, from pollution source identification to improved governance effectiveness.

[0050] Furthermore, in the water quality monitoring method based on big data analysis provided in this embodiment, step S100 includes: Step S110: Acquire water quality parameter data including chemical parameters, spectral characteristics and biological indicators from multiple types of sensors, classify and store the water quality parameter data through a preset acquisition protocol, and obtain an original data set.

[0051] The original dataset is: , In formula (1), represents the original dataset, Indicates the total number of sensors, Indicates the Chemical parameter data collected by sensors, Indicates the The spectral characteristic data collected by the sensor, Indicates the The biological indicator data collected by multiple sensors. Formula (1) describes the integration process of multi-type sensor data.

[0052] For example, for the scenario where multiple types of sensors obtain water quality parameter data, a lake water quality monitoring system can be imagined. The sensors include chemical sensors, spectrometers and biological sensors, which respectively collect chemical parameters such as pH value and dissolved oxygen, spectral characteristics such as ultraviolet absorption peaks, and biological indicators such as algae density.

[0053] The data collection protocol can be set to collect data every hour, with data stored categorized by sensor type. For example, chemical sensor data could be stored as pH 7.2 and dissolved oxygen 8.5 mg / L, spectral data could be stored as absorbance 0.15 at a wavelength of 254 nm, and biological data could be stored as algae density 5000 cells / mL. This categorized storage of raw data sets facilitates subsequent processing and ensures data traceability.

[0054] Step S120: Using a standardization algorithm to unify the format of the original data set, adjusting the data structure of chemical parameters, spectral characteristics and biological indicators based on a preset format template to obtain an initialized data set.

[0055] In one possible implementation, a normalization algorithm unifies the format of the raw data sets. Assume that different sensor data formats vary, where the pH value might be a string of "7.2" and the dissolved oxygen value might be a floating-point number of 8.5.

[0056] The normalization algorithm converts all values ​​to a unified floating-point format and adjusts the data structure based on a pre-set template. For example, the template requires chemical parameters to be stored in the "parameter name:value" format, such as "pH:7.2," spectral features to be stored as "wavelength:absorbance" key-value pairs, and biological indicators to be stored as "indicator:density." This creates a clear structure for the initial dataset, facilitating cross-system compatibility and subsequent analysis.

[0057] Step S130: Detect abnormal data points in the initialized data set using a statistical method. If the abnormal data points deviate from a preset threshold, use a median filtering algorithm to remove the abnormal data points to obtain a standardized data set.

[0058] Specifically, outlier data points are detected using statistical methods. For example, if the pH value sequence in the initial dataset is 7.2, 7.3, 9.8, and 7.1, a statistical method is used to calculate the mean (7.35) and standard deviation. If the threshold is set to the mean ± 2 times the standard deviation, 9.8 is a significant deviation. A median filter algorithm is used to remove outliers. The median value of the nearby points 7.2, 7.3, and 7.1 is calculated, and 9.8 is replaced with 7.2. This standardized dataset is therefore more accurate and reduces the interference of outliers in the analysis.

[0059] It should be noted that categorized storage of raw data sets improves data management efficiency. For example, monitoring personnel can quickly screen the pH value trend for a particular day without having to process messy data.

[0060] Standardization ensures a consistent data format, facilitating integration with data from other monitoring stations and improving regional water quality analysis capabilities. Removing outliers enhances data reliability. For example, after removing the erroneous pH value of 9.8, the pH value series better reflects the lake's true state, supporting accurate pollution control decisions.

[0061] In one example, a standardized dataset could be used for source analysis, assuming that discharge from a factory near a lake causes abnormal pH values. Monitoring personnel discovered a pH value as low as 6.8 during a specific period, and combined with increased absorbance at specific wavelengths in the spectral signature, they inferred that the pollutant might be acidic.

[0062] The impact of pollution was further verified by the decrease in algae density, a biological indicator. This multi-dimensional data analysis benefits from the integrity and reliability of standardized data sets, significantly improving the efficiency of pollution identification.

[0063] Preferably, the implementation of the above method can also reduce data processing costs.

[0064] A unified format reduces data cleaning time, and outlier removal prevents erroneous data from misleading decision-making. For example, removing the anomalous pH value of 9.8 eliminates the need to waste resources analyzing the unreasonable high alkalinity hypothesis. Collaborative analysis of data from multiple sensor types enhances the comprehensiveness of water quality monitoring, providing strong support for environmental protection.

[0065] It can be understood that the combination of classified storage, standardized processing and outlier elimination forms an efficient data processing chain.

[0066] Each step lays the foundation for subsequent analysis, ensuring data reliability from collection to application. For example, in lake water quality monitoring, standardized datasets can be directly input into machine learning models to predict pollution trends, facilitate early intervention, and significantly improve water quality management efficiency.

[0067] Furthermore, in the water quality monitoring method based on big data analysis provided in this embodiment, step S200 includes: Step S210: Acquire chemical parameters, spectral characteristics, and biological indicators from the standardized data set, and use the principal component analysis algorithm to calculate the eigenvectors of each principal component and the corresponding variance contribution rate to obtain a set of eigenvectors.

[0068] For example, for lake water quality monitoring scenarios, the standardized data set includes chemical parameters such as pH value and dissolved oxygen, spectral characteristics such as ultraviolet absorption peaks, and biological indicators such as algae density.

[0069] The principal component analysis algorithm is used to extract key features. For example, a data set includes pH 7.2, dissolved oxygen 8.5 mg / L, absorbance at 254 nm 0.15, and algae density 5000 cells / mL. Principal component analysis uses linear transformation to convert this high-dimensional data into a set of new feature vectors, or principal components. Each principal component is a linear combination of the original variables. The calculation results may generate several principal components, the first of which may primarily reflect changes in pH and dissolved oxygen, while the second principal component may be related to spectral features.

[0070] Step S220: For the set of feature vectors, a variance analysis method is used to calculate the variance contribution rate of each feature vector. If the variance contribution rate exceeds a preset threshold, the feature vector is retained to obtain a set of filtered feature vectors.

[0071] The judgment conditions for feature vector screening are: , In formula (2), represents the filtered feature vector set, Indicates the feature vectors, Indicates the The variance contribution rate of the eigenvectors, Indicates the preset threshold value, Represents the total number of original eigenvectors.

[0072] The variance contribution percentage indicates the proportion of the original data variation that each principal component explains. For example, the first principal component may contribute 60% and the second principal component may contribute 25%.

[0073] In one possible implementation, after calculating the variance contribution, a threshold, such as 20%, is set to select eigenvectors with higher contributions. If the contributions of the first and second principal components exceed the threshold, their eigenvectors are retained to form the filtered eigenvector set. Eigenvectors with lower contributions, such as the third principal component, which reflects only noise and has a contribution of only 5%, are eliminated. This ensures that the retained data dimensions focus on reflecting key changes in water quality.

[0074] Step S230: Based on the screened feature vector set, a matrix projection method is used to map the chemical parameters, spectral features, and biological indicators in the standardized data set to a low-dimensional space to generate a reduced-dimensional feature data set.

[0075] Specifically, the matrix projection method maps a standardized dataset into a lower-dimensional space. A set of selected eigenvectors is used as the projection basis, and the original data is projected into a two-dimensional space consisting of the first and second principal components through matrix operations. For example, data such as a pH value of 7.2 and a dissolved oxygen value of 8.5 mg / L are converted to two-dimensional coordinates, such as (3.5, 1.2). This dimensionality reduction simplifies the feature dataset, retaining only the key information for subsequent analysis.

[0076] It's important to note that the implementation of principal component analysis relies on the quality of data preprocessing. Normalizing the dataset ensures that the dimensions of each parameter are consistent. For example, if pH and absorbance values ​​vary significantly, they are normalized to a uniform scale. This ensures that the principal component analysis results accurately reflect the true associations between variables.

[0077] In one embodiment, a reduced-dimensional feature dataset can be used to classify water quality. For example, if the pH value in a certain area of ​​a lake is low, the projected two-dimensional data points may cluster in that specific area, separated from the normal water quality data points. This helps quickly identify polluted areas. The reduced-dimensional data can also be used as input for classification algorithms to determine the type of pollution.

[0078] The process of selecting eigenvectors using variance analysis improves computational efficiency. By removing low-contribution principal components, the data dimension is reduced, reducing the burden of subsequent processing. For example, original four-dimensional data can be reduced to two dimensions, significantly reducing computational effort while preserving essential information.

[0079] As you can understand, dimensionality-reduced feature datasets support multi-dimensional analysis. For example, monitoring personnel can visually observe changing trends in water quality parameters using two-dimensional scatter plots and verify the ecological impact of pollution using biological indicators. This analytical approach simplifies the interpretation of complex data and improves monitoring efficiency. For example, the absorbance information retained after dimensionality reduction can help track specific pollutants. If there is an abnormal increase in absorbance during a certain period, the reduced dimensionality data can be used to locate the source of the pollution and quickly implement remediation measures. The generation of dimensionality-reduced feature datasets ensures accurate and efficient analysis.

[0080] Furthermore, in the water quality monitoring method based on big data analysis provided in this embodiment, step S300 includes: Step S310: Obtain chemical parameters, spectral features, and biological indicators from the dimensionality reduction feature data set, and use a random forest algorithm to calculate the pollution index of each data point to obtain a pollution index set.

[0081] The pollution index of the data point is: , In formula (3), represents the pollution index of the i-th data point, represents the total number of decision trees in the random forest, represents the number of leaf nodes in each tree, Indicates the The prediction function of a decision tree, Indicates the The chemical parameter vector of data points, represents the spectral feature vector of the i-th data point, Represents the biological indicator vector of the i-th data point.

[0082] The pollution index set is: , In formula (4), represents the set of pollution indices, represents the total number of data points, Indicates the The pollution index of the data point, represents the random forest algorithm function, Indicates the Each data point contains the comprehensive feature vector of chemical parameter spectral characteristics and biological indicators.

[0083] For example, in a lake water quality monitoring scenario, the dimensionality reduction feature dataset includes chemical parameters such as pH value and dissolved oxygen, spectral features such as ultraviolet absorbance, and biological indicators such as algae density.

[0084] The random forest algorithm integrates multiple decision trees to assess the pollution index for each data point. In principle, the random forest algorithm takes dimensionality-reduced features as input and generates a pollution index based on their importance. For example, a data point with a pH of 6.5, dissolved oxygen of 7.8 mg / L, absorbance of 0.18, and an algae density of 6,000 cells / mL might generate a higher pollution index, such as 0.75, due to the low pH and high algae density. The random forest algorithm uses multiple sampling and random feature selection to ensure a stable index calculation that reflects the overall water quality.

[0085] Step S320: for the pollution index set, determine whether each pollution index exceeds the threshold when the preset threshold is exceeded, and if so, generate a high pollution mark to obtain a mark set.

[0086] The high pollution mark corresponding to the pollution index is: , In formula (5), Indicates the high pollution mark corresponding to the i-th pollution index, represents the i-th pollution index value, Indicates the preset pollution threshold. When the pollution index exceeds the threshold, it is marked as 1, otherwise it is marked as 0, thus achieving binary classification judgment of the pollution level.

[0087] The resulting set of tags is: , In formula (6), represents the generated tag set, Indicates the Markup elements, Indicates the pollution index, represents the threshold parameter, represents the threshold judgment function, Represents the total number of pollution indices. Formula (6) describes the complete mapping process from the pollution index set to the label set.

[0088] In one possible implementation, a set of pollution indices is categorized using a preset threshold. For example, if the threshold is 0.6, a pollution index of 0.75 exceeding the threshold is labeled as highly polluted; an index of 0.4 below the threshold is labeled as normal. The set of labels records the pollution status of each data point. For example, of 10 data points in a lake area, 3 are labeled as highly polluted and 7 are labeled as normal. This classification method facilitates rapid identification of polluted areas.

[0089] Step S330: Based on the marker set, the high-pollution markers are combined with the dimensionality-reduced feature data by a matrix projection method to generate low-dimensional space data containing the markers, thereby obtaining a labeled dataset.

[0090] The final output labeled dataset is: , In formula (7), represents the final output labeled dataset, represents the dimensionality reduction projection operator, represents the original data feature matrix, represents the label fusion coefficient, Represents the label information matrix.

[0091] Specifically, the matrix projection method combines highly contaminated labels with dimensionally reduced feature data to generate a labeled dataset.

[0092] The reduced feature dataset is originally a two-dimensional coordinate system, such as (3.2, 1.5). Through projection, high-pollution labels are attached to the corresponding coordinates, forming labeled data points, such as (3.2, 1.5, high pollution). For example, a data point may be labeled as highly polluted due to low pH and high algae density. Combining these coordinates with the label provides a visual representation of the pollution location. This approach preserves the spatial structure of the reduced data while highlighting the characteristics of pollution.

[0093] Step S340: Obtain high pollution labels and corresponding pollution indices from the labeled data set, and use a sorting method to arrange the labels according to the size of the pollution index to obtain a pollution assessment result.

[0094] Preferably, the sorting method arranges high-pollution markers according to the size of the pollution index. Assuming that the pollution indices of three highly polluted data points are 0.75, 0.82, and 0.65, respectively, they are sorted as 0.82, 0.75, and 0.65. The pollution assessment results show that the data point with an index of 0.82 is the most seriously polluted, which may correspond to an area with a pH value of 6.2 and an algae density of 7000 cells / mL. The sorted results make it easier for monitoring personnel to prioritize high-pollution points. For example, the point with the highest index may be located near the water inlet of a lake, indicating that the source of pollution may be from upstream discharge.

[0095] It should be noted that the random forest algorithm depends on the quality of feature data.

[0096] The dimensionality-reduced feature dataset has been processed through normalization and principal component analysis to ensure consistent dimensionality across all parameters and reduce noise. For example, the dimensionality differences between pH and absorbance have been eliminated, allowing the algorithm to accurately capture contamination associations between variables. In one embodiment, the labeled dataset supports visual analysis.

[0097] Monitoring personnel observed the distribution of high-pollution markers using a two-dimensional scatter plot and discovered that high-pollution points were concentrated in a certain area of ​​the lake, possibly related to industrial wastewater discharge. This intuitive presentation simplifies the location of pollution sources.

[0098] As you can imagine, pollution assessment results provide a basis for subsequent remediation efforts. The sorted ranking of high-pollution points guides monitoring personnel to prioritize the most heavily polluted areas. For example, an index of 0.82 may indicate that a specific pollutant exceeds the permitted limit. Combined with spectral signature analysis, this confirms that the source of pollution is organic matter emissions. This analytical approach improves the targeted nature of remediation efforts.

[0099] Furthermore, in the water quality monitoring method based on big data analysis provided in this embodiment, step S400 includes: Step S410: Obtain time series data from the pollution assessment results, use a convolutional neural network to perform feature extraction on the time series data, and obtain a time series feature set.

[0100] The results of the time series pollution assessment constructed from multiple pollutant concentration data are as follows: , In formula (8), represents the pollution assessment value at time t, represents the total number of pollutant types, Indicates the The weight coefficient of each pollutant, Indicates the time The concentration of pollutants, Indicates time The random error term.

[0101] The calculation process of single-layer feature extraction in convolutional neural network is: , In formula (9), represents the output feature map of the lth convolutional layer, represents the activation function, Indicates the The convolution kernel weight matrix of the layer, Indicates the The input data of the layer, represents the convolution operation, Indicates the Bias vector for layer l.

[0102] The mathematical representation of the complete time series feature set obtained after processing by the convolutional neural network is: , In formula (10), Represents the final extracted time series feature set, Indicates the The extracted feature vectors, Represents the total number of eigenvectors.

[0103] For example, in a lake water quality monitoring scenario, a pollution index set contains data from multiple time points, forming time series data. For example, a lake pollution index is collected once a day, and a time series of values ​​(such as 0.55, 0.58, and 0.62) is formed over 30 consecutive days, reflecting the changing trend of water quality.

[0104] Convolutional neural networks are used to extract time series features. Using convolution kernels, convolutional neural networks scan time series data, capturing local patterns, such as short-term fluctuations or sustained upward trends in the pollution index. Specifically, a convolution kernel might identify a pattern where the pollution index rises from 0.55 to 0.65 over five days, generating a set of time series features, such as short-term increases or periodic fluctuations. This method effectively extracts patterns of pollution variation over time.

[0105] Step S420: Analyze the rising trend of each feature of the time series feature set using a trend detection method to determine a continuously rising feature set.

[0106] In one possible implementation, a trend detection method analyzes the upward trend of a time series feature set. Trend detection can use moving averages or linear regression to determine whether a feature is consistently rising. For example, a feature indicating a pollution index rising from 0.60 to 0.75 over 10 days, demonstrating a steady upward trend, would be classified as a consistently rising feature set.

[0107] It’s important to note that trend detection must eliminate short-term fluctuations to identify long-term changes. For example, if the pollution index suddenly drops to 0.50 on a given day but then resumes its upward trend, the detection method will ignore this anomaly and focus on the overall trend.

[0108] Step S430: determine whether at least one feature in the continuously rising feature set exceeds a preset threshold; if so, generate a corresponding warning signal to obtain a warning signal set.

[0109] Specifically, if a feature in the continuously rising feature set exceeds a preset threshold, such as 0.70, an early warning signal is generated.

[0110] Suppose a characteristic of a lake reaches 0.73 on the 25th day, exceeding the threshold and generating an early warning signal, indicating that water quality may deteriorate. The early warning signal set records all characteristics exceeding the threshold and the time points. For example, three early warning signals may be generated within 30 days, corresponding to the 20th, 25th, and 28th days respectively. Preferably, the early warning signal set supports dynamic updates. If the characteristic subsequently falls below the threshold, the warning can be revoked, ensuring real-time monitoring.

[0111] Step S440: According to the warning signal set, the warning signal is combined with the time series feature through a matrix mapping method to output time series warning information.

[0112] The output timing warning information is: , In formula (11), Indicates the Time series warning information output at all times, represents the characteristic matrix of the p-th warning mode, represents the comprehensive feature vector at time t, Indicates the total number of warning modes, represents the importance weight of the p-th warning mode, represents the warning threshold bias parameter. Formula (11) maps the multimodal features to warning probability output through the sigmoid activation function.

[0113] In one embodiment, a matrix mapping method combines warning signals with time series features to generate time series warning information. Matrix mapping attaches warning signals to corresponding time points, forming structured data. For example, if the time series feature on day 25 is 0.73, the mapping results in a data point such as time 25, feature 0.73, and warning signal. This time series warning information intuitively reflects the moment of pollution deterioration and the intensity of its features.

[0114] Understandably, monitoring personnel can use time-series warning information to quickly pinpoint windows of increased pollution. For example, a warning on the 25th day might indicate a pollution discharge from an upstream factory, prompting further investigation. Combining the warning signal with spectral analysis might reveal an abnormally high UV absorbance on the 25th day, indicating organic pollutants. Time-series warning information can also be visualized, generating a line chart showing the pollution index over time, with warning signals highlighted in red. This approach allows monitoring personnel to intuitively understand water quality dynamics and quickly respond to potential pollution risks.

[0115] Furthermore, in the water quality monitoring method based on big data analysis provided in this embodiment, step S500 includes: Step S510: Obtain a data set from the time series warning information, compare the pollution assessment results with the historical pollution data through a distributed computing framework, and obtain a matching result set.

[0116] The data set structure obtained from the time series warning information is: , In formula (12), Indicates time Time series warning information data set, Indicates time No. pollution warning data points, Indicates time The total number of data points, Represents a collection of time series.

[0117] The overall matching degree between the pollution assessment and historical data in the distributed environment is calculated as: , In formula (13), Represents the matching result score under the distributed computing framework, Indicates the number of computing nodes, Indicates the The data shards processed by the computing nodes, Indicates the current pollution assessment results, represents historical pollution data, Represents the similarity function between the evaluation results and historical data.

[0118] The mathematical representation of the matching result set obtained after comparison is: , In formula (14), Represents the final set of matching results, represents the i-th pollution assessment result, Represents the corresponding historical pollution data, represents the matching weight, represents the difference measure between the two, represents the matching threshold, Represents the total number of matching pairs.

[0119] After obtaining data sets from time-series warning information, large amounts of data can be processed through a distributed computing framework. For example, in a lake monitoring system, the data set includes 30-day pollution indices, such as 0.55, 0.58, and 0.73, which are compared with historical pollution data.

[0120] The distributed computing framework distributes tasks across multiple nodes for parallel processing, quickly comparing current data with historical records. For example, if a pollution event in historical data shows an upward trend from 0.60 to 0.75, the current data is matched to this trend and a matching result set is generated, including time points and similarity information.

[0121] Step S520: For the set of matching results, a conditional judgment method is used. If at least one matching result points to an industrial emission area, the corresponding area identification data is extracted.

[0122] Conditional judgment methods are used to filter key information from the matching result set. Specifically, if a matching result points to an industrial emission area, such as an upstream factory discharge incident, the area identification data is extracted. For example, a matching result shows a pollution index of 0.73 on the 25th day, which aligns with historical industrial emission patterns, and the area identification is "Factory A." This method uses logical judgment to identify the pollution source, ensuring targeted analysis.

[0123] Step S530: According to the area identification data, the treatment data related to the emission area is searched through a pre-established mapping table to determine the treatment data set.

[0124] Using a mapping table to query remediation data based on regional identifier data is an efficient approach. In one possible implementation, a pre-established mapping table records the remediation measures corresponding to "Factory A," such as reducing wastewater discharge and installing filtration equipment. After querying, the remediation data set might include information such as a daily wastewater treatment volume of 10 tons and a filtration efficiency of 80%. This approach directly links pollution sources to remediation measures, facilitating subsequent processing.

[0125] Step S540: Use the random forest algorithm to jointly process the governance data set and the time series data to output the final governance data.

[0126] A random forest algorithm is used to jointly process the remediation data set and time series data to output the final remediation data. Preferably, the random forest algorithm uses multiple decision trees to analyze feature importance, combining pollution index trends in the time series data with the processing capacity in the remediation data to predict remediation effectiveness. For example, if the remediation data indicates that 10 tons of wastewater are processed daily, but the time series data shows that the pollution index has increased from 0.60 to 0.75, the algorithm may output a recommendation to increase the processing capacity to 15 tons.

[0127] Specifically, random forests can analyze multiple aspects, such as the operating time of treatment equipment and peak emission periods, to generate comprehensive treatment plans. In one example, suppose a pollution index of 0.73 on day 25 is associated with "Factory A," and historical data suggests that similar incidents can be mitigated by improving filtration efficiency.

[0128] After random forest analysis, it was recommended to increase the filtration efficiency from 80% to 90%, and it was predicted that the pollution index could be reduced to below 0.65.

[0129] It's important to note that this combined processing not only optimizes treatment plans but also dynamically adapts to pollution changes, providing real-time guidance. For example, based on time series data, the pollution index exceeded the standard multiple times on the 20th, 25th, and 28th days, with matching results pointing to the same area. The treatment data set provides multi-dimensional information, such as equipment status and treatment records.

[0130] Random Forest integrates these factors and outputs specific recommendations, such as adjusting emission periods and adding temporary purification equipment. This multifaceted approach ensures comprehensive and feasible solutions, enabling rapid responses to pollution risks.

[0131] Another aspect of the present invention relates to a water quality monitoring system based on big data analysis, which is used to implement the above-mentioned water quality monitoring method based on big data analysis. The water quality monitoring system based on big data analysis includes an acquisition module, a first generation module, a first output module, a second output module and a second generation module, wherein the acquisition module is used to acquire water quality parameter data collected by multiple types of sensors, and adopts a standardization algorithm to uniformly process the format of the water quality parameter data. If abnormal data points are detected, they are eliminated through a median filtering algorithm to obtain a standardized data set; the water quality parameter data includes chemical parameters, spectral characteristics and biological indicators; the first generation module is used to extract features of the chemical parameters, spectral characteristics and biological indicators for the standardized data set using a principal component analysis algorithm to obtain characteristics Eigenvector, if the variance contribution rate of the eigenvector exceeds the preset threshold, the eigenvector is retained to generate a reduced-dimensionality feature data set; the first output module is used to construct a pollution assessment model based on the reduced-dimensionality feature data set using the random forest algorithm, calculate the pollution index, and generate a high-pollution state mark if the pollution index exceeds the preset threshold, and output the pollution assessment result; the second output module is used to perform time series feature analysis based on the pollution assessment result using the convolutional neural network algorithm, and if it is detected that the pollution index shows a continuous upward trend, an early warning signal is generated and time series early warning information is output; the second generation module is used to match and analyze the pollution assessment results with historical pollution data using a distributed computing framework based on the time series early warning information, and if the matching result points to an industrial emission area, governance strategy data is generated.

[0132] Furthermore, the water quality monitoring system based on big data analysis provided in this embodiment has an acquisition module including a first acquisition unit, a second acquisition unit and a third acquisition unit, wherein the first acquisition unit is used to acquire water quality parameter data including chemical parameters, spectral characteristics and biological indicators from multiple types of sensors, and classify and store the water quality parameter data through a preset acquisition protocol to obtain an original data set; the second acquisition unit is used to use a standardized algorithm to unify the format of the original data set, and adjust the data structure of the chemical parameters, spectral characteristics and biological indicators based on a preset format template to obtain an initialized data set; the third acquisition unit is used to detect abnormal data points in the initialized data set through statistical methods. If the abnormal data points deviate from the preset threshold, the median filtering algorithm is used to eliminate the abnormal data points to obtain a standardized data set.

[0133] Preferably, the water quality monitoring system based on big data analysis provided in this embodiment, the first generation module includes a fourth acquisition unit, a fifth acquisition unit and a generation unit, wherein the fourth acquisition unit is used to obtain chemical parameters, spectral characteristics and biological indicators from the standardized data set, and use the principal component analysis algorithm to calculate the eigenvectors and corresponding variance contribution rates of each principal component to obtain a set of eigenvectors; the fifth acquisition unit is used to use the variance analysis method to calculate the variance contribution rate of each eigenvector in the set of eigenvectors, if the variance contribution rate exceeds a preset threshold, the eigenvector is retained to obtain a set of screened eigenvectors; the generation unit is used to map the chemical parameters, spectral characteristics and biological indicators in the standardized data set to a low-dimensional space based on the screened eigenvector set using a matrix projection method to generate a reduced-dimensional feature data set.

[0134] Furthermore, the water quality monitoring system based on big data analysis provided in this embodiment, the first output module includes a sixth acquisition unit, a seventh acquisition unit, an eighth acquisition unit and a ninth acquisition unit, wherein the sixth acquisition unit is used to obtain chemical parameters, spectral characteristics and biological indicators from the dimensionality reduction feature data set, and use the random forest algorithm to calculate the pollution index of each data point to obtain a pollution index set; the seventh acquisition unit is used to determine whether each pollution index exceeds the threshold value at a preset threshold value for the pollution index set, and if so, generate a high pollution mark to obtain a mark set; the eighth acquisition unit is used to combine the high pollution mark with the dimensionality reduction feature data through a matrix projection method according to the mark set to generate low-dimensional space data containing the mark to obtain a marked data set; the ninth acquisition unit is used to obtain the high pollution mark and the corresponding pollution index from the marked data set, and use a sorting method to arrange the marks according to the size of the pollution index to obtain a pollution assessment result.

[0135] Compared with the existing technology, the water quality monitoring method and system based on big data analysis provided in this embodiment collects water quality parameter data, including chemical parameters, spectral characteristics and biological indicators, through multiple types of sensors, standardizes the data and eliminates outliers, and then uses principal component analysis to extract features and reduce dimensionality. Based on the feature data set after dimensionality reduction, a random forest algorithm is used to construct a pollution assessment model, calculate the pollution index and generate an assessment result. Furthermore, a convolutional neural network is used to perform time series analysis on the pollution index, detect the upward trend and generate an early warning signal. Finally, the assessment results are matched and analyzed with historical data through a distributed computing framework to locate possible pollution sources and formulate governance strategies. This embodiment realizes real-time monitoring, assessment, early warning and tracing of water pollution, providing comprehensive technical support for water environment protection.

[0136] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they are aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the invention. Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the invention. Thus, the present invention is intended to include such changes and modifications as fall within the scope of the claims and their equivalents.

Claims

1. A water quality monitoring method based on big data analysis, characterized in that: The following steps are involved: Acquire water quality parameter data collected by multiple types of sensors, use a standardization algorithm to unify the format of the water quality parameter data, and remove any abnormal data points through a median filter algorithm to obtain a standardized data set; the water quality parameter data includes chemical parameters, spectral characteristics, and biological indicators; For the standardized data set, a principal component analysis algorithm is used to extract features of the chemical parameters, spectral features, and biological indicators to obtain a feature vector. If the variance contribution rate of the feature vector exceeds a preset threshold, the feature vector is retained to generate a reduced-dimensional feature data set. Based on the dimensionality-reduced feature dataset, a pollution assessment model is constructed using a random forest algorithm to calculate a pollution index. If the pollution index exceeds a preset threshold, a high pollution state flag is generated and a pollution assessment result is output; Based on the pollution assessment results, a convolutional neural network algorithm is used to perform time series feature analysis to obtain a pollution index. If it is detected that the pollution index shows a continuous upward trend, an early warning signal is generated and time series warning information is output; According to the time series warning information, a distributed computing framework is used to match and analyze the pollution assessment results with historical pollution data. If the matching result points to an industrial emission area, governance strategy data is generated.

2. The water quality monitoring method based on big data analysis according to claim 1, characterized in that: The steps of obtaining water quality parameter data collected by multiple types of sensors, formatting the water quality parameter data using a standardization algorithm, and removing abnormal data points using a median filtering algorithm to obtain a standardized data set include: Acquire water quality parameter data including chemical parameters, spectral characteristics, and biological indicators from multiple types of sensors, classify and store the water quality parameter data through a preset acquisition protocol, and obtain a raw data set; The raw data set is formatted uniformly using a standardized algorithm, and the data structures of chemical parameters, spectral characteristics, and biological indicators are adjusted based on a preset format template to obtain an initialized data set; Abnormal data points in the initialized data set are detected by statistical methods. If the abnormal data points deviate from a preset threshold, the median filtering algorithm is used to eliminate the abnormal data points to obtain a standardized data set.

3. The water quality monitoring method based on big data analysis according to claim 1, characterized in that: For the standardized data set, a principal component analysis algorithm is used to extract features of the chemical parameters, spectral features, and biological indicators to obtain a feature vector. If the variance contribution rate of the feature vector exceeds a preset threshold, the feature vector is retained. The steps of generating a dimensionality-reduced feature data set include: Chemical parameters, spectral characteristics and biological indicators are obtained from the standardized data set, and the principal component analysis algorithm is used to calculate the eigenvectors and corresponding variance contributions of each principal component to obtain a set of eigenvectors; For the set of feature vectors, a variance analysis method is used to calculate the variance contribution rate of each feature vector. If the variance contribution rate exceeds a preset threshold, the feature vector is retained to obtain a set of filtered feature vectors. According to the screening feature vector set, the chemical parameters, spectral features and biological indicators in the standardized data set are mapped to a low-dimensional space using a matrix projection method to generate a reduced-dimensional feature data set.

4. The water quality monitoring method based on big data analysis according to claim 1, characterized in that: Based on the dimensionality reduction feature data set, a pollution assessment model is constructed using a random forest algorithm to calculate a pollution index. If the pollution index exceeds a preset threshold, a high pollution state flag is generated. The steps of outputting the pollution assessment result include: Chemical parameters, spectral features, and biological indicators are obtained from the dimensionality-reduced feature dataset, and the pollution index of each data point is calculated using the random forest algorithm to obtain a pollution index set; For the pollution index set, determining whether each pollution index exceeds the threshold when a preset threshold is met, and if so, generating a high pollution mark to obtain a mark set; According to the marker set, the highly contaminated markers are combined with the dimensionality reduction feature data by a matrix projection method to generate low-dimensional spatial data containing the markers, thereby obtaining a labeled dataset; High pollution marks and corresponding pollution indexes are obtained from the labeled data set, and the marks are arranged according to the size of the pollution index using a sorting method to obtain a pollution assessment result.

5. The water quality monitoring method based on big data analysis according to claim 1, characterized in that: Based on the pollution assessment results, a convolutional neural network algorithm is used to perform time series feature analysis to obtain a pollution index. If the pollution index is detected to be continuously increasing, an early warning signal is generated. The steps of outputting the time series early warning information include: Acquire time series data from the pollution assessment results, and perform feature extraction on the time series data using a convolutional neural network to obtain a time series feature set; For the time series feature set, analyzing the rising trend of each feature by a trend detection method to determine a continuously rising feature set; Determine whether at least one feature in the continuously rising feature set exceeds a preset threshold, and if so, generate a corresponding warning signal to obtain a warning signal set; According to the warning signal set, the warning signal is combined with the time series feature through a matrix mapping method to output time series warning information.

6. The water quality monitoring method based on big data analysis according to claim 1, characterized in that: Based on the time series warning information, a distributed computing framework is used to match and analyze the pollution assessment results with historical pollution data. If the matching result points to an industrial emission area, the steps of generating control strategy data include: Acquire a data set from the time series warning information, and compare the pollution assessment results with historical pollution data through a distributed computing framework to obtain a matching result set; For the set of matching results, a conditional judgment method is adopted, and if at least one matching result points to an industrial emission area, the corresponding area identification data is extracted; According to the area identification data, querying the treatment data related to the emission area through a pre-established mapping table to determine the treatment data set; The random forest algorithm is used to jointly process the governance data set and the time series data to output the final governance data.

7. A water quality monitoring system based on big data analysis, used to implement the water quality monitoring method based on big data analysis according to any one of claims 1 to 6, characterized in that: The water quality monitoring system based on big data analysis includes: An acquisition module is used to obtain water quality parameter data collected by multiple types of sensors, and use a standardization algorithm to unify the format of the water quality parameter data. If any abnormal data points are detected, they are removed through a median filtering algorithm to obtain a standardized data set; the water quality parameter data includes chemical parameters, spectral characteristics and biological indicators; A first generation module is configured to extract features of the chemical parameters, spectral features, and biological indicators using a principal component analysis algorithm for the standardized data set to obtain a feature vector. If the variance contribution rate of the feature vector exceeds a preset threshold, the feature vector is retained to generate a reduced-dimensional feature data set. A first output module is configured to construct a pollution assessment model using a random forest algorithm based on the dimensionality reduction feature dataset, calculate a pollution index, generate a high pollution state flag if the pollution index exceeds a preset threshold, and output a pollution assessment result; A second output module is configured to perform time series feature analysis using a convolutional neural network algorithm based on the pollution assessment results to obtain a pollution index. If a continuous upward trend in the pollution index is detected, an early warning signal is generated and time series warning information is output; The second generation module is used to match and analyze the pollution assessment results with historical pollution data based on the time series warning information using a distributed computing framework, and generate governance strategy data if the matching result points to an industrial emission area.

8. The water quality monitoring system based on big data analysis according to claim 7, characterized in that: The acquisition module includes: A first acquisition unit is configured to acquire water quality parameter data including chemical parameters, spectral characteristics, and biological indicators from multiple types of sensors, and classify and store the water quality parameter data using a preset acquisition protocol to obtain an original data set; a second acquisition unit, configured to perform format unification processing on the original data set using a standardized algorithm, and adjust the data structure of chemical parameters, spectral characteristics, and biological indicators based on a preset format template to obtain an initialized data set; The third acquisition unit is used to detect abnormal data points in the initialized data set by a statistical method. If the abnormal data points deviate from a preset threshold, the median filtering algorithm is used to eliminate the abnormal data points to obtain a standardized data set.

9. The water quality monitoring system based on big data analysis according to claim 7, characterized in that: The first generation module includes: a fourth acquisition unit, configured to acquire chemical parameters, spectral characteristics, and biological indicators from the standardized data set, and calculate the eigenvectors and corresponding variance contributions of each principal component using a principal component analysis algorithm to obtain a set of eigenvectors; a fifth acquiring unit, configured to calculate, for the set of feature vectors, a variance contribution rate of each feature vector using a variance analysis method, and retain the feature vector if the variance contribution rate exceeds a preset threshold, thereby obtaining a set of filtered feature vectors; The generating unit is used to map the chemical parameters, spectral features and biological indicators in the standardized data set to a low-dimensional space using a matrix projection method according to the screening feature vector set to generate a reduced-dimensional feature data set.

10. The water quality monitoring system based on big data analysis according to claim 7, characterized in that: The first output module includes: a sixth acquisition unit, configured to acquire chemical parameters, spectral features, and biological indicators from the dimensionally reduced feature data set, and calculate the pollution index of each data point using a random forest algorithm to obtain a pollution index set; a seventh obtaining unit, configured to determine, for the pollution index set, whether each pollution index exceeds a threshold value when the threshold value is set, and generate a high pollution mark if the pollution index exceeds the threshold value, thereby obtaining a mark set; an eighth acquisition unit, configured to combine the high-pollution markers with the dimensionality reduction feature data by a matrix projection method according to the marker set, to generate low-dimensional spatial data containing the markers, and obtain a labeled data set; The ninth acquisition unit is used to obtain high pollution marks and corresponding pollution indexes from the labeled data set, and arrange the marks according to the size of the pollution index using a sorting method to obtain a pollution assessment result.

Citation Information

Cited By

  • Sewage treatment method and system based on deep learning

    CN121030476A

  • Tea moisture measurement error compensation method and system

    CN121207911A

  • Online spectrum detection compensation method, equipment and system based on standard liquid dilution curve

    CN121409898A

  • Method and system for testing rigidity of main shaft of gang tool machine

    CN121917175A