A multi-source water environment data processing method and system based on intelligent Internet of Things

By adopting a multi-source water environment data processing method based on smart IoT in water environment monitoring, the problems of multi-source heterogeneous data integration and sharing are solved, efficient data storage and real-time analysis are realized, and the efficiency and accuracy of water environment monitoring are improved.

CN119474077BActive Publication Date: 2025-05-02BEIJING UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510052583.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-02
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

There is a problem that multi-source heterogeneous data is difficult to effectively integrate and share in water environment monitoring. The data is characterized by multi-dimensional, high noise and nonlinearity. Traditional data processing methods are difficult to meet the requirements of real-time, accuracy and comprehensiveness.

Method used

The multi-source water environment data processing method based on smart IoT is adopted, and multi-source heterogeneous data is converted into a unified format through a data adapter, wavelet analysis and principal component analysis are performed for dimensionality reduction, data storage is adopted for data storage, and streaming data processing and machine learning methods are used for real-time analysis and prediction.

Benefits of technology

It realizes efficient storage, rapid retrieval, real-time analysis and intelligent early warning of multi-source water environment data, solves the problems of heterogeneity, redundancy and poor real-time performance of data, and improves the efficiency and accuracy of water environment monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119474077B_ABST
    Figure CN119474077B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-source water environment data processing method and system based on intelligent Internet of Things, which relates to the field of data processing technology, including: standardizing the acquired multi-source heterogeneous water environment monitoring data, obtaining the standardized data, performing multi-scale decomposition, and obtaining low-dimensional data; using a distributed storage architecture for low-dimensional data, dispersively storing the data on several nodes, establishing a data index and metadata management mechanism, and performing data retrieval and access; for the low-dimensional data, using a stream data processing method to clean, aggregate and analyze, extract key indicators and events, and construct water quality abnormality warning rules by setting water quality parameter thresholds and change rate thresholds in combination with a decision tree method to determine whether there is water quality abnormality; and constructing a water quality prediction model by a support vector machine and a random forest machine learning method to predict water quality trends. The present invention provides strong technical support for water environment monitoring and management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a multi-source water environment data processing method and system based on smart Internet of Things. Background Art

[0002] In water environment monitoring, there is a problem that multi-source heterogeneous data is difficult to effectively integrate and share. The data format, acquisition frequency, and transmission protocol between different monitoring stations are different, which makes it difficult to interconnect and integrate data for analysis. At the same time, water environment monitoring data has the characteristics of multi-dimensionality, high noise, and nonlinearity. How to achieve data dimensionality reduction and feature extraction while ensuring information integrity is a technical problem that needs to be solved urgently.

[0003] In addition, the amount of water environment monitoring data is huge, often reaching TB or even PB levels. How to achieve efficient data transmission and storage within limited transmission bandwidth and storage space is also a thorny issue. Traditional data processing methods are difficult to adapt to the real-time, accuracy and comprehensive requirements of water environment monitoring. A new solution based on smart IoT is urgently needed to break through the barriers of multi-source data, achieve seamless integration and intelligent analysis of global data, and provide a reliable basis for water environment management decisions. Summary of the invention

[0004] In order to solve the technical problems existing in the above-mentioned prior art, the present invention proposes a multi-source water environment data processing method and system based on smart Internet of Things, which solves the problems of heterogeneity, redundancy, and poor real-time performance of water environment monitoring data, and realizes efficient storage, rapid retrieval, real-time analysis and intelligent early warning of data.

[0005] On the one hand, to achieve the above-mentioned purpose, the present invention provides a multi-source water environment data processing method based on smart Internet of Things, comprising:

[0006] The acquired multi-source heterogeneous water environment monitoring data are standardized to obtain standardized data, and the standardized data are multi-scale decomposed to obtain low-dimensional data;

[0007] A distributed storage architecture is used for the low-dimensional data, the data is dispersedly stored on several nodes, and a data index and metadata management mechanism is established to perform data retrieval and access;

[0008] For the low-dimensional data, a stream data processing method is used to clean, aggregate and analyze it, extract key indicators and events, and build water quality abnormality warning rules by setting water quality parameter thresholds and change rate thresholds combined with a decision tree method to determine whether there is water quality abnormality;

[0009] Through support vector machine and random forest machine learning methods, a water quality prediction model is constructed to predict water quality trends.

[0010] On the other hand, to achieve the above-mentioned purpose, the present invention also provides a multi-source water environment data processing system based on smart Internet of Things, comprising:

[0011] The data adaptation module is used to convert the multi-source heterogeneous water environment monitoring data into a unified data model and format using a data adapter method to obtain standardized data;

[0012] The wavelet analysis module is used to perform multi-scale decomposition on the standardized data by using the wavelet analysis method, and to remove high-frequency noise interference by using a preset threshold filtering method to obtain low-dimensional data;

[0013] A distributed storage module is used to adopt a distributed storage architecture to disperse and store the low-dimensional data on a number of nodes, establish a data index and metadata management mechanism, and realize fast retrieval and access of data;

[0014] A stream data processing module is used to process, clean, aggregate and analyze the low-dimensional data using a stream data processing method, extract key indicators and events, and build early warning rules by setting water quality parameter thresholds and change rate thresholds in combination with a decision tree algorithm to determine whether there is water quality anomaly;

[0015] The machine learning prediction module is used to establish a water quality prediction model and predict water quality trends using support vector machine and random forest machine learning methods.

[0016] Compared with the prior art, the present invention has the following advantages and technical effects:

[0017] (1) The present invention discloses a multi-source water environment data processing method and system based on intelligent Internet of Things. For multi-source heterogeneous water environment monitoring data, a data adapter model is used to convert it into a unified format to eliminate data barriers. Then, dimensionality reduction processing is performed through wavelet analysis and principal component analysis to extract key features. The data after dimensionality reduction is compressed and stored using a distributed storage architecture, and an index mechanism is established to achieve fast retrieval.

[0018] (2) The present invention uses stream data processing technology to perform real-time analysis on compressed and stored data, and combines the decision tree model to build warning rules to achieve real-time warning of water quality anomalies. At the same time, a machine learning model is trained based on historical data to predict water quality trends. This method solves the problems of heterogeneity, redundancy, and poor real-time performance of water environment monitoring data, and achieves efficient storage, rapid retrieval, real-time analysis, and intelligent warning of data, providing powerful technical support for water environment monitoring and management. The present invention significantly improves the efficiency and accuracy of water environment monitoring, and is of great significance to protecting water resources and improving the water environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0020] Figure 1 This is a flow chart of a multi-source water environment data processing method based on smart Internet of Things according to an embodiment of the present invention;

[0021] Figure 2 This is a structural schematic diagram of a multi-source water environment data processing system based on smart Internet of Things according to an embodiment of the present invention. DETAILED DESCRIPTION

[0022] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0023] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0024] The present invention proposes a multi-source water environment data processing method based on smart Internet of Things. Figure 1 ,include:

[0025] The acquired multi-source heterogeneous water environment monitoring data are standardized to obtain standardized data, and the standardized data are decomposed at multiple scales to obtain low-dimensional data;

[0026] A distributed storage architecture is used for low-dimensional data, which is stored on several nodes in a dispersed manner. A data index and metadata management mechanism is established to perform data retrieval and access.

[0027] For low-dimensional data, we use streaming data processing methods to clean, aggregate and analyze it, extract key indicators and events, and build water quality anomaly warning rules by setting water quality parameter thresholds and change rate thresholds combined with decision tree methods to determine whether there is water quality anomaly.

[0028] Through support vector machine and random forest machine learning methods, a water quality prediction model is constructed to predict water quality trends.

[0029] Furthermore, the standardized data are obtained, including:

[0030] For data from different sources and formats, a pre-established data adapter model is used to convert the data into a unified data model and format;

[0031] The converted data is quality checked to determine whether the data meets the preset data quality standards. If not, the data is cleaned and repaired to obtain the standardized data.

[0032] Specifically, for data from different sources and formats, a pre-established data adapter model is used to convert the data into a unified data model and format; the converted data is quality checked to determine whether the data meets the preset data quality standards. If not, the data is cleaned and repaired to obtain high-quality standardized data.

[0033] This embodiment also includes classifying the data using a clustering algorithm according to the characteristic attributes of the standardized data, dividing the data with similar characteristics into the same category to form different data sets; for each data set, using an association rule mining algorithm to discover the association relationship and rules between the data and obtain the association rules within the data set; comparing and integrating the association rules of different data sets to find the association relationship between different data sources, forming association rules for cross-source data, and realizing data interconnection and interoperability.

[0034] Take a city as an example, and obtain data from the Municipal Environmental Protection Bureau, Water Resources Bureau and third-party monitoring agencies at the same time. These data formats are different. The Environmental Protection Bureau provides XML format, the Water Resources Bureau uses CSV format, and the third-party agency uses JSON format. To this end, it is necessary to establish corresponding data adapters for different data sources to convert heterogeneous data into a unified standard format. Convert all data into structured data containing fields such as monitoring points, time, indicators, and values. Data quality inspection is a key step to ensure the reliability of subsequent analysis. Taking the dissolved oxygen indicator as an example, the normal range should be between 0-20mg / L. If obvious outliers such as -5mg / L or 100mg / L are found in the data, data cleaning is required. Median filling, linear interpolation and other methods can be used to correct outliers to ensure the accuracy and continuity of the data.

[0035] Further, obtaining the low-dimensional data includes:

[0036] The wavelet analysis method is used to perform multi-scale decomposition on the standardized data, extract data features of different frequency bands, and use a preset threshold filtering method to remove high-frequency noise interference to obtain the low-dimensional data.

[0037] Specifically, the data after standardization is obtained, and the wavelet analysis algorithm is used to perform multi-scale decomposition on the data according to the preset wavelet basis function to obtain the wavelet coefficients of different frequency bands. For the wavelet coefficients obtained by wavelet decomposition, the characteristic vectors of each frequency band are determined according to the energy distribution characteristics of each frequency band to obtain a multi-scale feature representation. According to the characteristic vectors of each frequency band, a clustering algorithm is used to perform cluster analysis on the characteristic vectors to obtain the cluster centers of different frequency bands as the representative features of the frequency band. For the representative features of each frequency band, by calculating the difference between each feature and a preset threshold, it is determined whether the feature is a high-frequency noise feature. If the difference is greater than the threshold, it is determined to be a noise feature. For the frequency band determined to be a noise feature, a wavelet threshold filtering method is used to set the wavelet coefficients of the frequency band to zero to achieve the removal of high-frequency noise. According to the wavelet coefficients after removing the high-frequency noise, a wavelet reconstruction algorithm is used to inversely transform the wavelet coefficients of each frequency band to obtain the denoised time domain data. For the denoised time domain data, the principal component analysis algorithm is used to perform dimensionality reduction processing, extract the low-dimensional feature representation of the data, and obtain smooth low-dimensional data as input for subsequent analysis.

[0038] Furthermore, a wavelet analysis method is used to perform multi-scale decomposition on the standardized data, including:

[0039] According to the preset wavelet basis function, the wavelet analysis algorithm is used to perform multi-scale decomposition on the data to obtain wavelet coefficients in different frequency bands;

[0040] For the wavelet coefficients obtained by wavelet decomposition, the feature vectors of each frequency band are determined according to the energy distribution characteristics of each frequency band, and a multi-scale feature representation is obtained;

[0041] According to the feature vectors of each frequency segment, a clustering algorithm is used to perform cluster analysis on the feature vectors, to obtain cluster centers of different frequency segments respectively, and to use different cluster centers as representative features of the corresponding frequency segments;

[0042] For the representative features of each frequency band, by calculating the difference between different representative features and a preset threshold, it is determined whether the corresponding feature is a high-frequency noise feature. If the difference is greater than the threshold, it is determined to be a noise feature;

[0043] For the frequency bands determined to be noise characteristics, the wavelet threshold filtering method is used to set the wavelet coefficients of the corresponding frequency bands to zero to remove high-frequency noise;

[0044] According to the wavelet coefficients after removing high-frequency noise, the wavelet coefficients of each frequency band are inversely transformed using the wavelet reconstruction algorithm to obtain the denoised time domain data;

[0045] For the denoised time domain data, principal component analysis method is used to reduce the dimension and obtain low-dimensional data.

[0046] Specifically, wavelet analysis is a powerful signal processing tool that can perform multi-scale decomposition on data. In this embodiment, Daubechies wavelet is selected as the basis function, and 5 layers of wavelet decomposition are performed on indicators such as dissolved oxygen and pH value. In this way, wavelet coefficients of different frequency bands can be obtained, reflecting the different scale characteristics of water quality changes. For the decomposed wavelet coefficients, the energy distribution of each frequency band can be calculated. For example, the low frequency band may reflect the long-term trend of water quality changes, while the high frequency band may contain short-term fluctuations and noise information. By analyzing the energy distribution, the characteristic vectors of each frequency band can be extracted to form a multi-scale feature representation of the data. In order to further extract representative features, the characteristic vectors can be clustered. In this embodiment, taking the K-means algorithm as an example, the characteristic vectors of each frequency band are clustered into 3-5 categories, and the cluster center can be used as the representative feature of the frequency band. This can greatly reduce the data dimension and highlight the key features. Next, it is necessary to identify and remove high-frequency noise. A threshold can be set in advance, and the frequency band with an energy share of less than 5% is regarded as potential noise. By calculating the difference between the characteristics of each frequency band and the threshold, it can be determined whether it is a noise feature. For the high-frequency bands judged as noise, the soft threshold method is used to set their wavelet coefficients to zero to achieve denoising. After denoising, the wavelet reconstruction algorithm is used to inversely transform the processed wavelet coefficients to obtain the denoised time domain data. Taking dissolved oxygen data as an example, the denoised data curve will become smoother, removing short-term random fluctuations and better reflecting the changing trend of dissolved oxygen. In order to further reduce the data dimension, the denoised data can be subjected to principal component analysis. Assuming that the original data contains 10 water quality indicators, through principal component analysis, it may be found that the first three principal components can explain more than 90% of the data variance. In this way, these three principal components can be used as a low-dimensional representation of the data, greatly simplifying the subsequent analysis process.

[0047] Furthermore, the data index and metadata management mechanism is established to perform data retrieval and access, including:

[0048] According to a preset distributed storage architecture, the low-dimensional data is dispersedly stored on a plurality of storage nodes;

[0049] For low-dimensional data stored in a decentralized manner, an inverted index technology is used to establish a data index, the index content of which includes a data identifier, a storage node identifier, and a timestamp, and the data index is stored in an index database;

[0050] For decentralized low-dimensional data, extract data attribute information, including data volume, data format, acquisition time and acquisition device, generate metadata, and store the metadata in a metadata database;

[0051] If a data search request is received, the request content is parsed, search conditions are extracted, and the index database is searched according to the search conditions to obtain data identifiers and storage node identifier information that meet the conditions;

[0052] According to the index retrieval results, the node address of the target data storage is obtained, and the corresponding storage node is accessed through the distributed file system to obtain the target monitoring data;

[0053] The target monitoring data is aggregated and integrated, and according to the request requirements, a correlation analysis algorithm is used to analyze and mine the monitoring data to generate an analysis result report.

[0054] Specifically, according to the preset distributed storage architecture, the low-dimensional data is distributed and stored on multiple server nodes. In this embodiment, the data is distributed and stored on three server nodes, and each node is responsible for storing a part of the data. The advantage of this is that the data load can be distributed to multiple nodes, improving the efficiency of data storage and access, while enhancing the reliability and scalability of the system. For low-dimensional data stored in a decentralized manner, an inverted index technology is used to establish a data index. For example, for dissolved oxygen data, an index can be established, and the index content includes data identification (such as data number), storage node identification (such as node 1, node 2, node 3), timestamp (such as October 26, 2023 10:00), etc., and the index is stored in a dedicated index database. When it is necessary to find dissolved oxygen data for a certain period of time, the storage node where the data is located can be quickly located through the index. For low-dimensional data stored in a decentralized manner, data attribute information is extracted and metadata is generated. Metadata records the basic information of the data. For example, you can extract information such as data volume (such as 100MB), data format (such as table), collection time (such as October 26, 2023), collection equipment (such as multi-parameter water quality analyzer), generate metadata, and store the metadata in a dedicated metadata database. Metadata can help users understand the basic situation of the data and facilitate data management and use. Receive a data retrieval request, such as "query for data with dissolved oxygen greater than 5mg / L at all monitoring points in a lake from October 26, 2023 to October 28, 2023". Parse the request content and extract the retrieval conditions, including time range, monitoring indicators, monitoring points, and indicator thresholds. According to these retrieval conditions, search in the index database to obtain information such as data identifiers and storage node identifiers that meet the conditions. According to the index retrieval results, you can know which nodes the target data is stored on, such as node 1 and node 3. Through the distributed file system, access these two nodes to obtain the target monitoring data. According to the requirements of the request, a correlation analysis algorithm, such as the Pearson correlation coefficient, is used to analyze and mine the dissolved oxygen data and other indicators (such as water temperature and pH value) to generate an analysis result report. For example, the correlation between dissolved oxygen and water temperature can be analyzed to determine whether there is a negative correlation, and a report can be generated to illustrate the size and significance level of the correlation coefficient. The generated analysis result report is returned to the requesting terminal, such as the user's computer or mobile phone. On the interface of the requesting terminal, the report content is displayed in the form of charts, tables, etc. For example, a line chart is used to show the trend of dissolved oxygen over time, and a scatter plot is used to show the relationship between dissolved oxygen and water temperature. In this way, the characteristics and rules of the data can be intuitively understood, and the data retrieval and access process can be completed.

[0055] Furthermore, the decision tree method is combined to construct water quality abnormality warning rules to determine whether there is water quality abnormality, including:

[0056] Performing data cleaning on the low-dimensional data to remove noise and outliers;

[0057] Aggregate the cleaned data according to the preset aggregation rules;

[0058] Use data mining methods to analyze aggregated data and extract key indicators and key events;

[0059] According to the water quality parameter threshold and change rate threshold, combined with the decision tree model, the water quality abnormality warning rules are constructed;

[0060] Inputting the extracted key indicators and key events into a decision tree model to determine whether there is water quality abnormality;

[0061] If the decision tree model determines that there is abnormal water quality, it will trigger real-time warning and intelligent decision-making, and notify relevant departments to take corresponding measures.

[0062] Specifically, stream data processing technology can realize the continuous collection of low-dimensional data, such as using the Internet of Things sensor network to collect water temperature, pH value, dissolved oxygen and other indicator data every 5 minutes. Data cleaning is a key step to ensure data quality. By setting a reasonable threshold range, obviously abnormal data can be filtered out. In this embodiment, if the pH value exceeds the range of 3-10, it may be caused by sensor failure or sudden emission of pollutants, which requires further verification.

[0063] Data mining methods can extract valuable information from massive amounts of data. Through time series analysis, the periodic changes in water quality indicators can be discovered; through association rule mining, the correlation between different pollutants can be found. For example, it was found that an increase in ammonia nitrogen concentration is often accompanied by a decrease in dissolved oxygen, which may indicate the emission of organic pollutants.

[0064] In this embodiment, it is necessary to combine expert experience and historical data to construct the abnormal water quality warning rules. An absolute threshold value of a single indicator can be set, such as a total phosphorus concentration exceeding 0.2 mg / L is abnormal; a change rate threshold value can also be set, such as a pH value that changes by more than 1 unit within 30 minutes triggers an early warning. These rules should be personalized according to the characteristics of different water bodies. The decision tree model is an intuitive and effective classification method. By training historical data, a decision tree that can quickly judge the water quality status can be constructed. In this embodiment, first determine whether the dissolved oxygen is lower than 5 mg / L. If so, then determine whether the COD is higher than 40 mg / L, and so on, and finally conclude whether the water quality is abnormal. Once the water quality is abnormal, the system will automatically trigger the early warning mechanism. Through real-time monitoring and rapid response, water quality problems can be controlled in the bud to avoid causing greater environmental impact. At the same time, the long-term accumulated data also provides a scientific basis for the formulation of water environment protection policies, which helps to achieve sustainable management of water resources.

[0065] Furthermore, data mining methods are used to analyze the aggregated data and extract key indicators and key events, including:

[0066] Convert low-dimensional data into a format suitable for mining algorithm input, and use data aggregation methods to aggregate the converted data according to specific dimensions;

[0067] Input the summarized data into the selected data mining method for training to obtain the mining results;

[0068] The key indicators and the key events are extracted respectively according to the mining results.

[0069] Specifically, according to business needs, determine the data mining algorithm, such as association rules, clustering algorithms or decision tree algorithms, to discover hidden patterns and laws from the monitoring data. Input the aggregated data into the selected data mining algorithm for training, and use the algorithm to automatically discover patterns such as association rules, clustering results or decision trees from the data. Identify business-related key indicators from the mining results, such as abnormal indicators or indicators exceeding the threshold, as the focus of business monitoring and early warning. Extract business-related key events from the mining results, such as equipment failures, sudden increases in traffic and other abnormal situations, as key events for business monitoring and disposal. Form reports or visualization results with the extracted key indicators and key events to provide intuitive data analysis results for business personnel to assist them in making decisions.

[0070] Furthermore, a water quality prediction model is constructed, including:

[0071] Obtain historical water quality data, clean and preprocess the data, remove outliers and missing values, and normalize the data to obtain a standardized water quality data set;

[0072] According to the characteristics of water quality data and prediction objectives, support vector machine, random forest and gradient boosting tree machine learning methods were selected to build an initial water quality prediction model;

[0073] Dividing the standardized water quality data set into a training set and a test set, using the training set data to train the initial water quality prediction model, and obtaining a trained water quality prediction model by adjusting model parameters and iterative optimization;

[0074] The trained water quality prediction model is evaluated using the test set data, and the prediction accuracy, recall rate and F1 value evaluation indicators of the model are calculated using the cross-validation method. The model is optimized according to the evaluation results to obtain a water quality prediction model.

[0075] Specifically, historical water quality data is obtained, the data is cleaned and preprocessed, outliers and missing values ​​are removed, and data normalization is performed to obtain a standardized water quality data set. According to the characteristics of water quality data and prediction targets, machine learning algorithms such as support vector machines, random forests, and gradient boosting trees are selected to construct an initial water quality prediction model. The standardized water quality data set is divided into a training set and a test set. The initial water quality prediction model is trained using the training set data. By adjusting the model parameters and iterative optimization, a water quality prediction model with excellent performance is obtained. The test set data is used to evaluate the water quality prediction model with excellent performance. The cross-validation method is used to calculate the model's prediction accuracy, recall rate, F1 value and other evaluation indicators, and the model is further optimized based on the evaluation results. The optimized water quality prediction model is applied to real-time water quality data. The water quality trend changes in the future period are obtained through model prediction and compared with the preset threshold. If the predicted value exceeds the normal range, the early warning mechanism is triggered.

[0076] When the water quality prediction results are abnormal, the system automatically generates warning information and notifies relevant personnel through SMS, email, etc. At the same time, combined with the specific indicators of abnormal water quality, corresponding emergency disposal suggestions are given. This embodiment also includes regularly retraining and optimizing the water quality prediction model, using the latest water quality historical data to continuously update the model parameters, improve the adaptability and prediction accuracy of the model, and ensure the timeliness and reliability of water quality prediction and abnormal warning.

[0077] This embodiment also provides a multi-source water environment data processing system based on smart IoT, such as Figure 2 ,include:

[0078] The data adaptation module is used to convert the multi-source heterogeneous water environment monitoring data into a unified data model and format using a data adapter method to obtain standardized data;

[0079] The wavelet analysis module is used to perform multi-scale decomposition on the standardized data by using the wavelet analysis method, and to remove high-frequency noise interference by using a preset threshold filtering method to obtain low-dimensional data;

[0080] A distributed storage module is used to adopt a distributed storage architecture to disperse and store the low-dimensional data on a number of nodes, establish a data index and metadata management mechanism, and realize fast retrieval and access of data;

[0081] A stream data processing module is used to process, clean, aggregate and analyze the low-dimensional data using a stream data processing method, extract key indicators and events, and build early warning rules by setting water quality parameter thresholds and change rate thresholds in combination with a decision tree algorithm to determine whether there is water quality anomaly;

[0082] The machine learning prediction module is used to establish a water quality prediction model and predict water quality trends using support vector machine and random forest machine learning methods.

[0083] The above are only preferred specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A multi-source water environment data processing method based on smart Internet of Things, characterized in that: include: The acquired multi-source heterogeneous water environment monitoring data are standardized to obtain standardized data, and the standardized data are multi-scale decomposed to obtain low-dimensional data; The method for obtaining the low-dimensional data is to perform multi-scale decomposition on the standardized data using a wavelet analysis method, specifically including: According to the preset wavelet basis function, the wavelet analysis algorithm is used to perform multi-scale decomposition on the data to obtain wavelet coefficients of different frequency bands, extract data features of different frequency bands, and use the preset threshold filtering method to remove high-frequency noise interference. According to the wavelet coefficients after removing the high-frequency noise, the wavelet reconstruction algorithm is used to inversely transform the wavelet coefficients of each frequency band to obtain the denoised time domain data; for the denoised time domain data, the principal component analysis method is used to perform dimensionality reduction processing to obtain the low-dimensional data; A distributed storage architecture is used for the low-dimensional data, the data is dispersedly stored on several nodes, a data index and metadata management mechanism is established, and data retrieval and access are performed; wherein, for the dispersedly stored low-dimensional data, an inverted index technology is used to establish a data index, the index content includes a data identifier, a storage node identifier and a timestamp, and the data index is stored in an index database; for the dispersedly stored low-dimensional data, data attribute information is extracted, including data volume, data format, acquisition time and acquisition device, metadata is generated, and the metadata is stored in the metadata database; For the low-dimensional data, a stream data processing method is used to clean, aggregate and analyze it, extract key indicators and events, and build water quality abnormality warning rules by setting water quality parameter thresholds and change rate thresholds combined with a decision tree method to determine whether there is water quality abnormality; Acquire historical water quality data, clean and preprocess the data, remove outliers and missing values, and normalize the data to obtain a standardized water quality data set; based on the characteristics of water quality data and prediction objectives, select support vector machine, random forest and gradient boosting tree machine learning methods to build a water quality prediction model and predict water quality trends.

2. The multi-source water environment data processing method based on intelligent Internet of Things according to claim 1 is characterized in that: Obtaining the standardized data includes: For data from different sources and formats, a pre-established data adapter model is used to convert the data into a unified data model and format; The converted data is quality checked to determine whether the data meets the preset data quality standards. If not, the data is cleaned and repaired to obtain the standardized data.

3. The multi-source water environment data processing method based on intelligent Internet of Things according to claim 1 is characterized in that: Removing the high frequency noise interference includes: For the wavelet coefficients obtained by wavelet decomposition, the feature vectors of each frequency band are determined according to the energy distribution characteristics of each frequency band, and a multi-scale feature representation is obtained; According to the feature vectors of each frequency segment, a clustering algorithm is used to perform cluster analysis on the feature vectors, to obtain cluster centers of different frequency segments respectively, and to use different cluster centers as representative features of the corresponding frequency segments; For the representative features of each frequency band, by calculating the difference between different representative features and a preset threshold, it is determined whether the corresponding feature is a high-frequency noise feature. If the difference is greater than the threshold, it is determined to be a noise feature; For the frequency bands determined to be noise characteristics, the wavelet threshold filtering method is used to set the wavelet coefficients of the corresponding frequency bands to zero to remove high-frequency noise.

4. The multi-source water environment data processing method based on intelligent Internet of Things according to claim 1 is characterized in that: Establish the data index and metadata management mechanism to perform data retrieval and access, including: According to a preset distributed storage architecture, the low-dimensional data is dispersedly stored on a plurality of storage nodes; If a data search request is received, the request content is parsed, search conditions are extracted, and the index database is searched according to the search conditions to obtain data identifiers and storage node identifier information that meet the conditions; According to the index retrieval results, the node address of the target data storage is obtained, and the corresponding storage node is accessed through the distributed file system to obtain the target monitoring data; The target monitoring data is aggregated and integrated, and according to the request requirements, a correlation analysis algorithm is used to analyze and mine the monitoring data to generate an analysis result report.

5. The multi-source water environment data processing method based on intelligent Internet of Things according to claim 1 is characterized in that: Combined with the decision tree method, water quality abnormality warning rules are constructed to determine whether there is water quality abnormality, including: Performing data cleaning on the low-dimensional data to remove noise and outliers; Aggregate the cleaned data according to the preset aggregation rules; Use data mining methods to analyze aggregated data and extract key indicators and key events; According to the water quality parameter threshold and change rate threshold, combined with the decision tree model, the water quality abnormality warning rules are constructed; Inputting the extracted key indicators and key events into a decision tree model to determine whether there is water quality abnormality; If the decision tree model determines that there is abnormal water quality, it will trigger real-time warning and intelligent decision-making, and notify relevant departments to take corresponding measures.

6. The multi-source water environment data processing method based on intelligent Internet of Things according to claim 5 is characterized in that: Use data mining methods to analyze the aggregated data and extract key indicators and key events, including: Convert the low-dimensional data into a format suitable for mining algorithm input, and use a data aggregation method to aggregate the format-converted data according to a specific dimension; Input the summarized data into the selected data mining method for training to obtain the mining results; The key indicators and the key events are extracted respectively according to the mining results.

7. The multi-source water environment data processing method based on intelligent Internet of Things according to claim 1 is characterized in that: Constructing the water quality prediction model includes: Dividing the standardized water quality data set into a training set and a test set, using the training set data to train the initial water quality prediction model, and obtaining a trained water quality prediction model by adjusting model parameters and iterative optimization; The trained water quality prediction model is evaluated using the test set data, and the prediction accuracy, recall rate and F1 value evaluation indicators of the model are calculated using a cross-validation method. The model is optimized based on the evaluation results to obtain the water quality prediction model.

8. The multi-source water environment data processing method based on intelligent Internet of Things according to claim 7 is characterized in that: Predict water quality trends, including: The water quality prediction model is applied to real-time water quality data, and a predicted value is obtained through model prediction. The predicted value is compared with a preset threshold value. If the predicted value exceeds the normal range, an early warning mechanism is triggered, and early warning information is automatically generated. Relevant personnel are notified via SMS or email, and corresponding emergency disposal suggestions are given in combination with specific indicators of abnormal water quality.

9. A multi-source water environment data processing system based on intelligent Internet of Things, characterized in that: include: The data adaptation module is used to convert the multi-source heterogeneous water environment monitoring data into a unified data model and format using a data adapter method to obtain standardized data; The wavelet analysis module is used to perform multi-scale decomposition on the standardized data by using the wavelet analysis method, and to remove high-frequency noise interference by using a preset threshold filtering method to obtain low-dimensional data; The method for obtaining the low-dimensional data is to perform multi-scale decomposition on the standardized data using a wavelet analysis method, specifically including: According to the preset wavelet basis function, the wavelet analysis algorithm is used to perform multi-scale decomposition on the data to obtain wavelet coefficients of different frequency bands, extract data features of different frequency bands, and use the preset threshold filtering method to remove high-frequency noise interference. According to the wavelet coefficients after removing the high-frequency noise, the wavelet reconstruction algorithm is used to inversely transform the wavelet coefficients of each frequency band to obtain the denoised time domain data; for the denoised time domain data, the principal component analysis method is used to perform dimensionality reduction processing to obtain the low-dimensional data; A distributed storage module is used to adopt a distributed storage architecture to disperse and store the low-dimensional data on a number of nodes, establish a data index and metadata management mechanism, and realize fast retrieval and access of data; Among them, for the low-dimensional data stored in a decentralized manner, an inverted index technology is used to establish a data index, the index content includes a data identifier, a storage node identifier and a timestamp, and the data index is stored in an index database; for the low-dimensional data stored in a decentralized manner, data attribute information is extracted, including data volume, data format, collection time and collection device, metadata is generated, and the metadata is stored in a metadata database; A stream data processing module is used to process, clean, aggregate and analyze the low-dimensional data using a stream data processing method, extract key indicators and events, and build early warning rules by setting water quality parameter thresholds and change rate thresholds in combination with a decision tree algorithm to determine whether there is water quality anomaly; The machine learning prediction module is used to obtain historical water quality data, clean and preprocess the data, remove outliers and missing values, and normalize the data to obtain a standardized water quality data set; according to the characteristics of water quality data and prediction targets, support vector machine, random forest and gradient boosting tree machine learning methods are selected to build a water quality prediction model and predict water quality trends.

Citation Information

Patent Citations

  • Distributed indexing method and system for large-scale high-dimensional data

    CN108090182A

  • Data distributed storage method and system based on cloud platform

    CN116521087A

  • Sewage treatment effect evaluation method and system based on data analysis

    CN119226976A