Water pollutant monitoring method and system based on big data and storage medium

By rationally setting up monitoring points and data collection cycles within water bodies, and utilizing information entropy calculation and time series models, the problems of information redundancy and omission in water pollution monitoring have been solved, enabling efficient and accurate water quality data monitoring and anomaly prediction, and supporting pollution control and emergency response.

CN120870484AInactive Publication Date: 2025-10-31河南省地质研究院
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510793333.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-10-31
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In existing technologies, water pollution monitoring methods suffer from problems such as unreasonable settings for monitoring points and data collection cycles, leading to information redundancy or omissions, and inaccurate and untimely monitoring.

Method used

By dividing the target water area into multiple monitoring zones, setting reasonable monitoring points and data collection cycles, and using information entropy calculation and time series models to optimize the number of monitoring points and data collection cycles, a first model and a second model are generated to achieve real-time monitoring and anomaly prediction of water quality data.

Benefits of technology

It improves the accuracy and timeliness of water pollutant monitoring, reduces data redundancy and omissions, lowers monitoring costs, ensures timely response in the event of pollution incidents, and provides scientific evidence for pollution control and emergency response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120870484A_ABST
    Figure CN120870484A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of water body monitoring, and discloses a water body pollutant monitoring method and system based on big data and a storage medium. The method comprises the steps of obtaining target data capable of representing the pollution degree based on water quality data, calculating a first value and a second value of a monitoring area based on the target data, and setting a monitoring point for a target water area based on the second value; setting a plurality of different collection periods based on the initial collection period and the maximum collection period, calculating third values of the different collection periods based on the target data, and obtaining an optimal collection period based on the third values; collecting time sequence water quality data based on the optimal collection period, and training to generate a first model and a second model based on the time sequence water quality data; and acquiring real-time water quality data, inputting the real-time water quality data into the first model, inputting the output value of the first model and the real-time water quality data into the second model, and predicting the time point when the abnormality occurs by the second model. According to the invention, the accuracy and timeliness of water pollutant monitoring can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of water body monitoring technology, and in particular to water pollutant monitoring methods, systems and storage media based on big data. Background Technology

[0002] With rapid industrialization and urbanization, water pollution has become increasingly serious. Traditional methods for monitoring water pollutants mainly rely on manual sampling and laboratory analysis, which suffer from limitations such as limited sampling points, low monitoring frequency, and delayed data analysis, making it difficult to reflect the water pollution status in a timely and comprehensive manner. Therefore, developing an efficient, accurate, and real-time method for monitoring water pollutants is of significant practical importance.

[0003] A similar prior art patent application, CN119474157A, provides a control method for a fishery aquatic environment monitoring device, comprising: obtaining a primary monitoring area and a secondary monitoring area; obtaining a primary monitoring device array and a secondary monitoring device array; setting a monitoring bandwidth; obtaining a primary monitoring index set sequence and a secondary monitoring index set sequence; obtaining the water state coefficient of the primary monitoring area and the water state coefficient of the secondary monitoring area; and performing aquatic environment monitoring on the primary and secondary monitoring areas according to the adjusted monitoring bandwidth.

[0004] Similar prior art includes Chinese patent application CN118312738A, which provides a big data-based marine water pollution monitoring system, comprising: a marine water sampling module, a water pollution detection module, a pollution source tracing module, and a zoned collaborative filtering module; the marine water sampling module is used to extract samples of different depths from the marine water within a region; the water pollution detection module is used to detect pollution in the sample data and obtain the degree of pollution; the pollution source tracing module is used to locate the pollution source and the time of pollution occurrence in the polluted water sample; and the zoned collaborative filtering module is used to perform collaborative filtering calculations on other regions based on big data technology to find similar pollution sources and the time of pollution occurrence in other regions.

[0005] However, neither of the above two technical solutions took into account the problem of information redundancy or omission caused by unreasonable settings of monitoring points and water quality data collection cycles, which in turn led to inaccurate and untimely water pollution monitoring.

[0006] Therefore, the present invention provides a method, system and storage medium for monitoring water pollutants based on big data. Summary of the Invention

[0007] This application provides a method, system, and storage medium for monitoring water pollutants based on big data. By setting reasonable monitoring points and water quality data collection cycles, the accuracy and timeliness of water pollutant monitoring can be improved.

[0008] In a first aspect, this application provides a method for monitoring water pollutants based on big data, the method comprising:

[0009] Step S1: Divide the target water area into multiple different monitoring areas, collect water quality data of the monitoring areas, obtain target data that can represent the degree of pollution based on the water quality data, calculate the first value of the monitoring area based on the target data, calculate the second value based on the first value, and set up monitoring points for the target water area based on the second value.

[0010] Step S2: Set the initial collection period, obtain the maximum collection period, set multiple different collection periods based on the initial collection period and the maximum collection period, calculate the third value of different collection periods based on the target data, and obtain the optimal collection period based on the third value.

[0011] Step S3: Collect time-series water quality data of each monitoring point in the target water area based on the optimal collection period, and train and generate the first model and the second model based on the time-series water quality data.

[0012] Step S4: Obtain real-time water quality data, input the real-time water quality data into the first model, obtain the output value of the first model, input the output value of the first model and the real-time water quality data into the second model, and use the second model to predict the time point of the anomaly.

[0013] In conjunction with the first aspect, in the first implementation of the first aspect of this application, calculating the first value of the monitoring area includes:

[0014] The target data for each monitoring area is obtained, the target data of all monitoring areas are added together to obtain the first result value, the ratio of the corresponding target data to the first result value for each monitoring area is calculated to obtain the second result value, the second result value is used as the probability of the pollutant appearing in the corresponding monitoring area, and the second result value of each monitoring area is substituted into the information entropy calculation formula to obtain the corresponding first value.

[0015] All monitoring areas are sorted according to the direction of water flow. For any two adjacent monitoring areas, the first value of the first monitoring area located upstream of the water flow is called the first representative value, and the first value of the second monitoring area located downstream of the water flow is called the second representative value. The second value is calculated using the first formula: Second value = 1 - Second representative value / First representative value.

[0016] In conjunction with the first aspect, in the second implementation of the first aspect of this application, monitoring points are set up for the target water area, including:

[0017] Multiple different data intervals are preset, and a different number of monitoring points are set for each different data interval. For each monitoring area, the data interval to which the corresponding second value belongs is determined. Based on the data interval, the number of corresponding monitoring points is determined, and a corresponding number of monitoring points are set for the corresponding monitoring area.

[0018] For each monitoring area, determine whether it belongs to a specific monitoring area. If it does, add a preset number of monitoring points to the monitoring area based on the original monitoring points.

[0019] In conjunction with the first aspect, in the third implementation of the first aspect of this application, step S2 includes:

[0020] Between the initial collection period and the maximum collection period, several collection periods are set with a preset first time difference. Starting from the time when the target data first reaches the detectable level, a preset first time interval is set, and the corresponding target data is acquired every first time interval. The collection of target data stops at the time point corresponding to the maximum collection period. Based on the collected target data, the third value of each collection period is calculated, and the minimum collection period with the third value greater than the preset first threshold is taken as the optimal collection period.

[0021] In conjunction with the first aspect, in the fourth implementation of the first aspect of this application, the calculation of a third value based on the target data for different collection periods includes:

[0022] For each collection cycle, multiple target data are acquired within the collection cycle. The target data obtained within the collection cycle is divided into several intervals. The number of times the target data appears in each interval is counted. The number of times the target data appears in each interval is divided by the total number of times the target data appears in the corresponding interval. The probability of the target data appearing in the corresponding interval is substituted into the information entropy calculation formula to obtain the third value corresponding to the collection cycle.

[0023] In conjunction with the first aspect, in the fifth implementation of the first aspect of this application, the first model is generated, including:

[0024] Collect historical water quality data, calculate the pollution level corresponding to each historical water quality data point, label each time series water quality data point of the historical water quality data, and divide the labeled time series water quality data into training set and validation set;

[0025] The training set is input into different classification algorithms to train different models. The accuracy of different models is evaluated using the validation set, and the model with the highest accuracy is selected as the first model.

[0026] In conjunction with the first aspect, in the sixth implementation of the first aspect of this application, a second model is generated, including:

[0027] Obtain the first time series data, label the pollution level of the first time series data, and construct the second model based on the first time series data and the corresponding pollution level.

[0028] In conjunction with the first aspect, in the seventh implementation of the first aspect of this application, the pollution level corresponding to each historical water quality data point is calculated, including:

[0029] The system obtains the evaluation value of each data element in the water quality data, presets several pollution levels, obtains the relevant value of each data element for each pollution level, and calculates the pollution level corresponding to each historical water quality data based on the evaluation value and the relevant value.

[0030] Secondly, this application provides a water pollutant monitoring system based on big data, the system comprising:

[0031] The first calculation module is used to divide the target water area into multiple different monitoring areas, collect water quality data of the monitoring areas, obtain target data that can represent the degree of pollution based on the water quality data, calculate the first value of the monitoring area based on the target data, calculate the second value based on the first value, and set monitoring points for the target water area based on the second value.

[0032] The second calculation module is used to set the initial collection period, obtain the maximum collection period, set multiple different collection periods based on the initial collection period and the maximum collection period, calculate the third value of different collection periods based on the target data, and obtain the optimal collection period based on the third value.

[0033] The model generation module is used to collect time-series water quality data from each monitoring point in the target water area based on the optimal collection period, and to train and generate the first and second models based on the time-series water quality data.

[0034] The anomaly monitoring module is used to acquire real-time water quality data. The real-time water quality data is input into the first model to obtain the output value of the first model. The output value of the first model and the real-time water quality data are input into the second model, which predicts the time point when the anomaly will occur.

[0035] A third aspect of this application provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the above-described big data-based water pollutant monitoring method.

[0036] Compared with the prior art, the beneficial effects of the present invention are at least as follows:

[0037] The technical solution provided in this application, by rationally setting monitoring points and optimizing the data collection cycle, can more accurately capture the dynamic changes of pollutants in water bodies, reduce data redundancy and omissions, and thus significantly improve the accuracy of monitoring results. The optimized data collection cycle and monitoring point layout ensure that key data can be obtained in a timely manner when pollution events occur, enabling rapid response to abnormal situations and buying valuable time for pollution control and emergency response. By scientifically calculating and determining the optimal number of monitoring points and data collection cycle, unnecessary waste of monitoring resources is avoided, monitoring costs are reduced, and monitoring efficiency is improved. The first and second models generated by training time-series water quality data can simulate the pollution evolution process of water quality data and predict future trends in pollution levels, providing a scientific basis for taking measures in advance. Attached Figure Description

[0038] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0039] Figure 1 This is a schematic diagram of an embodiment of the water pollutant monitoring method based on big data in this application.

[0040] Figure 2 This is a flowchart illustrating the setting up of monitoring points in the target water area in this application embodiment;

[0041] Figure 3 This is a flowchart illustrating the process of obtaining the optimal collection period in an embodiment of this application;

[0042] Figure 4 This is a schematic diagram of an embodiment of a water pollutant monitoring system based on big data in this application. Detailed Implementation

[0043] This application provides a method, system, and storage medium for monitoring water pollutants based on big data. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0044] For ease of understanding, the specific process of the embodiments of this application is described below. Please refer to [link / reference]. Figure 1 One embodiment of the water pollutant monitoring method based on big data in this application includes:

[0045] Step S1: Divide the target water area into multiple different monitoring areas, collect water quality data of the monitoring areas, obtain target data that can represent the degree of pollution based on the water quality data, calculate the first value of the monitoring area based on the target data, calculate the second value based on the first value, and set up monitoring points for the target water area based on the second value.

[0046] Specifically, when monitoring water pollutants, sensors need to be set up in the target water area to collect water quality-related data. Traditional monitoring point and collection cycle settings rely heavily on experience, which may lead to redundant or missing data.

[0047] To improve the rationality and scientific nature of monitoring points, the target water area is divided into multiple monitoring zones. To obtain suitable monitoring points, the area of ​​each monitoring zone is relatively small. Water quality data is collected from these zones, and target data representing the degree of pollution is obtained based on this data. For example, assuming the target water area is next to a chemical plant, and the plant primarily discharges heavy metal pollutants, with cadmium being the most abundant, cadmium concentration is used as the target data. Based on this target data, a first value is calculated for each monitoring zone. This first value quantifies the uncertainty of the target data across the monitoring zone, i.e., the uniformity of its distribution. A higher first value indicates higher uncertainty and a more uniform distribution of the target data; a lower first value indicates lower uncertainty and a more concentrated distribution of the target data, requiring more monitoring points. A second value is then calculated based on the first value. This second value represents the information transmission efficiency between the monitoring zone and adjacent monitoring zones. A higher second value indicates higher information transmission efficiency, which may reduce the number of monitoring points. Subsequently, monitoring points are set for the target water area based on the second value. By using quantitative calculation methods to set appropriate monitoring points in reasonable locations within the target water area, the accuracy of water pollutant monitoring is improved.

[0048] Step S2: Set the initial collection period, obtain the maximum collection period, set multiple different collection periods based on the initial collection period and the maximum collection period, calculate the third value of different collection periods based on the target data, and obtain the optimal collection period based on the third value.

[0049] Specifically, to set a reasonable collection period, first, set an initial collection period. This initial period is usually short, such as 10, 20, or 30 minutes. Taking an initial period of 10 minutes as an example, then determine the maximum collection period. The maximum collection period is typically set as the time elapsed from when the target data first reaches a detectable level to when it reaches its highest value. Assuming the maximum collection period is 60 minutes, setting multiple different collection periods based on the initial and maximum periods means selecting multiple different collection periods between the initial and maximum periods, such as from 10 minutes to 60 minutes. Multiple different collection periods were selected, namely 15 minutes, 20 minutes, 30 minutes, 40 minutes and 50 minutes. Target data was then collected for each collection period. A third value was calculated based on the target data for each collection period. The third value refers to the information transmission efficiency of the target data from the initial collection period to other collection periods in the same monitoring area. The optimal collection period was then set based on the third value. The larger the third value, the higher the information transmission efficiency from the initial collection period to the corresponding collection period. The optimal collection period was selected based on the third value to ensure that the most representative water quality data information was obtained through a reasonable number of data collections.

[0050] Step S3: Collect time-series water quality data of each monitoring point in the target water area based on the optimal collection period, and train and generate the first model and the second model based on the time-series water quality data.

[0051] Specifically, a first model and a second model are trained based on time-series water quality data. The specific process of generating the first and second models will be explained in detail later. The first model can obtain the pollution level of the water quality data, and the second model can obtain the pollution evolution process of the water quality data based on real-time numerical data and pollution level.

[0052] Step S4: Obtain real-time water quality data, input the real-time water quality data into the first model, obtain the output value of the first model, input the output value of the first model and the real-time water quality data into the second model, and use the second model to predict the time point of the anomaly.

[0053] Specifically, real-time water quality data is acquired and input into a first model to obtain the output value of the first model, which refers to the pollution level of the real-time water quality data. The output value of the first model and the real-time water quality data are then input into a second model, which predicts the time point when the anomaly will occur. The second model can perform pollution evolution analysis on the water quality data based on the pollution level and real-time water quality data to obtain the specific time point when the anomaly occurs (i.e., the pollution level deteriorates significantly), which facilitates relevant personnel to take measures before the pollution worsens and prevents the pollution from expanding further.

[0054] In one specific embodiment, calculating a first value for the monitoring area and calculating a second value based on the first value specifically includes the following steps:

[0055] The target data for each monitoring area is obtained, and the target data of all monitoring areas are added together to obtain the first result value. The ratio of the corresponding target data to the first result value for each monitoring area is calculated to obtain the second result value. The second result value is used as the probability of the pollutant appearing in the corresponding monitoring area. The second result value of each monitoring area is substituted into the information entropy calculation formula to obtain the corresponding first value.

[0056] Specifically, to obtain reasonable monitoring points, it is necessary to first calculate the first value for the monitoring area. To calculate the first value, the target data needs to be transformed into a probability distribution. The target data for each monitoring area at the same time point is obtained, and the target data for all monitoring areas are summed to obtain the first result value. For each monitoring area, the ratio of the corresponding target data to the first result value is calculated to obtain the second result value. The ratio of its corresponding target data to the sum of the target data for all monitoring areas is then calculated. This ratio, i.e., the second result value, is used as the probability of pollutants appearing in the corresponding monitoring area. Substituting the second result value into the information entropy calculation formula yields the corresponding first value H(x). i The formula for calculating information entropy is: Wherein, P(x i ) is the second result value of the i-th monitoring area, and N is the total number of monitoring areas.

[0057] All monitoring areas are sorted according to the direction of water flow. For any two adjacent monitoring areas, the first value of the first monitoring area located upstream of the water flow is called the first representative value, and the first value of the second monitoring area located downstream of the water flow is called the second representative value. The second value is calculated using the first formula: Second value = 1 - Second representative value / First representative value.

[0058] Specifically, in order to calculate the second value, all monitoring areas are sorted according to the direction of water flow. For any two adjacent monitoring areas, the first value of the first monitoring area located upstream of the water flow is taken as the first representative value, and the first value of the second monitoring area located downstream of the water flow is taken as the second representative value. The second value is calculated using the first formula mentioned above. The second value is used as the information transmission efficiency of the first monitoring area in the spatial dimension. Based on the information transmission efficiency, monitoring points are set for the monitoring area to ensure that the monitoring points are set reasonably and scientifically, so as to avoid both data redundancy and data omission.

[0059] In one specific embodiment, setting up monitoring points for the target water area includes the following steps:

[0060] Multiple different data intervals are preset, and a different number of monitoring points are set for each different data interval. For each monitoring area, the data interval to which the corresponding second value belongs is determined, and the number of monitoring points corresponding to the data interval is determined. The corresponding number of monitoring points is set for the corresponding monitoring area.

[0061] For each monitoring area, determine whether it belongs to a specific monitoring area. If it does, add a preset number of monitoring points to the monitoring area based on the original monitoring points.

[0062] Specifically, such as Figure 2 The diagram shows a flowchart for setting up monitoring points in a target water area. A higher second value indicates a greater variation in target data from the upstream monitoring area to the adjacent downstream monitoring area; therefore, more monitoring points should be set up for the corresponding monitoring area. A lower second value indicates less variation in target data from the upstream monitoring area to the adjacent downstream monitoring area; the upstream monitoring area may already contain most of the information, so fewer monitoring points will be set up. Therefore, based on the second value, multiple different data intervals are preset, and a different number of monitoring points are set up for each different data interval. For example, the second value typically ranges from 0 to 1, so five intervals are preset: (0, 0.2], (0.2, 0.4], (0.2, 0.6], (0.6, 0.2], (0.6, 0.2], (0.2, 0.4], (0.2, 0.6 ... [8], (0.8, 0.1], the corresponding number of monitoring points is set to 0.33, 0.5, 1, 2, 4. If the second value of the monitoring area belongs to the interval (0, 0.2], then a monitoring point is set every 3 monitoring areas. If the second value of the monitoring area belongs to the interval (0.2, 0.4], then a monitoring point is set every 2 monitoring areas. If the second value of the monitoring area belongs to the interval (0.4, 0.6], then 1 monitoring point is set for each monitoring area. If the second value of the monitoring area belongs to the interval (0.6, 0.8], then 2 monitoring points are set for each monitoring area. If the second value of the monitoring area belongs to the interval (0.8, 1], then 4 monitoring points are set for each monitoring area.

[0063] It should be noted that for some monitoring areas, if these monitoring areas belong to specific monitoring areas, which refer to environmentally sensitive points such as water sources and ecological protection areas, a preset number of monitoring points will be added to the monitoring area based on the original monitoring points. For example, if the second value of a certain monitoring area belongs to the interval (0, 0.2], and originally only one monitoring point is set every 3 monitoring areas, then two monitoring points will be added to this monitoring area.

[0064] In one specific embodiment, step S2 includes the following steps:

[0065] Between the initial collection period and the maximum collection period, several collection periods are set with a preset first time difference. Starting from the time when the target data first reaches the detectable level, a preset first time interval is set, and the corresponding target data is acquired every first time interval. The collection of target data stops at the time point corresponding to the maximum collection period. Based on the collected target data, the third value of each collection period is calculated, and the minimum collection period with the third value greater than the preset first threshold is taken as the optimal collection period.

[0066] Specifically, such as Figure 3 The flowchart shown illustrates the process of obtaining the optimal collection cycle. To set a reasonable collection cycle that ensures the collected water quality monitoring data is neither redundant nor unrepresentative of changes in water quality, several collection cycles are set between the initial and maximum collection cycles, based on a preset first time difference. For example, if the initial collection cycle is 10 minutes, the maximum collection cycle is 90 minutes, and the first time difference is 20 minutes, then the set collection cycles could be 30 minutes, 50 minutes, or 70 minutes. A preset first time interval is established from the initial collection cycle, for example, 5 minutes. Assuming the first time point at the start of monitoring is 0:00, target data is acquired every 5 minutes until 1:30, resulting in 27 target data acquisitions. Based on the collected target data, the calculation for each collection cycle is performed. The third value represents the uncertainty of the target data distribution within the corresponding collection period. The larger the third value, the higher the uncertainty of the target data distribution, and the shorter the corresponding collection period should be. A first threshold is preset. If the third value is less than or equal to the first threshold, it means that the target data does not change much within the corresponding collection period and the distribution is relatively uniform. Therefore, the collection period can be increased further. If the third threshold is greater than the first threshold, it means that the target data changes significantly within the corresponding collection period. Therefore, the smallest collection period among the collection periods with a third value greater than the first threshold is taken as the optimal collection period. The above method quantifies the uncertainty of the target data between different collection periods, sets the optimal collection period, and ensures that the dynamic changes of pollutants can be captured in a timely and accurate manner when a pollution event occurs.

[0067] In one specific embodiment, calculating a third value for each collection cycle based on multiple collected target data includes the following steps:

[0068] For each collection cycle, multiple target data are acquired within the collection cycle. The target data obtained within the collection cycle is divided into several intervals. The number of times the target data appears in each interval is counted. The number of times the target data appears in each interval is divided by the total number of times the target data appears in the corresponding interval. The probability of the target data appearing in the corresponding interval is substituted into the information entropy calculation formula to obtain the third value corresponding to the collection cycle.

[0069] Specifically, for each collection cycle, such as a 30-minute collection cycle, seven target data points are acquired within the collection cycle, assuming they are {10, 15, 10, 20, 15, 25, 30}. These target data points are divided into several data intervals, assuming three intervals: (0-10], (10-20], and (20-30]. The number of times the target data appears in each interval is counted. The target data {10, 10} appears twice in the (0-10) interval, {15, 20, 15} appears three times in the (10-20) interval, and {25, 30} appears twice in the (20-30) interval. Therefore, the probabilities of the target data appearing in each interval are 2 / 7 = 0.29, 3 / 7 = 0.42, and 2 / 7 = 0.29, respectively. Substituting the probabilities of each interval into the information entropy calculation formula yields the third value H(y) corresponding to the collection cycle. k The formula for calculating information entropy here is: Where P(y) k ) represents the probability of the target data appearing in the k-th interval, and M represents the number of target data included in the current collection period.

[0070] In one specific embodiment, generating the first model includes the following steps:

[0071] Collect historical water quality data, calculate the pollution level corresponding to each historical water quality data point, label each time series water quality data point of the historical water quality data, and divide the labeled time series water quality data into training set and validation set;

[0072] The training set is input into different classification algorithms to train different models. The accuracy of different models is evaluated using the validation set, and the model with the highest accuracy is selected as the first model.

[0073] Specifically, the first model is used to obtain the pollution level of water quality data. To improve the accuracy of the first model, historical water quality data is collected, including polluted water quality data and normal water quality data corresponding to historical water pollution events. The pollution level corresponding to each historical water quality data is calculated. Labeling each time series water quality data in the historical water quality data means labeling the corresponding pollution level for each time series water quality data. The labeled time series water quality data is divided into training set and validation set. The training set is input into different classification methods, such as support vector machine and random forest, to obtain multiple different classification models. For each classification model, the accuracy of different models is evaluated using the validation set, and the model with the highest accuracy is selected as the first model.

[0074] In one specific embodiment, generating the second model includes the following steps:

[0075] Obtain the first time series data, label the pollution level of the first time series data, and construct the second model based on the first time series data and the corresponding pollution level.

[0076] Specifically, the second model can simulate the pollution evolution process of water quality data based on water quality data and corresponding pollution levels, and predict the future trend of pollution level changes. In order to build the second model, first time series data is obtained. The first time series data refers to the time series data corresponding to the occurrence of pollution events. The pollution level is labeled for each time series data in the first data series data. The first time series data and the corresponding pollution level are used as learning data to build the second model.

[0077] In one specific embodiment, the pollution level corresponding to each historical water quality data point is calculated, which specifically includes the following steps:

[0078] The system obtains the evaluation value of each data element in the water quality data, presets several pollution levels, obtains the relevant value of each data element for each pollution level, and calculates the pollution level corresponding to each historical water quality data based on the evaluation value and the relevant value.

[0079] Specifically, the assessment value refers to the expert's evaluation of each data element relative to other data elements. For example, the assessment value of the data element pH relative to pH itself is 1, and the assessment value of the data element pH relative to the data element temperature is 2, indicating that the data element pH is more important than temperature in assessing the degree of pollution. The correlation value refers to the probability of judging the current pollution level when there is only one data element. Assuming there are four pollution levels, namely 1, 2, 3, and 4, representing no pollution, light pollution, moderate pollution, and heavy pollution, respectively, when the cadmium concentration is 0.001, the probability of judging the pollution level as 1 is 0.9, and the probability of judging the pollution level as 4 is 0.1. The pollution level corresponding to each historical water quality data is calculated based on the assessment value and the correlation value.

[0080] In one specific embodiment, the pollution level corresponding to each historical water quality data point is calculated based on the assessment value and the correlation value, specifically including the following steps:

[0081] The degree of pollution is calculated using the second formula: PL=(r1,r2,...r o (v1,v2,...,v) o ), where r j =(s 1j ,s 2j ,...,s zj ) T , where v o It is the evaluation value corresponding to the o-th data element, s 1jLet $\frac{j}{j}$ be the correlation value of the j-th data element for each of the $z$ different contamination levels, $o$ be the number of data elements, and $z$ be the number of contamination levels. Using the second formula, we obtain $o$ result values ​​{c1, c2, ..., c$ for each data element corresponding to the $z$ contamination levels. m The pollution level corresponding to the maximum value among the o results is taken as the pollution level of the corresponding water quality data.

[0082] Specifically, in order to quickly calculate the pollution level of water quality data, based on the evaluation value of each data element and the evaluation value set for each pollution level for the data element, the result value corresponding to each water quality data is calculated using the second formula mentioned above. The pollution level corresponding to the maximum value is taken as the pollution level of the water quality data. The above method combines the evaluation value and the relevant value to quantify the pollution level and provides a scientific calculation method for calculating the pollution level.

[0083] The above describes the water pollutant monitoring method based on big data in the embodiments of this application. The following describes the water pollutant monitoring system based on big data in the embodiments of this application. Please refer to [link / reference]. Figure 4 One embodiment of the water pollutant monitoring system based on big data in this application includes:

[0084] The first calculation module is used to divide the target water area into multiple different monitoring areas, collect water quality data of the monitoring areas, obtain target data that can represent the degree of pollution based on the water quality data, calculate the first value of the monitoring area based on the target data, calculate the second value based on the first value, and set monitoring points for the target water area based on the second value.

[0085] The second calculation module is used to set the initial collection period, obtain the maximum collection period, set multiple different collection periods based on the initial collection period and the maximum collection period, calculate the third value of different collection periods based on the target data, and obtain the optimal collection period based on the third value.

[0086] The model generation module is used to collect time-series water quality data from each monitoring point in the target water area based on the optimal collection period, and to train and generate the first and second models based on the time-series water quality data.

[0087] The anomaly monitoring module is used to acquire real-time water quality data. The real-time water quality data is input into the first model to obtain the output value of the first model. The output value of the first model and the real-time water quality data are input into the second model, which predicts the time point when the anomaly will occur.

[0088] This application also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the steps of a big data-based water pollutant monitoring method.

[0089] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0090] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0091] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A water pollutant monitoring method based on big data, characterized in that, The method includes: Step S1: Divide the target water area into multiple different monitoring areas, collect water quality data of the monitoring areas, obtain target data that can represent the degree of pollution based on the water quality data, calculate the first value of the monitoring area based on the target data, calculate the second value based on the first value, and set up monitoring points for the target water area based on the second value. Step S2: Set the initial collection period, obtain the maximum collection period, set multiple different collection periods based on the initial collection period and the maximum collection period, calculate the third value of different collection periods based on the target data, and obtain the optimal collection period based on the third value. Step S3: Collect time-series water quality data of each monitoring point in the target water area based on the optimal collection period, and train and generate the first model and the second model based on the time-series water quality data. Step S4: Obtain real-time water quality data, input the real-time water quality data into the first model, obtain the output value of the first model, input the output value of the first model and the real-time water quality data into the second model, and use the second model to predict the time point of the anomaly.

2. The method according to claim 1, characterized in that, Calculate a first value for the monitored area, and then calculate a second value based on the first value, including: The target data for each monitoring area is obtained, the target data of all monitoring areas are added together to obtain the first result value, the ratio of the corresponding target data to the first result value for each monitoring area is calculated to obtain the second result value, the second result value is used as the probability of the pollutant appearing in the corresponding monitoring area, and the second result value of each monitoring area is substituted into the information entropy calculation formula to obtain the corresponding first value. All monitoring areas are sorted according to the direction of water flow. For any two adjacent monitoring areas, the first value of the first monitoring area located upstream of the water flow is called the first representative value, and the first value of the second monitoring area located downstream of the water flow is called the second representative value. The second value is calculated using the first formula: Second value = 1 - Second representative value / First representative value.

3. The method according to claim 1, characterized in that, Establish monitoring points in the target water area, including: Multiple different data intervals are preset, and a different number of monitoring points are set for each different data interval. For each monitoring area, the data interval to which the corresponding second value belongs is determined. Based on the data interval, the number of corresponding monitoring points is determined, and a corresponding number of monitoring points are set for the corresponding monitoring area. For each monitoring area, determine whether it belongs to a specific monitoring area. If it does, add a preset number of monitoring points to the monitoring area based on the original monitoring points.

4. The method according to claim 1, characterized in that, Step S2 includes: Between the initial collection period and the maximum collection period, several collection periods are set with a preset first time difference. Starting from the time when the target data first reaches the detectable level, a preset first time interval is set, and the corresponding target data is acquired every first time interval. The collection of target data stops at the time point corresponding to the maximum collection period. Based on the collected target data, the third value of each collection period is calculated, and the minimum collection period with the third value greater than the preset first threshold is taken as the optimal collection period.

5. The method according to claim 1, characterized in that, The third value is calculated based on the target data for different collection periods, including: For each collection cycle, multiple target data are acquired within the collection cycle. The target data obtained within the collection cycle is divided into several intervals. The number of times the target data appears in each interval is counted. The number of times the target data appears in each interval is divided by the total number of times the target data appears in the corresponding interval. The probability of the target data appearing in the corresponding interval is substituted into the information entropy calculation formula to obtain the third value corresponding to the collection cycle.

6. The method according to claim 1, characterized in that, Generate the first model, including: Collect historical water quality data, calculate the pollution level corresponding to each historical water quality data point, label each time series water quality data point of the historical water quality data, and divide the labeled time series water quality data into training set and validation set; The training set is input into different classification algorithms to train different models. The accuracy of different models is evaluated using the validation set, and the model with the highest accuracy is selected as the first model.

7. The method according to claim 1, characterized in that, Generate a second model, including: Obtain the first time series data, label the pollution level of the first time series data, and construct the second model based on the first time series data and the corresponding pollution level.

8. The method according to claim 1, characterized in that, Calculate the pollution level corresponding to each historical water quality data point, including: The system obtains the evaluation value of each data element in the water quality data, presets several pollution levels, obtains the relevant value of each data element for each pollution level, and calculates the pollution level corresponding to each historical water quality data based on the evaluation value and the relevant value.

9. A water pollutant monitoring system based on big data, used to implement the water pollutant monitoring method based on big data as described in any one of claims 1-8, characterized in that, The system includes: The first calculation module is used to divide the target water area into multiple different monitoring areas, collect water quality data of the monitoring areas, obtain target data that can represent the degree of pollution based on the water quality data, calculate the first value of the monitoring area based on the target data, calculate the second value based on the first value, and set monitoring points for the target water area based on the second value. The second calculation module is used to set the initial collection period, obtain the maximum collection period, set multiple different collection periods based on the initial collection period and the maximum collection period, calculate the third value of different collection periods based on the target data, and obtain the optimal collection period based on the third value. The model generation module is used to collect time-series water quality data from each monitoring point in the target water area based on the optimal collection period, and to train and generate the first and second models based on the time-series water quality data. The anomaly monitoring module is used to acquire real-time water quality data. The real-time water quality data is input into the first model to obtain the output value of the first model. The output value of the first model and the real-time water quality data are input into the second model, which predicts the time point when the anomaly will occur.

10. A computer-readable storage medium storing instructions thereon, characterized in that, When the instructions are executed by the processor, they implement the water pollutant monitoring method based on big data as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Marine water body pollution monitoring system based on big data

    CN118312738A

  • Fishery water body environment monitoring equipment control method

    CN119474157A