Regional watershed sudden environmental event early warning method based on big data
By using big data-based methods and leveraging negative binomial distribution and multi-source monitoring data, a gridded early warning threshold matrix was established, overcoming the limitations of traditional early warning methods and enabling precise early warning and dynamic response to watershed environmental events, thereby improving the timeliness and accuracy of early warnings.
Patent Information
- Application Number
- CN202511127404.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-11
AI Technical Summary
Traditional early warning methods rely on data from a single monitoring station, which cannot fully reflect the environmental conditions of the watershed. Static thresholds cannot adapt to the spatiotemporal changes of pollution events, resulting in delayed early warnings and affecting the effectiveness of decision-making and response measures.
Based on big data methods, historical pollution event data of the target watershed is acquired, statistically analyzed by spatial grid and time segmentation, and a gridded early warning threshold matrix is established using negative binomial distribution. Multi-source monitoring data is collected in real time for comparison to trigger graded early warnings. The early warning thresholds are optimized through confidence verification and dynamic parameter updates.
It achieves accurate spatiotemporal risk characterization, rapidly triggers tiered early warnings, reduces false alarm rates, adapts to changes in watershed pollution patterns, and provides scientific and dynamic decision support.
Smart Images

Figure CN120932423A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and more specifically, to a method for early warning of sudden environmental events in regional watersheds based on big data. Background Technology
[0002] In the field of regional watershed environmental management, water pollution and chemical spills are particularly prominent. They often break out suddenly without warning, and the scope of harm can quickly spread from a local area to the entire watershed. Moreover, the resulting ecological damage and health harm have far-reaching and lasting effects, and constantly and seriously threaten the ecological security of the region and the health of residents. Therefore, it is necessary to issue early warnings for sudden environmental incidents.
[0003] Traditional early warning methods have several limitations. Firstly, their data sources are often limited, relying heavily on data from single monitoring stations, resulting in a limited monitoring scope and an inability to comprehensively and accurately reflect the overall environmental condition of the watershed. They also fail to capture the heterogeneity of localized pollution risks arising from differences in hydrological and geographical characteristics within the watershed. Secondly, traditional early warning systems employ static threshold models. Watershed pollution patterns are significantly influenced by seasonal hydrological variations and human activities; the probability and characteristics of pollution events constantly change over time and space. Static thresholds cannot promptly reflect these spatiotemporal evolutions, leading to delayed early warnings and impacting the effectiveness of environmental management decisions and response measures.
[0004] No effective solutions have yet been proposed to address the problems in the relevant technologies. Summary of the Invention
[0005] In view of this, the present invention provides a method for early warning of sudden environmental events in regional watersheds based on big data, in order to solve the problems mentioned above.
[0006] To solve the above problems, the specific technical solution adopted by the present invention is as follows:
[0007] A method for early warning of sudden environmental events in regional watersheds based on big data includes the following steps:
[0008] S1. Obtain historical pollution event data for the target watershed and perform pollution event frequency statistics according to spatial grids and time segments; based on the statistical results, establish a gridded early warning threshold matrix using negative binomial distribution;
[0009] S2. Collect multi-source monitoring data of the target watershed in real time, and count the frequency of new pollution events according to spatiotemporal grids; compare the statistical results with the warning threshold matrix of the corresponding grid in real time, and trigger a graded warning if the frequency of new pollution events exceeds the warning threshold.
[0010] S3. Verify the confidence level of newly occurring events that have not triggered a warning, and dynamically recalculate the negative binomial distribution parameters of the verified data;
[0011] S4. Based on the recalculated negative binomial distribution parameters, optimize the gridded early warning threshold matrix and return to step S2 to achieve iterative monitoring of sudden environmental events.
[0012] Preferably, the steps of acquiring historical pollution event data for the target watershed and statistically analyzing the frequency of pollution events by spatial grid and time segmentation, and establishing a gridded early warning threshold matrix based on the statistical results using a negative binomial distribution, include the following steps:
[0013] S11. Collect historical pollution event data of the target watershed, and perform data cleaning and standardization on the historical pollution event data;
[0014] S12. Based on hydrogeographic characteristics, the target watershed is divided into spatial grids, and time segments are divided according to seasonal cycles to obtain several spatiotemporal units.
[0015] S13. Statistically analyze the frequency of historical pollution events and safe operating time for each spatiotemporal unit, and use the maximum likelihood estimation method to calculate the probability of events occurring in each spatiotemporal unit.
[0016] S14. Based on the event occurrence probability of each spatiotemporal unit, calculate the early warning threshold through the quantile function of the negative binomial distribution, and establish a gridded early warning threshold matrix.
[0017] Preferably, the step of dividing the target watershed into spatial grids based on hydrogeographic features and dividing the time into segments according to seasonal cycles to obtain several spatiotemporal units includes the following steps:
[0018] S121. Obtain the geospatial basic data of the target watershed and divide the target watershed into grids according to the preset spatial grid;
[0019] S122. Perform hydrological consistency verification on the divided spatial grids, and remove grids that fail the hydrological consistency verification as invalid grids to obtain valid grids.
[0020] S123. Divide the time period according to the preset seasonal cycle, and map the data of each historical pollution event to the corresponding time period according to the occurrence time, and form a spatiotemporal unit by combining the effective grid.
[0021] Preferably, the step of statistically analyzing the historical pollution event frequency and safe operating time of each spatiotemporal unit, and calculating the event probability of each spatiotemporal unit using the maximum likelihood estimation method includes the following steps:
[0022] S131. For each spatiotemporal unit, count the total number of historical pollution events, and based on the total number of monitoring days in that spatiotemporal unit, calculate the number of days in which no pollution events occurred as the safe operating time.
[0023] S132. Based on the total number of historical pollution events and the safe operating time, the probability of pollution events occurring within the spatiotemporal unit is calculated using the maximum likelihood estimation method.
[0024] Preferably, the step of calculating the probability of a pollution event occurring within the spatiotemporal unit using the maximum likelihood estimation method based on the total number of historical pollution events and the safe operating time includes the following steps:
[0025] S1321. Based on the total number of historical pollution events and the safe operating time, construct a likelihood function with the probability of pollution events occurring as the unknown.
[0026] S1322. Take the logarithm of the likelihood function to obtain the log-likelihood function, and differentiate the probability of pollution event occurrence with respect to the log-likelihood function to construct the maximum likelihood estimation equation.
[0027] S1323. Solve the maximum likelihood estimation equation to obtain the probability of pollution events occurring within the spatiotemporal unit.
[0028] Preferably, the step of calculating the early warning threshold based on the event occurrence probability of each spatiotemporal unit using the quantile function of the negative binomial distribution and establishing a gridded early warning threshold matrix includes the following steps:
[0029] S141. Based on the event occurrence probability of each spatiotemporal unit, calculate the quantiles at different confidence levels using the quantile function of the negative binomial distribution.
[0030] S142. Based on the confidence level of the early warning threshold, construct multi-level early warning rules for each spatiotemporal unit;
[0031] S143. Integrate the warning thresholds of all spatiotemporal units and summarize them according to grid number and time period to form a structured gridded warning threshold matrix.
[0032] Preferably, the step of calculating the warning threshold using the quantile function of the negative binomial distribution based on the event occurrence probability of each spatiotemporal unit includes the following steps:
[0033] S1411. Set the success probability parameter of the negative binomial distribution based on the event occurrence probability of the spatiotemporal unit, and determine the success number threshold in combination with the early warning requirements.
[0034] S1412. Based on prior knowledge, estimate the shape and scale parameters of the negative binomial distribution; and call the quantile function of the negative binomial distribution at multiple confidence levels to calculate the corresponding event frequency threshold.
[0035] S1413. Use the quantile results of each spatiotemporal unit at different confidence levels as the warning threshold, and mark the confidence level of the warning threshold.
[0036] Preferably, the real-time acquisition of multi-source monitoring data of the target watershed, the statistical counting of new pollution events according to spatiotemporal grids, and the real-time comparison of the statistical results with the warning threshold matrix of the corresponding grid, triggering a graded warning if the frequency of new pollution events exceeds the warning threshold, includes the following steps:
[0037] S21. Collect multi-source monitoring data of the target watershed in real time, and map the multi-source monitoring data to the corresponding spatial grid according to coordinates and monitoring time;
[0038] S22. Count the frequency of new pollution events in each spatiotemporal grid according to the time period, and generate a dynamic pollution event frequency matrix;
[0039] S23. The dynamic pollution event frequency matrix is compared with the corresponding grid's early warning threshold matrix in real time. If the pollution event frequency exceeds the corresponding level threshold, the graded early warning response mechanism is triggered.
[0040] Preferably, the step of comparing the dynamic pollution event frequency matrix with the corresponding grid's early warning threshold matrix in real time, and triggering a graded early warning response mechanism if the pollution event frequency exceeds the corresponding level threshold, includes the following steps:
[0041] S231. Extract the statistical frequency of each grid within the current time window from the dynamic pollution event frequency matrix, and load the warning threshold matrix of the corresponding grid.
[0042] S232. Compare the frequency of pollution events with the thresholds of each level in the early warning threshold matrix in real time, and determine the level of exceeding the limit;
[0043] S233. Determine the corresponding graded early warning response mechanism based on the level of exceeding the limit.
[0044] Preferably, the step of verifying the confidence level of newly occurring events that have not triggered an early warning, and dynamically recalculating the negative binomial distribution parameters of the verified data, includes the following steps:
[0045] S31. Extract data on spatiotemporal units that have not triggered warnings and their newly emerging pollution events, and conduct preliminary validity screening of the data;
[0046] S32. Based on multi-dimensional verification rules, perform confidence verification on newly emerging events after preliminary screening, and save the newly emerging events that pass the confidence verification.
[0047] S33. Dynamically update the negative binomial distribution parameters of the corresponding spatiotemporal unit using monitoring data of newly occurring events verified by confidence level.
[0048] The beneficial effects of this invention are as follows:
[0049] This invention achieves precise spatiotemporal risk characterization by dividing spatiotemporal units according to hydrogeographic features and constructing a gridded early warning threshold matrix using historical data. By collecting multi-source data in real time and dynamically comparing the threshold matrix, a tiered early warning response mechanism can be quickly triggered, improving the timeliness of event handling. Furthermore, by introducing a confidence verification mechanism to control the quality of newly occurring events that have not triggered early warnings, it ensures that only high-confidence data participates in the dynamic updating of the negative binomial distribution parameters. This avoids false alarms interfering with accuracy, continuously adapts to changes in watershed pollution patterns, and reduces the false alarm rate while ensuring environmental safety. This provides scientific, dynamic, and reliable decision support for regional watershed environmental risk management. Attached Figure Description
[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly described below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:
[0051] Figure 1 This is a flowchart of a regional watershed emergency environmental event early warning method based on big data according to an embodiment of the present invention. Detailed Implementation
[0052] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.
[0053] According to an embodiment of the present invention, a method for early warning of sudden environmental events in regional watersheds based on big data is provided.
[0054] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments, such as... Figure 1 As shown, the regional watershed emergency environmental event early warning method based on big data according to an embodiment of the present invention includes the following steps:
[0055] S1. Obtain historical pollution event data for the target watershed and perform pollution event frequency statistics according to spatial grids and time segments; based on the statistical results, establish a gridded early warning threshold matrix using negative binomial distribution;
[0056] As a preferred embodiment, the steps of acquiring historical pollution event data of the target watershed and statistically analyzing the frequency of pollution events by spatial grid and time segmentation, and establishing a gridded early warning threshold matrix based on the statistical results using a negative binomial distribution, include the following steps:
[0057] S11. Collect historical pollution event data of the target watershed, and perform data cleaning and standardization on the historical pollution event data;
[0058] It should be noted that historical pollution event data for the target watershed can be obtained through environmental protection department monitoring records, enterprise pollution discharge reports, news media reports, historical disaster records, etc. Data types include the time and location (latitude and longitude coordinates) of the pollution event, the type and concentration of pollutants, the duration, and the scope of impact.
[0059] Data cleaning includes removing duplicate records, filling in missing values, correcting errors, and handling outliers. Data standardization includes format standardization, content standardization, and numerical standardization.
[0060] S12. Based on hydrogeographic characteristics, the target watershed is divided into spatial grids, and time segments are divided according to seasonal cycles to obtain several spatiotemporal units.
[0061] As a preferred embodiment, the step of dividing the target watershed into spatial grids based on hydrogeographic features and dividing the time into segments according to seasonal cycles to obtain several spatiotemporal units includes the following steps:
[0062] S121. Obtain the geospatial basic data of the target watershed and divide the target watershed into grids according to the preset spatial grid;
[0063] It should be noted that geospatial basic data includes digital elevation models (DEM), water system distribution maps, administrative boundaries, land use types, soil types, distribution of industrial tailings ponds, distribution of highways, distribution of oil and gas pipelines, etc.
[0064] Furthermore, during the classification process, tailings ponds of industrial enterprises, high-frequency road sections for hazardous chemical transport vehicles, and areas where oil and gas pipelines are distributed are designated as high-risk areas. A densified grid division strategy is then adopted: based on the density of tailings ponds, the frequency of hazardous chemical transport routes, and the distribution density (or total kilometers) of oil and gas pipelines, the grid granularity is refined in high-risk areas, giving them higher spatiotemporal resolution than ordinary areas. This enables more refined modeling and risk assessment of key areas, improving the ability to identify and respond to sudden environmental events.
[0065] S122. Perform hydrological consistency verification on the divided spatial grids, and remove grids that fail the hydrological consistency verification as invalid grids to obtain valid grids.
[0066] Among them, the hydrological consistency check is to determine whether the grid has hydrological significance, that is, whether the geographical features (such as topography, soil and vegetation) within the grid are similar, and whether the grid conforms to the hydrological characteristics of the watershed (such as confluence relationships). Grids with no water area at all or grids with too small an area (such as less than 0.1% of the total area) are removed.
[0067] S123. Divide the time period according to the preset seasonal cycle, and map the data of each historical pollution event to the corresponding time period according to the occurrence time, and form a spatiotemporal unit by combining the effective grid.
[0068] Furthermore, when dividing time periods, peak electricity consumption periods for industrial production or residential life, peak traffic periods, and periods of frequent extreme weather (such as typhoon season and periods of high seismic activity) can be designated as special time periods. These special time periods can be further subdivided using a more intensive division method to improve the sensitivity and resolution of spatiotemporal unit division. This helps to more accurately capture the high incidence of pollution events.
[0069] S13. Statistically analyze the frequency of historical pollution events and safe operating time for each spatiotemporal unit, and use the maximum likelihood estimation method to calculate the probability of events occurring in each spatiotemporal unit.
[0070] As a preferred embodiment, the step of statistically analyzing the frequency of historical pollution events and safe operating time for each spatiotemporal unit, and calculating the probability of event occurrence for each spatiotemporal unit using the maximum likelihood estimation method, includes the following steps:
[0071] S131. For each spatiotemporal unit, count the total number of historical pollution events, and based on the total number of monitoring days in that spatiotemporal unit, calculate the number of days in which no pollution events occurred as the safe operating time.
[0072] S132. Based on the total number of historical pollution events and the safe operating time, the probability of pollution events occurring within the spatiotemporal unit is calculated using the maximum likelihood estimation method.
[0073] As a preferred embodiment, the step of calculating the probability of a pollution event occurring within the spatiotemporal unit using the maximum likelihood estimation method based on the total number of historical pollution events and the safe operating time includes the following steps:
[0074] S1321. Based on the total number of historical pollution events and the safe operating time, construct a likelihood function with the probability of pollution events occurring as the unknown.
[0075] It should be noted that when constructing the likelihood function with the probability of a pollution event as the unknown, we need to assume that within a certain spatiotemporal unit, the total number of days in the observation period is n, the number of days with pollution events is k, and the number of days without events is nk; let the probability of a pollution event be p; and treat each day as an independent Bernoulli trial (whether an event occurs or not), for a total of n trials; then the number of pollution events follows a binomial distribution B(n,p), and the corresponding likelihood function is:
[0076]
[0077] In the formula, p is an unknown, and n and k are known statistical values.
[0078] S1322. Take the logarithm of the likelihood function to obtain the log-likelihood function, and differentiate the probability of pollution event occurrence with respect to the log-likelihood function to construct the maximum likelihood estimation equation.
[0079] S1323. Solve the maximum likelihood estimation equation to obtain the probability of pollution events occurring within the spatiotemporal unit.
[0080] S14. Based on the event occurrence probability of each spatiotemporal unit, calculate the early warning threshold through the quantile function of the negative binomial distribution, and establish a gridded early warning threshold matrix.
[0081] In a preferred embodiment, the step of calculating the early warning threshold based on the event occurrence probability of each spatiotemporal unit using the quantile function of the negative binomial distribution and establishing a gridded early warning threshold matrix includes the following steps:
[0082] S141. Based on the event occurrence probability of each spatiotemporal unit, calculate the quantiles at different confidence levels using the quantile function of the negative binomial distribution.
[0083] In a preferred embodiment, the step of calculating the warning threshold using the quantile function of the negative binomial distribution based on the event occurrence probability of each spatiotemporal unit includes the following steps:
[0084] S1411. Set the success probability parameter of the negative binomial distribution based on the event occurrence probability of the spatiotemporal unit, and determine the success number threshold in combination with the early warning requirements.
[0085] It should be noted that the negative binomial distribution describes "the distribution of the number of failures before reaching the r-th success in independent trials with a success probability of p". In pollution incident early warning, it can be defined as:
[0086] A successful event is the occurrence of a contamination event (probability p); a failed event is the absence of contamination (probability 1-p).
[0087] Therefore, we calculate the threshold for the total number of events (success + failure) that may occur before reaching r pollution events.
[0088] S1412. Based on prior knowledge, estimate the shape and scale parameters of the negative binomial distribution; and call the quantile function of the negative binomial distribution at multiple confidence levels to calculate the corresponding event frequency threshold.
[0089] Specifically, the shape parameter of the negative binomial distribution is the success rate threshold r, and the scale parameter is the success probability p. If historical data is insufficient, it can be set using prior knowledge: for example, if the long-term pollution probability of a certain area is stable at p = 0.01, then this value can be directly used as the scale parameter. The shape parameter r is determined by the early warning requirements (e.g., r = 1 or r = 3).
[0090] Then, statistical tools are used to calculate quantile thresholds at different confidence levels: for example, at a 95% confidence level, a threshold for "the total number of events before r pollution events" is calculated, such that the probability of the actual number of events exceeding this threshold is only 5%. By adjusting the confidence levels (such as 90%, 95%, 99%), multi-level thresholds can be generated to balance the risks of false alarms and false negatives.
[0091] S1413. Use the quantile results of each spatiotemporal unit at different confidence levels as the warning threshold, and mark the confidence level of the warning threshold.
[0092] Specifically, if the threshold for a certain area at a 90% confidence level is 14 times / day, at 95% it is 29 times / day, and at 99% it is 99 times / day, then the confidence level should be labeled according to the requirements:
[0093] Level 1 warning (high risk) is set to the threshold of ≥99% of events (immediate action required);
[0094] Level 2 warning (medium risk) is set at 95% threshold ≤ number of events < 99% threshold (enhanced monitoring);
[0095] Level 3 early warning (low risk) is set to 90% threshold ≤ number of events < 95% threshold (monitoring trends).
[0096] S142. Based on the confidence level of the early warning threshold, construct multi-level early warning rules for each spatiotemporal unit;
[0097] S143. Integrate the warning thresholds of all spatiotemporal units and summarize them according to grid number and time period to form a structured gridded warning threshold matrix.
[0098] S2. Collect multi-source monitoring data of the target watershed in real time, and count the frequency of new pollution events according to spatiotemporal grids; compare the statistical results with the warning threshold matrix of the corresponding grid in real time, and trigger a graded warning if the frequency of new pollution events exceeds the warning threshold.
[0099] As a preferred implementation, the real-time acquisition of multi-source monitoring data of the target watershed, the statistical counting of new pollution events according to spatiotemporal grids, and the real-time comparison of the statistical results with the warning threshold matrix of the corresponding grid, triggering a graded warning if the frequency of new pollution events exceeds the warning threshold, includes the following steps:
[0100] S21. Collect multi-source monitoring data of the target watershed in real time, and map the multi-source monitoring data to the corresponding spatial grid according to coordinates and monitoring time;
[0101] Specifically, when collecting multi-source monitoring data of a target watershed, multi-source data of the target watershed can be obtained in real time through IoT sensors (such as water quality monitoring buoys, drone inspections, satellite remote sensing, etc.), including but not limited to pollution indicators such as chemical oxygen demand (COD), ammonia nitrogen concentration, pH value, dissolved oxygen, and flow velocity, as well as sensor location coordinates (latitude and longitude) and collection timestamps.
[0102] S22. Count the frequency of new pollution events in each spatiotemporal grid according to the time period, and generate a dynamic pollution event frequency matrix;
[0103] S23. The dynamic pollution event frequency matrix is compared with the corresponding grid's early warning threshold matrix in real time. If the pollution event frequency exceeds the corresponding level threshold, the graded early warning response mechanism is triggered.
[0104] In a preferred embodiment, the step of comparing the dynamic pollution event frequency matrix with the corresponding grid's early warning threshold matrix in real time, and triggering a graded early warning response mechanism if the pollution event frequency exceeds the corresponding level threshold, includes the following steps:
[0105] S231. Extract the statistical frequency of each grid within the current time window from the dynamic pollution event frequency matrix, and load the warning threshold matrix of the corresponding grid.
[0106] S232. Compare the frequency of pollution events with the thresholds of each level in the early warning threshold matrix in real time, and determine the level of exceeding the limit;
[0107] S233. Determine the corresponding graded early warning response mechanism based on the level of exceeding the limit.
[0108] It should be noted that the tiered early warning response mechanism specifically includes: Level 1 early warning: immediately triggering an emergency response (such as closing the sewage outlet and activating emergency treatment facilities), and notifying the environmental protection department, enterprise leaders, and downstream residents; Level 2 early warning: increasing the monitoring frequency (such as changing from every hour to every 15 minutes), and dispatching patrol personnel to conduct on-site verification; Level 3 early warning: automatically recording and generating trend reports for subsequent analysis.
[0109] S3. Verify the confidence level of newly occurring events that have not triggered a warning, and dynamically recalculate the negative binomial distribution parameters of the verified data;
[0110] In a preferred embodiment, the confidence verification of newly occurring events that have not triggered an early warning, and the dynamic recalculation of the negative binomial distribution parameters of the verified data, includes the following steps:
[0111] S31. Extract data on spatiotemporal units that have not triggered warnings and their newly emerging pollution events, and conduct preliminary validity screening of the data;
[0112] It should be noted that for spatiotemporal units that have not triggered warnings, data on newly occurring pollution events in that unit are extracted, including the event occurrence time, pollution index values (such as COD concentration), and sensor locations. Then, data with abnormal monitoring times, data located outside the grid boundaries, or data with drifting sensor coordinates are removed to obtain valid data.
[0113] S32. Based on multi-dimensional verification rules, perform confidence verification on newly emerging events after preliminary screening, and save the newly emerging events that pass the confidence verification.
[0114] It should be noted that confidence verification includes spatiotemporal consistency verification, indicator correlation verification, and so on. Spatiotemporal consistency verification checks whether the event matches the known distribution of pollution sources or historically high-incidence areas; indicator correlation verification verifies the logical relationship between pollution indicators. For example, high COD is usually accompanied by low dissolved oxygen. If an event shows that COD exceeds the standard but dissolved oxygen is normal, the confidence level is low.
[0115] S33. Dynamically update the negative binomial distribution parameters of the corresponding spatiotemporal unit using monitoring data of newly occurring events verified by confidence level.
[0116] Specifically, dynamically updating the negative binomial distribution parameters of the corresponding spatiotemporal units using monitoring data of newly occurring events verified by confidence means using these reliable data as the latest observation samples, combining them with existing historical event statistics, and recalculating key parameters such as the success probability and number of successes of pollution events through maximum likelihood estimation or Bayesian update methods. This achieves adaptive optimization of the negative binomial distribution, enabling the early warning threshold system to continuously evolve with spatiotemporal changes and environmental fluctuations, thereby improving the sensitivity and accuracy of response to emergencies.
[0117] S4. Based on the recalculated negative binomial distribution parameters, optimize the gridded early warning threshold matrix and return to step S2 to achieve iterative monitoring of sudden environmental events.
[0118] Specifically, optimizing the gridded early warning threshold matrix based on the recalculated negative binomial distribution parameters involves re-introducing dynamically updated parameters such as pollution event probabilities into the negative binomial distribution model, recalculating the quantiles at each confidence level, and thus generating an early warning threshold matrix that more closely reflects the current pollution characteristics of the watershed, replacing the original matrix. The optimized matrix is then used for subsequent comparison of pollution event frequencies and tiered early warning, enabling iterative identification and dynamic adaptation to sudden environmental events.
[0119] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, optical storage, etc.) containing computer-usable program code.
[0120] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for early warning of sudden environmental events in regional watersheds based on big data, characterized in that, Includes the following steps: S1. Obtain historical pollution event data for the target watershed and perform pollution event frequency statistics according to spatial grid and time segmentation; Based on statistical results, a gridded early warning threshold matrix is established using the negative binomial distribution; S2. Real-time collection of multi-source monitoring data of the target watershed, and statistical analysis of the frequency of new pollution events according to spatiotemporal grids; The statistical results are compared with the warning threshold matrix of the corresponding grid in real time. If the frequency of new pollution events exceeds the warning threshold, a graded warning is triggered. S3. Verify the confidence level of newly occurring events that have not triggered a warning, and dynamically recalculate the negative binomial distribution parameters of the verified data; S4. Based on the recalculated negative binomial distribution parameters, optimize the gridded early warning threshold matrix and return to step S2 to achieve iterative monitoring of sudden environmental events.
2. The method for early warning of sudden environmental events in regional watersheds based on big data according to claim 1, characterized in that, The process involves acquiring historical pollution event data for the target watershed and statistically analyzing the frequency of pollution events by spatial grid and time segmentation. Based on statistical results, the following steps are taken to establish a gridded early warning threshold matrix using the negative binomial distribution: S11. Collect historical pollution event data of the target watershed, and perform data cleaning and standardization on the historical pollution event data; S12. Based on hydrogeographic characteristics, the target watershed is divided into spatial grids, and time segments are divided according to seasonal cycles to obtain several spatiotemporal units. S13. Statistically analyze the frequency of historical pollution events and safe operating time for each spatiotemporal unit, and use the maximum likelihood estimation method to calculate the probability of events occurring in each spatiotemporal unit. S14. Based on the event occurrence probability of each spatiotemporal unit, calculate the early warning threshold through the quantile function of the negative binomial distribution, and establish a gridded early warning threshold matrix.
3. The method for early warning of sudden environmental events in regional watersheds based on big data according to claim 2, characterized in that, The process of dividing the target watershed into spatial grids based on hydrogeographic features and further dividing it into time segments according to seasonal cycles to obtain several spatiotemporal units includes the following steps: S121. Obtain the geospatial basic data of the target watershed and divide the target watershed into grids according to the preset spatial grid; S122. Perform hydrological consistency verification on the divided spatial grids, and remove grids that fail the hydrological consistency verification as invalid grids to obtain valid grids. S123. Divide the time period according to the preset seasonal cycle, and map the data of each historical pollution event to the corresponding time period according to the occurrence time, and form a spatiotemporal unit by combining the effective grid.
4. The method for early warning of sudden environmental events in a regional watershed based on big data according to claim 2, characterized in that, The process of statistically analyzing the frequency of historical pollution events and safe operating time for each spatiotemporal unit, and calculating the probability of events occurring in each spatiotemporal unit using the maximum likelihood estimation method, includes the following steps: S131. For each spatiotemporal unit, count the total number of historical pollution events, and based on the total number of monitoring days in that spatiotemporal unit, calculate the number of days in which no pollution events occurred as the safe operating time. S132. Based on the total number of historical pollution events and the safe operating time, the probability of pollution events occurring within the spatiotemporal unit is calculated using the maximum likelihood estimation method.
5. The method for early warning of sudden environmental events in a regional watershed based on big data according to claim 4, characterized in that, The method of calculating the probability of a pollution event occurring within a spatiotemporal unit using the maximum likelihood estimation method, based on the total number of historical pollution events and the safe operating time, includes the following steps: S1321. Based on the total number of historical pollution events and the safe operating time, construct a likelihood function with the probability of pollution events occurring as the unknown. S1322. Take the logarithm of the likelihood function to obtain the log-likelihood function, and differentiate the probability of pollution event occurrence with respect to the log-likelihood function to construct the maximum likelihood estimation equation. S1323. Solve the maximum likelihood estimation equation to obtain the probability of pollution events occurring within the spatiotemporal unit.
6. The method for early warning of sudden environmental events in a regional watershed based on big data according to claim 2, characterized in that, The process of calculating the early warning threshold based on the event occurrence probability of each spatiotemporal unit using the quantile function of the negative binomial distribution and establishing a gridded early warning threshold matrix includes the following steps: S141. Based on the event occurrence probability of each spatiotemporal unit, calculate the quantiles at different confidence levels using the quantile function of the negative binomial distribution. S142. Based on the confidence level of the early warning threshold, construct multi-level early warning rules for each spatiotemporal unit; S143. Integrate the warning thresholds of all spatiotemporal units and summarize them according to grid number and time period to form a structured gridded warning threshold matrix.
7. The method for early warning of sudden environmental events in a regional watershed based on big data according to claim 6, characterized in that, The calculation of the early warning threshold based on the event occurrence probability of each spatiotemporal unit using the quantile function of the negative binomial distribution includes the following steps: S1411. Set the success probability parameter of the negative binomial distribution based on the event occurrence probability of the spatiotemporal unit, and determine the success number threshold in combination with the early warning requirements. S1412. Based on prior knowledge, estimate the shape and scale parameters of the negative binomial distribution; and call the quantile function of the negative binomial distribution at multiple confidence levels to calculate the corresponding event frequency threshold. S1413. Use the quantile results of each spatiotemporal unit at different confidence levels as the warning threshold, and mark the confidence level of the warning threshold.
8. The method for early warning of sudden environmental events in regional watersheds based on big data according to claim 1, characterized in that, The process of collecting multi-source monitoring data of the target watershed in real time, counting the frequency of new pollution events according to spatiotemporal grids, and comparing the statistical results with the warning threshold matrix of the corresponding grid in real time, and triggering a graded warning if the frequency of new pollution events exceeds the warning threshold, includes the following steps: S21. Collect multi-source monitoring data of the target watershed in real time, and map the multi-source monitoring data to the corresponding spatial grid according to coordinates and monitoring time; S22. Count the frequency of new pollution events in each spatiotemporal grid according to the time period, and generate a dynamic pollution event frequency matrix; S23. The dynamic pollution event frequency matrix is compared with the corresponding grid's early warning threshold matrix in real time. If the pollution event frequency exceeds the corresponding level threshold, the graded early warning response mechanism is triggered.
9. A method for early warning of sudden environmental events in a regional watershed based on big data, as described in claim 8, is characterized in that... The step of comparing the dynamic pollution event frequency matrix with the corresponding grid's early warning threshold matrix in real time, and triggering a graded early warning response mechanism if the pollution event frequency exceeds the corresponding level threshold, includes the following steps: S231. Extract the statistical frequency of each grid within the current time window from the dynamic pollution event frequency matrix, and load the warning threshold matrix of the corresponding grid. S232. Compare the frequency of pollution events with the thresholds of each level in the early warning threshold matrix in real time, and determine the level of exceeding the limit; S233. Determine the corresponding graded early warning response mechanism based on the level of exceeding the limit.
10. A method for early warning of sudden environmental events in a regional watershed based on big data, as described in claim 1, characterized in that, The process of performing confidence verification on newly occurring events that have not triggered warnings, and dynamically recalculating the negative binomial distribution parameters based on the verified data, includes the following steps: S31. Extract data on spatiotemporal units that have not triggered warnings and their newly emerging pollution events, and conduct preliminary validity screening of the data; S32. Based on multi-dimensional verification rules, perform confidence verification on newly emerging events after preliminary screening, and save the newly emerging events that pass the confidence verification. S33. Dynamically update the negative binomial distribution parameters of the corresponding spatiotemporal unit using monitoring data of newly occurring events verified by confidence level.