Training method and apparatus and regional rainwater runoff pollution load prediction method and apparatus

By screening the correlation between historical rainwater runoff pollution load data and characteristic data, a prediction model was established, which solved the problems of long cycle, high cost and low accuracy of rainwater runoff pollution load calculation in the existing technology, and achieved efficient and accurate rainwater runoff pollution load prediction.

WO2025200277A1PCT designated stage Publication Date: 2025-10-02THREE GORGES GROUP IND DEVELOPMENT (BEIJING) CO LTD +1

Patent Information

Application Number
PCT/CN2024/114937
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-27
Filing Date
2024-08-27
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

The existing method for calculating rainwater runoff pollution load relies on a large amount of real-time monitoring data, which has the characteristics of long cycle, high cost, low accuracy and lack of screening of abnormal data, resulting in low model prediction efficiency and unable to meet the timeliness and accuracy requirements of water environment governance.

Method used

By obtaining the historical rainwater runoff pollution load data and historical characteristic data of the region, the characteristic data to be trained is screened based on the correlation value, a prediction model is established, invalid data is reduced, and the accuracy and efficiency of the training data are improved.

Benefits of technology

It improves the accuracy of the prediction model, reduces the workload of model training, meets the timeliness requirements of water environment governance, and enhances the universality and accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024114937_02102025_PF_FP_ABST
    Figure CN2024114937_02102025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of water environment treatment, and provides a training method and apparatus and a regional rainwater runoff pollution load prediction method and apparatus. The prediction model training method comprises: acquiring historical rainwater runoff pollution load data of a region and historical feature data of the region, wherein the historical feature data comprises feature data representing attribute data categories and feature data representing meteorological data categories; on the basis of correlation values between the historical rainwater runoff pollution load data and the historical feature data, screening the historical feature data to obtain feature data to be trained; and training an initial prediction model on the basis of the feature data to be trained and the historical rainwater runoff pollution load data to obtain a prediction model for regional rainwater runoff pollution load prediction. The solution can improve the accuracy of training data, thereby obtaining a prediction model having high accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Training method, regional rainwater runoff pollution load prediction method and device

[0001] This application claims priority to a Chinese patent application filed with the Patent Office of China on March 27, 2024, with application number 202410354280.X and application name “Training Method, Regional Rainwater Runoff Pollution Load Prediction Method and Device,” the entire contents of which are incorporated herein by reference. Technical Field

[0002] The present application relates to the field of water environment management technology, and in particular to a training method, a method and a device for predicting regional rainwater runoff pollution load. Background Art

[0003] Calculating the pollution load of rainwater runoff is fundamental to water environment management. Existing methods typically rely on statistical analysis and calculation of large amounts of field data (at least dozens of times) over long periods of time (more than one year). This approach requires real-time, on-site sampling and monitoring of the pollution load of rainwater runoff, resulting in long cycles and high costs. Furthermore, rainfall is seasonal and random each year, leading to large errors in the results obtained during actual operation, which cannot meet the timeliness and accuracy requirements of water environment management project design.

[0004] Existing methods also propose using models to predict rainwater runoff pollution loads, thereby reducing sampling cycles and costs. However, existing models still rely on a large amount of monitoring data from the area to be measured. Due to the uncertainty of rainfall monitoring, there are often a large number of abnormal and invalid data in the existing monitoring data. Existing models still lack screening and training for abnormal and invalid data, which ultimately affects the accuracy of model predictions. The processing of abnormal and invalid data will lead to low model processing efficiency.

[0005] Summary of the Invention

[0006] The present application provides a training method, a method and an apparatus for predicting regional rainwater runoff pollution load, which are used to solve the problems of large data processing volume and low accuracy of the prediction model.

[0007] In a first aspect, the present application provides a prediction model training method, comprising: obtaining historical rainwater runoff pollution load data of a region and historical characteristic data of the region; wherein the historical characteristic data includes characteristic data representing attribute data categories and characteristic data representing meteorological data categories; based on the correlation value between the historical rainwater runoff pollution load data and the historical characteristic data, screening the historical characteristic data to obtain characteristic data to be trained; training an initial prediction model based on the characteristic data to be trained and the historical rainwater runoff pollution load data to obtain a prediction model for regional rainwater runoff pollution load prediction.

[0008] In one implementable manner, the historical characteristic data includes underlying surface data of the multiple areas; obtaining historical rainwater runoff pollution load data of the area includes: obtaining single rainwater runoff pollution load data of the underlying surface corresponding to the underlying surface data; and obtaining historical rainwater runoff pollution load data characterizing the rainwater runoff pollution load of the area based on the area of ​​the underlying surface in the underlying surface data and the single rainwater runoff pollution load data.

[0009] In one implementable manner, there are multiple regions; based on the correlation value between the historical rainwater runoff pollution load data and the historical feature data, the historical feature data is screened to obtain the feature data to be trained, including: clustering processing based on the historical feature data of each region to obtain the regional category corresponding to each region; based on the correlation value between the historical rainwater runoff pollution load data and the historical feature data in each regional category, the historical feature data in each regional category is screened to obtain the feature data to be trained corresponding to each regional category, so as to train a prediction model for the corresponding regional category based on the feature data to be trained and the historical rainwater runoff pollution load data of each regional category.

[0010] In one implementable manner, the historical feature data includes feature data corresponding to multiple feature categories; the feature data to be trained is screened from the historical feature data based on the correlation value between the historical rainwater runoff pollution load data and the historical feature data, including: respectively obtaining the correlation value between the historical rainwater runoff pollution load data and each feature data; and based on the size of the correlation value corresponding to each feature data, screening the feature data to be trained from the feature data corresponding to the multiple feature categories.

[0011] In a second aspect, the present application provides a method for predicting regional rainwater runoff pollution load, comprising: obtaining target feature data of a target area; wherein the target feature data includes feature data representing attribute data categories and feature data representing meteorological data categories; obtaining preset feature categories based on the target feature data, filtering the target area feature data based on the preset feature categories, and inputting the filtered target feature data into a pre-trained prediction model to obtain rainwater runoff pollution load data of the target area output by the prediction model; wherein the preset feature categories are obtained based on the correlation value between historical rainwater runoff pollution load data and the feature data corresponding to each feature category in the historical feature data.

[0012] In one implementable manner, the prediction model includes multiple models, different prediction models correspond to different area categories, and different area categories correspond to different preset feature categories; the preset feature category is obtained based on the target feature data, the target area feature data is screened based on the preset feature category, and the screened target feature data is input into a pre-trained prediction model to obtain the rainwater runoff pollution load data of the target area output by the prediction model, including: classifying the target area based on the target feature data to obtain the target area category of the target area; based on the preset feature category in the target area category, the feature data corresponding to the feature category is screened in the target area data to obtain the screened target area feature data; the screened target area feature data is input into the trained training model corresponding to the target area category to obtain the rainwater runoff pollution load data of the target area.

[0013] In a third aspect, the present application provides a prediction model training device, comprising: a historical data acquisition module, configured to acquire historical rainwater runoff pollution load data of a region and historical characteristic data of the region; wherein the historical characteristic data includes characteristic data representing attribute data categories and characteristic data representing meteorological data categories; a data processing module, configured to filter and obtain feature data to be trained from the historical feature data based on the correlation value between the historical rainwater runoff pollution load data and the historical characteristic data; a model training module, configured to train an initial prediction model based on the feature data to be trained and the historical rainwater runoff pollution load data to obtain a prediction model for regional rainwater runoff pollution load prediction.

[0014] In a fourth aspect, the present application provides a regional rainwater runoff pollution load prediction device, comprising: a target data acquisition module, configured to obtain target feature data of a target area; wherein the target data includes feature data representing attribute data categories and feature data representing meteorological data categories; a prediction module, configured to obtain preset feature categories based on the target feature data, to filter the target area feature data based on the preset feature categories, and to input the filtered target area feature data into a pre-trained prediction model to obtain the rainwater runoff pollution load data of the target area output by the prediction model.

[0015] In a fifth aspect, the present application provides an electronic device, comprising: a processor and a memory; the memory is used to store instructions; the processor is used to execute the instructions in the memory, so that the electronic device executes the training method of the prediction model as described in the first aspect or the regional rainwater runoff pollution load prediction method as described in the second aspect.

[0016] In a sixth aspect, the present application provides a computer-readable storage medium, which stores computer execution instructions. When the computer execution instructions are executed by a processor, they are used to implement the training method of the prediction model as described in the first aspect or the regional rainwater runoff pollution load prediction method as described in the second aspect.

[0017] In a seventh aspect, the present application provides a computer program product, comprising a computer program, which, when executed by a processor, implements the prediction model training method as described in the first aspect or the regional rainwater runoff pollution load prediction method as described in the second aspect.

[0018] The training method, regional rainwater runoff pollution load prediction method and device provided in the present application obtain historical feature data that characterizes attributes and meteorology in the region and obtain historical rainwater runoff pollution load data, and calculate the correlation value between the historical feature data and the historical rainwater runoff pollution load data, and screen the feature data to be trained that has a greater correlation with the historical pollution load data, thereby removing invalid data in the historical feature data, improving the accuracy of the training data, improving the accuracy of the prediction model obtained by training, and reducing the volume of the feature data to be trained used to train the prediction model, reducing the workload of model training, and improving the efficiency of model training. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0020] FIG1 is a diagram showing an implementation environment structure of an exemplary embodiment of the present application;

[0021] FIG2 is a flow chart of a prediction model training method according to an exemplary embodiment of the present application;

[0022] FIG3 is a schematic diagram of cluster analysis results shown in an exemplary embodiment of the present application;

[0023] FIG4 is a schematic diagram of the fitting results of predicted values ​​and observed values ​​shown in an exemplary embodiment of the present application;

[0024] FIG5 is a flow chart of a method for predicting regional rainwater runoff pollution load according to an exemplary embodiment of the present application;

[0025] FIG6 is a diagram showing principal coordinate analysis results according to an exemplary embodiment of the present application;

[0026] FIG7 is a diagram showing a comparison between measured values ​​and predicted values ​​according to an exemplary embodiment of the present application;

[0027] FIG8 is a schematic diagram of the structure of a prediction model training device according to an exemplary embodiment of the present application;

[0028] FIG9 is a schematic structural diagram of a device for predicting regional rainwater runoff pollution load according to an exemplary embodiment of the present application;

[0029] FIG10 is a block diagram of an electronic device according to an exemplary embodiment of the present application.

[0030] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION

[0031] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0032] First, let’s explain the terms involved in this application:

[0033] EMC: refers to the average concentration of rainfall runoff, which is used to evaluate and calculate the large amount of pollutants carried in surface rainfall runoff. Pollutants can include different pollution factors such as COD (organic pollutants), SS (suspended solids), NH3-N (ammonia nitrogen), TN (total nitrogen) and TP (total phosphorus).

[0034] In this application, rainwater runoff pollution load data generally refers to EMC-related data. Specifically, because data from a single rainfall event is incidental, meaningless, and inaccurate, historical rainwater runoff pollution load data can be the average EMC data for a specific time unit (e.g., a year, a quarter, a three-year period, etc.), that is, averaged over a number of EMC data within that time unit. Specifically, it is obtained by taking a weighted average of the average runoff concentrations of individual rainfall events within that time unit. For example, if the time unit is a year, the corresponding historical rainwater runoff is the weighted average of the EMC data from a number of single rainfall events within that year.

[0035] Therefore, the rainwater runoff pollution load data of the target area generally output by the prediction model in this application is also the average EMC data under the time unit. Of course, in some embodiments, the EMC data reflects the concentration of pollutants. When treating rainwater pollution, the general pollution quantity data can more intuitively display the sewage situation in the area, thereby specifying the corresponding treatment method. The pollution quantity data is generally obtained by multiplying the EMC data by the runoff flow. Specifically, for a single rainfall, the pollution quantity data of the single rainfall is the EMC data of the single rainfall multiplied by the runoff flow / rainfall corresponding to the single rainfall. As for the pollution quantity data under a certain time unit, the pollution quantity data under the time unit is the sum of the products of the EMC data of each rainfall in the time unit and the runoff flow of the corresponding rainfall session, that is, the pollution quantity data under the time unit is The sum of the pollution data of each single rainfall in the time unit. Therefore, in some embodiments, a set of data can be added when the prediction model is used (the runoff flow of the target area in the time unit, and the runoff flow refers to the sum of the runoff flow of all rainfall sessions in the time unit). When the prediction model is processed, the average EMC data in the time unit can be obtained first, and then the average EMC data in the time unit and the input rainwater runoff pollution load data are multiplied by the average EMC data and the runoff flow in the time unit to obtain the pollution data of the target area in the time unit. The prediction model inputs the pollution data of the target area in the time unit. In this way, it can also be regarded as that in some special embodiments, the rainwater runoff pollution load data of the target area output by the prediction model is the pollution data of the target area in the time unit.

[0036] The pollution problem of rainwater runoff caused by rainfall is becoming increasingly serious. The calculation of regional rainwater runoff pollution load is the basic basis for water environment governance. However, due to the complex pollution sources and many influencing factors, the calculation and prediction of rainwater runoff pollution load is difficult, and there are common problems such as complex monitoring, low accuracy, and lack of timeliness.

[0037] Currently, the calculation of rainwater runoff pollution load prediction usually uses real-time sampling and monitoring methods. After calculation, the event mean concentration (EMC) of the event runoff is obtained. For each underlying surface, real-time monitoring of at least 15-20 representative rainfall runoff data throughout the year is required to obtain a relatively accurate weighted average concentration of pollutants. Therefore, this method requires repeated data measurement for different underlying surfaces and different regions. This is not only time-consuming and labor-intensive, with a long monitoring cycle, but also difficult to control the sampling standardization and large errors during actual operation.

[0038] Rainwater runoff pollution load prediction also uses model methods, or modifies the model based on real-time monitoring data. Models such as SWMM (dynamic precipitation-runoff simulation model), SWAT (watershed scale model), and HSPF (hydrological simulation model) are calibrated and modified based on regional conditions, basic data, and monitoring data. However, such methods still have high requirements for basic data, requiring a large amount of data on meteorology, hydrology, land use, pollutant accumulation and flushing, and drainage network layout. These basic data are generally obtained from historical data monitored over the years, or new monitoring is added according to the needs of the model to obtain relevant basic data to meet the needs of model calculation and parameter calibration, and ensure the accuracy of model calculations.

[0039] Therefore, the existing regional rainwater runoff pollution load prediction technology, regardless of the model method or the actual measurement method, relies on a large amount of monitoring data in the target area. According to the different target areas, a large amount of data needs to be monitored repeatedly. It is impossible to mine the common and characteristic factors, common and characteristic data between different regions, and there is no mining, screening and application of the monitoring data disclosed by existing research, so as to reduce the workload of repeated monitoring. It can be seen that the defects of regional rainwater runoff pollution load prediction technology include three aspects: First, when relying on a large amount of real-time monitoring data, or relying on historical monitoring data, there is generally incomplete data, and real-time monitoring data needs to be supplemented. The supplementary monitoring data is large in volume and low in efficiency. At the same time, due to the seasonality and randomness of rainfall, it takes a long time to obtain a large amount of actual monitoring data. Secondly, the accuracy of on-site monitoring data such as hydrology and water quality is a big challenge due to the influence of monitoring environment conditions and randomness of sampling operations. Both real-time monitoring data and historical monitoring data contain a large amount of abnormal and invalid data. Existing methods generally lack data screening and scientific elimination of invalid or interfering data, which affects the accuracy of the results. Thirdly, existing models are all established for specific target areas and rely on a large amount of data in the target area. They cannot be used for other target areas (such as different cities) and need to retrain a model based on a large amount of data from other target areas. Therefore, the universality and pertinence of the methods are not strong, and the workload of monitoring and calculation is relatively large.

[0040] To address the above issues, this embodiment proposes a prediction model training method and a regional rainwater runoff pollution load prediction method. This prediction model training method fully utilizes historical data from multiple regions, not just the target region, significantly increasing the volume of monitoring data. Through specific clustering, it enables the scientific utilization of historical data from different regions. Furthermore, this prediction model training method can filter high-quality training data, reducing the training data volume and thus improving the accuracy of the prediction model. Therefore, the rainwater runoff pollution load prediction method established based on this prediction model can efficiently obtain highly accurate regional rainwater runoff pollution load data.

[0041] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0042] It should be noted that the rainwater runoff pollution proposed in this embodiment generally refers to the pollution caused by initial rainwater. In actual operation, the pollution caused by initial rainwater is relatively large and has research and monitoring value. However, in the later stages of rainfall, the resulting rainwater runoff pollution load will decrease and has no research value. The historical rainwater runoff pollution load data and the rainwater runoff pollution load data for the target area proposed in this embodiment generally refer to the pollution load under a specific time unit, such as the pollution load of one year, one quarter, or two quarters. The pollution load of a single rainfall is accidental and has little guiding significance for water environment governance projects.

[0043] The prediction model training method and regional rainwater runoff pollution load prediction method provided in this application can both be applied to the implementation environment shown in Figure 1. As shown in Figure 1, the implementation environment includes: a terminal 1 and a server 2, wherein the terminal 1 and the server 2 are communicatively connected. Of course, the number of terminals 1 and servers 2 shown in Figure 1 is for illustration only. In other embodiments, the implementation environment may also include other numbers of terminals 1 and servers 2, and this is not specifically limited here.

[0044] The terminal 1 in this embodiment can be understood as various electronic devices, such as wired or wireless terminals, including smartphones, computers, smart home appliances, and in-vehicle terminals, and is not specifically limited here. The server 2 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server, and is not specifically limited here.

[0045] When the prediction model training method in this embodiment is applied to the implementation environment of Figure 1, the terminal 1 can obtain the historical rainwater runoff pollution load data of the region and the historical feature data of the region, wherein the historical feature data includes feature data representing the attribute data category and feature data representing the meteorological data category; and send the historical rainwater runoff pollution load data of the region and the historical feature data of the region to the server end 2, so that after the server end 2 obtains the historical rainwater runoff pollution load data of the region and the historical feature data of the region, based on the correlation value between the historical rainwater runoff pollution load data and the historical feature data, it filters the feature data to be trained in the historical feature data; based on the feature data to be trained and the historical rainwater runoff pollution load data, the initial prediction model is trained to obtain a prediction model for regional rainwater runoff pollution load prediction; accordingly, the prediction model training device is set in the server end 2 to complete the training of the prediction model.

[0046] When the regional rainwater runoff pollution load prediction method in this embodiment is applied to the implementation environment of Figure 1, the terminal 1 can obtain the target feature data of the target area; wherein the target feature data includes feature data representing the attribute data category and feature data representing the meteorological data category; and send the target feature data of the target area to the server end 2, so that after obtaining the target feature data of the target area, the server end 2 obtains the preset feature category based on the target feature data, filters the target area feature data based on the preset feature category, and inputs the filtered target feature data into the pre-trained prediction model to obtain the rainwater runoff pollution load data of the target area output by the prediction model; wherein the preset feature category is obtained based on the correlation value between the historical rainwater runoff pollution load data and the feature data corresponding to each feature category in the historical feature data; accordingly, the regional rainwater runoff pollution load prediction device is set in the server end 2 to complete the rainwater runoff pollution load prediction of the target area.

[0047] Of course, the implementation environment shown in Figure 1 in this embodiment is only exemplary. In other embodiments, the prediction model training method and the regional rainwater runoff pollution load prediction method can also be applied to other implementation environments, and the implementation environments of the prediction model training method and the regional rainwater runoff pollution load prediction method can be the same or different.

[0048] FIG2 is a prediction model training method according to an embodiment of the present application. As shown in FIG2 , the prediction model training method may include steps S210 to S250, which are described in detail as follows:

[0049] Step S210: Acquire historical rainwater runoff pollution load data and historical characteristic data of the region.

[0050] In this embodiment, the historical characteristic data includes characteristic data representing the attribute data category and characteristic data representing the meteorological data category. The historical rainwater runoff pollution load data represents the average EMC data of the region in a certain time unit, that is, it is obtained by weighted averaging the single rainwater runoff pollution load data. For example, if one year is a time unit, the historical rainwater runoff pollution load data represents the annual average EMC data, that is, it is obtained by weighted averaging the single rainwater runoff pollution load data in that year.

[0051] It should be noted that when calculating the historical rainwater runoff pollution load data, this embodiment does not use all the single rainwater runoff pollution load data under the time unit. It can first be screened to obtain representative single rainwater runoff pollution load data, and then the historical rainwater runoff pollution load data under the time unit can be calculated.

[0052] Among them, the characteristic data of the attribute data category include characteristic data such as the terrain and economic level of the region. Among them, the terrain of the region can include the underlying surface data, average slope data, etc. of the region. The underlying surface data can also be refined into data such as the underlying surface type, the impermeability of the underlying surface, the slope of the underlying surface, and the area of ​​the underlying surface; the characteristic data of the economic level can include characteristic data such as the annual per capita income of the region that may affect EMC.

[0053] The characteristic data of meteorological data category include annual precipitation, annual average temperature, secondary precipitation, secondary precipitation duration, secondary average rainfall intensity, dry period before secondary rainfall and other characteristic data that may affect EMC. Among them, single meteorological data can correspond to rainfall underlying surface data. For example, for each underlying surface, there are characteristic data such as secondary precipitation and secondary precipitation duration under the underlying surface scenario.

[0054] Specifically, historical characteristic data include multiple data categories, such as attribute data categories and meteorological data categories, and each data category includes multiple characteristic data. The characteristic data are distinguished by characteristic categories, and the characteristic category can be regarded as the smallest category unit in the historical characteristic data. For example, the impermeability of the underlying surface and the average temperature mentioned above can be regarded as two characteristic categories, that is, the historical characteristic data can also be regarded as multiple characteristic categories under each data category, each characteristic category corresponds to characteristic data, and multiple characteristic data constitute the historical characteristic data of the region.

[0055] The historical characteristic data of the region in this embodiment is the historical monitoring data of the region, which can be obtained through local inquiries such as literature databases, statistical yearbooks or statistical bulletins. In addition, the historical characteristic data and historical rainwater runoff pollution load data of the region obtained in this embodiment should have a relatively long time span, such as a time span of 5 years, 6 years, or 10 years, so as to ensure that the historical characteristic data and historical rainwater runoff pollution load data of the region obtained have sufficient data volume and time representativeness for prediction model training.

[0056] This time span can be regarded as the time for obtaining historical characteristic data and historical rainwater runoff pollution load data. If the data acquisition time span proposed above is 5 years, the historical characteristic data and historical rainwater runoff pollution load data of the past five years are obtained, and then the mean of the average historical characteristic data and the mean of the historical rainwater runoff pollution load data in a certain time unit are calculated, represented by the time unit, and the mean is used as the historical characteristic data and historical rainwater runoff pollution load data corresponding to the region.

[0057] The time unit is a preset value to reduce the randomness of the rainfall pollution load data. The time unit is generally longer, such as year, quarter, etc. If the time unit is year, then taking a time span of 5 years as an example, the corresponding historical characteristic data and historical rainwater runoff pollution load data of the region are the annual average historical characteristic data and historical rainwater runoff pollution load data. In this way, the historical rainwater runoff pollution load data and the annual average data in the historical characteristic data of the region (such as annual precipitation, annual average temperature and other data are regarded as 5-year annual average data), and the corresponding historical rainwater runoff pollution load data are the data obtained by averaging the 5-year rainwater runoff pollution load data. Of course, in some embodiments, quarterly representation can also be used for data acquisition. At this time, the annual precipitation, annual average temperature and other data obtained are correspondingly modified to quarterly precipitation, quarterly average temperature, etc. There is no restriction on the data acquisition method here.

[0058] In a specific embodiment, a method for calculating historical rainwater runoff pollution load data is proposed, including step S10: obtaining single rainwater runoff pollution load data of the underlying surface corresponding to the underlying surface data; step S12: based on the area of ​​the underlying surface in the underlying surface data and the single rainwater runoff pollution load data, obtaining historical rainwater runoff pollution load data representing the regional rainwater runoff pollution load.

[0059] Specifically, for each underlying surface in the area, the single rainwater runoff pollution load data of the corresponding underlying surface is obtained. The single rainwater runoff pollution load data can be obtained at the same time as the historical characteristic data is obtained, or after obtaining the historical characteristic data, the single rainwater runoff pollution load data can be obtained based on the underlying surface data. That is, at this time, the single rainwater runoff pollution load data under each underlying surface data and the corresponding data such as the precipitation amount, precipitation duration, average rainfall intensity, and dry period before the rain can be obtained.

[0060] In this embodiment, based on the area of ​​the underlying surface in the underlying surface data and the single rainwater runoff pollution load data, historical rainwater runoff pollution load data representing the regional rainwater runoff pollution load is obtained, and the historical rainwater runoff pollution load data can be obtained by performing a weighted average calculation on the single rainwater runoff pollution load data.

[0061] Specifically, based on the underlying surface data and the single rainwater runoff pollution load data, the average rainwater runoff pollution load data under the corresponding underlying surface can be calculated. For example, for historical rainwater runoff pollution load data with years as the time unit, multiple single rainwater runoff pollution load data under a certain underlying surface in a certain year can be obtained. By taking the average, the historical rainwater runoff pollution load data under the underlying surface in that year can be obtained. If the time span is 5 years, the rainwater runoff pollution load data under the underlying surface for each of the 5 years will be calculated separately.

[0062] Then, based on the size of the underlying surface in the region, the weights of different underlying surfaces in the region are obtained. The weight of the underlying surface with a larger area is higher, and the weight of the smaller area is lower. The historical rainwater runoff pollution load data is calculated using the weight and the historical rainwater runoff pollution load data of each underlying surface. The historical rainwater runoff pollution load data obtained at this time is the historical rainwater runoff pollution load data under the corresponding time unit. The prediction model obtained by subsequent training using this data also predicts the historical rainwater runoff pollution load data of the target area under the time unit, such as the rainwater runoff pollution load data of the target area in a certain year.

[0063] Of course, in order to ensure the accuracy of the training data, after obtaining the historical rainwater runoff pollution load data under the time unit, the average value can be calculated based on the time span to obtain the average value of the historical rainwater runoff pollution load data under the time span and the time unit. Later, when the model is trained, for a set of data input into the model, the historical rainwater runoff pollution load data under the time unit corresponds to the feature data to be trained, and the average value of the historical rainwater runoff pollution load data under the time span and the time unit corresponds to the average value of the feature data to be trained under the time span. For example, in one embodiment, the time span is 5 years and the time unit is year, then the historical rainwater runoff pollution load data of each year and the feature data to be trained of the corresponding year are input into the model as a set of training data for training, and the average value of the historical rainwater runoff pollution load data obtained by averaging the historical rainwater runoff pollution load data for 5 years and the average value of the data to be trained for 5 years are input into the model as a set of training data for training. Alternatively, the average value of the historical rainwater runoff pollution load data of any two or more years in 5 years and the average value of the data to be trained in the corresponding year can be input into the model as a set of training data for training.

[0064] In this embodiment, for an area, there may be multiple different types of underlying surfaces in the area, and there may be multiple single rainwater runoff pollution load data under each underlying surface. Especially for those with a long time span, the corresponding single rainwater runoff pollution load data under each underlying surface that needs to be processed may be more, and the amount of data to be processed is large. At this time, the single rainwater runoff pollution load data under each underlying surface can be screened first, and the representative single rainwater runoff pollution load data under each underlying surface can be selected, and then the historical rainwater runoff pollution load data can be calculated based on the screened data. The screening method can be manual screening based on the description of historical literature; or it can be input into the machine model for different underlying surfaces, and the machine model can be used to calculate the historical rainwater runoff pollution load data. The characteristics of each single rainwater runoff pollution load data are analyzed by the machine model, so that representative single rainwater runoff pollution load data can be screened out. The screening of single rainwater runoff pollution load data by the machine model can be based on the correlation analysis between the single rainwater runoff pollution load data of the machine model, so as to screen the single rainwater runoff pollution load data with high correlation as the representative single rainwater runoff pollution load data, or the underlying surface data is learned by the machine model and then the single rainwater runoff pollution load data are scored, and the single rainwater runoff pollution load data with high scores are screened out as the representative single rainwater runoff pollution load data. There is no specific restriction here. Of course, the single rainwater runoff pollution load data of the underlying surface can also be randomly selected, and there is no specific restriction here.

[0065] Step S230: Based on the correlation value between the historical rainwater runoff pollution load data and the historical feature data, the feature data to be trained is obtained by screening the historical feature data.

[0066] In this embodiment, the historical feature data has a large data volume due to its wide range and long time span, and there are some invalid feature data that are irrelevant to the historical rainwater runoff pollution load. The invalid feature data will cause the model training rate to slow down and affect the results of the model training. Therefore, it can be seen that based on the correlation between each feature data in the historical feature data and the historical rainwater runoff pollution load data, the training feature data is screened and irrelevant or low-correlation feature data is excluded, thereby improving the efficiency of model training; the correlation value represents the degree of correlation between the historical feature data and the historical rainwater runoff pollution data. The larger the value, the higher the correlation.

[0067] In this way, the correlation values ​​between the historical rainwater runoff pollution load data and each feature data can be obtained respectively; based on the size of the correlation value corresponding to each feature data, the feature data corresponding to multiple feature categories are screened to obtain the feature data to be trained. The feature data to be trained is the feature data with a greater correlation with the historical rainwater runoff pollution load data of the corresponding area. Therefore, invalid data in the historical rainwater runoff pollution load data can be excluded, the data volume can be reduced, and the reliability of the data to be trained can be improved.

[0068] Of course, in this embodiment, when training the prediction model, historical rainwater runoff pollution load data and historical characteristic data corresponding to multiple regions can be obtained, and multiple regions can be classified, so that a prediction model of the corresponding regional category can be trained based on the data of the corresponding regional category, so that when the rainwater runoff pollution load data of the target area is subsequently predicted, the prediction module of the regional category of the target area can be directly used for prediction. The universality of this prediction module is high, which can reduce the workload of rainwater runoff pollution load prediction.

[0069] Specifically, for at least one area of ​​the same area type, the historical rainwater load data and historical characteristic data of all areas in the area type are used as the historical rainwater load data and historical characteristic data of the area type, and then the correlation data of each historical rainwater load data in the area type and each characteristic data in the historical characteristic data of the corresponding area are calculated. The degree of correlation between the characteristic data of each area in the area type and the historical rainwater load data of the corresponding area can be obtained, thereby screening the feature data to be trained for the corresponding area type.

[0070] It is understandable that the classification method in this embodiment may be a cluster analysis method, a similarity matching method, a regression method, etc., to perform classification based on the historical characteristic data and historical rainwater runoff pollution load data of each region, and no specific limitation is made here.

[0071] Step S250: training the initial prediction model based on the feature data to be trained and the historical rainwater runoff pollution load data to obtain a prediction model for regional rainwater runoff pollution load prediction.

[0072] In this embodiment, an initial prediction model may be established first, and then the initial prediction model may be trained based on the feature data to be trained and the historical rainwater runoff pollution load data to obtain a prediction model for regional rainwater runoff pollution load prediction.

[0073] In this embodiment, a corresponding initial prediction model can be established for a region. Of course, for universality, an initial prediction model corresponding to a regional category can also be established for a regional category, and then a prediction model corresponding to the regional category can be trained. Subsequently, when making predictions for the target area, the regional category of the target area can be detected, and the prediction model corresponding to the regional category can be used to predict the rainwater runoff pollution load of the target area.

[0074] Specifically, the initial prediction model can be a random forest model, a linear regression model, a neural network model, etc., and there is no specific limitation here.

[0075] In this embodiment, the feature data to be trained is used as the independent variable, and the historical rainwater runoff pollution load data is used as the dependent variable. The feature data to be trained and the historical rainwater runoff pollution load data corresponding to each initial prediction model are divided into a training set and a test set (the data volume ratio is 4:1). The training set is used to train the initial prediction model, and the trained prediction model is used to predict the samples in the test set.

[0076] In a specific embodiment, the initial prediction model is a random forest model. The random forest regression model is established using software or languages ​​such as RStudio, SPSS, and Python. The number of decision trees ntree in the model is set to 500. The number of candidate variables mtry in the case of decision tree splitting is tried one by one according to the actual situation until the ideal value is found. The remaining parameters can be set to the default values ​​in the model. In model evaluation and optimization, the square error, mean absolute error, determination coefficient, etc. can be used to evaluate the accuracy and generalization ability of the model, and the random forest parameters are adjusted according to the evaluation results to optimize the performance of the model.

[0077] In this embodiment, a prediction model training method is proposed, which fully utilizes historical rainwater runoff pollution load data and historical characteristic data as basic database information, and by establishing a correlation between historical rainwater runoff pollution load data and historical characteristic data, non-major influencing factors on historical rainwater load data in historical characteristic data are eliminated, and characteristic data with a high correlation with historical rainwater load data are screened out, thereby greatly simplifying the conditions for subsequent model training, reducing the data volume, reducing the data processing time cost for training, and meeting the timeliness requirements of engineering management. In addition, the data to be trained is more reliable and accurate, which can improve the accuracy of model training.

[0078] At the same time, this application also classifies multiple regions according to historical rainwater runoff pollution load data and historical feature data through cluster analysis, and trains the initial prediction model of the corresponding regional type using the feature data to be trained in the regional type obtained by classification to obtain a prediction model for the corresponding regional category. Then, the corresponding prediction model can be obtained based on the regional category of the target area, which can improve the universality of the regional rainwater runoff pollution load. There is no need to repeat the monitoring data acquisition, model training and other steps for each region, which greatly reduces the monitoring workload and repetitive work, and especially solves the defect that accurate prediction models cannot be trained in areas with little monitoring data and incomplete information.

[0079] In one embodiment, a method for calculating correlation values ​​in a region category to obtain feature data to be trained is also proposed. Specifically, there are multiple regions. In this case, step S230 may include steps S21 to S23, which are specifically described as follows:

[0080] Step S21: performing clustering processing based on the historical characteristic data of each region to obtain the region category corresponding to each region.

[0081] In this embodiment, the historical characteristic data of each region and the historical rainwater runoff pollution load data are used to classify all the involved regions according to their similarity by adopting a system cluster analysis method.

[0082] The number of clusters is determined by the "tree diagram method" (by drawing a tree structure diagram and finding the optimal number of clusters by cutting a part of the tree structure); the distance metric is determined according to the final clustering effect, and different distances such as Euclidean distance, Manhattan distance, and Canberra distance can be selected. Data analysis can be completed through data analysis software such as Spss, Past, and RStudio.

[0083] Based on the results of cluster analysis, we can analyze the typical characteristics of various types of regions (such as terrain, climate, economy, pollution, etc.) and sort out the reasons for classification.

[0084] Step S23: Based on the correlation value between the historical rainwater runoff pollution load data and the historical feature data in each regional category, the historical feature data in each regional category is screened to obtain the feature data to be trained corresponding to each regional category, so as to train the prediction model of the corresponding regional category based on the feature data to be trained and the historical rainwater runoff pollution load data of each regional category.

[0085] In this embodiment, for a regional category, the correlation value between the historical rainwater runoff pollution load data of each region in the regional category and the feature data of each corresponding feature category is calculated, so that the feature data with higher correlation data is used as the training data for each region in the regional category, and the training data for each region in a regional category is used as the training data for the regional category, wherein the training data for the regional category corresponds to the historical rainwater runoff pollution load data of a certain region in the regional category.

[0086] In one embodiment, the present application further proposes a method for screening feature data to be trained, wherein the historical feature data includes feature data corresponding to multiple feature categories, which includes steps S31 to S32, as described in detail as follows:

[0087] Step S31: respectively obtaining correlation values ​​between historical rainwater runoff pollution load data and each characteristic data.

[0088] In this embodiment, for a region or a region type, correlation analysis and partial correlation analysis methods can be used to screen out feature categories that have a greater impact on historical rainwater runoff pollution load data. Specifically, the correlation value between the historical rainwater runoff pollution load data and the feature data of each feature category in the corresponding region is calculated. The correlation value can be the Pearson correlation coefficient.

[0089] In one embodiment, the historical rainwater runoff pollution load data may include historical rainwater runoff pollution load data of pollution factors such as COD, SS, NH3-N, TN and TP. Therefore, calculating the correlation value between the rainwater runoff pollution load data and each characteristic data may be calculating the correlation value between the historical rainwater runoff pollution load data of each pollution factor and each characteristic data.

[0090] Step S32: based on the magnitude of the correlation value corresponding to each feature data, feature data corresponding to a plurality of feature categories are screened to obtain feature data to be trained.

[0091] In this embodiment, a threshold value can be set. When the correlation value is greater than the threshold value, the correlation between the feature data of the corresponding feature category and the historical rainwater runoff pollution ancient river data is not significant, and the feature data can be eliminated. If the correlation value is not greater than the threshold value, the correlation between the feature data of the corresponding feature category and the historical rainwater runoff pollution load data is significant, and the feature data can be used as feature data to be trained.

[0092] This embodiment proposes to screen the feature data to be trained based on the correlation value, thereby reducing the amount of data for model training and improving the efficiency of model training. At the same time, by eliminating irrelevant data, the reliability of the feature data to be trained is improved, further improving the accuracy of the prediction model.

[0093] Based on the above-mentioned prediction model training method, this embodiment discloses a specific embodiment. By querying and obtaining literature related to rainwater runoff pollution for D1, D2, D3, D4, D5, D6, D7, D8, D9, D10, D11, D12, and D13 through a literature database, 584 sets of single-event rainwater runoff pollution load data (including COD, SS, TN, and TP pollution factors) for single rainfall on different underlying surfaces are sorted out, and the weighted average of the historical rainwater runoff pollution load data for all rainfall events in each region is calculated. In addition, corresponding historical characteristic data (such as precipitation, precipitation duration, average rainfall intensity and dry period before rain, underlying surface type, impermeability and slope, average slope, annual precipitation, annual average temperature, and economic level) are obtained from databases and yearbooks. The economic level can be reflected by data such as GDP (gross domestic product, the GDP of the corresponding region). The above data spans 15 years to ensure sufficient data volume and annual representativeness.

[0094] Specifically, the data in Table 1 below can be obtained. Of course, Table 1 only shows some historical rainwater runoff pollution load data and some historical characteristic data. Among them, COD, SS, TN and TP in Table 1 are historical rainwater runoff pollution load data of corresponding pollution factors.

[0095] Table 1: Examples of historical stormwater runoff pollution load data and historical characteristic data

[0096] The above 13 regions were classified by systematic clustering method. Combined with the data in Table 1, cluster analysis was completed by RStudio software, and the cluster analysis results were obtained as shown in Figure 3. The number of clusters was determined by "tree diagram method". It can be clearly seen from Figure 3 that the 13 regions can be divided into 3 regional categories. According to the cluster analysis results, the typical characteristics of various types of regions can be analyzed (such as terrain, climate, economy and pollution, etc.), and the characteristic data of some typical characteristic categories of the corresponding regional categories can be obtained, as shown in Table 2. Of course, the data in Table 2 are also exemplary and do not represent the characteristic data of all characteristic categories of each regional category.

[0097] Among them, SS in Table 2 ave TN ave TP ave and COD ave Historical rainwater runoff pollution load data corresponding to different pollution factors.

[0098] Table 2: Typical characteristics of regional categories Feature data of the category

[0099] For the three regional categories, correlation analysis was used to analyze the correlation data between the historical rainwater runoff pollution load data and the feature data of the corresponding feature categories. In the three regional categories in this embodiment, annual precipitation and single precipitation were negatively correlated with the historical rainwater runoff pollution load data, while the single average rainfall intensity, dry period before rain, and underlying surface impermeability were positively correlated with the historical rainwater runoff pollution load data. Per capita GDP (gross domestic product) was only positively correlated with the historical rainwater runoff pollution load data of COD. Therefore, the feature data of the above feature categories were selected as key feature data affecting the rainwater surface runoff pollution load and were selected as the feature data to be trained. The average slope, average temperature, precipitation duration, and slope had no correlation with the historical rainwater runoff pollution load data and were eliminated in the subsequent model training.

[0100] After obtaining the training feature data for each feature category, a random forest model was used to model and predict the rainwater runoff pollution load prediction data for different regional categories during a single rainfall event. The random forest model was implemented using the RandomForest package in the RStudio software platform. The training feature data selected through correlation analysis served as the independent variable, and the historical rainwater runoff pollution load data served as the dependent variable. Each data set was divided into a training set and a test set (4:1) to establish a random forest regression model. After debugging, the number of decision trees in the model was set to 500, and the number of candidate variables in the decision tree splitting case, mtry, was set to 3 for optimal results. The remaining parameters were set to the default values ​​in the model, resulting in prediction models corresponding to the three regional categories.

[0101] Furthermore, the prediction models corresponding to the three regional categories were evaluated, and the change trends of the predicted values ​​and observed values ​​of the pollution factors (COD, SS, TN and TP) of the prediction models corresponding to the three regional categories were obtained respectively, and a schematic diagram of the fitting results of the predicted values ​​and the observed values ​​was obtained as shown in Figure 4. In Figure 4, the horizontal axis is the observed value obtained by actually observing the rainwater runoff pollution load data in the area of ​​the corresponding regional category, and the vertical axis is the predicted value calculated according to the prediction model of the corresponding regional category. Each point is the fitting point corresponding to the observed value and the predicted value, and the straight line is the fitting line between the observed value and the predicted value. It can be seen that the predicted values ​​of the prediction models corresponding to the three regional categories fit the observed values ​​well, and the fitting degree of the predicted value and the observed value (R 2 ) is 0.68, and most of the fitting degrees are above 0.75, indicating that the prediction models corresponding to the three regional categories are generally reasonable.

[0102] FIG5 is a method for predicting regional rainwater runoff pollution load according to an embodiment of the present application. As shown in FIG5 , the method for predicting regional rainwater runoff pollution load may include steps S510 to S530, which are described in detail as follows:

[0103] Step S510: Acquire target feature data of the target area.

[0104] In this embodiment, if it is necessary to predict the regional rainwater runoff pollution load of the target area, the target characteristic data of the target area can be obtained. The target characteristic data includes characteristic data representing the attribute data category and characteristic data representing the meteorological data category, and the specific reference can be made to the historical characteristic data.

[0105] Of course, in some embodiments, if the target area is the area in FIG. 2 , the historical feature data of the corresponding area may be directly selected for prediction.

[0106] It should be noted that, in some embodiments, the target characteristic data may refer to historical characteristic data. In other embodiments, since the rainwater runoff pollution load data of the target area is predicted, unlike the historical characteristic data, the precipitation index-related characteristics and underlying surface data such as precipitation, precipitation duration, average rainfall intensity and pre-rain dry period of the target area may not be obtained. In other words, the target characteristic data may only include data such as average slope, multi-year precipitation, multi-year average temperature, and economic level.

[0107] Step S530: Obtain preset feature categories based on the target feature data to filter the target area feature data based on the preset feature categories, and input the filtered target feature data into a pre-trained prediction model to obtain rainwater runoff pollution load data of the target area output by the prediction model.

[0108] In this embodiment, the preset feature categories are obtained based on correlation values ​​between historical rainwater runoff pollution load data and feature data corresponding to each feature category in the historical feature data.

[0109] Specifically, the prediction models include multiple ones, different prediction models correspond to different regional categories, and different regional categories correspond to different preset feature categories. The preset feature category can be obtained by recording the corresponding regional category when screening the feature data to be trained, that is, when obtaining the feature data to be trained through the correlation value, the feature category corresponding to the feature data to be trained is used as the preset feature category corresponding to the regional category.

[0110] In a specific embodiment, step S530 may include steps S41 to S43, specifically:

[0111] Step S41: classifying the target area based on the target feature data to obtain a target area category of the target area.

[0112] In this embodiment, the target area is classified based on the area category corresponding to the existing prediction model to determine the target area category of the target area. Specifically, principal coordinate analysis is performed based on the historical feature data and target feature data of the area where the prediction model is trained in Figure 2 to determine the target category to which the target area belongs in the existing area category.

[0113] Principal coordinate analysis can be performed using software such as RStudio. The distance metric is selected based on the clustering effect, and a numerical standardization is performed to reduce the impact caused by the difference in numerical magnitude between variables. Of course, principal coordinate analysis is exemplary, and other classification methods can be used in other embodiments, which are not specifically limited here.

[0114] After classification, analysis can be performed based on the target feature data and historical feature data in the same target area category to determine whether the classification of the target area is reasonable.

[0115] In other embodiments, the classification of the target area may not be restricted to existing area categories. If the distance between the target area and the existing area categories is too large, other areas other than the area in Figure 2 may be obtained, and prediction models of other area categories may be trained based on the steps in Figure 2 until the target area category of the target area and the prediction model corresponding to the target area category are obtained.

[0116] Of course, the classification of the target area here means that the target area is not the area used when the prediction model is trained. If the target area is the area used when the prediction model is trained, the area category and data of the target area during the prediction model training can be directly based on the area category and data of the target area during the prediction model training.

[0117] Step S42: Based on the preset feature categories in the target area category, feature data corresponding to the feature categories are filtered in the target area data to obtain filtered target area feature data.

[0118] The prediction feature category in this embodiment is the feature category of the feature data to be trained in the corresponding regional category. For a regional category, the feature data of the corresponding prediction feature category has a high correlation with the rainwater runoff pollution load data in the area of ​​the corresponding regional category. In this way, the target feature data can be filtered through the preset feature category, reducing the data processing volume of the rainwater runoff pollution load prediction in the target area and improving the prediction efficiency.

[0119] Of course, in some embodiments, the target feature data may not contain feature data of the preset feature category. In this case, it is sufficient to only screen the target feature data for feature data of the predicted feature category.

[0120] Step S43: inputting the filtered target area feature data into the trained model corresponding to the target area category to obtain the rainwater runoff pollution load data of the target area.

[0121] The prediction of the rainwater runoff pollution load in the target area is also based on the training model of the target area category, thereby improving the accuracy of the prediction of the rainwater runoff pollution load in the target area.

[0122] It should be noted that the rainwater runoff pollution load data output by the prediction model in this embodiment is rainwater runoff pollution load data in time units, which is the time unit used when the prediction model is trained. For example, if the time unit is year, the target feature data is the relevant data of the target area in a certain year that needs to be predicted. The output of the corresponding prediction model is the rainwater runoff pollution load data of the target area in that year. If it is necessary to calculate the pollution amount data of the target area in that year, it can be obtained by multiplying the rainwater runoff pollution load data by the rainfall or runoff flow in the target area in that year.

[0123] In this embodiment, based on the three prediction models proposed in the above embodiment (i.e., the prediction models corresponding to the three regional categories A, B and C corresponding to the schematic diagram of the predicted value and observed value fitting results shown in Figure 4), a specific implementation method is proposed here.

[0124] In this embodiment, the target area is D9. Since the target area is the area used in the training process of the prediction models corresponding to the three area categories A, B and C, the prediction model of the corresponding area category is directly selected.

[0125] Of course, this embodiment can also use principal coordinate analysis to determine whether the division of D9 into the three regional categories of A, B, and C is reasonable. The Euclidean distance is selected as the distance metric, and a numerical standardization is performed to reduce the influence caused by the difference in numerical magnitude between the variables. The principal coordinate analysis result shown in Figure 6 is obtained, wherein the area circled by the solid circle is the area of ​​the C regional category, the area circled by the segmented dotted circle is the area of ​​the A regional category, and the area circled by the dotted dotted circle is the area of ​​the B regional category. The variance contribution rate of the first principal coordinate (i.e., the horizontal coordinate in Figure 6) is 51.12%, the variance contribution rate of the second principal coordinate is 31.99% (i.e., the vertical coordinate in Figure 6), and the cumulative variance contribution rate is 83.11%, indicating that the information dimensionality reduction effect is good and the regional categories obtained by this regional classification are relatively reasonable.

[0126] Then, D9 can be selected for field sampling. A total of three sampling points are set up, namely asphalt pavement, cement pavement and green space. Initial rainwater is collected in the rainwater well. For the target characteristic data corresponding to the three sampling points, such as rainfall of 20.4 mm, which is a moderate rain level, average rainfall intensity of 0.028 mm / min, and dry period before rain of 15 days, prediction is made through the prediction model to obtain the predicted value. Water samples are collected at the rainwater well mouth within 0, 3, 6, 9, 12, 15, 20, 25, 30, and 60 minutes after the runoff begins to form on the road surface. The rainwater runoff pollution load data of COD, SS, TN and TP related pollution factors are tested. Then, the single rainwater runoff pollution load data of a single rainfall is calculated according to the following formula, that is, the measured value is obtained:

[0127] Among them, C j is the pollutant concentration measured at the jth time period (mg / L, milligrams per liter); V j is the runoff flow in the jth period (m 3 , cubic meters); n is the number of sampling time periods.

[0128] The comparison results between the measured values ​​and the predicted values ​​are shown in FIG7 . As can be seen from FIG7 , the error between the predicted values ​​of the D9 rainwater runoff pollution load data obtained using the prediction model proposed in this embodiment and the measured values ​​on three different underlying surfaces does not exceed 15%, indicating that the accuracy of this prediction model is relatively high.

[0129] Figure 8 is a prediction model training device shown in an embodiment of the present application. As shown in Figure 8, the prediction model training device 800 may specifically include: a historical data acquisition module 810, configured to acquire historical rainwater runoff pollution load data and historical feature data of a region; wherein the historical feature data includes feature data representing attribute data categories and feature data representing meteorological data categories; a data processing module 830, configured to filter and obtain feature data to be trained in the historical feature data based on the correlation value between the historical rainwater runoff pollution load data and the historical feature data; a model training module 850, configured to train the initial prediction model based on the feature data to be trained and the historical rainwater runoff pollution load data to obtain a prediction model for regional rainwater runoff pollution load prediction.

[0130] In one implementable manner, the historical data acquisition module 810 includes: an underlying surface data acquisition unit, configured to obtain single rainwater runoff pollution load data of the underlying surface corresponding to the underlying surface data; a historical rainwater runoff pollution load data acquisition unit, configured to obtain historical rainwater runoff pollution load data representing the regional rainwater runoff pollution load based on the area of ​​the underlying surface in the underlying surface data and the single rainwater runoff pollution load data.

[0131] In one implementable manner, there are multiple regions; the data processing module 830 includes: a clustering unit, configured to perform clustering processing based on the historical feature data of each region to obtain the region category corresponding to each region; a first data processing unit, configured to screen the historical feature data in each region category based on the correlation value between the historical rainwater runoff pollution load data and the historical feature data in each region category to obtain the feature data to be trained corresponding to each region category, so as to train a prediction model for the corresponding region category based on the feature data to be trained and the historical rainwater runoff pollution load data of each region category.

[0132] In one implementable manner, the historical feature data includes feature data corresponding to multiple feature categories; the data processing module 830 includes: a correlation processing unit, configured to respectively obtain the correlation values ​​between the historical rainwater runoff pollution load data and each feature data; a second data processing unit, configured to screen the feature data corresponding to multiple feature categories based on the size of the correlation value corresponding to each feature data to obtain the feature data to be trained.

[0133] The prediction model training device provided in this embodiment can be used to execute the above-mentioned prediction model training method. Its implementation principle and technical effects are similar and will not be described in detail in this embodiment.

[0134] Figure 9 is a regional rainwater runoff pollution load prediction device shown in an embodiment of the present application. As shown in Figure 9, the regional rainwater runoff pollution load prediction device 900 may specifically include: a target data acquisition module 910, configured to obtain target feature data of the target area; wherein the target data includes feature data representing the attribute data category and feature data representing the meteorological data category; a prediction module 930, configured to obtain preset feature categories based on the target feature data, to filter the target area feature data based on the preset feature categories, and input the filtered target area feature data into a pre-trained prediction model to obtain the rainwater runoff pollution load data of the target area output by the prediction model.

[0135] In one implementable manner, the prediction model includes multiple models, different prediction models correspond to different area categories, and different area categories correspond to different preset feature categories; the prediction module 930 includes: a target area category acquisition unit, configured to classify the target area based on the target feature data to obtain the target area category of the target area; a target area feature data acquisition unit, configured to filter feature data corresponding to the feature category in the target area data based on the preset feature category in the target area category to obtain the filtered target area feature data; a prediction unit, configured to input the filtered target area feature data into the trained training model corresponding to the target area category to obtain the rainwater runoff pollution load data of the target area.

[0136] The regional rainwater runoff pollution load prediction device provided in this embodiment can be used to execute the above-mentioned regional rainwater runoff pollution load prediction method. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.

[0137] Figure 10 is a block diagram of an electronic device according to an exemplary embodiment. Please refer to Figure 10. The electronic device 1000 may include: a processor 101 and a memory 102, wherein the processor 101 and the memory 102 can communicate; exemplarily, the processor 101 and the memory 102 communicate through a communication bus 103, the memory 102 is used to store instructions, and the processor 101 is used to call the instructions in the memory to execute the prediction model training method or the regional rainwater runoff pollution load prediction method shown in any of the above method embodiments.

[0138] The processor may be a central processing unit (CPU), or other general-purpose processor, a digital signal processor (DSP), or an application-specific integrated circuit (ASIC). The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in this application may be directly implemented as being executed by a hardware processor, or may be implemented by a combination of hardware and software modules in the processor.

[0139] The present application provides a computer-readable storage medium having computer-executable instructions stored thereon; when the computer-executable instructions are executed by a processor, they are used to implement a prediction model training method or a regional rainwater runoff pollution load prediction method as described in any of the above embodiments.

[0140] An embodiment of the present application provides a computer program product, which includes instructions. When the instructions are executed, the computer executes the above-mentioned prediction model training method or regional rainwater runoff pollution load prediction method.

[0141] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the application disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, and the true scope and spirit of the present application are indicated by the following claims.

[0142] It should be understood that the present application is not limited to the exact structure described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A prediction model training method, characterized in that: include: Acquire historical rainwater runoff pollution load data of a region and historical characteristic data of the region; wherein the historical characteristic data includes characteristic data representing attribute data categories and characteristic data representing meteorological data categories; Based on the correlation value between the historical rainwater runoff pollution load data and the historical feature data, filtering the historical feature data to obtain feature data to be trained; An initial prediction model is trained based on the feature data to be trained and the historical rainwater runoff pollution load data to obtain a prediction model for regional rainwater runoff pollution load prediction.

2. The method according to claim 1, characterized in that The historical characteristic data includes underlying surface data of the area; the historical rainwater runoff pollution load data of the acquired area includes: Obtaining single rainwater runoff pollution load data of the underlying surface corresponding to the underlying surface data; Based on the area of ​​the underlying surface in the underlying surface data and the single rainwater runoff pollution load data, historical rainwater runoff pollution load data representing the rainwater runoff pollution load in the area is obtained.

3. The method according to claim 1 or 2, characterized in that The number of the regions is multiple; the feature data to be trained is obtained by screening the historical feature data based on the correlation value between the historical rainwater runoff pollution load data and the historical feature data, including: Perform clustering based on the historical characteristic data of each region to obtain the corresponding regional category of each region; Based on the correlation value between the historical rainwater runoff pollution load data and the historical characteristic data in each regional category, the historical characteristic data in each regional category are screened to obtain the feature data to be trained corresponding to each regional category, and the prediction model of the corresponding regional category is trained based on the feature data to be trained and the historical rainwater runoff pollution load data of each regional category.

4. The method according to any one of claims 1 to 3, characterized in that The historical feature data includes feature data corresponding to a plurality of feature categories; the feature data to be trained is obtained by screening the historical feature data based on the correlation value between the historical rainwater runoff pollution load data and the historical feature data, including: respectively obtaining correlation values ​​between the historical rainwater runoff pollution load data and each characteristic data; Based on the magnitude of the correlation value corresponding to each feature data, the feature data to be trained is obtained by screening the feature data corresponding to the plurality of feature categories.

5. A method for predicting regional rainwater runoff pollution load, characterized in that: include: Acquire target feature data of the target area; wherein the target feature data includes feature data representing the attribute data category and feature data representing the meteorological data category; Based on the target feature data, a preset feature category is obtained to filter the target area feature data based on the preset feature category, and the filtered target feature data is input into a pre-trained prediction model to obtain rainwater runoff pollution load data of the target area output by the prediction model; wherein the preset feature category is obtained based on the correlation value between the historical rainwater runoff pollution load data and the feature data corresponding to each feature category in the historical feature data.

6. The method according to claim 5, characterized in that The prediction models include multiple ones, different prediction models correspond to different area categories, and different area categories correspond to different preset feature categories; the preset feature categories are obtained based on the target feature data, the target area feature data are screened based on the preset feature categories, and the screened target feature data are input into the pre-trained prediction model to obtain the rainwater runoff pollution load data of the target area output by the prediction model, including: classifying the target area based on the target feature data to obtain a target area category of the target area; Based on a preset feature category in the target area category, filtering feature data corresponding to the feature category in the target area data to obtain filtered target area feature data; The filtered target area characteristic data is input into the trained training model corresponding to the target area category to obtain the rainwater runoff pollution load data of the target area.

7. A prediction model training device, characterized in that: include: A historical data acquisition module is configured to acquire historical rainwater runoff pollution load data of a region and historical characteristic data of the region; wherein the historical characteristic data includes characteristic data representing attribute data categories and characteristic data representing meteorological data categories; a data processing module configured to filter the historical feature data to obtain feature data to be trained based on a correlation value between the historical rainwater runoff pollution load data and the historical feature data; The model training module is configured to train an initial prediction model based on the feature data to be trained and the historical rainwater runoff pollution load data to obtain a prediction model for regional rainwater runoff pollution load prediction.

8. A regional rainwater runoff pollution load prediction device, characterized in that: include: A target data acquisition module is configured to acquire target feature data of a target area; wherein the target data includes feature data representing a category of attribute data and feature data representing a category of meteorological data; The prediction module is configured to obtain a preset feature category based on the target feature data, to filter the target area feature data based on the preset feature category, and to input the filtered target area feature data into a pre-trained prediction model to obtain the rainwater runoff pollution load data of the target area output by the prediction model.

9. An electronic device, characterized in that: The electronic device includes: a processor and a memory; the memory is used to store instructions; the processor is used to execute the instructions in the memory, so that the electronic device executes the training method of the prediction model as described in any one of claims 1 to 4 or the regional rainwater runoff pollution load prediction method as described in any one of claims 5 to 6.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the training method of the prediction model as described in any one of claims 1 to 4 or the regional rainwater runoff pollution load prediction method as described in any one of claims 5 to 6.

11. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the method for training a prediction model according to any one of claims 1 to 4 or the method for predicting regional rainwater runoff pollution load according to any one of claims 5 to 6.

Citation Information

Patent Citations

  • Power load prediction method and device

    CN113592138A

  • CNN-LSTM-BP multi-mode air pollutant prediction method in combination with meteorological characteristics

    CN115951014A

  • Runoff pollution load calculation model construction method and runoff pollution load calculation method

    CN116822366A

  • Non-point source pollution load prediction method and device, electronic equipment and readable storage medium

    CN117291296A

  • Training method and regional rainwater runoff pollution load prediction method and device

    CN117951531A

Cited By

  • Intelligent control method and control system for sewage treatment based on Internet of Things

    CN121721952A

  • Method and device for predicting target runoff pollutant load, equipment and medium

    CN122022533A

  • A runoff change attribution analysis method based on underlying surface spatial characteristics

    CN122451371A